Spaces:
Sleeping
title: CommunityPulse — NGO Multi-Agent Coordination Environment
emoji: 🔄
colorFrom: blue
colorTo: purple
sdk: docker
app_port: 7860
pinned: false
NGO Multi-Agent Coordination Environment
OpenEnv-compatible RL environment for training multi-agent negotiation strategies
Meta PyTorch OpenEnv Hackathon — Theme #1: Multi-Agent Interactions
Links
- GitHub Repository: https://github.com/Sayali5115/communitypulse-env
- Hugging Face Space: https://huggingface.co/spaces/CommunityPulse/communitypulse-env
- Colab Notebook: https://colab.research.google.com/drive/1kIgEIoCZavHsFVhUVfTYTVXbZ8qnoVMD?usp=sharing
- Blog Post: https://huggingface.co/spaces/CommunityPulse/communitypulse-env/blob/main/blog_post.md
The Problem
When disasters strike, multiple NGOs must coordinate limited volunteers and resources. Each organization has different priorities:
- Some prioritize speed (deploy quickly)
- Others prioritize efficiency (conserve resources)
- Others prioritize equity (fair distribution)
The Challenge: How do these organizations negotiate, form coalitions, and make decisions when their goals conflict?
Current coordination is manual and inefficient. We need AI systems that can learn optimal multi-agent negotiation strategies.
The Solution
A proper Reinforcement Learning environment where agents learn through trial and error:
- Not a simulation — Agents start random and gradually discover strategies
- Real learning — Epsilon-greedy exploration + policy gradient updates
- Multi-agent interactions — Cooperation, competition, negotiation, coalition formation
- OpenEnv compatible — Works with Stable-Baselines3, RLlib, CleanRL
How It Works
Environment Design
3 Agents with Different Strategies:
- Cooperative Agent — Maximizes total impact (allocates 60% of resources)
- Competitive Agent — Maximizes individual score (allocates 40% of resources)
- Negotiator Agent — Finds balanced compromises (allocates 50% of resources)
4 Multi-Agent Scenarios:
- Cooperation Task — Agents coordinate to maximize shared benefit
- Competition Task — Agents compete while maintaining efficiency
- Negotiation Task — Agents negotiate Pareto-optimal allocations
- Coalition Task — Agents form strategic alliances dynamically
Observation Space (What Agents See)
{
'urgency': Discrete(10), # How critical is the situation (1-10)
'available_resources': Box(0, 100), # Volunteers available
'people_affected': Box(0, 300), # People needing help
'other_agents_actions': Box(0, 100), # What others did last step
'coalition_status': MultiBinary(3), # Who's in the coalition
'communication_channel': Box(0, 1) # Cooperation signals
}
Action Space (What Agents Do)
Box(3) = [allocation_percentage, cooperation_signal, negotiation_bid]
- Allocation % (0–1): How much of resources to deploy
- Cooperation signal (0–1): Willingness to cooperate (0=compete, 1=cooperate)
- Negotiation bid (0–1): Offer for negotiation
Reward Function
Rewards depend on the task type and agent interactions:
Cooperation Task:
- High reward for coordinated allocations (low variance;
std < 0.15triggers a 3.0 bonus) cooperation_reward = (total_allocation / resources) × urgency × 2.0when within budget- Penalty for over-allocation (multiplied down to
× 1.0)
Competition Task:
- Individual rewards:
individual_impact = alloc × resources × (urgency / 10.0) - Penalty of
−2.0if relative allocation exceeds 1.5× the mean (greediness check) - Nash equilibrium emerges as agents converge to balanced strategies
Negotiation Task:
reward = (fairness × 5.0) + (efficiency × 2.0) + (negotiation_quality × 3.0)- Fairness =
1 / (1 + std(allocations)); rewards equal splits - Efficiency =
min(total_alloc, 1.0) × urgency
Coalition Task:
- When ≥ 2 agents allocate > 50%:
coalition_reward = mean_alloc × resources × urgency × 1.5 - Otherwise falls back to
× 0.8multiplier - Agents learn when forming alliances outperforms going alone
Global Bonuses (all tasks):
progress_bonus = (episode_count / 100) × 2.0— rewards cumulative learningstep_bonus = (1 − step / max_steps) × 1.0— rewards efficient early decisions- Floor:
total_reward = max(total_reward, 1.0)
Learning Mechanism
Real RL Learning (Not Hardcoded!):
Epsilon-Greedy Exploration
- Start: ε = 1.0 (100% random actions)
- End: ε = 0.05 (5% random, 95% learned policy)
- Decay: 0.995 per episode
Policy Gradient Updates
gradient = reward × action × 0.01theta += lr × mean(gradient)- Actions with high rewards are reinforced; low-reward actions are discouraged
Progress Bonus
- Simulates cumulative learning effects
- Later episodes have a higher baseline reward
Colab Notebook
The notebook CommunityPulse_Colab.ipynb mirrors the full training pipeline and runs end-to-end in ~2 minutes with no API key needed.
Cells overview:
| Cell | Contents |
|---|---|
| 1 | Install dependencies (gymnasium, numpy, matplotlib, scipy) |
| 2 | NGOCoordinationEnv — full environment definition (ngo_coordination_env.py) |
| 3 | SimpleRLAgent — epsilon-greedy agent with policy gradient updates (train.py) |
| 4 | train_multi_agent_system() — 100-episode training loop |
| 5 | Execute training (num_episodes=100, max_steps=50, num_agents=3) |
| 6 | plot_learning_curves() — generates all 3 result visualizations inline |
| 7 | Save training_results.json and print summary |
| 8 | Download outputs (Google Colab auto-download) |
To run: Open the notebook in Colab → Runtime → Run all
Training Results
Performance Improvement
Total Episodes: 100
First 10 Episodes Avg: 2353.89
Last 10 Episodes Avg: 5103.71
Improvement: 116.8%
Learning Curves
Overall Learning Progression:
Agents start with random behavior (2300 reward) and gradually learn optimal strategies (5100 reward). The dashed blue trend line confirms consistent upward improvement across all 5,000 steps.
Task-by-Task Performance:
Each of the 4 multi-agent scenarios shows clear learning progression. Cooperation (blue) stabilises quickly; Competition (green) remains noisier due to individual incentives; Negotiation (orange) shows periodic resets as bids are refined; Coalition (red) achieves the highest absolute rewards.
Average Rewards by Task:
Coalition formation tasks achieve by far the highest average rewards (275.48), followed by Negotiation (26.81), Competition (18.65), and Cooperation (7.47), demonstrating emergent strategic behaviour.
Key Findings
- Agents learn cooperation — Variance in allocations decreases over time; coordination bonus triggers more frequently
- Nash equilibrium emerges — In competition tasks, agents converge to balanced strategies avoiding the greediness penalty
- Coalition formation — Agents learn when to form alliances (≥ 2 agents > 50% allocation) vs. act independently
- Theory-of-mind reasoning — Agents adapt behaviour based on
other_agents_actionsin the observation
Quick Start
Installation
pip install -r requirements.txt
Run Training
python train.py
Output:
- 100 episodes in ~30 seconds
- 3 graphs saved to
results/ - Training logs in JSON format
Test Environment
python test_env.py
Verify Submission
python verify_submission.py
Use with Your RL Algorithm
Basic Usage
from ngo_coordination_env import NGOCoordinationEnv
# Create environment
env = NGOCoordinationEnv(num_agents=3, max_steps=50)
# Reset
observation, info = env.reset()
# Step
actions = your_agent.select_action(observation) # Shape: (3, 3)
observation, reward, terminated, truncated, info = env.step(actions)
With Stable-Baselines3
from stable_baselines3 import PPO
from ngo_coordination_env import NGOCoordinationEnv
env = NGOCoordinationEnv()
model = PPO("MultiInputPolicy", env, verbose=1)
model.learn(total_timesteps=10000)
model.save("ngo_coordination_ppo")
With RLlib
from ray.rllib.algorithms.ppo import PPOConfig
from ngo_coordination_env import NGOCoordinationEnv
config = PPOConfig().environment(NGOCoordinationEnv)
algo = config.build()
algo.train()
Project Structure
ngo-coordination-env/
├── ngo_coordination_env.py # Main RL environment (gymnasium.Env)
├── train.py # Training script with simple RL agents
├── CommunityPulse_Colab.ipynb # End-to-end Colab notebook
├── test_env.py # Test script
├── requirements.txt # Dependencies
├── README.md # This file
└── results/ # Generated outputs
├── overall_learning_curve.png
├── task_comparison.png
├── task_progression.png
└── training_results.json
Technical Details
Built with OpenEnv
- Framework: gymnasium 1.3.0 (latest OpenEnv release)
- Proper gym.Env interface:
reset(),step(), observation/action spaces - Compatible with: Stable-Baselines3, RLlib, CleanRL, any PyTorch RL library
Requirements
gymnasium>=0.29.0
numpy>=1.24.0
matplotlib>=3.7.0
scipy>=1.10.0
Performance
- Training time: ~30 seconds (100 episodes via script) / ~2 minutes (Colab notebook)
- Memory usage: <100 MB
- CPU: Works on 2 cores
- Total steps: 5,000 (100 episodes × 50 steps)
Real-World Applications
This environment can train AI systems for:
- Disaster Response Coordination — Multiple NGOs coordinating relief efforts
- Hospital Resource Allocation — Hospitals sharing medical supplies during crises
- Budget Negotiations — Departments negotiating limited budgets
- Supply Chain Partnerships — Companies forming logistics coalitions
- LLM Multi-Agent Training — Teaching language models to handle multi-agent scenarios
Citation
@misc{ngo_coordination_env_2025,
title = {NGO Multi-Agent Coordination Environment},
author = {Meta PyTorch OpenEnv Hackathon Submission},
year = {2025},
note = {Theme #1: Multi-Agent Interactions}
% url = https://huggingface.co/spaces/CommunityPulse/communitypulse-env
}


