Spaces:
Sleeping
Sleeping
| # LogSentinel v2: Building a Multi-Agent SOC War-Room for RLVR Training | |
| **Author:** Surya-sj | |
| **HF Space:** [Surya-sj/logsentinel](https://huggingface.co/spaces/Surya-sj/logsentinel) | |
| **Colab:** [Open Notebook](https://colab.research.google.com/drive/1N-We4n7g9vtH1A1Emtjndqpc8X0CH8zl) | |
| --- | |
| ## The Problem I Set Out to Solve | |
| Modern incident response is not a single-agent, single-turn task. When a production system goes down at 3am, four people are on a call simultaneously: | |
| - The **Incident Commander** coordinating the response | |
| - The **App SRE** watching pod logs and nginx errors | |
| - The **DB SRE** staring at replication lag and connection pool metrics | |
| - The **Security Analyst** hunting for injection patterns and exfiltration signals | |
| Each of them sees a *different slice* of reality. They have to talk to each other, negotiate severity, and act fast β without stepping on each other's toes. | |
| Current LLM benchmarks test none of this. They ask a single agent to classify logs or answer MCQs. That's not incident response. That's a quiz. | |
| I built **LogSentinel v2** to fix that. | |
| --- | |
| ## What I Built | |
| LogSentinel v2 is an OpenEnv-compliant reinforcement learning environment that simulates a realistic SOC (Security Operations Center) war-room. It supports: | |
| - **4 agent roles** with different log views | |
| - **5-phase episode lifecycle** driven by actions, not step counts | |
| - **Composable verifiable rewards** with 5 components and 4 anti-hacking penalties | |
| - **Procedural scenario generation** with seeds for reproducibility | |
| - **Adaptive curriculum** that adjusts difficulty based on rolling success rate | |
| - **GRPO training pipeline** built on TRL + Unsloth | |
| --- | |
| ## Environment Design | |
| ### The 5-Phase Lifecycle | |
| ``` | |
| [DETECT] β [TRIAGE] β [MITIGATE] β [VERIFY] β [FINAL_REPORT] | |
| ``` | |
| Phase transitions are **action-driven**. The agent can't skip phases β it has to actually do the work of each phase before advancing. Detecting an incident transitions to triage. Executing a mitigation transitions to verify. This forces long-horizon reasoning. | |
| ### Partial Observability | |
| Each role sees only the logs from sources relevant to their domain: | |
| ``` | |
| incident_commander β all sources (but gets fewer metrics) | |
| app_sre β nginx, app-server-1, app-server-2, k8s | |
| db_sre β postgres-primary, app-server-1 | |
| security_analyst β nginx, app-server-1, waf, audit-log | |
| ``` | |
| This means a DB SRE running a query about replication lag sees things the Security Analyst cannot β and vice versa. Agents must use **handoff actions** to share findings via a shared board that all roles can read. | |
| ### The Shared Board | |
| When a DB SRE discovers replication lag, they write to the shared board: | |
| ```json | |
| { | |
| "action_type": "request_handoff", | |
| "agent_role": "db_sre", | |
| "handoff_to": "incident_commander", | |
| "handoff_note": "Replication lag 15.2s on pg-replica-1. WAL 256MB behind. Recommend read failover." | |
| } | |
| ``` | |
| Every agent sees this in their next observation. This is the coordination mechanism β explicit, structured, and gradeable. | |
| --- | |
| ## Reward Engineering | |
| The reward formula is: | |
| ``` | |
| R_total = 0.40 Γ R_outcome | |
| + 0.20 Γ R_detection_f1 | |
| + 0.15 Γ R_severity_accuracy | |
| + 0.10 Γ R_efficiency | |
| + 0.15 Γ R_teamwork | |
| β penalties | |
| ``` | |
| Every component is logged in `state.reward_breakdown` so you can plot each one during training. | |
| ### Why this formula works for RLVR | |
| Each component requires **real work** to earn: | |
| - **R_outcome** only goes up when `service_health > 0.7` β which only happens after evidence-backed mitigations | |
| - **R_detection_f1** uses precision + recall, so spamming all incident types tanks your precision | |
| - **R_severity_accuracy** gives partial credit for near-misses (P2 when truth is P1 = 0.5, not 0) | |
| - **R_efficiency** rewards finishing in fewer steps, so the agent learns to be decisive | |
| - **R_teamwork** rewards useful handoffs with non-trivial notes β empty handoffs don't count | |
| ### Anti-Hacking Mechanisms | |
| I spent a lot of time thinking about how an RL agent would try to game this environment. Here's what I found and how I stopped it: | |
| | Attack | What the agent would do | How I stop it | | |
| |--------|------------------------|---------------| | |
| | Incident spam | Propose every incident type for max recall | F1 precision term collapses | | |
| | Blind mitigation | Execute mitigations without looking at logs | Penalty + no world state change | | |
| | Instant report | Submit report without detecting anything | `report_before_detection` penalty | | |
| | Noop farming | Repeat `observe_logs` to burn steps safely | Noop counter β efficiency drop | | |
| --- | |
| ## Scenario Generation | |
| Scenarios are procedurally generated from a `ScenarioConfig`: | |
| ```python | |
| ScenarioConfig( | |
| num_incidents=3, | |
| difficulty=DifficultyLevel.HARD, | |
| attack_subtlety=0.8, # attacker traffic looks like normal POSTs | |
| observability_quality=0.7, # 30% of normal logs are dropped | |
| confounding_noise_ratio=0.5, | |
| seed=42, # fully reproducible | |
| ) | |
| ``` | |
| On **hard** difficulty, there are 3 simultaneous incidents: | |
| 1. A DB replication lag cascade causing nginx 502s | |
| 2. A memory leak causing OOMKills | |
| 3. A SQL injection + data exfiltration attack that initially looks like normal search traffic | |
| The security incident is the hardest to catch β the first log from the attacker is a normal-looking `POST /api/search 200 OK`. You need to correlate the UNION SELECT pattern, the 4.2MB response anomaly, and the rate limit breach across multiple sources to identify it. | |
| --- | |
| ## Adaptive Curriculum | |
| The `adaptive_curriculum` task adjusts difficulty automatically: | |
| ``` | |
| rolling_avg > 0.70 β promote to harder difficulty | |
| rolling_avg < 0.35 β demote to easier difficulty | |
| window = last 20 episodes | |
| ``` | |
| This means a model that's mastered easy single-incident scenarios gets automatically challenged with multi-incident hard scenarios β without any manual tuning. | |
| --- | |
| ## Training Pipeline | |
| The GRPO training script uses TRL's `GRPOTrainer` with Unsloth for 4-bit quantization: | |
| ```python | |
| model, tokenizer = FastLanguageModel.from_pretrained( | |
| model_name="unsloth/Qwen2.5-7B-Instruct-bnb-4bit", | |
| max_seq_length=2048, | |
| load_in_4bit=True, | |
| ) | |
| ``` | |
| The reward function passed to GRPO is the environment's own `grade_action` + episode-level `compute_episode_reward` β the same signal used during evaluation. This is what makes it RLVR: the reward is **verifiable** against ground truth, not a learned value function. | |
| The Colab notebook walks through: | |
| 1. Installing deps | |
| 2. Running baseline (heuristic) evaluation | |
| 3. GRPO training | |
| 4. Post-training evaluation | |
| 5. Plotting baseline vs trained reward curves | |
| --- | |
| ## Results | |
| ### Baseline (Heuristic Agent) | |
| The baseline agent uses simple rules: classify logs by level, propose incident based on keyword matching, assign severity heuristically. | |
| ``` | |
| Avg total reward: 0.45 | |
| Success rate: ~60% | |
| Avg detection F1: 0.51 | |
| Avg severity accuracy: 0.72 | |
| Avg teamwork score: 0.00 β handoffs never used | |
| ``` | |
| The heuristic never uses handoffs (teamwork = 0) and never verifies recovery (efficiency suffers). It gets lucky on severity because the mapping is simple. | |
| ### Real LLM Benchmark Results | |
| We benchmarked two Qwen3 models against the heuristic baseline β no fine-tuning, zero-shot: | |
| | Metric | Heuristic Baseline | Qwen3-0.6B (local) | Qwen3-32B (Groq) | | |
| |--------|-------------------|---------------------|------------------| | |
| | Avg Total Reward | 0.41 | **0.49** | **0.49** | | |
| | Efficiency Score | 0.00 | **0.81** | **0.81** | | |
| | Avg Steps to Resolve | 45 | **8** | **8** | | |
| | Success Rate | 66.7% | 66.7% | 66.7% | | |
| The standout result: **efficiency**. The heuristic agent grinds through 45 steps because it follows a fixed sequence. Both Qwen3 models resolve incidents in 8 steps β they read the phase, output the right JSON action, and advance. The reward gap comes almost entirely from not wasting steps. | |
| Teamwork score (handoffs) remains 0.0 for all zero-shot models β this is exactly the gap that GRPO training is designed to close. A trained agent learns that `request_handoff` actions earn `R_teamwork` reward, creating emergent coordination behaviour that no prompt engineering alone achieves. | |
| --- | |
| ## What I Learned | |
| **1. Phase design matters more than reward weights.** | |
| Early versions had step-count phase transitions. The agent learned to stall in easy phases. Switching to action-driven transitions immediately fixed this β the agent can only advance by doing the right thing. | |
| **2. Anti-hacking is a first-class design concern.** | |
| Every reward component I added, I immediately asked: how would an RL agent game this? The answer was almost always "spam the easiest action." Designing against that upfront saved a lot of debugging later. | |
| **3. Partial observability creates emergent coordination pressure.** | |
| When each role sees different logs, a single agent acting as all roles simultaneously has to explicitly decide what to share. This pressure is what makes handoffs meaningful rather than just a formality. | |
| **4. Verifiable rewards are worth the engineering effort.** | |
| Ground truth is pre-computed at scenario generation time. Grading is deterministic and fast (<1ms per step). This makes the training loop extremely tight compared to learned reward models. | |
| --- | |
| ## Try It | |
| ```bash | |
| # Reset a hard SOC scenario | |
| curl -X POST https://Surya-sj-logsentinel.hf.space/reset \ | |
| -H "Content-Type: application/json" \ | |
| -d '{"task_name": "soc_warroom_hard", "seed": 42}' | |
| # See all tasks | |
| curl https://Surya-sj-logsentinel.hf.space/tasks | |
| # Full API docs | |
| https://Surya-sj-logsentinel.hf.space/docs | |
| ``` | |
| Or open the [Colab notebook](https://colab.research.google.com/drive/1N-We4n7g9vtH1A1Emtjndqpc8X0CH8zl) and run the full baseline evaluation β no GPU needed for that part. | |
| --- | |
| ## Links | |
| - **HF Space:** https://huggingface.co/spaces/Surya-sj/logsentinel | |
| - **Colab Notebook:** https://colab.research.google.com/drive/1N-We4n7g9vtH1A1Emtjndqpc8X0CH8zl | |
| - **API Docs:** https://Surya-sj-logsentinel.hf.space/docs | |