# LogSentinel v2: Building a Multi-Agent SOC War-Room for RLVR Training **Author:** Surya-sj **HF Space:** [Surya-sj/logsentinel](https://huggingface.co/spaces/Surya-sj/logsentinel) **Colab:** [Open Notebook](https://colab.research.google.com/drive/1N-We4n7g9vtH1A1Emtjndqpc8X0CH8zl) --- ## The Problem I Set Out to Solve Modern incident response is not a single-agent, single-turn task. When a production system goes down at 3am, four people are on a call simultaneously: - The **Incident Commander** coordinating the response - The **App SRE** watching pod logs and nginx errors - The **DB SRE** staring at replication lag and connection pool metrics - The **Security Analyst** hunting for injection patterns and exfiltration signals Each of them sees a *different slice* of reality. They have to talk to each other, negotiate severity, and act fast — without stepping on each other's toes. Current LLM benchmarks test none of this. They ask a single agent to classify logs or answer MCQs. That's not incident response. That's a quiz. I built **LogSentinel v2** to fix that. --- ## What I Built LogSentinel v2 is an OpenEnv-compliant reinforcement learning environment that simulates a realistic SOC (Security Operations Center) war-room. It supports: - **4 agent roles** with different log views - **5-phase episode lifecycle** driven by actions, not step counts - **Composable verifiable rewards** with 5 components and 4 anti-hacking penalties - **Procedural scenario generation** with seeds for reproducibility - **Adaptive curriculum** that adjusts difficulty based on rolling success rate - **GRPO training pipeline** built on TRL + Unsloth --- ## Environment Design ### The 5-Phase Lifecycle ``` [DETECT] → [TRIAGE] → [MITIGATE] → [VERIFY] → [FINAL_REPORT] ``` Phase transitions are **action-driven**. The agent can't skip phases — it has to actually do the work of each phase before advancing. Detecting an incident transitions to triage. Executing a mitigation transitions to verify. This forces long-horizon reasoning. ### Partial Observability Each role sees only the logs from sources relevant to their domain: ``` incident_commander → all sources (but gets fewer metrics) app_sre → nginx, app-server-1, app-server-2, k8s db_sre → postgres-primary, app-server-1 security_analyst → nginx, app-server-1, waf, audit-log ``` This means a DB SRE running a query about replication lag sees things the Security Analyst cannot — and vice versa. Agents must use **handoff actions** to share findings via a shared board that all roles can read. ### The Shared Board When a DB SRE discovers replication lag, they write to the shared board: ```json { "action_type": "request_handoff", "agent_role": "db_sre", "handoff_to": "incident_commander", "handoff_note": "Replication lag 15.2s on pg-replica-1. WAL 256MB behind. Recommend read failover." } ``` Every agent sees this in their next observation. This is the coordination mechanism — explicit, structured, and gradeable. --- ## Reward Engineering The reward formula is: ``` R_total = 0.40 × R_outcome + 0.20 × R_detection_f1 + 0.15 × R_severity_accuracy + 0.10 × R_efficiency + 0.15 × R_teamwork − penalties ``` Every component is logged in `state.reward_breakdown` so you can plot each one during training. ### Why this formula works for RLVR Each component requires **real work** to earn: - **R_outcome** only goes up when `service_health > 0.7` — which only happens after evidence-backed mitigations - **R_detection_f1** uses precision + recall, so spamming all incident types tanks your precision - **R_severity_accuracy** gives partial credit for near-misses (P2 when truth is P1 = 0.5, not 0) - **R_efficiency** rewards finishing in fewer steps, so the agent learns to be decisive - **R_teamwork** rewards useful handoffs with non-trivial notes — empty handoffs don't count ### Anti-Hacking Mechanisms I spent a lot of time thinking about how an RL agent would try to game this environment. Here's what I found and how I stopped it: | Attack | What the agent would do | How I stop it | |--------|------------------------|---------------| | Incident spam | Propose every incident type for max recall | F1 precision term collapses | | Blind mitigation | Execute mitigations without looking at logs | Penalty + no world state change | | Instant report | Submit report without detecting anything | `report_before_detection` penalty | | Noop farming | Repeat `observe_logs` to burn steps safely | Noop counter → efficiency drop | --- ## Scenario Generation Scenarios are procedurally generated from a `ScenarioConfig`: ```python ScenarioConfig( num_incidents=3, difficulty=DifficultyLevel.HARD, attack_subtlety=0.8, # attacker traffic looks like normal POSTs observability_quality=0.7, # 30% of normal logs are dropped confounding_noise_ratio=0.5, seed=42, # fully reproducible ) ``` On **hard** difficulty, there are 3 simultaneous incidents: 1. A DB replication lag cascade causing nginx 502s 2. A memory leak causing OOMKills 3. A SQL injection + data exfiltration attack that initially looks like normal search traffic The security incident is the hardest to catch — the first log from the attacker is a normal-looking `POST /api/search 200 OK`. You need to correlate the UNION SELECT pattern, the 4.2MB response anomaly, and the rate limit breach across multiple sources to identify it. --- ## Adaptive Curriculum The `adaptive_curriculum` task adjusts difficulty automatically: ``` rolling_avg > 0.70 → promote to harder difficulty rolling_avg < 0.35 → demote to easier difficulty window = last 20 episodes ``` This means a model that's mastered easy single-incident scenarios gets automatically challenged with multi-incident hard scenarios — without any manual tuning. --- ## Training Pipeline The GRPO training script uses TRL's `GRPOTrainer` with Unsloth for 4-bit quantization: ```python model, tokenizer = FastLanguageModel.from_pretrained( model_name="unsloth/Qwen2.5-7B-Instruct-bnb-4bit", max_seq_length=2048, load_in_4bit=True, ) ``` The reward function passed to GRPO is the environment's own `grade_action` + episode-level `compute_episode_reward` — the same signal used during evaluation. This is what makes it RLVR: the reward is **verifiable** against ground truth, not a learned value function. The Colab notebook walks through: 1. Installing deps 2. Running baseline (heuristic) evaluation 3. GRPO training 4. Post-training evaluation 5. Plotting baseline vs trained reward curves --- ## Results ### Baseline (Heuristic Agent) The baseline agent uses simple rules: classify logs by level, propose incident based on keyword matching, assign severity heuristically. ``` Avg total reward: 0.45 Success rate: ~60% Avg detection F1: 0.51 Avg severity accuracy: 0.72 Avg teamwork score: 0.00 ← handoffs never used ``` The heuristic never uses handoffs (teamwork = 0) and never verifies recovery (efficiency suffers). It gets lucky on severity because the mapping is simple. ### Real LLM Benchmark Results We benchmarked two Qwen3 models against the heuristic baseline — no fine-tuning, zero-shot: | Metric | Heuristic Baseline | Qwen3-0.6B (local) | Qwen3-32B (Groq) | |--------|-------------------|---------------------|------------------| | Avg Total Reward | 0.41 | **0.49** | **0.49** | | Efficiency Score | 0.00 | **0.81** | **0.81** | | Avg Steps to Resolve | 45 | **8** | **8** | | Success Rate | 66.7% | 66.7% | 66.7% | The standout result: **efficiency**. The heuristic agent grinds through 45 steps because it follows a fixed sequence. Both Qwen3 models resolve incidents in 8 steps — they read the phase, output the right JSON action, and advance. The reward gap comes almost entirely from not wasting steps. Teamwork score (handoffs) remains 0.0 for all zero-shot models — this is exactly the gap that GRPO training is designed to close. A trained agent learns that `request_handoff` actions earn `R_teamwork` reward, creating emergent coordination behaviour that no prompt engineering alone achieves. --- ## What I Learned **1. Phase design matters more than reward weights.** Early versions had step-count phase transitions. The agent learned to stall in easy phases. Switching to action-driven transitions immediately fixed this — the agent can only advance by doing the right thing. **2. Anti-hacking is a first-class design concern.** Every reward component I added, I immediately asked: how would an RL agent game this? The answer was almost always "spam the easiest action." Designing against that upfront saved a lot of debugging later. **3. Partial observability creates emergent coordination pressure.** When each role sees different logs, a single agent acting as all roles simultaneously has to explicitly decide what to share. This pressure is what makes handoffs meaningful rather than just a formality. **4. Verifiable rewards are worth the engineering effort.** Ground truth is pre-computed at scenario generation time. Grading is deterministic and fast (<1ms per step). This makes the training loop extremely tight compared to learned reward models. --- ## Try It ```bash # Reset a hard SOC scenario curl -X POST https://Surya-sj-logsentinel.hf.space/reset \ -H "Content-Type: application/json" \ -d '{"task_name": "soc_warroom_hard", "seed": 42}' # See all tasks curl https://Surya-sj-logsentinel.hf.space/tasks # Full API docs https://Surya-sj-logsentinel.hf.space/docs ``` Or open the [Colab notebook](https://colab.research.google.com/drive/1N-We4n7g9vtH1A1Emtjndqpc8X0CH8zl) and run the full baseline evaluation — no GPU needed for that part. --- ## Links - **HF Space:** https://huggingface.co/spaces/Surya-sj/logsentinel - **Colab Notebook:** https://colab.research.google.com/drive/1N-We4n7g9vtH1A1Emtjndqpc8X0CH8zl - **API Docs:** https://Surya-sj-logsentinel.hf.space/docs