logsentinel / docs /demo_script.md
Surya-sj's picture
feat: LogSentinel v2 — Multi-Agent SOC War-Room
f7d58be
|
Raw
History Blame Contribute Delete
4.57 kB
# LogSentinel v2 — Demo Script
## 5-Minute Hackathon Demo
### Setup (1 min)
```bash
# Clone and start server
git clone https://github.com/YOUR_USERNAME/logsentinel
cd logsentinel
pip install -r requirements.txt
uvicorn server:app --host 0.0.0.0 --port 7860
```
### Part 1: Environment Overview (1 min)
```bash
# Check health
curl http://localhost:7860/health
# → {"status": "healthy"}
# List available tasks
curl http://localhost:7860/tasks | python -m json.tool
# Shows: log_classification, incident_detection, full_triage,
# soc_warroom_easy/medium/hard, adaptive_curriculum
```
### Part 2: Multi-Agent Episode (2 min)
**Start a hard SOC war-room scenario:**
```bash
curl -X POST http://localhost:7860/reset \
-H "Content-Type: application/json" \
-d '{"task_name": "soc_warroom_hard", "seed": 42}'
```
*Show the response:*
- Agent role: `incident_commander`
- Phase: `detect`
- Logs are role-filtered (only sees sources relevant to IC role)
- `shared_board` is empty (no handoffs yet)
**DB SRE detects replication lag:**
```bash
curl -X POST http://localhost:7860/step \
-H "Content-Type: application/json" \
-d '{
"action": {
"action_type": "propose_incident",
"agent_role": "db_sre",
"incident_type": "outage",
"evidence_indices": [17, 18, 19]
}
}'
```
**DB SRE hands off to IC:**
```bash
curl -X POST http://localhost:7860/step \
-H "Content-Type: application/json" \
-d '{
"action": {
"action_type": "request_handoff",
"agent_role": "db_sre",
"handoff_to": "incident_commander",
"handoff_note": "Replication lag 15.2s on pg-replica-1. WAL 256MB behind. Recommend read query failover immediately."
}
}'
```
**Show the shared board is updated:**
```bash
curl http://localhost:7860/state | python -m json.tool
# Shows: shared_board with handoff_1 entry
```
**IC assigns severity:**
```bash
curl -X POST http://localhost:7860/step \
-H "Content-Type: application/json" \
-d '{
"action": {
"action_type": "assign_severity",
"agent_role": "incident_commander",
"incident_type": "outage",
"severity": "P1"
}
}'
```
**Execute mitigation:**
```bash
curl -X POST http://localhost:7860/step \
-H "Content-Type: application/json" \
-d '{
"action": {
"action_type": "execute_mitigation",
"agent_role": "app_sre",
"mitigation_id": "failover_read_queries_to_primary",
"evidence_indices": [17, 18, 19, 20]
}
}'
```
**Submit joint report:**
```bash
curl -X POST http://localhost:7860/step \
-H "Content-Type: application/json" \
-d '{
"action": {
"action_type": "submit_joint_report",
"agent_role": "incident_commander",
"report": {
"incidents": [
{"type": "outage", "severity": "P1", "description": "DB replication lag cascade"}
],
"severity": "P1",
"summary": "Cascading outage from DB replication lag. Read queries failed over to primary. Service health restored."
}
}
}'
```
### Part 3: Reward Breakdown (1 min)
```bash
curl http://localhost:7860/state | python -c "
import json, sys
state = json.load(sys.stdin)
rb = state['reward_breakdown']
print('--- REWARD BREAKDOWN ---')
print(f'R_outcome: {rb[\"r_outcome\"]:.3f} (×0.40)')
print(f'R_detection_f1: {rb[\"r_detection_f1\"]:.3f} (×0.20)')
print(f'R_severity_acc: {rb[\"r_severity_accuracy\"]:.3f} (×0.15)')
print(f'R_efficiency: {rb[\"r_efficiency\"]:.3f} (×0.10)')
print(f'R_teamwork: {rb[\"r_teamwork\"]:.3f} (×0.15)')
print(f'Penalties: {rb[\"penalty_spam\"]+rb[\"penalty_unsafe\"]:.3f}')
print(f'TOTAL: {rb[\"total\"]:.3f}')
"
```
### Key Talking Points
1. **Partial observability**: Each agent role sees different logs — DB SRE sees postgres, Security Analyst sees nginx + WAL
2. **Phase lifecycle**: Detect → Triage → Mitigate → Verify → Final Report (not step-count based)
3. **Composable rewards**: 5 components with weights, all logged for RLVR training
4. **Anti-hacking**: Unsafe mitigation (no evidence) gets penalised; spam proposals reduce F1
5. **Curriculum**: `adaptive_curriculum` task auto-adjusts difficulty based on rolling success rate
6. **Deterministic**: Same seed → same episode (fully reproducible)
## Running Tests
```bash
pytest tests/ -v
```
## Running Baseline Eval
```bash
python training/eval_baseline_vs_trained.py \
--mode baseline \
--tasks soc_warroom_easy soc_warroom_medium soc_warroom_hard \
--episodes 50 \
--plot
```