File size: 25,777 Bytes
e04abc7
234641b
6daf142
e04abc7
6daf142
e04abc7
6daf142
 
 
e04abc7
 
 
1607c63
6daf142
1607c63
 
6daf142
234641b
 
912886b
 
 
 
 
 
1607c63
6daf142
8d396e9
 
 
234641b
1607c63
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
c5afe7c
 
 
 
 
 
 
 
 
d40c62f
c5afe7c
 
 
 
 
1607c63
 
 
 
 
 
 
6daf142
 
1607c63
 
 
 
 
6daf142
 
 
1607c63
6daf142
1607c63
 
 
 
 
 
 
 
 
 
 
 
 
6daf142
 
1607c63
38d66e6
1607c63
 
 
 
 
 
 
 
 
 
 
 
6daf142
 
1607c63
 
 
 
 
6daf142
 
 
de57d9d
 
 
 
 
 
6daf142
1607c63
 
 
 
 
 
 
 
 
 
 
 
 
6daf142
 
de57d9d
 
6daf142
de57d9d
 
 
 
6daf142
 
 
 
1607c63
 
 
6daf142
 
 
1607c63
de57d9d
 
6daf142
 
 
1607c63
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
6daf142
 
 
 
 
 
 
de57d9d
6daf142
de57d9d
6daf142
de57d9d
6daf142
de57d9d
6daf142
 
 
 
 
1607c63
 
 
6daf142
 
 
de57d9d
 
 
 
 
 
6daf142
1607c63
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
38d66e6
1607c63
 
 
 
6daf142
29fef1e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
c1e9eef
 
6daf142
 
 
 
 
 
 
 
 
 
 
29fef1e
 
6daf142
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
6193264
1607c63
234641b
6193264
1607c63
38d66e6
1607c63
 
38d66e6
6193264
38d66e6
6193264
1607c63
 
 
6193264
234641b
6193264
 
 
 
 
 
 
 
 
 
 
 
234641b
6193264
234641b
 
 
6193264
234641b
6193264
 
 
 
 
 
 
 
 
38d66e6
d40c62f
1607c63
2ce0305
234641b
1607c63
234641b
1607c63
 
 
 
 
 
 
 
 
 
 
 
6daf142
1607c63
6daf142
1607c63
 
38d66e6
 
 
234641b
2ce0305
6193264
38d66e6
 
8d396e9
38d66e6
6193264
 
6daf142
38d66e6
 
 
 
 
 
 
 
 
 
6daf142
 
 
 
 
de57d9d
 
1607c63
de57d9d
1607c63
de57d9d
 
1607c63
 
 
6daf142
de57d9d
1607c63
 
 
 
 
 
 
6daf142
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
---
title: Enterprise Contract Guardian
emoji: πŸ“‹
colorFrom: blue
colorTo: green
sdk: docker
app_port: 7860
tags:
  - openenv
pinned: false
---

# Enterprise Contract Guardian β€” OpenEnv Environment

> **Meta PyTorch OpenEnv Hackathon Γ— Scaler School of Technology β€” Grand Finale Submission**
> **Theme #3.1**: World Modeling β†’ Professional Tasks Β· ⭐ **Scaler AI Labs bonus track**: Multi-App RL Environment for Enterprise Workflows

> πŸ“› **Naming**: **Enterprise Contract Guardian** is the product name. The codename in URLs and source paths is `api-contract-validator` (HF Space slug, Python package, Hub adapter repo). Both refer to the same artifact β€” the HF Spaces title now reflects the product name; the URLs were not changed because they would invalidate every linked artefact.

> πŸ”— **Two URLs explained** β€” for judges:
> - **Repo URL** (browse files, README, notebook, code, results): https://huggingface.co/spaces/pushpam14/api-contract-validator
> - **App endpoint** (the running OpenEnv server, used by `/reset`, `/step`, `/state`): https://pushpam14-api-contract-validator.hf.space
>
> They are the same Space β€” HuggingFace exposes the file browser at `huggingface.co/spaces/<user>/<name>` and the running app at `<user>-<name>.hf.space`. **For automated graders**, point at the app endpoint. **For browsing the codebase, notebook, or training artefacts**, use the repo URL. All training artefacts (`results/reward_curve.png`, `results/before_after.png`, `trained_scores.json`, the Colab notebook, the full training log, the trained LoRA adapter at [`pushpam14/api-contract-validator-grpo-7b`](https://huggingface.co/pushpam14/api-contract-validator-grpo-7b)) are committed to **both** the GitHub repo and the HF Space repo.

An OpenEnv RL environment that trains agents to do what senior platform engineers do when an API breaks in production: **detect the violation, trace which downstream services are affected, propose a backward-compatible fix, and verify the fix doesn't cascade**.

> πŸ“– **Read these first**:
> - [`ENTERPRISE_CONTRACT_GUARDIAN_STORY.md`](ENTERPRISE_CONTRACT_GUARDIAN_STORY.md) β€” product narrative + 20-section walkthrough, mermaid diagrams, two worked incident examples
> - [`TECHNICAL_ARCHITECTURE.md`](TECHNICAL_ARCHITECTURE.md) β€” every tool documented, build flow, runtime sequence diagrams, why-each-decision interview Q&A, reproducibility checklist

## The Story

> An engineer ships a "small" change to the Users API on Friday evening. It passes local tests. On Monday, **four downstream teams break** β€” the Orders service, the Billing pipeline, the Notification worker, and the Analytics ETL. The root cause: a single field renamed in one spec, with no awareness of who consumed it.
>
> This environment teaches agents the full workflow β€” not just "find the bug," but **reason about blast radius, propose fixes that preserve compatibility, and verify the migration across every consumer.**

## Why This Environment Matters (Theme #3.1 Alignment)

Per `themes.md` Theme #3.1: *"environments that require real interaction with tools, APIs, or dynamic systems where the model is expected to do real hard work instead of exploiting short-cuts."*

- βœ… **Real tools/APIs**: OpenAPI specs, payloads, consumer service graphs
- βœ… **Partially observable world**: agent discovers the consumer graph through queries
- βœ… **Persistent state**: violations found, consumers traced, fixes proposed build up across steps
- βœ… **Multi-step orchestration**: `detect β†’ trace β†’ propose β†’ validate`
- βœ… **Enterprise workflow nuance**: versioning, deprecation, backward compatibility
- βœ… **Verifiable reward**: every step has a deterministic, objective grader

## What Makes This Submission Strong

This is not a static prompt benchmark. It is a runnable OpenEnv environment with hidden ground truth, stateful episodes, objective rewards, real training, and public proof artifacts.

| Area | What is included |
|---|---|
| **Environment depth** | 9 tasks across detection, downstream impact tracing, and backward-compatible fix verification |
| **Episode/data mix** | Seeded synthetic enterprise API scenarios: OpenAPI specs, payloads, version diffs, consumer service graphs, and migration candidates |
| **Reward richness** | 14 independent reward signals covering correct findings, proximity, duplicates, false positives, missed consumers, malformed patches, broken consumers, and anti-spam |
| **Training evidence** | 300 GRPO steps on Qwen2.5-7B + LoRA, public WandB report, reward curve, training state JSON, full logs, and trained adapter on Hugging Face |
| **Before/after evaluation** | Three-way comparison: untrained Qwen2.5-72B, untrained Qwen2.5-7B, and trained Qwen2.5-7B + LoRA |
| **Headline improvement** | `detect_breaking_changes`: 0.01 β†’ 0.67 for the same 7B base model after GRPO |

The environment uses generated, deterministic scenarios rather than a scraped external dataset. That is intentional: every episode has known ground truth, which makes the reward signal auditable and lets judges reproduce the same task with a fixed `seed`.

## Architecture: Phase 1 β†’ Phase 2 β†’ Phase 3

| Phase | What the agent does | Task examples |
|---|---|---|
| **Phase 1 β€” Detection** (inherited from Round 1) | Read one OpenAPI spec + payload, report violations | `find_type_mismatches`, `validate_nested_objects`, `detect_breaking_changes` |
| **Phase 2 β€” Impact Tracing** | Given a detected breaking change, identify all downstream consumers whose contracts are violated | `trace_downstream_blast_radius` |
| **Phase 3 β€” Fix & Verify** | Propose a backward-compatible migration; verify against every consumer spec | `propose_backward_compat_fix`, `multi_service_cascade_fix` |

**Real-world applications:**
- CI/CD contract gate that blocks a PR with predicted downstream impact
- Automated migration-plan generator for API versioning
- Enterprise API gateway pre-deployment safety check
- SDK compatibility auditor across microservices
- OAuth2/auth schema change impact analysis

## How It Works

Each episode places the agent inside a **simulated enterprise** with 3–5 microservices, each owning an OpenAPI spec and declaring which other services consume it.

```
Enterprise Service Graph
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ UsersService β”‚ ─────▢ β”‚ OrdersService β”‚ ─────▢ β”‚BillingServiceβ”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜        β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜        β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
       β”‚                        β”‚
       β–Ό                        β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ NotificationsSvc  β”‚    β”‚  AnalyticsETL    β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
```

### Episode Flow (Phase 2/3 tasks)

```
reset()
  β†’  Agent receives: a changed spec (producer) + service graph with consumer declarations.

Phase 1 β€” Detection
  step(violation_report)    β†’  Correct? +1.0 | Proximity +0.3 | False positive -0.3 | Duplicate -0.1

Phase 2 β€” Impact Tracing
  step(trace_impact)        β†’  For each consumer correctly flagged: +reward; missed consumer: penalty

Phase 3 β€” Fix Proposal
  step(propose_fix)         β†’  Fix validates against ALL consumers: +big reward | breaks β‰₯1 consumer: penalty
  step(validate_fix)        β†’  Deterministic cross-spec check confirms/rejects the fix

step(DONE)  β†’  Completeness bonus = 0.5 Γ— (correct_violations / total) Γ— (consumers_traced / total) Γ— fix_valid
```

Phase 1 tasks retain the simple single-spec flow (used as curriculum starters β€” `help_guide.md Β§6`).

## Tasks

### Phase 1 β€” Detection (curriculum starters, inherited from Round 1)

| Task | Difficulty | Violations | Max Steps | What the Agent Must Find |
|------|-----------|------------|-----------|--------------------------|
| `find_type_mismatches` | Easy | 4 | 10 | Type mismatches, missing required fields, invalid enums at the top level. Sampled from a pool of 12 β€” 495 unique episode combinations |
| `validate_nested_objects` | Medium | 7 | 15 | Violations inside nested objects and arrays β€” requires traversing deep structures. 2 variants: Order Service / Event Booking |
| `detect_breaking_changes` | Hard | 9 | 20 | Breaking changes between two API spec versions β€” type changes, removed fields, narrowed enums, new required fields |
| `validate_response_schema` | Expert | 10 | 25 | Subtle format errors in an API response: invalid date formats, pattern mismatches, out-of-range numerics, bad enum values. 2 variants |
| `validate_cross_field_constraints` | Expert | 7 | 18 | Cross-field arithmetic and date ordering on Invoice API β€” line totals, subtotal sum, tax calculation, discount rules for trial accounts |
| `validate_auth_request` | Expert | 6 | 14 | OAuth2 token and API key management violations β€” invalid grant types, bad scopes, MFA token patterns, IP format, rate limits. 2 variants |

### Phase 2 β€” Impact Tracing (finale β€” multi-service)

| Task | Difficulty | Max Steps | What the Agent Must Do |
|---|---|---|---|
| `trace_downstream_blast_radius` | Hard | 20 | Given a breaking change in a producer spec + a consumer service graph, identify every downstream service whose contract is violated. Graded on precision + recall against ground-truth consumer impact. |

### Phase 3 β€” Fix & Verify (finale β€” full workflow)

| Task | Difficulty | Max Steps | What the Agent Must Do |
|---|---|---|---|
| `propose_backward_compat_fix` | Expert | 25 | Given a detected breaking change, propose a migration (aliasing, deprecation, version bump). Graded by whether the fix validates against all consumer specs. |
| `multi_service_cascade_fix` | Expert | 40 | Full workflow: `detect β†’ trace β†’ propose β†’ validate` in one episode, across 3–5 services. Sparse reward with per-phase sub-rewards. |

### Randomised Episode Generation

All tasks support seed-based randomisation, making the environment suitable for
**training** (varied seeds) as well as **evaluation** (fixed seeds):

- `find_type_mismatches` β€” samples 4 from a pool of 12 violations (495 unique combinations)
- `validate_nested_objects` β€” 2 complete scenario variants (Order Service / Event Booking)
- `validate_response_schema` β€” 2 complete scenario variants with different violation sets
- `validate_auth_request` β€” 2 complete scenario variants (OAuth2 / API key management)
- Pass `seed` in the `reset()` call to select a deterministic episode

## Action Space

Each step the agent submits a `ValidatorAction`. The action type it sends depends on the current episode phase.

### Detection actions (Phase 1)

| Field | Type | Description |
|-------|------|-------------|
| `action_type` | `str` | `report_violation` |
| `field_path` | `str` | Dot-notation path to the violated field (e.g. `customer.email`, `items[1].quantity`). Special values: `DONE` to end episode, `HINT` for a location clue |
| `violation_type` | `str` | One of: `type_mismatch`, `missing_required`, `invalid_enum`, `format_error`, `extra_field`, `breaking_change`, `cross_field_constraint` |
| `description` | `str` | Human-readable explanation of the violation |
| `suggested_fix` | `str` | Optional suggested correction |

### Impact tracing actions (Phase 2 β€” finale)

| Field | Type | Description |
|---|---|---|
| `action_type` | `str` | `trace_impact` |
| `affected_services` | `list[str]` | Names of downstream services the agent believes are impacted |
| `reasoning` | `str` | Brief justification for each entry |

### Fix-proposal actions (Phase 3 β€” finale)

| Field | Type | Description |
|---|---|---|
| `action_type` | `str` | `propose_fix` or `validate_fix` |
| `fix_strategy` | `str` | One of: `field_alias`, `version_bump`, `deprecation_window`, `dual_write`, `consumer_patch` |
| `spec_patch` | `dict` | JSON patch to apply to the producer spec |
| `rationale` | `str` | Why this preserves backward compatibility |

## Observation Space

After each step the agent receives a `ValidatorObservation`:

| Field | Type | Description |
|-------|------|-------------|
| `task_name` | `str` | Current task identifier |
| `task_description` | `str` | Natural-language instructions for the agent |
| `api_spec` | `dict` | The OpenAPI specification (or version diff for hard task) |
| `payload` | `dict` | The API payload to validate |
| `violations_found` | `list[dict]` | Violations correctly identified so far |
| `violations_remaining` | `int` | Number of planted violations still undetected |
| `feedback` | `str` | Result of the last submitted report |
| `max_steps` | `int` | Step budget for this episode |
| `done` | `bool` | Whether the episode has ended |
| `reward` | `float` | Reward for the last action |

## Reward Function

Multiple **independent** reward signals (per `help_guide.md Β§7`) β€” reduces reward-hacking risk, provides rich training signal.

### Detection rewards (Phase 1)

| Event | Reward | Rationale |
|-------|--------|-----------|
| Correct violation (path + type match) | **+1.0** | Primary incentive |
| Proximity match (right path, wrong type) | **+0.3** | Encourages finding the right field first |
| HINT requested | **βˆ’0.5** | Informative but expensive |
| Duplicate report | **βˆ’0.1** | Light penalty β€” track what you already found |
| False positive | **βˆ’0.3** | Penalises guessing |
| DONE signal | **+0.5 Γ— (found/total)** | Completeness bonus |

### Impact-tracing rewards (Phase 2)

| Event | Reward | Rationale |
|---|---|---|
| Correctly identified affected consumer | **+0.8** | Reward recall |
| Missed affected consumer | **βˆ’0.5** | Penalise under-reporting |
| False-flag unaffected consumer | **βˆ’0.4** | Penalise over-reporting |

### Fix-proposal rewards (Phase 3)

| Event | Reward | Rationale |
|---|---|---|
| Fix validates against ALL consumers | **+2.0** | Major incentive β€” this is the goal |
| Fix breaks 1+ consumer | **βˆ’1.0** | Must be backward compatible |
| Malformed spec patch | **βˆ’0.5** | Format compliance |
| Invalid strategy for this violation class | **βˆ’0.3** | Encourages strategy selection |

### Cross-cutting signals

| Signal | Reward | Rationale (`help_guide.md Β§7`) |
|---|---|---|
| **Step budget** | Hard max-step limit per task | Discourages padding and forces concise analysis |
| **Format compliance** | βˆ’0.2 for malformed actions | Enforces schema |
| **Anti-hacking (spam)** | βˆ’1.0 if > 3Γ— total violations reported | Prevents "report everything" exploit |

**Final episode score** = weighted blend of phase scores; see `server/rewards.py`.

## How To Run And See Output

If you are reviewing this submission, the fastest path is:

1. **Try the hosted environment**:

   ```bash
   curl https://pushpam14-api-contract-validator.hf.space/health

   curl -X POST https://pushpam14-api-contract-validator.hf.space/reset \
        -H "Content-Type: application/json" \
        -d '{"task_name":"trace_downstream_blast_radius","seed":1}'
   ```

2. **Inspect the committed training outputs**:

   - Reward curve: [`results/reward_curve.png`](results/reward_curve.png)
   - Before/after chart: [`results/before_after.png`](results/before_after.png)
   - Training proof: [`results/TRAINING_RUN_PROOF.md`](results/TRAINING_RUN_PROOF.md)
   - Full training log: [`results/training_full_log.txt`](results/training_full_log.txt)
   - Score files: [`../baseline_72b_v2_scores.json`](../baseline_72b_v2_scores.json), [`../baseline_7b_scores.json`](../baseline_7b_scores.json), [`../trained_scores.json`](../trained_scores.json)

3. **Re-run locally if desired**:

   ```bash
   git clone https://github.com/kumarpushpam17-personal/Hackathon
   cd Hackathon/api_contract_validator
   pip install -e .
   uvicorn server.app:app --host 0.0.0.0 --port 7860
   ```

   In another terminal:

   ```bash
   curl http://localhost:7860/health
   openenv validate
   ```

4. **Re-run inference or training**:

   - Baseline/trained inference: [`inference.py`](inference.py)
   - Colab training notebook: [`training/grpo_colab.ipynb`](training/grpo_colab.ipynb)
   - HF Jobs launcher used for the submitted run: [`training/run_in_hf_jobs.py`](training/run_in_hf_jobs.py)
   - Full training instructions: [`training/README.md`](training/README.md)

   If Hugging Face's notebook viewer shows a blank/white page for the `.ipynb`, open the notebook in Google Colab or download it and open with Jupyter. The real training metrics are also committed directly in [`results/reward_curve.png`](results/reward_curve.png), [`results/training_state.json`](results/training_state.json), and [`results/training_full_log.txt`](results/training_full_log.txt), so the notebook viewer is not required to verify the run.

## Setup

### Prerequisites

- Python 3.10+
- Docker (for containerised deployment)
- `openenv-core` (`pip install openenv-core`)

### Local Development

```bash
git clone https://github.com/kumarpushpam17-personal/Hackathon
cd Hackathon/api_contract_validator
pip install -e .

uvicorn server.app:app --host 0.0.0.0 --port 7860 --reload
```

### Docker

```bash
docker build -t api-contract-validator .
docker run -p 7860:7860 api-contract-validator
```

### Run Inference

```bash
export API_BASE_URL="https://router.huggingface.co/v1"
export MODEL_NAME="Qwen/Qwen2.5-72B-Instruct"
export HF_TOKEN="your-token-here"

python inference.py
```

## Validate Submission

```bash
openenv validate
```

## Training Results β€” Three-way comparison (before vs after)

> **Training**: GRPO via TRL + Unsloth Β· **Hardware**: HuggingFace Jobs L4 (24 GB) Β· **Steps**: 300 Β· **Wall-time**: 1 h 56 min
> **Inference**: Same temperature (0.7) for all three columns β€” the comparison is sampling-fair.

### Reward Curve (training progress)

![Reward Curve](results/reward_curve.png)

*Mean episode reward across 300 GRPO training steps on Qwen2.5-7B-Instruct (4-bit + LoRA r=16). The curve plateaus around 1.0–1.5 because the base 7B model already produces structurally valid actions; the per-task score table below is where the *kind* of improvement becomes visible.*

### Apples-to-apples per-task comparison

![Before vs After](results/before_after.png)

*Three-bar chart: Qwen2.5-72B baseline (dark grey), Qwen2.5-7B baseline (light grey, **same base as trained**), Qwen2.5-7B + LoRA after GRPO (green). The headroom task `detect_breaking_changes` is the standout β€” both untrained models score 0.01; the trained adapter scores 0.67.*

| Task | Phase | 72B baseline | **7B baseline** | **7B + LoRA (trained)** | Ξ” vs 7B base |
|---|---|---|---|---|---|
| `find_type_mismatches` | 1 | 0.75 | 0.75 | 0.75 | = |
| `validate_nested_objects` | 1 | 0.99 | 0.57 | 0.57 | = |
| **`detect_breaking_changes`** | 1 | **0.01** | **0.01** | **0.67** | **+0.66** 🎯 |
| `validate_response_schema` | 1 | 0.99 | 0.70 | 0.30 | -0.40 |
| `validate_cross_field_constraints` | 1 | 0.99 | 0.43 | 0.29 | -0.14 |
| `validate_auth_request` | 1 | 0.99 | 0.83 | 0.33 | -0.50 |
| `trace_downstream_blast_radius` | 2 | 0.67 | 0.99 | 0.99 | = |
| `propose_backward_compat_fix` | 3 | 0.99 | 0.99 | 0.99 | = |
| `multi_service_cascade_fix` | 2+3 | 0.99 | 0.99 | 0.99 | = |
| **Mean** |   | 0.82 | 0.70 | 0.65 | -0.05 |

Score files: [`../baseline_72b_v2_scores.json`](../baseline_72b_v2_scores.json) (72B), [`../baseline_7b_scores.json`](../baseline_7b_scores.json) (untrained 7B β€” apples-to-apples baseline), [`../trained_scores.json`](../trained_scores.json) (7B + LoRA after GRPO).

### What the comparison shows

The middle column (untrained Qwen-7B) is the fair baseline β€” same base model as the trained one, no adapter. **Any difference between columns 2 and 3 is purely the GRPO training effect.**

**The headline win**:

> **`detect_breaking_changes` went from 0.01 β†’ 0.67.** Both untrained models β€” *including the 10Γ— larger 72B* β€” scored 0.01 on this task. Both models knew where the breaking changes were (they earned the +0.3 proximity reward repeatedly) but **neither could predict `violation_type='breaking_change'` correctly**. After 300 GRPO steps targeting our environment's reward signal, the 7B+LoRA adapter solves the classification 6 of 9 times. **This is RL training value, not model size.**

**The trade-off (we're being honest here)**:

Three Phase 1 tasks regressed (`validate_response_schema`, `validate_cross_field_constraints`, `validate_auth_request`). Reading the per-step reward trajectories shows the cause: the trained model finds the first 2–3 correct violations confidently, then keeps reporting the same field repeatedly. This is the **classic RL fine-tuning trade-off** β€” GRPO heavily reinforced specific high-reward action patterns from training (Phase 2/3 episodes give +2.0 fix rewards vs Phase 1's +1.0 per violation). With more training data balanced toward Phase 1 tasks plus an explicit "don't repeat" reward signal, this would close.

**Phase 2 / Phase 3 tasks** maintain their scores in the trained adapter β€” the model didn't *forget* multi-service reasoning while learning the breaking-change classification.

| Phase | WandB Report | Notebook |
|---|---|---|
| GRPO main run (Qwen-7B, 300 steps) | https://wandb.ai/pushpamsubscriptions-inn/openenv-contract-guardian/reports/Enterprise-Contract-Guardian-GRPO-training-Qwen-7B-LoRA-300-steps---VmlldzoxNjY3MTAxMA?accessToken=3dhumexjta1umyk04rq6dx47iww4t25utt3j0x7063b7pvzzibp8jah29grhlwpb | [`training/grpo_colab.ipynb`](training/grpo_colab.ipynb) |
| Trained adapter | https://huggingface.co/pushpam14/api-contract-validator-grpo-7b | β€” |

See [`training/README.md`](training/README.md) for the three ways to run the pipeline (HF Jobs / Colab / local).

## Why This Matters

API contract violations are the **#1 cause of production incidents in microservice architectures**. Every platform team deals with this weekly. No existing RL environment teaches agents to reason about multi-service contract impact.

**Who benefits from an agent trained on this environment:**
- Platform / API gateway teams β€” pre-merge contract safety checks
- CI/CD pipelines β€” automated impact analysis before deploy
- API versioning toolchains β€” backward-compat migration planning
- Any engineering org operating β‰₯ 3 microservices

This is a genuinely underexplored domain in RL/LLM training β€” no prior benchmarks exist for multi-service API contract reasoning. A model trained here would be publishable as a research artifact.

## Links

| Resource | URL |
|---|---|
| HuggingFace Space (live env) | https://huggingface.co/spaces/pushpam14/api-contract-validator |
| Live env endpoint | https://pushpam14-api-contract-validator.hf.space |
| Health check | https://pushpam14-api-contract-validator.hf.space/health |
| Trained LoRA adapter | https://huggingface.co/pushpam14/api-contract-validator-grpo-7b |
| WandB training run (300 steps) | https://wandb.ai/pushpamsubscriptions-inn/openenv-contract-guardian/reports/Enterprise-Contract-Guardian-GRPO-training-Qwen-7B-LoRA-300-steps---VmlldzoxNjY3MTAxMA?accessToken=3dhumexjta1umyk04rq6dx47iww4t25utt3j0x7063b7pvzzibp8jah29grhlwpb |
| Training proof + full logs | [`results/TRAINING_RUN_PROOF.md`](results/TRAINING_RUN_PROOF.md) |
| Training Notebook (Colab) | [`training/grpo_colab.ipynb`](training/grpo_colab.ipynb) |
| Story + Technical Guide | [`ENTERPRISE_CONTRACT_GUARDIAN_STORY.md`](ENTERPRISE_CONTRACT_GUARDIAN_STORY.md) |
| Tech Architecture & Build Flow | [`TECHNICAL_ARCHITECTURE.md`](TECHNICAL_ARCHITECTURE.md) |
| GitHub repo | https://github.com/kumarpushpam17-personal/Hackathon |
| HF mini-blog writeup (separate MD in Space) | [`BLOG.md`](BLOG.md) |
| Trained adapter model card | https://huggingface.co/pushpam14/api-contract-validator-grpo-7b |

Quick test:

```bash
curl https://pushpam14-api-contract-validator.hf.space/health
# {"status":"healthy"}

curl -X POST https://pushpam14-api-contract-validator.hf.space/reset \
     -H "Content-Type: application/json" -d '{}'
```

## Project Structure

```
api_contract_validator/
β”œβ”€β”€ openenv.yaml              # OpenEnv manifest
β”œβ”€β”€ pyproject.toml            # Python project metadata
β”œβ”€β”€ Dockerfile                # Container definition
β”œβ”€β”€ inference.py              # Baseline inference script (OpenAI client, phase-aware)
β”œβ”€β”€ README.md                 # This file
β”œβ”€β”€ models.py                 # Pydantic models (Action, Observation, State β€” all 3 phases)
β”œβ”€β”€ client.py                 # WebSocket client (EnvClient subclass)
β”œβ”€β”€ __init__.py               # Package exports
β”œβ”€β”€ results/                  # Training plots (.png) β€” committed, embedded above
β”œβ”€β”€ tests/
β”‚   └── test_environment.py   # 28 tests across all 3 phases
└── server/
    β”œβ”€β”€ app.py                # FastAPI wiring (create_app)
    β”œβ”€β”€ environment.py        # Core environment β€” multi-phase orchestration
    β”œβ”€β”€ logging_setup.py      # Structured JSON episode logging
    β”œβ”€β”€ spec_generator.py     # Phase 1 β€” task scenarios with planted violations
    β”œβ”€β”€ service_graph.py      # Phase 2 β€” simulated enterprise service graph
    β”œβ”€β”€ impact_tracer.py      # Phase 2 β€” ground-truth consumer-impact computation
    β”œβ”€β”€ fix_validator.py      # Phase 3 β€” cross-spec fix verification
    └── rewards.py            # Composable reward rubrics (multi-phase, independent signals)
```

## License

BSD-style β€” see LICENSE file.