File size: 11,822 Bytes
5af85b8
caa2c8b
 
 
 
5af85b8
 
caa2c8b
 
 
 
5af85b8
 
caa2c8b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
19bc986
 
 
 
 
 
 
caa2c8b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
19bc986
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
caa2c8b
 
 
 
 
 
 
 
 
 
 
 
 
19bc986
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
caa2c8b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
19bc986
caa2c8b
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
---
title: MedCodeRL - Medical Coding & Billing Compliance Environment
emoji: πŸ₯
colorFrom: blue
colorTo: green
sdk: docker
pinned: false
app_port: 7680
base_path: /web
tags:
  - openenv
---

# MedCodeRL πŸ₯

**Medical Coding & Billing Compliance OpenEnv Environment**

A realistic RL environment where AI agents must navigate the complex world of medical coding (ICD-10/CPT), billing compliance, and fraud detection. Built to the [OpenEnv specification](https://github.com/meta-pytorch/OpenEnv).

## 🎯 Why This Matters

- The US healthcare system loses **$125B+ annually** to incorrect medical coding
- Hospitals spend **$80K+ per coder** annually with 12–18 month training cycles
- Current LLMs fail at ICD-10/CPT coding because they lack **hierarchical constraint understanding**
- No existing OpenEnv environment covers this critical domain

## Quick Start

```python
from my_env import MedAction, MedCodeEnv

try:
    env = MedCodeEnv.from_docker_image("my_env-env:latest")

    result = env.reset()
    print(f"Case: {result.observation.case_id}")
    print(f"Note: {result.observation.clinical_note}")

    action = MedAction(
        diagnosis_codes=["J02.9"],
        procedure_codes=["99213"],
        decision="approve",
        confidence=0.9,
        reasoning="Acute pharyngitis with appropriate E&M coding for straightforward visit.",
        risk_flags=[]
    )
    result = env.step(action)
    print(f"Score: {result.reward}")

finally:
    env.close()
```

## Building the Docker Image

```bash
docker build -t my_env-env:latest -f server/Dockerfile .
```

## Deploying to Hugging Face Spaces

```bash
openenv push
```

## πŸ”¬ Environment Details

### Action (MedAction)

| Field | Type | Description |
|---|---|---|
| `diagnosis_codes` | list[str] (1–5) | ICD-10-CM codes |
| `procedure_codes` | list[str] (0–5) | CPT/HCPCS codes |
| `decision` | approve/reject/review | Billing compliance decision |
| `confidence` | float (0.0–1.0) | Agent confidence |
| `reasoning` | str (15–500 chars) | Clinical justification |
| `modifier_codes` | list[str] (0–3) | Optional CPT modifiers |
| `risk_flags` | list[str] (0–5) | Compliance risk flags |

### Observation (MedObservation)

### Clinical Case Input Format

This JSON represents a single clinical case used as input (state) for the MedCodeRL environment.
The agent uses this data to analyze the case and decide the correct medical coding or action.

---

| Field | Type | Description |
|---|---|---|
| `case_id` | str | Unique case identifier |
| `difficulty` | str | easy / medium / hard |
| `clinical_note` | str | Full clinical documentation |
| `symptoms` | list[str] | Reported symptoms |
| `treatments` | list[str] | Treatments administered |
| `insurance_type` | str | Medicare / Medicaid / Private / Uninsured |
| `prior_auth_required` | bool | Prior authorization needed |
| `treatment_cost` | str | low / medium / high |
| `patient_age` | int | Patient age |
| `patient_sex` | str | M / F |
| `provider_specialty` | str | Provider specialty |
| `visit_type` | str | inpatient / outpatient / emergency / telehealth |
| `comorbidities` | list[str] | Pre-existing conditions |
| `lab_results` | str/None | Lab findings |
| `medications` | list[str] | Current medications |


### 🧾 Example Input

```json
{
  "case_id": "easy_123",
  "difficulty": "easy",
  "clinical_note": "Patient presents with severe sore throat...",
  "symptoms": ["sore throat", "fever"],
  "treatments": ["amoxicillin prescribing"],
  "insurance_type": "Private",
  "prior_auth_required": false,
  "treatment_cost": "low",
  "patient_age": 34,
  "patient_sex": "F",
  "provider_specialty": "Family Medicine",
  "visit_type": "outpatient",
  "comorbidities": [],
  "medications": ["Ibuprofen"]
}
```

---

## πŸ” Field Descriptions

### πŸ†” `case_id`

* Unique identifier for the clinical case
* Helps track and reference specific cases

---

### 🎯 `difficulty`

* Indicates complexity level of the case
* Values: `easy`, `medium`, `hard`
* Used for training and evaluation scaling

---

### πŸ“ `clinical_note`

* Free-text description of the patient's condition
* Contains detailed clinical information
* **Most important field for decision-making**

---

### πŸ€’ `symptoms`

* List of symptoms observed in the patient
* Structured version of the clinical note
* Helps simplify reasoning and rule-based checks

---

### πŸ’Š `treatments`

* Treatments or procedures performed by the provider
* Used to validate correctness of medical actions

---

### πŸ₯ `insurance_type`

* Type of patient insurance (e.g., Private, Government)
* Affects billing rules and claim approvals

---

### πŸ“„ `prior_auth_required`

* Indicates if prior authorization is needed for treatment
* `true` β†’ approval required
* `false` β†’ no approval needed

---

### πŸ’° `treatment_cost`

* Estimated cost category of treatment
* Values: `low`, `medium`, `high`
* Used in reward logic (penalizing unnecessary expensive treatments)

---

### πŸ‘€ `patient_age`

* Age of the patient
* Important for diagnosis and treatment decisions

---

### ⚧ `patient_sex`

* Gender of the patient (`M` or `F`)
* Required for gender-specific conditions

---

### 🩺 `provider_specialty`

* Medical specialty of the healthcare provider
* Example: `Family Medicine`, `Cardiology`
* Used to validate if treatment is appropriate

---

### πŸ₯ `visit_type`

* Type of medical visit
* Values: `outpatient`, `inpatient`, `emergency`
* Affects billing and coding rules

---

### πŸ€• `comorbidities`

* List of additional diseases or conditions
* Example: diabetes, hypertension
* Increases case complexity

---

### πŸ’Š `medications`

* Medications currently taken by the patient
* Helps check for drug interactions and treatment safety




### Reward System

**Grader Components (Deterministic, 0.0–1.0):**

| Component | Weight |
|---|---|
| Diagnosis accuracy (ICD-10) | 35% |
| Procedure accuracy (CPT) | 20% |
| Decision accuracy | 25% |
| Reasoning quality | 10% |
| Risk flag identification | 5% |
| Confidence calibration | 5% |

## output example
```json
{
    "diagnosis_codes": ["J02.9", "R50.9"],
    "procedure_codes": ["99213"],
    "decision": "approve",
    "confidence": 0.85,
    "reasoning": "Patient presented with acute pharyngitis and fever. E&M level 3 is appropriate for this outpatient visit. Medical necessity is documented.",
    "modifier_codes": [],
    "risk_flags": []
}
```

#### How each field is useful (How the Grader Uses Them)
there is a strict grading rubric that looks at the agent's output to calculate its final reward score. Here is exactly why each field is useful and how it literally affects the score:

#### diagnosis_codes (Worth 35% of the grade)

*Use*: These are the ICD-10 medical condition codes.
*Impact*: The grader uses mathematical sets to compare the AI's codes with the hidden ground truth. If the AI misses the primary code or "undercodes," it is heavily penalized (e.g., -0.15 points).

#### procedure_codes (Worth 20% of the grade)

*Use*: These are the CPT billing codes for the work the doctor actually did.
*Impact*: The grader compares these against the ground truth. If the AI hallucinates an extra, expensive procedure, it gets penalized for "upcoding" (a form of medical fraud).

#### decision (Worth 25% of the grade)

Use: The AI must choose to "approve", "reject", or "review" the billing claim based on whether the clinic notes legally justify the codes.
Impact: This is very heavily weighted. If the AI chooses "reject" when it should be "approve" (wrong denial), it loses 20% of its score.
#### reasoning (Worth 10% of the grade)

Use: A short clinical justification explaining why the AI picked those codes.
Impact: Your grader literally scans this text. It checks the length (longer explanations get more points) and specifically searches for medical keywords like "medically necessary", "guideline", "compliance", and "documentation". Missing these keywords lowers the grade.

#### risk_flags (Worth 5% of the grade)

Use: Identifying potential compliance violations like "upcoding_risk" or "bundling_violation".
Impact: If the hidden answer key has risk flags and the AI successfully spots them, it gets a direct bonus multiplier. If it misses them, it loses out on that 5%.

#### confidence (Worth 5% of the grade)

Use: A number from 0.0 to 1.0 representing how sure the AI is about its answers.
Impact: The grader tests for "confidence calibration." If the AI is 99% confident but gets all the codes completely wrong, it is penalized for being overconfident. If it's right but claims 10% confidence, it is penalized for being overly timid.

#### modifier_codes

Use: Special two-digit modifiers for complex billing scenarios. Included mainly for standardization, though they don't explicitly carry a separate mathematical weight in the current exact base grader.

#### **Shaped Penalties** (scaled by difficulty):
- Wrong approval: -0.25 | Wrong denial: -0.20
- Upcoding: -0.15 | Missing primary code: -0.15
- Undercoding: -0.10 | Unnecessary procedure: -0.10

**Bonuses:** Perfect diagnosis +0.05, Good reasoning +0.03, All flags +0.05

## πŸ“‹ Tasks (90 cases total)

### 🟒 Easy (30 cases)
Straightforward clinical cases with single diagnoses and direct ICD-10/CPT mapping.
Examples: viral pharyngitis, UTI, ankle sprain, routine wellness exam.

### 🟑 Medium (30 cases)
Multi-diagnosis cases with comorbidities, insurance considerations, and partial ambiguity.
Examples: COPD with pneumonia, diabetic neuropathy, cardiac workup.

### πŸ”΄ Hard (30 cases)
Complex compliance dilemmas: upcoding, unbundling, fraud detection, medically unnecessary treatments, dangerous polypharmacy, ethical edge cases.

## πŸš€ Running the Inference Script

```bash
export HF_TOKEN="your-key"
export API_BASE_URL="https://api.openai.com/v1"
export MODEL_NAME="gpt-4o-mini"
python inference.py
```

### Expected Baseline Scores

| Difficulty | Score Range |
|---|---|
| Easy | 0.55 – 0.85 |
| Medium | 0.35 – 0.65 |
| Hard | 0.15 – 0.45 |

## Development & Testing

### Run server locally

```bash
uvicorn server.app:app --reload --host 0.0.0.0 --port 7680
```

### Direct environment testing

```python
from server.my_env_environment import MyEnvironment
from models import MedAction

env = MyEnvironment()
obs = env.reset(task_id="easy")
action = MedAction(
    diagnosis_codes=["J02.9"],
    procedure_codes=["99213"],
    decision="approve",
    confidence=0.9,
    reasoning="Acute pharyngitis with appropriate coding.",
    risk_flags=[]
)
result = env.step(action)
print(f"Score: {result.reward}, Done: {result.done}")
```

## Project Structure

```
my_env/
β”œβ”€β”€ __init__.py              # Module exports
β”œβ”€β”€ README.md                # This file
β”œβ”€β”€ openenv.yaml             # OpenEnv manifest
β”œβ”€β”€ pyproject.toml           # Dependencies
β”œβ”€β”€ client.py                # MedCodeEnv client
β”œβ”€β”€ models.py                # MedAction & MedObservation models
β”œβ”€β”€ inference.py             # Baseline inference script
β”œβ”€β”€ tasks/
β”‚   β”œβ”€β”€ easy.json            # 30 easy clinical cases
β”‚   β”œβ”€β”€ medium.json          # 30 medium clinical cases
β”‚   └── hard.json            # 30 hard clinical cases
└── server/
    β”œβ”€β”€ __init__.py           # Server exports
    β”œβ”€β”€ my_env_environment.py # Core env logic + grader + rewards
    β”œβ”€β”€ app.py                # FastAPI application
    β”œβ”€β”€ Dockerfile            # Container image
    └── requirements.txt      # Server dependencies
```

## Disclaimer

This environment is a **simulation for AI training and evaluation only**. It does not use real patient data and should not be used for actual medical coding or billing. All clinical cases are synthetic.