File size: 20,718 Bytes
734d48a
 
 
 
 
 
 
 
 
 
 
d980cf6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2f9aaf8
 
d980cf6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
a87e8f5
d980cf6
a87e8f5
d980cf6
a87e8f5
d980cf6
0c8b295
d980cf6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
a87e8f5
0c8b295
a87e8f5
d980cf6
a87e8f5
d980cf6
a87e8f5
d980cf6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
443c45a
d980cf6
a87e8f5
0c8b295
443c45a
0c8b295
443c45a
d980cf6
 
 
0c8b295
d980cf6
 
 
 
443c45a
0c8b295
443c45a
d980cf6
 
 
443c45a
d980cf6
 
 
443c45a
d980cf6
0c8b295
d980cf6
0c8b295
443c45a
d980cf6
443c45a
0c8b295
443c45a
d980cf6
443c45a
d980cf6
a87e8f5
d980cf6
 
 
a87e8f5
d980cf6
 
 
 
 
 
a87e8f5
d980cf6
443a6c2
0c8b295
 
d980cf6
0c8b295
d980cf6
0c8b295
d980cf6
 
 
 
 
0c8b295
d980cf6
 
 
 
0c8b295
d980cf6
0c8b295
 
 
d980cf6
 
 
 
 
 
 
 
 
 
0c8b295
d980cf6
0c8b295
d980cf6
0c8b295
d980cf6
0c8b295
d980cf6
 
 
 
 
 
0c8b295
 
 
d980cf6
 
 
 
 
 
0c8b295
d980cf6
0c8b295
d980cf6
0c8b295
d980cf6
a87e8f5
d980cf6
 
a87e8f5
d980cf6
 
 
 
 
 
 
 
0c8b295
 
 
d980cf6
 
 
 
 
 
 
 
 
 
0c8b295
d980cf6
0c8b295
d980cf6
0c8b295
d980cf6
0c8b295
d980cf6
 
 
 
0c8b295
d980cf6
 
 
 
 
 
 
 
0c8b295
 
 
d980cf6
 
 
 
 
 
 
 
0c8b295
d980cf6
0c8b295
d980cf6
 
 
 
 
0c8b295
d980cf6
0c8b295
d980cf6
0c8b295
d980cf6
0c8b295
d980cf6
0c8b295
d980cf6
 
 
 
 
 
 
 
 
 
 
 
 
 
0c8b295
d980cf6
0c8b295
 
d980cf6
0c8b295
 
 
d980cf6
 
 
 
 
 
 
0c8b295
d980cf6
0c8b295
d980cf6
0c8b295
d980cf6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
0c8b295
d980cf6
 
0c8b295
d980cf6
0c8b295
d980cf6
0c8b295
d980cf6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
0c8b295
 
 
d980cf6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
0c8b295
d980cf6
0c8b295
d980cf6
 
 
 
 
0c8b295
d980cf6
 
 
 
 
 
0c8b295
 
d980cf6
0c8b295
d980cf6
 
 
 
0c8b295
 
 
 
d980cf6
 
 
 
0c8b295
 
 
 
 
 
d980cf6
443a6c2
d980cf6
a87e8f5
d980cf6
0c8b295
d980cf6
0c8b295
d980cf6
 
 
 
 
0c8b295
d980cf6
 
 
 
 
 
 
 
0c8b295
d980cf6
 
 
 
 
 
 
0c8b295
 
d980cf6
0c8b295
d980cf6
0c8b295
d980cf6
443a6c2
d980cf6
 
 
 
 
 
 
 
 
0c8b295
d980cf6
0c8b295
d980cf6
a87e8f5
d980cf6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
a87e8f5
d980cf6
 
a87e8f5
d980cf6
a87e8f5
d980cf6
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
0c8b295
d980cf6
 
 
 
 
 
 
 
 
 
 
 
 
0c8b295
d980cf6
0c8b295
 
 
d980cf6
 
 
 
 
 
 
 
 
 
 
 
 
0c8b295
d980cf6
0c8b295
d980cf6
0c8b295
d980cf6
0c8b295
d980cf6
 
 
 
 
 
 
0c8b295
d980cf6
 
 
 
 
 
0c8b295
d980cf6
 
 
 
 
 
0c8b295
d980cf6
 
 
 
 
a87e8f5
0c8b295
 
 
d980cf6
0c8b295
d980cf6
 
 
0c8b295
d980cf6
 
 
 
0c8b295
 
 
d980cf6
0c8b295
d980cf6
0c8b295
 
a87e8f5
d980cf6
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
---
title: Data Cleaning OpenEnv
emoji: 🧹
colorFrom: blue
colorTo: green
sdk: docker
sdk_version: "3.10"
app_file: main.py
pinned: false
---

# 🧹 CleanifyAI β€” Data Cleaning OpenEnv

<div align="center">

[![HuggingFace Space](https://img.shields.io/badge/πŸ€—%20HuggingFace-Space-blue)](https://huggingface.co/spaces/cleanify-ai/Data-cleaning)
[![GitHub](https://img.shields.io/badge/GitHub-ReverseCoder1%2FCleanifyAI-black?logo=github)](https://github.com/ReverseCoder1/CleanifyAI)
[![OpenEnv](https://img.shields.io/badge/OpenEnv-Compliant-green)](https://huggingface.co/spaces/cleanify-ai/Data-cleaning)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](LICENSE)
[![Python 3.10](https://img.shields.io/badge/Python-3.10-blue?logo=python)](https://python.org)
[![FastAPI](https://img.shields.io/badge/FastAPI-0.104.1-009688?logo=fastapi)](https://fastapi.tiangolo.com)

**A reinforcement-learning environment where AI agents learn to clean real-world messy datasets β€” step by step.**

*Scaler Γ— OpenEnv Hackathon Submission*

[πŸš€ Live API](https://thorodin103-data-cleaning-openenv.hf.space) Β· [πŸ“– Swagger Docs](https://thorodin103-data-cleaning-openenv.hf.space/docs) Β· [πŸ€— HuggingFace](https://huggingface.co/spaces/cleanify-ai/Data-cleaning)

</div>

---

## πŸ“‹ Table of Contents

- [Overview](#-overview)
- [Project Structure](#-project-structure)
- [Setup & Installation](#-setup--installation)
- [Tasks](#-tasks)
- [Operations Reference](#-operations-reference)
- [Reward & Scoring System](#-reward--scoring-system)
- [API Reference](#-api-reference)
- [Inference Script](#-inference-script)
- [Data Models](#-data-models)
- [Datasets](#-datasets)
- [Baseline Scores](#-baseline-scores)
- [License](#-license)

---

## 🌟 Overview

**CleanifyAI** is a fully OpenEnv-compliant environment that challenges AI agents to autonomously clean messy, real-world datasets through a sequence of structured operations. It mimics professional data engineering pipelines and rewards agents that apply operations in the correct, logical order.

| Property | Value |
|---|---|
| **Environment Name** | `data-cleaning-openenv` |
| **Tasks** | 4 (Easy, Medium, Hard, Expert) |
| **Operations** | 9 (dedup, fill, dtype fix, outlier removal, rename, validate, finish) |
| **Scoring** | Weighted multi-component, strictly in `(0, 1)` |
| **API** | OpenEnv-compliant REST via FastAPI |
| **Framework** | Python 3.10, FastAPI, Pandas, NumPy |
| **Inference** | OpenAI-compatible LLM client |
| **Deployed at** | `https://thorodin103-data-cleaning-openenv.hf.space` |

> ⚠️ **Score Constraint**: The Scaler platform rejects scores of exactly `0.0` or `1.0`. All scoring paths in this codebase clamp strictly to `(0.0001, 0.9999)`.

---

## πŸ“ Project Structure

```
CleanifyAI/
β”‚
β”œβ”€β”€ inference.py                  # πŸ€– LLM agent β€” emits [START]/[STEP]/[END] stdout lines
β”œβ”€β”€ environment.py                # πŸ‹οΈ Core OpenEnv environment & reward computation
β”œβ”€β”€ models.py                     # πŸ“¦ Pydantic models: Action, Observation, Reward, StepResult
β”œβ”€β”€ main.py                       # 🌐 FastAPI server with all REST endpoints
β”œβ”€β”€ Dockerfile                    # 🐳 Python 3.10-slim container, port 7860
β”œβ”€β”€ openenv.yaml                  # πŸ“„ OpenEnv spec manifest
β”œβ”€β”€ pyproject.toml                # πŸ“¦ Python dependency config
β”œβ”€β”€ uv.lock                       # πŸ”’ Locked dependency versions
β”‚
β”œβ”€β”€ datasets/
β”‚   β”œβ”€β”€ task_metadata.json        # βš™οΈ Per-task config (steps, operations, scoring weights)
β”‚   β”œβ”€β”€ easy/
β”‚   β”‚   β”œβ”€β”€ dirty.csv             # πŸ—‘οΈ Employee dataset with duplicates + bad column names
β”‚   β”‚   └── gold.csv              # βœ… Gold standard cleaned version
β”‚   β”œβ”€β”€ medium/
β”‚   β”‚   β”œβ”€β”€ dirty.csv             # πŸ—‘οΈ Customer dataset with missing values + wrong dtypes
β”‚   β”‚   └── gold.csv              # βœ… Gold standard
β”‚   β”œβ”€β”€ hard/
β”‚   β”‚   β”œβ”€β”€ dirty.csv             # πŸ—‘οΈ Orders dataset requiring full pipeline
β”‚   β”‚   └── gold.csv              # βœ… Gold standard
β”‚   └── expert/
β”‚       β”œβ”€β”€ dirty.csv             # πŸ—‘οΈ Sales dataset β€” strict operation order required
β”‚       └── gold.csv              # βœ… Gold standard
β”‚
β”œβ”€β”€ static/
β”‚   └── index.html                # πŸ–₯️ Web UI for interactive exploration
β”‚
└── server/
    └── app.py                    # πŸ”§ Server initialization module
```

---

## πŸš€ Setup & Installation

### Prerequisites

- Python 3.10+
- Docker (for containerized deployment)
- A Hugging Face account (`HF_TOKEN`)
- An OpenAI-compatible API endpoint and model

---

### Local Development

**1. Clone the repository**
```bash
git clone https://github.com/ReverseCoder1/CleanifyAI.git
cd CleanifyAI
```

**2. Install dependencies**
```bash
pip install fastapi==0.104.1 uvicorn==0.24.0 pydantic==2.5.0 \
            pandas==2.1.3 numpy==1.26.2 openai>=2.7.2 \
            pyyaml==6.0.1 python-dotenv==1.0.0
```

**3. Create a `.env` file**
```env
API_BASE_URL=https://api.openai.com/v1
MODEL_NAME=gpt-4o-mini
HF_TOKEN=your_hugging_face_token_here
```

**4. Start the FastAPI server**
```bash
uvicorn main:app --host 0.0.0.0 --port 7860 --reload
```

- **API:** http://localhost:7860
- **Swagger UI:** http://localhost:7860/docs

---

### Docker Deployment

```bash
# Build
docker build -t cleanify-ai .

# Run
docker run -p 7860:7860 \
  -e HF_TOKEN=your_token \
  -e MODEL_NAME=gpt-4o-mini \
  -e API_BASE_URL=https://api.openai.com/v1 \
  cleanify-ai
```

---

### Run the Inference Agent

```bash
python inference.py
```

Runs the LLM agent across all 3 hackathon tasks and streams hackathon-spec log lines to stdout.

---

## 🎯 Tasks

Four progressively complex tasks. The hackathon evaluates **easy**, **medium**, and **hard**. Expert is available for extended benchmarking.

| Task ID | Difficulty | Max Steps | Key Operations | Scoring |
|---|---|---|---|---|
| `easy_dedup_rename` | ⭐ Easy | 10 | `remove_duplicates`, `rename_columns` | dup 50% + schema 50% |
| `medium_missing_dtype` | ⭐⭐ Medium | 15 | `fill_missing_*`, `fix_dtype` | missing 50% + dtype 50% |
| `hard_full_pipeline` | ⭐⭐⭐ Hard | 20 | Full pipeline | 20% Γ— 5 components |
| `expert_sales_pipeline` | ⭐⭐⭐⭐ Expert | 25 | All 9 ops in strict order | Weighted (schema 25%) |

---

### ⭐ Easy β€” `easy_dedup_rename`

**Dataset:** Employee records (`emp_id`, `emp_name`, `dept`, `salary`, `age`)

**Dirty conditions:**
- Duplicate rows
- Column names with spaces and inconsistent casing (`EMP ID`, `DEP T`, `SAL ARY`)

**Optimal sequence:**
```
remove_duplicates β†’ rename_columns β†’ finish
```

**Scoring:** `duplicate_score Γ— 0.5 + schema_score Γ— 0.5`

---

### ⭐⭐ Medium β€” `medium_missing_dtype`

**Dataset:** Customer records (`customer_id`, `age`, `salary`, `gender`, `purchases`, `region`, `joined_date`)

**Dirty conditions:**
- NaN values in `age`, `salary`, `gender`, `region`
- `salary` stored as `object` instead of `float`

**Optimal sequence:**
```
fill_missing_mean  (numeric columns)
fill_missing_mode  (categorical columns)
fix_dtype          β†’ finish
```

**Scoring:** `missing_score Γ— 0.5 + dtype_score Γ— 0.5`

---

### ⭐⭐⭐ Hard β€” `hard_full_pipeline`

**Dataset:** Orders (`order_id`, `product`, `quantity`, `price`, `customer_id`, `status`, `order_date`, `rating`)

**Dirty conditions:**
- Duplicate order entries
- Missing `quantity` and `price` values
- `quantity` stored as object type
- Extreme price outliers

**Optimal sequence:**
```
remove_duplicates β†’ fill_missing_* β†’ fix_dtype β†’ remove_outliers β†’ validate_schema β†’ finish
```

**Scoring:** `duplicate Γ— 0.2 + missing Γ— 0.2 + dtype Γ— 0.2 + outlier Γ— 0.2 + schema Γ— 0.2`

---

### ⭐⭐⭐⭐ Expert β€” `expert_sales_pipeline`

**Dataset:** Sales transactions β€” highest complexity, penalises out-of-order operations heavily.

**Optimal sequence (strictly enforced):**
```
remove_duplicates β†’ rename_columns β†’ fill_missing_mode β†’ fix_dtype β†’ remove_outliers β†’ validate_schema β†’ finish
```

**Scoring:** `duplicate Γ— 0.15 + missing Γ— 0.20 + dtype Γ— 0.20 + outlier Γ— 0.20 + schema Γ— 0.25`

---

## πŸ”§ Operations Reference

All operations are invoked via JSON actions sent to `POST /step/{task_id}`.

### `remove_duplicates`
```json
{"operation": "remove_duplicates", "parameters": {}}
```
Drops exact duplicate rows using `pandas.drop_duplicates()`. Resets the index after removal.
- Optional parameter: `"subset": ["col1", "col2"]` β€” deduplicate on specific columns only

---

### `fill_missing_mean`
```json
{"operation": "fill_missing_mean", "parameters": {}}
```
Fills NaN values in numeric columns with the column mean. Skips non-numeric columns to avoid type errors.
- Optional parameter: `"column": "col_name"` β€” target a single column

---

### `fill_missing_mode`
```json
{"operation": "fill_missing_mode", "parameters": {}}
```
Fills NaN values with the most frequent value (mode). Works for both numeric and categorical columns.
- Optional parameter: `"column": "col_name"`

---

### `fill_missing_median`
```json
{"operation": "fill_missing_median", "parameters": {}}
```
Fills NaN values in numeric columns with the column median. More robust to outliers than mean.
- Optional parameter: `"column": "col_name"`

---

### `fix_dtype`
```json
{"operation": "fix_dtype", "parameters": {"dtype": "auto"}}
```
Attempts to convert columns to the most appropriate type.
- `"dtype": "auto"` β€” tries `int` then `float`, skips if conversion fails
- `"dtype": "int"` β€” convert to integer
- `"dtype": "float"` β€” convert to float
- `"dtype": "str"` β€” convert to string
- Optional parameter: `"column": "col_name"`

---

### `remove_outliers`
```json
{"operation": "remove_outliers", "parameters": {"method": "iqr"}}
```
Removes rows where numeric values fall outside the outlier fence.
- `"method": "iqr"` β€” IQR method: removes values outside `[Q1 βˆ’ 1.5Γ—IQR, Q3 + 1.5Γ—IQR]`
- `"method": "zscore"` β€” Z-score method: removes values beyond Β±3Οƒ
- Optional parameter: `"column": "col_name"` β€” target a single numeric column

---

### `rename_columns`
```json
{"operation": "rename_columns", "parameters": {}}
```
Auto-renames all columns to `snake_case` (lowercase, spaces β†’ underscores).
- Optional parameter: `"mapping": {"Old Name": "new_name"}` β€” explicit rename map

---

### `validate_schema`
```json
{"operation": "validate_schema", "parameters": {}}
```
Compares current column names against the gold dataset schema.
- Returns missing columns (in gold but not current)
- Returns extra columns (in current but not in gold)
- Returns a success message if schemas match perfectly

---

### `finish`
```json
{"operation": "finish", "parameters": {}}
```
Signals the agent is done. Triggers final reward computation and ends the episode immediately. **Always call this when cleaning is complete.**

---

## πŸ“Š Reward & Scoring System

Reward is computed after every step and returned as a `Reward` object. The total is a weighted sum of components minus penalties, **clamped strictly to `(0.0001, 0.9999)`**.

### Score Components

| Component | What It Measures | How It's Calculated |
|---|---|---|
| `duplicate_score` | Row count vs gold dataset | Proportional to excess/deficit rows |
| `missing_score` | Missing values filled vs gold | Fraction of needed fills completed |
| `dtype_score` | Column types match gold | Matched columns Γ· total columns |
| `outlier_score` | Numeric values within 3Οƒ of gold mean | Per-column average, then mean across columns |
| `schema_score` | Column names match gold schema | Matched column names Γ· gold column count |
| `penalty` | Step efficiency + operation order | See sequence penalty below |

---

### Sequence Penalty

The optimal operation order is:
```
remove_duplicates β†’ fix_dtype β†’ fill_missing_* β†’ remove_outliers β†’ validate_schema
```

Penalties for deviations:

| Violation | Penalty |
|---|---|
| Out-of-order operation | βˆ’0.08 |
| Repeated operation (non-fill/outlier) | βˆ’0.02 |
| Unknown operation | βˆ’0.01 |
| Using >80% of allowed steps | βˆ’0.05 |
| **Maximum total penalty** | **βˆ’0.25** |

---

### Score Clamping (Critical)

The Scaler grader rejects scores of exactly `0.0` or `1.0`. The following clamping is enforced at every level:

```python
# environment.py β€” _compute_reward()
def _sc(v):
    return round(max(0.0001, min(0.9999, float(v))), 4)

# Applied to ALL Reward fields: total, duplicate_score, missing_score, etc.
return Reward(
    total=_sc(total),
    duplicate_score=_sc(dup_score),
    ...
)
```

```python
# inference.py β€” every printed reward
def _clamp(v: float) -> float:
    return max(0.01, min(0.99, float(v)))

# [STEP] and [END] lines both use _clamp() before formatting
```

---

## 🌐 API Reference

**Base URL:** `https://thorodin103-data-cleaning-openenv.hf.space`

| Method | Endpoint | Description |
|---|---|---|
| `POST` | `/reset` | Reset environment (body: `{"task_id": "..."}`) |
| `POST` | `/reset/{task_id}` | Reset specific task environment |
| `POST` | `/step` | Take action (body: `{"task_id": "...", "operation": "...", "parameters": {}}`) |
| `POST` | `/step/{task_id}` | Take action in specific task |
| `GET` | `/state` | Get current environment state |
| `GET` | `/state/{task_id}` | Get state for specific task |
| `GET` | `/tasks` | List all tasks with full metadata |
| `GET` | `/validate` | Run OpenEnv spec validation across all tasks |
| `GET` | `/health` | Health check |
| `GET` | `/docs` | Interactive Swagger UI |
| `POST` | `/leaderboard/submit` | Submit a score entry |
| `GET` | `/leaderboard` | Get current leaderboard rankings |

---

### Example: Reset a task
```bash
curl -X POST https://thorodin103-data-cleaning-openenv.hf.space/reset/easy_dedup_rename
```
```json
{
  "observation": {
    "task_id": "easy_dedup_rename",
    "step": 0,
    "columns": ["EMP ID", "EMP NAME", "DEP T", "SAL ARY", "AGE"],
    "duplicate_count": 5,
    "missing_values": {"EMP ID": 0, "EMP NAME": 0, ...},
    "message": "Environment reset. Start cleaning!"
  },
  "reward": {"total": 0.0001},
  "done": false
}
```

---

### Example: Take a step
```bash
curl -X POST https://thorodin103-data-cleaning-openenv.hf.space/step/easy_dedup_rename \
  -H "Content-Type: application/json" \
  -d '{"operation": "remove_duplicates", "parameters": {}}'
```
```json
{
  "observation": {"step": 1, "duplicate_count": 0, "message": "Removed 5 duplicate rows. Rows: 20 -> 15"},
  "reward": {"total": 0.4821, "duplicate_score": 0.9999, "schema_score": 0.0001},
  "done": false
}
```

---

### Example: Validate the environment
```bash
curl https://thorodin103-data-cleaning-openenv.hf.space/validate
```
```json
{
  "openenv_valid": true,
  "tasks": {
    "easy_dedup_rename":    {"status": "passed"},
    "medium_missing_dtype": {"status": "passed"},
    "hard_full_pipeline":   {"status": "passed"},
    "expert_sales_pipeline":{"status": "passed"}
  }
}
```

---

## πŸ€– Inference Script

`inference.py` is the hackathon submission entry point. It runs an LLM agent across all tasks and emits structured stdout lines that the platform parser reads.

### Required Stdout Format

> The format below is **mandatory**. The platform parser reads these exact line types.

```
[START] task=<task_name> env=<benchmark> model=<model_name>
[STEP]  step=<n> action=<action_str> reward=<0.00> done=<true|false> error=<msg|null>
[END]   success=<true|false> steps=<n> score=<0.00> rewards=<r1,r2,...,rn>
```

**Rules:**
- One `[START]` line at episode begin
- One `[STEP]` line per step, immediately after `env.step()` returns
- One `[END]` line after episode end β€” **always emitted, even on exception** (via `finally` block)
- `reward` and `rewards` formatted to **2 decimal places**
- `done` and `success` are lowercase: `true` or `false`
- `score=` field in `[END]` is **mandatory** β€” its absence causes Task Validation failure
- `error` is the raw error string, or `null` if none

**Example output:**
```
[START] task=easy_dedup_rename env=data-cleaning-openenv model=gpt-4o-mini
[STEP] step=1 action=remove_duplicates reward=0.48 done=false error=null
[STEP] step=2 action=rename_columns reward=0.96 done=false error=null
[STEP] step=3 action=finish reward=0.96 done=true error=null
[END] success=true steps=3 score=0.96 rewards=0.48,0.96,0.96
```

---

### Agent Loop

For each task the agent follows this loop:

1. Call `env.reset()` to initialise the episode
2. Build a prompt from the observation (shape, columns, missing values, dtypes, sample rows)
3. Send prompt to LLM via OpenAI-compatible client
4. Parse the JSON response into an `Action`
5. Call `env.step(action)` and record the reward
6. Emit a `[STEP]` line
7. Repeat until `done=true` or `MAX_STEPS` (20) reached
8. Compute `score = average(rewards)`, clamped to `(0, 1)`
9. Emit `[END]` line via `finally` block

---

### Environment Variables

| Variable | Default | Description |
|---|---|---|
| `API_BASE_URL` | `https://api.openai.com/v1` | OpenAI-compatible API endpoint |
| `MODEL_NAME` | `gpt-4o-mini` | Model identifier |
| `HF_TOKEN` | *(required)* | Hugging Face / API key |

---

## πŸ“¦ Data Models

### `Action`
```json
{
  "operation": "remove_duplicates",
  "parameters": {}
}
```
- `operation` β€” one of the 9 valid operations
- `parameters` β€” operation-specific options (`column`, `strategy`, `method`, `dtype`, `mapping`, `subset`)

---

### `Observation`
```json
{
  "task_id": "easy_dedup_rename",
  "step": 1,
  "dataset_info": {"total_rows": 15, "has_duplicates": false, "has_missing": false},
  "columns": ["emp_id", "emp_name", "dept", "salary", "age"],
  "shape": [15, 5],
  "missing_values": {"emp_id": 0, "emp_name": 0},
  "dtypes": {"emp_id": "int64", "emp_name": "object"},
  "duplicate_count": 0,
  "sample_rows": [{"emp_id": 101, "emp_name": "Alice", ...}],
  "available_operations": ["remove_duplicates", "rename_columns", "finish"],
  "task_description": "Clean an employee dataset by...",
  "message": "Removed 5 duplicate rows."
}
```

---

### `Reward`
```json
{
  "total": 0.4821,
  "duplicate_score": 0.9999,
  "missing_score": 0.0001,
  "dtype_score": 0.0001,
  "outlier_score": 0.0001,
  "schema_score": 0.0001,
  "penalty": 0.0
}
```
All values are clamped to `(0.0001, 0.9999)`.

---

### `StepResult`
```json
{
  "observation": { ... },
  "reward": { ... },
  "done": false,
  "info": {
    "step": 1,
    "operation": "remove_duplicates",
    "reward_history": [0.4821]
  }
}
```

---

## πŸ—ƒοΈ Datasets

Each task has a paired `dirty.csv` and `gold.csv`. The dirty file is loaded at reset; the gold file is used as the scoring reference throughout the episode.

### Easy β€” Employee Dataset
| Property | Value |
|---|---|
| Dirty columns | `EMP ID`, `EMP NAME`, `DEP T`, `SAL ARY`, `AGE` |
| Gold columns | `emp_id`, `emp_name`, `dept`, `salary`, `age` |
| Issues | Duplicate rows, space-separated column names |
| Rows | ~20 dirty β†’ ~15 gold after dedup |

### Medium β€” Customer Dataset
| Property | Value |
|---|---|
| Columns | `customer_id`, `age`, `salary`, `gender`, `purchases`, `region`, `joined_date` |
| Issues | NaN in `age`, `salary`, `gender`, `region`; `salary` as `object` instead of `float` |
| Rows | ~30, no duplicates |

### Hard β€” Orders Dataset
| Property | Value |
|---|---|
| Columns | `order_id`, `product`, `quantity`, `price`, `customer_id`, `status`, `order_date`, `rating` |
| Issues | Duplicate orders, missing `quantity`/`price`, wrong dtypes, price outliers |
| Rows | ~50 dirty, full pipeline required |

### Expert β€” Sales Dataset
| Property | Value |
|---|---|
| Issues | All of the above plus column naming problems |
| Unique challenge | Operations must be applied in strict optimal order β€” out-of-order is penalised βˆ’0.08 per violation |

---


## πŸ“ˆ Baseline Scores

Baseline agent: **gpt-4o-mini** (from `openenv.yaml`)

| Task | Score |
|---|---|
| `easy_dedup_rename` | **0.9900** |
| `medium_missing_dtype` | **0.7000** |
| `hard_full_pipeline` | **0.6636** |
| **Average** | **0.7845** |

---

## πŸ“„ License

MIT License β€” free to use, modify, and distribute.

---

<div align="center">

Built for the **Scaler Γ— OpenEnv Hackathon**

πŸ”— [GitHub](https://github.com/ReverseCoder1/CleanifyAI) Β· [HuggingFace Space](https://huggingface.co/spaces/cleanify-ai/Data-cleaning) Β· [Live API Docs](https://thorodin103-data-cleaning-openenv.hf.space/docs)

*CleanifyAI β€” making data clean, one step at a time.*

</div>