Spaces:
Sleeping
title: Data Cleaning OpenEnv
emoji: π§Ή
colorFrom: blue
colorTo: green
sdk: docker
sdk_version: '3.10'
app_file: main.py
pinned: false
π§Ή CleanifyAI β Data Cleaning OpenEnv
A reinforcement-learning environment where AI agents learn to clean real-world messy datasets β step by step.
Scaler Γ OpenEnv Hackathon Submission
π Table of Contents
- Overview
- Project Structure
- Setup & Installation
- Tasks
- Operations Reference
- Reward & Scoring System
- API Reference
- Inference Script
- Data Models
- Datasets
- Baseline Scores
- License
π Overview
CleanifyAI is a fully OpenEnv-compliant environment that challenges AI agents to autonomously clean messy, real-world datasets through a sequence of structured operations. It mimics professional data engineering pipelines and rewards agents that apply operations in the correct, logical order.
| Property | Value |
|---|---|
| Environment Name | data-cleaning-openenv |
| Tasks | 4 (Easy, Medium, Hard, Expert) |
| Operations | 9 (dedup, fill, dtype fix, outlier removal, rename, validate, finish) |
| Scoring | Weighted multi-component, strictly in (0, 1) |
| API | OpenEnv-compliant REST via FastAPI |
| Framework | Python 3.10, FastAPI, Pandas, NumPy |
| Inference | OpenAI-compatible LLM client |
| Deployed at | https://thorodin103-data-cleaning-openenv.hf.space |
β οΈ Score Constraint: The Scaler platform rejects scores of exactly
0.0or1.0. All scoring paths in this codebase clamp strictly to(0.0001, 0.9999).
π Project Structure
CleanifyAI/
β
βββ inference.py # π€ LLM agent β emits [START]/[STEP]/[END] stdout lines
βββ environment.py # ποΈ Core OpenEnv environment & reward computation
βββ models.py # π¦ Pydantic models: Action, Observation, Reward, StepResult
βββ main.py # π FastAPI server with all REST endpoints
βββ Dockerfile # π³ Python 3.10-slim container, port 7860
βββ openenv.yaml # π OpenEnv spec manifest
βββ pyproject.toml # π¦ Python dependency config
βββ uv.lock # π Locked dependency versions
β
βββ datasets/
β βββ task_metadata.json # βοΈ Per-task config (steps, operations, scoring weights)
β βββ easy/
β β βββ dirty.csv # ποΈ Employee dataset with duplicates + bad column names
β β βββ gold.csv # β
Gold standard cleaned version
β βββ medium/
β β βββ dirty.csv # ποΈ Customer dataset with missing values + wrong dtypes
β β βββ gold.csv # β
Gold standard
β βββ hard/
β β βββ dirty.csv # ποΈ Orders dataset requiring full pipeline
β β βββ gold.csv # β
Gold standard
β βββ expert/
β βββ dirty.csv # ποΈ Sales dataset β strict operation order required
β βββ gold.csv # β
Gold standard
β
βββ static/
β βββ index.html # π₯οΈ Web UI for interactive exploration
β
βββ server/
βββ app.py # π§ Server initialization module
π Setup & Installation
Prerequisites
- Python 3.10+
- Docker (for containerized deployment)
- A Hugging Face account (
HF_TOKEN) - An OpenAI-compatible API endpoint and model
Local Development
1. Clone the repository
git clone https://github.com/ReverseCoder1/CleanifyAI.git
cd CleanifyAI
2. Install dependencies
pip install fastapi==0.104.1 uvicorn==0.24.0 pydantic==2.5.0 \
pandas==2.1.3 numpy==1.26.2 openai>=2.7.2 \
pyyaml==6.0.1 python-dotenv==1.0.0
3. Create a .env file
API_BASE_URL=https://api.openai.com/v1
MODEL_NAME=gpt-4o-mini
HF_TOKEN=your_hugging_face_token_here
4. Start the FastAPI server
uvicorn main:app --host 0.0.0.0 --port 7860 --reload
- API: http://localhost:7860
- Swagger UI: http://localhost:7860/docs
Docker Deployment
# Build
docker build -t cleanify-ai .
# Run
docker run -p 7860:7860 \
-e HF_TOKEN=your_token \
-e MODEL_NAME=gpt-4o-mini \
-e API_BASE_URL=https://api.openai.com/v1 \
cleanify-ai
Run the Inference Agent
python inference.py
Runs the LLM agent across all 3 hackathon tasks and streams hackathon-spec log lines to stdout.
π― Tasks
Four progressively complex tasks. The hackathon evaluates easy, medium, and hard. Expert is available for extended benchmarking.
| Task ID | Difficulty | Max Steps | Key Operations | Scoring |
|---|---|---|---|---|
easy_dedup_rename |
β Easy | 10 | remove_duplicates, rename_columns |
dup 50% + schema 50% |
medium_missing_dtype |
ββ Medium | 15 | fill_missing_*, fix_dtype |
missing 50% + dtype 50% |
hard_full_pipeline |
βββ Hard | 20 | Full pipeline | 20% Γ 5 components |
expert_sales_pipeline |
ββββ Expert | 25 | All 9 ops in strict order | Weighted (schema 25%) |
β Easy β easy_dedup_rename
Dataset: Employee records (emp_id, emp_name, dept, salary, age)
Dirty conditions:
- Duplicate rows
- Column names with spaces and inconsistent casing (
EMP ID,DEP T,SAL ARY)
Optimal sequence:
remove_duplicates β rename_columns β finish
Scoring: duplicate_score Γ 0.5 + schema_score Γ 0.5
ββ Medium β medium_missing_dtype
Dataset: Customer records (customer_id, age, salary, gender, purchases, region, joined_date)
Dirty conditions:
- NaN values in
age,salary,gender,region salarystored asobjectinstead offloat
Optimal sequence:
fill_missing_mean (numeric columns)
fill_missing_mode (categorical columns)
fix_dtype β finish
Scoring: missing_score Γ 0.5 + dtype_score Γ 0.5
βββ Hard β hard_full_pipeline
Dataset: Orders (order_id, product, quantity, price, customer_id, status, order_date, rating)
Dirty conditions:
- Duplicate order entries
- Missing
quantityandpricevalues quantitystored as object type- Extreme price outliers
Optimal sequence:
remove_duplicates β fill_missing_* β fix_dtype β remove_outliers β validate_schema β finish
Scoring: duplicate Γ 0.2 + missing Γ 0.2 + dtype Γ 0.2 + outlier Γ 0.2 + schema Γ 0.2
ββββ Expert β expert_sales_pipeline
Dataset: Sales transactions β highest complexity, penalises out-of-order operations heavily.
Optimal sequence (strictly enforced):
remove_duplicates β rename_columns β fill_missing_mode β fix_dtype β remove_outliers β validate_schema β finish
Scoring: duplicate Γ 0.15 + missing Γ 0.20 + dtype Γ 0.20 + outlier Γ 0.20 + schema Γ 0.25
π§ Operations Reference
All operations are invoked via JSON actions sent to POST /step/{task_id}.
remove_duplicates
{"operation": "remove_duplicates", "parameters": {}}
Drops exact duplicate rows using pandas.drop_duplicates(). Resets the index after removal.
- Optional parameter:
"subset": ["col1", "col2"]β deduplicate on specific columns only
fill_missing_mean
{"operation": "fill_missing_mean", "parameters": {}}
Fills NaN values in numeric columns with the column mean. Skips non-numeric columns to avoid type errors.
- Optional parameter:
"column": "col_name"β target a single column
fill_missing_mode
{"operation": "fill_missing_mode", "parameters": {}}
Fills NaN values with the most frequent value (mode). Works for both numeric and categorical columns.
- Optional parameter:
"column": "col_name"
fill_missing_median
{"operation": "fill_missing_median", "parameters": {}}
Fills NaN values in numeric columns with the column median. More robust to outliers than mean.
- Optional parameter:
"column": "col_name"
fix_dtype
{"operation": "fix_dtype", "parameters": {"dtype": "auto"}}
Attempts to convert columns to the most appropriate type.
"dtype": "auto"β triesintthenfloat, skips if conversion fails"dtype": "int"β convert to integer"dtype": "float"β convert to float"dtype": "str"β convert to string- Optional parameter:
"column": "col_name"
remove_outliers
{"operation": "remove_outliers", "parameters": {"method": "iqr"}}
Removes rows where numeric values fall outside the outlier fence.
"method": "iqr"β IQR method: removes values outside[Q1 β 1.5ΓIQR, Q3 + 1.5ΓIQR]"method": "zscore"β Z-score method: removes values beyond Β±3Ο- Optional parameter:
"column": "col_name"β target a single numeric column
rename_columns
{"operation": "rename_columns", "parameters": {}}
Auto-renames all columns to snake_case (lowercase, spaces β underscores).
- Optional parameter:
"mapping": {"Old Name": "new_name"}β explicit rename map
validate_schema
{"operation": "validate_schema", "parameters": {}}
Compares current column names against the gold dataset schema.
- Returns missing columns (in gold but not current)
- Returns extra columns (in current but not in gold)
- Returns a success message if schemas match perfectly
finish
{"operation": "finish", "parameters": {}}
Signals the agent is done. Triggers final reward computation and ends the episode immediately. Always call this when cleaning is complete.
π Reward & Scoring System
Reward is computed after every step and returned as a Reward object. The total is a weighted sum of components minus penalties, clamped strictly to (0.0001, 0.9999).
Score Components
| Component | What It Measures | How It's Calculated |
|---|---|---|
duplicate_score |
Row count vs gold dataset | Proportional to excess/deficit rows |
missing_score |
Missing values filled vs gold | Fraction of needed fills completed |
dtype_score |
Column types match gold | Matched columns Γ· total columns |
outlier_score |
Numeric values within 3Ο of gold mean | Per-column average, then mean across columns |
schema_score |
Column names match gold schema | Matched column names Γ· gold column count |
penalty |
Step efficiency + operation order | See sequence penalty below |
Sequence Penalty
The optimal operation order is:
remove_duplicates β fix_dtype β fill_missing_* β remove_outliers β validate_schema
Penalties for deviations:
| Violation | Penalty |
|---|---|
| Out-of-order operation | β0.08 |
| Repeated operation (non-fill/outlier) | β0.02 |
| Unknown operation | β0.01 |
| Using >80% of allowed steps | β0.05 |
| Maximum total penalty | β0.25 |
Score Clamping (Critical)
The Scaler grader rejects scores of exactly 0.0 or 1.0. The following clamping is enforced at every level:
# environment.py β _compute_reward()
def _sc(v):
return round(max(0.0001, min(0.9999, float(v))), 4)
# Applied to ALL Reward fields: total, duplicate_score, missing_score, etc.
return Reward(
total=_sc(total),
duplicate_score=_sc(dup_score),
...
)
# inference.py β every printed reward
def _clamp(v: float) -> float:
return max(0.01, min(0.99, float(v)))
# [STEP] and [END] lines both use _clamp() before formatting
π API Reference
Base URL: https://thorodin103-data-cleaning-openenv.hf.space
| Method | Endpoint | Description |
|---|---|---|
POST |
/reset |
Reset environment (body: {"task_id": "..."}) |
POST |
/reset/{task_id} |
Reset specific task environment |
POST |
/step |
Take action (body: {"task_id": "...", "operation": "...", "parameters": {}}) |
POST |
/step/{task_id} |
Take action in specific task |
GET |
/state |
Get current environment state |
GET |
/state/{task_id} |
Get state for specific task |
GET |
/tasks |
List all tasks with full metadata |
GET |
/validate |
Run OpenEnv spec validation across all tasks |
GET |
/health |
Health check |
GET |
/docs |
Interactive Swagger UI |
POST |
/leaderboard/submit |
Submit a score entry |
GET |
/leaderboard |
Get current leaderboard rankings |
Example: Reset a task
curl -X POST https://thorodin103-data-cleaning-openenv.hf.space/reset/easy_dedup_rename
{
"observation": {
"task_id": "easy_dedup_rename",
"step": 0,
"columns": ["EMP ID", "EMP NAME", "DEP T", "SAL ARY", "AGE"],
"duplicate_count": 5,
"missing_values": {"EMP ID": 0, "EMP NAME": 0, ...},
"message": "Environment reset. Start cleaning!"
},
"reward": {"total": 0.0001},
"done": false
}
Example: Take a step
curl -X POST https://thorodin103-data-cleaning-openenv.hf.space/step/easy_dedup_rename \
-H "Content-Type: application/json" \
-d '{"operation": "remove_duplicates", "parameters": {}}'
{
"observation": {"step": 1, "duplicate_count": 0, "message": "Removed 5 duplicate rows. Rows: 20 -> 15"},
"reward": {"total": 0.4821, "duplicate_score": 0.9999, "schema_score": 0.0001},
"done": false
}
Example: Validate the environment
curl https://thorodin103-data-cleaning-openenv.hf.space/validate
{
"openenv_valid": true,
"tasks": {
"easy_dedup_rename": {"status": "passed"},
"medium_missing_dtype": {"status": "passed"},
"hard_full_pipeline": {"status": "passed"},
"expert_sales_pipeline":{"status": "passed"}
}
}
π€ Inference Script
inference.py is the hackathon submission entry point. It runs an LLM agent across all tasks and emits structured stdout lines that the platform parser reads.
Required Stdout Format
The format below is mandatory. The platform parser reads these exact line types.
[START] task=<task_name> env=<benchmark> model=<model_name>
[STEP] step=<n> action=<action_str> reward=<0.00> done=<true|false> error=<msg|null>
[END] success=<true|false> steps=<n> score=<0.00> rewards=<r1,r2,...,rn>
Rules:
- One
[START]line at episode begin - One
[STEP]line per step, immediately afterenv.step()returns - One
[END]line after episode end β always emitted, even on exception (viafinallyblock) rewardandrewardsformatted to 2 decimal placesdoneandsuccessare lowercase:trueorfalsescore=field in[END]is mandatory β its absence causes Task Validation failureerroris the raw error string, ornullif none
Example output:
[START] task=easy_dedup_rename env=data-cleaning-openenv model=gpt-4o-mini
[STEP] step=1 action=remove_duplicates reward=0.48 done=false error=null
[STEP] step=2 action=rename_columns reward=0.96 done=false error=null
[STEP] step=3 action=finish reward=0.96 done=true error=null
[END] success=true steps=3 score=0.96 rewards=0.48,0.96,0.96
Agent Loop
For each task the agent follows this loop:
- Call
env.reset()to initialise the episode - Build a prompt from the observation (shape, columns, missing values, dtypes, sample rows)
- Send prompt to LLM via OpenAI-compatible client
- Parse the JSON response into an
Action - Call
env.step(action)and record the reward - Emit a
[STEP]line - Repeat until
done=trueorMAX_STEPS(20) reached - Compute
score = average(rewards), clamped to(0, 1) - Emit
[END]line viafinallyblock
Environment Variables
| Variable | Default | Description |
|---|---|---|
API_BASE_URL |
https://api.openai.com/v1 |
OpenAI-compatible API endpoint |
MODEL_NAME |
gpt-4o-mini |
Model identifier |
HF_TOKEN |
(required) | Hugging Face / API key |
π¦ Data Models
Action
{
"operation": "remove_duplicates",
"parameters": {}
}
operationβ one of the 9 valid operationsparametersβ operation-specific options (column,strategy,method,dtype,mapping,subset)
Observation
{
"task_id": "easy_dedup_rename",
"step": 1,
"dataset_info": {"total_rows": 15, "has_duplicates": false, "has_missing": false},
"columns": ["emp_id", "emp_name", "dept", "salary", "age"],
"shape": [15, 5],
"missing_values": {"emp_id": 0, "emp_name": 0},
"dtypes": {"emp_id": "int64", "emp_name": "object"},
"duplicate_count": 0,
"sample_rows": [{"emp_id": 101, "emp_name": "Alice", ...}],
"available_operations": ["remove_duplicates", "rename_columns", "finish"],
"task_description": "Clean an employee dataset by...",
"message": "Removed 5 duplicate rows."
}
Reward
{
"total": 0.4821,
"duplicate_score": 0.9999,
"missing_score": 0.0001,
"dtype_score": 0.0001,
"outlier_score": 0.0001,
"schema_score": 0.0001,
"penalty": 0.0
}
All values are clamped to (0.0001, 0.9999).
StepResult
{
"observation": { ... },
"reward": { ... },
"done": false,
"info": {
"step": 1,
"operation": "remove_duplicates",
"reward_history": [0.4821]
}
}
ποΈ Datasets
Each task has a paired dirty.csv and gold.csv. The dirty file is loaded at reset; the gold file is used as the scoring reference throughout the episode.
Easy β Employee Dataset
| Property | Value |
|---|---|
| Dirty columns | EMP ID, EMP NAME, DEP T, SAL ARY, AGE |
| Gold columns | emp_id, emp_name, dept, salary, age |
| Issues | Duplicate rows, space-separated column names |
| Rows | ~20 dirty β ~15 gold after dedup |
Medium β Customer Dataset
| Property | Value |
|---|---|
| Columns | customer_id, age, salary, gender, purchases, region, joined_date |
| Issues | NaN in age, salary, gender, region; salary as object instead of float |
| Rows | ~30, no duplicates |
Hard β Orders Dataset
| Property | Value |
|---|---|
| Columns | order_id, product, quantity, price, customer_id, status, order_date, rating |
| Issues | Duplicate orders, missing quantity/price, wrong dtypes, price outliers |
| Rows | ~50 dirty, full pipeline required |
Expert β Sales Dataset
| Property | Value |
|---|---|
| Issues | All of the above plus column naming problems |
| Unique challenge | Operations must be applied in strict optimal order β out-of-order is penalised β0.08 per violation |
π Baseline Scores
Baseline agent: gpt-4o-mini (from openenv.yaml)
| Task | Score |
|---|---|
easy_dedup_rename |
0.9900 |
medium_missing_dtype |
0.7000 |
hard_full_pipeline |
0.6636 |
| Average | 0.7845 |
π License
MIT License β free to use, modify, and distribute.
Built for the Scaler Γ OpenEnv Hackathon
π GitHub Β· HuggingFace Space Β· Live API Docs
CleanifyAI β making data clean, one step at a time.