meta-com-openenv-agent / CHECKLIST.MD
Monike123's picture
Initial OpenEnv Agent Submission
d72844a
Raw
History Blame Contribute Delete
9.48 kB
# Project Execution Checklist
> **How to use:** Mark items `[x]` when complete. Add date + initials next to each completed item.
> Never delete a completed item β€” this log is permanent.
> If a phase is blocked, add `[!]` and explain why in `docs/error_and_solving_chat.md`.
---
## Phase 1 β€” Environment Foundation
**Goal:** A fully runnable, sandboxed OpenEnv-compatible environment.
**Success:** `python tests/test_env.py` executes without error and verifies all 3 API methods.
- [ ] Create root project directory (`Meta_com/`)
- [ ] Create subdirectory structure:
- [ ] `skills/` β€” agent skill definition files
- [ ] `docs/` β€” idea log and error log
- [ ] `env/` β€” environment implementation (Python)
- [ ] `agent/` β€” agent modules
- [ ] `tests/` β€” automated test suite
- [ ] `tasks/` β€” task repository folders
- [ ] Set up Python virtual environment (`.venv`)
- [ ] Create `requirements.txt` (initial: `pytest`, `ast`, `gitpython`)
- [ ] Implement OpenEnv-compatible class structure in `env/conflict_env.py`:
- [ ] Class inherits or follows OpenEnv interface
- [ ] Constructor accepts `task_name` parameter to select difficulty
- [ ] Stores internal state dict
- [ ] Implement `reset()` API:
- [ ] Copies task repo to a temp sandbox directory
- [ ] Returns initial state dict
- [ ] State includes: file contents, conflict count, syntax status, test status, done flag
- [ ] Implement `step(action)` API:
- [ ] Accepts action dict with `file` and `resolution` keys
- [ ] Applies resolution to sandbox repo
- [ ] Calls reward engine to compute reward
- [ ] Returns `(next_state, reward, done, info)` tuple
- [ ] Implement `state()` API:
- [ ] Reads current repo sandbox state without modifying it
- [ ] Returns identical schema to `reset()` output
- [ ] Define reward signal structure in `env/reward_engine.py`:
- [ ] Conflict marker count function
- [ ] AST syntax check function
- [ ] Test runner function
- [ ] Semantic duplication detector (stretch goal)
- [ ] Aggregation function combining all components
- [ ] Add deterministic conflict detection:
- [ ] Regex scan for `<<<<<<<`, `=======`, `>>>>>>>`
- [ ] Count blocks (not lines)
- [ ] Return list of conflict blocks with line numbers
---
## Phase 2 β€” Task Design
**Goal:** Three task repositories covering easy / medium / hard conflict scenarios.
**Success:** Each task can be `reset()` and the resulting state contains valid conflict markers.
- [ ] **Task 1 β€” Single-File Conflict** (`tasks/easy_conflict/`)
- [ ] Create a simple Python file (e.g. `calculator.py`) with 1–2 conflict blocks
- [ ] Conflict should involve a simple arithmetic function (HEAD: `return a + b`, incoming: `return a + b + 1`)
- [ ] Include a unit test file `test_calculator.py` with 2+ tests
- [ ] Add conflict markers manually to simulate a real `git merge` conflict
- [ ] Validate: `env.reset()` returns `conflict_count >= 1`
- [ ] **Task 2 β€” Multi-File Dependency Conflict** (`tasks/medium_conflict/`)
- [ ] File A: `utils.py` β€” function signature changed (e.g. added a parameter)
- [ ] File B: `main.py` β€” imports and calls `utils.py` with old signature
- [ ] Conflict markers present in `utils.py`
- [ ] `main.py` should break (ImportError or TypeError) until the agent resolves both files
- [ ] Include tests that exercise the cross-file call chain
- [ ] **Task 3 β€” Semantic Logic Merge** (`tasks/hard_conflict/`)
- [ ] Both HEAD and incoming branch added different logic to the same function body
- [ ] The combined logic must include changes from **both** branches (not just pick one)
- [ ] Naive "take mine" or "take theirs" solutions should fail tests
- [ ] Requires the agent to understand intent and recombine logic safely
- [ ] Include tests that verify the merged behavior
- [ ] Create `tasks/README.md` explaining each task format
- [ ] Add `conflict_markers_validator.py` script that validates task files contain markers
- [ ] Add `task_schema.json` describing expected task directory structure
- [ ] Add validation scripts for each task:
- [ ] `tasks/easy_conflict/validate.py`
- [ ] `tasks/medium_conflict/validate.py`
- [ ] `tasks/hard_conflict/validate.py`
---
## Phase 3 β€” Agent Skills
**Goal:** Modular, testable skill functions the agent uses to process conflicts.
**Success:** Each skill function passes its unit tests and can be called independently.
- [ ] **Skill 1 β€” Conflict Marker Detection** (`agent/conflict_detector.py`)
- [ ] `detect_conflicts(file_content: str) β†’ list[ConflictBlock]`
- [ ] Returns list of `ConflictBlock` objects with: `head_lines`, `incoming_lines`, `start_line`, `end_line`
- [ ] Handles edge cases: nested blocks, missing end marker, empty files
- [ ] Unit tests: `tests/test_conflict_detector.py`
- [ ] **Skill 2 β€” Syntax Validation** (`agent/validator.py`)
- [ ] `validate_syntax(file_content: str) β†’ tuple[bool, str]`
- [ ] Uses `ast.parse()` for Python files
- [ ] Returns `(is_valid, error_message)` tuple
- [ ] Handles SyntaxError, IndentationError gracefully
- [ ] Unit tests: `tests/test_validator.py`
- [ ] **Skill 3 β€” Multi-File Reasoning** (`agent/dependency_resolver.py`)
- [ ] `find_call_sites(files: dict[str, str], changed_function: str) β†’ list[CallSite]`
- [ ] Scans all files for references to a changed function signature
- [ ] Returns list of file + line number pairs to update
- [ ] Unit tests: `tests/test_dependency_resolver.py`
- [ ] **Skill 4 β€” Semantic Merge Resolution** (`agent/resolver.py`)
- [ ] `resolve_conflict(block: ConflictBlock, context: str) β†’ str`
- [ ] Combines HEAD and incoming logic minimizing duplication
- [ ] Optional: LLM-assisted resolution for hard tasks
- [ ] Unit tests: `tests/test_resolver.py`
- [ ] **Skill 5 β€” Test Execution** (`agent/test_runner.py`)
- [ ] `run_tests(task_path: str) β†’ TestResult`
- [ ] Runs `pytest` in subprocess and parses output
- [ ] Returns `TestResult` with `passed`, `failed`, `errors` counts
- [ ] Unit tests: `tests/test_test_runner.py`
---
## Phase 4 β€” Evaluation & Reward
**Goal:** Complete, calibrated reward engine and scoring pipeline.
**Success:** Total episode reward is deterministic and matches expected values from manual evaluation.
- [ ] **Marker Removal Scoring**
- [ ] +0.15 awarded only when block is *fully* removed (no partial markers)
- [ ] Penalize if markers remain but file is reported as resolved
- [ ] Test: manually place 3 blocks, resolve 2, verify reward = 0.30
- [ ] **Syntax Reward**
- [ ] +0.20 for passing `ast.parse()` after each step
- [ ] βˆ’0.10 for failing `ast.parse()` after each step
- [ ] Reward is per-file (if multi-file task, sum across all files)
- [ ] **Unit Test Reward**
- [ ] +0.30 for all tests passing (binary β€” partial pass = 0)
- [ ] Stretch: +0.10 per test for partial credit
- [ ] **AST Duplication Penalty**
- [ ] βˆ’0.20 if semantic analysis detects duplicated function bodies
- [ ] Use AST node comparison, not string comparison
- [ ] Stretch: detect near-duplicates via tree edit distance
- [ ] **Final Reward Aggregation** (`env/reward_engine.py`)
- [ ] `compute_reward(state_before, state_after, action) β†’ float`
- [ ] All reward components summed into scalar
- [ ] Clamped to range `[βˆ’1.0, +1.0]`
- [ ] Document reward edge cases in `docs/idea_and_development_chat.md`
---
## Phase 5 β€” Deployment
**Goal:** Public Hugging Face Space with working demo UI.
**Success:** Anyone can visit the Space URL, select a conflict task, run the agent, and see the resolution + reward score.
- [ ] **Hugging Face Space Setup**
- [ ] Create new Space under HF account
- [ ] Choose `Gradio` SDK
- [ ] Add `README.md` with Space metadata (title, description, tags)
- [ ] Confirm Space shows "Running" status
- [ ] **Demo Interface** (`app.py`)
- [ ] Task selector dropdown (easy / medium / hard)
- [ ] Code display: before (with conflict markers) vs after (resolved)
- [ ] Diff view using `difflib`
- [ ] Reward breakdown table (per component)
- [ ] "Run Agent" button triggers full episode
- [ ] **Test Repository Upload**
- [ ] All 3 task folders zip-packaged and accessible
- [ ] Users can optionally upload their own conflicted repo (stretch)
- [ ] **Result Viewer**
- [ ] Step-by-step episode log shown
- [ ] Each step: action taken, reward received, state after
- [ ] Final episode summary: total reward, success/fail
- [ ] **Deploy Public Demo**
- [ ] Push to Hugging Face Space main branch
- [ ] Verify demo loads in browser
- [ ] Test all 3 tasks from UI
- [ ] Share link in `README.md`
---
## Phase 6 β€” Documentation
**Goal:** All documentation current, consistent, and useful for any external developer.
**Success:** A new contributor can understand and run the project using docs alone.
- [ ] Update `HISTORY.md` β€” all changes since v0.3 documented
- [ ] Update all `skills/` files to reflect final implementations
- [ ] Maintain `docs/idea_and_development_chat.md` β€” log all major design decisions
- [ ] Maintain `docs/error_and_solving_chat.md` β€” log all bugs and their solutions
- [ ] Write `CONTRIBUTING.md` β€” guide for contributors
- [ ] Write `AGENTS.md` β€” describes how to add a new agent type
- [ ] Add docstrings to all Python modules
- [ ] Add type hints to all public functions
- [ ] Generate API reference with `pdoc` or `sphinx`