Spaces:
Sleeping
Idea and Development Chat
Purpose: Log all architecture decisions, design explorations, open questions, and planning discussions here. Treat this as a searchable reasoning journal β write down WHY you made a decision, not just what it is. Format entries with a date and topic heading.
How to Use This File
When you are:
- Designing a new module β write the options you considered and why you chose one
- Changing an existing design β write what changed and what triggered the change
- Unsure about an approach β write the open question and leave it for later resolution
- Discussing with a collaborator or AI β paste key excerpts and conclusions
[2026-04-07] β Project Concept Selection
Problem Statement
Git merge conflicts are one of the most universally painful experiences in software development. When two branches touch the same lines:
- Git inserts conflict markers (
<<<<<<<,=======,>>>>>>>) into the file - The developer must manually inspect both versions and decide how to combine them
- In large codebases with hundreds of files, this is painstaking and error-prone
- No existing tool can autonomously resolve conflicts with high accuracy
Why This is a Good AI Benchmark Task
- Structured input: Conflict markers provide exact boundaries of disagreement
- Verifiable output: AST parsing and unit tests give objective success/fail signal
- Multi-skill: Requires reading, reasoning, editing, and validation all in one episode
- Scalable difficulty: Easy (trivial merge) β Hard (semantic reasoning required)
- Real-world relevance: Every developer team faces this problem
Platform Decision: OpenEnv
OpenEnv was selected over custom environments because:
- Standard
reset() / step() / state()API is familiar to RL researchers - Compatible with multiple agent frameworks (RL policies, LLM tool-use, hybrid)
- Easy to wrap into a Gymnasium-compatible environment later
- Deployment on Hugging Face is straightforward
Alternative Considered: Pure LLM Benchmark
A simpler alternative would be to just prompt an LLM with a conflict and check the output. This was rejected because:
- No sandboxed execution β can't verify syntax or tests
- No episode structure β can't study multi-step resolution
- Can't support RL training loops
- Doesn't fit the "agent operating in environment" paradigm
[2026-04-07] β Task Difficulty Design
Design Question
How many difficulty levels, and what differentiates them?
Decision: 3 Levels
| Level | Why It Exists |
|---|---|
| Easy | Baseline β any agent should pass this. Confirms environment works. |
| Medium | Tests cross-file reasoning. More realistic real-world scenario. |
| Hard | Tests semantic understanding. Requires genuine comprehension, not just marker removal. |
Easy Task Design Rationale
A single Python file with 1β2 conflict blocks where:
- HEAD block adds simple logic
- Incoming block also adds simple logic
- The correct merge is to include BOTH additions (not pick one)
Or: HEAD and incoming change a return value differently β tests define which is correct.
The unit tests are the ground truth. The agent wins if tests pass.
Medium Task Design Rationale
Two files: utils.py and main.py.
utils.py gets a new function parameter in the incoming branch. The conflict sits in utils.py. After the agent resolves it by accepting the new signature, main.py now breaks β because it still calls the old signature.
The agent must:
- Notice the signature changed
- Find all call sites (in
main.py) - Update them to match
This tests cross-file awareness β something LLMs frequently fail at in practice.
Hard Task Design Rationale
Both HEAD and incoming branches independently added useful logic to the same function body:
- HEAD: added input validation
- Incoming: added a new computation mode
A naive "take mine" or "take theirs" strategy fails the tests, because the tests verify BOTH behaviors. The agent must understand what each branch was trying to do and merge them.
This tests semantic reasoning β the hardest and most valuable capability.
[2026-04-07] β Reward Function Design
Design Question
Should the reward be sparse (only on success) or dense (every step)?
Decision: Dense Reward
Reasons:
- Sparse reward (only +1 on full success) makes learning very hard β especially on hard tasks
- Dense reward gives signal even for partial progress (removed 1 of 3 blocks)
- Mixed reward (dense during episode + bonus on success) is the ideal approach for RL
Reward Component Breakdown
Why +0.15 for marker removal?
- Conflict removal is a necessary but not sufficient condition for success
- A lower value than syntax/test rewards because removing markers is easy (just delete them) without actually fixing the code
- The syntax and test rewards ensure the removal was valid
Why β0.10 for syntax error?
- The agent must be penalized for making the file worse
- A small negative prevents the agent from making random edits
- Not too large so that the agent can recover from mistakes
Why +0.30 for tests passing?
- This is the hardest to achieve and the most meaningful signal β worth the most
- Tests passing = end goal achieved (from a developer's perspective)
Why β0.20 for semantic duplication?
- Duplicated logic is a silent bug β tests may still pass but the code quality degrades
- This teaches the agent to produce clean merges, not just technically correct ones
Max Episode Reward Calculation
For a 3-block easy task:
3 blocks Γ 0.15 = 0.45 (all markers removed)
+ 0.20 (syntax valid)
+ 0.30 (tests pass)
= 0.95 max reward
For a multi-step episode, rewards accumulate across steps. The agent can also receive negative rewards if it introduces syntax errors along the way.
[2026-04-07] β Agent Architecture Decisions
Question: Rule-based vs LLM-based resolver?
Option A β Pure Rule-Based
- Pros: fast, deterministic, no API calls
- Cons: fails on semantic tasks, brittle to edge cases
- Valid for: easy task only
Option B β Pure LLM
- Pros: handles semantic reasoning, flexible
- Cons: slow, costly, non-deterministic
- Valid for: hard task
Decision: Hybrid (rule-first, LLM fallback)
if conflict is trivially resolvable (one side is subset of other):
use rule-based resolver
elif conflict can be resolved by AST merge:
use AST-based resolver
else:
use LLM resolver with task context as prompt
This gives performance on easy tasks and capability on hard tasks.
Question: How does the agent receive state?
The agent receives a state dict from the environment:
{
"files": {"main.py": "<content>"},
"conflict_count": 3,
"syntax_valid": False,
"tests_passing": False,
"done": False
}
The agent's action is:
{
"file": "main.py",
"resolution": "<full resolved content>"
}
One action = one file resolved. Multi-file tasks require multiple steps.
[2026-04-07] β Deployment Architecture
Hugging Face Space Design
User visits Space URL
βββΊ Selects task (Easy / Medium / Hard)
βββΊ Clicks "Run Agent"
βββΊ Backend: env.reset() called
βββΊ Agent runs episode
βββΊ Results shown: before/after diff + reward table
Why Gradio?
- Built into HF Spaces natively
- Fast to prototype interactive demos
- supports code display, diffs, tables natively
Single-file app: app.py (everything in one file for Space simplicity)
Open Questions (To Be Resolved)
| Question | Status |
|---|---|
| Should hard task use GPT-4 or a local LLM? | π² Open |
| Should multi-file actions be batched or sequential? | π² Open |
| Should the environment support Java/JS files, or Python only? | π² Open β Python only for v1 |
| Should episode length be capped? (max N steps) | π² Open β suggest cap at 10 steps |
| Should we persist episode logs to disk for analysis? | π² Open |