File size: 6,683 Bytes
0d566ba
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
---
title: EIM+ Real Verification Engineer
emoji: 🧪
colorFrom: blue
colorTo: indigo
sdk: gradio
sdk_version: 4.44.0
python_version: 3.12.12
app_file: chat_app.py
pinned: false
---

# EIM+ Real Verification Engineer

A Gradio coding assistant with an iterative repair engine, actual child-process test execution, persistent correction/repair memory (optional private Hugging Face Dataset sync), and a repeatable benchmark harness.

## Files

- `app.py` — core engine: model providers, constrained subprocess execution, scoring, and experience memory.
- `eim_plus.py` — candidate generation, independent test verification, holdout checks, weak-test/mutation checks, repair logs, retries, and result cache controls.
- `chat_app.py` — bilingual chat UI and file-aware assistant.
- `terminal_verify.py` — stand-alone real terminal runner with stdout/stderr/exit/timeout evidence, explicit opt-ins for higher-risk operations, and JSON reports.
- `correction_memory.py` — explicit user-correction memory in Arabic/English, plus optional Hugging Face Dataset sync.
- `eval_eim.py` — 25-task model benchmark; stores labeled runs and compares rates, confidence intervals, attempts, restarts, regressions, and time.

## Deploy on Hugging Face Spaces

1. Create a **Gradio** Space and push all files from this folder to its repository.
2. Select **ZeroGPU** hardware in the Space settings. On a free personal account, Hugging Face currently requires a verified email and an account older than 30 days; eligible accounts can host up to two ZeroGPU Spaces. Current official docs list a **5 GPU-minute daily quota** for Free accounts, shared at account level; this is not unlimited compute. Details can change: [ZeroGPU official documentation](https://huggingface.co/docs/hub/en/spaces-zerogpu).
3. This app automatically chooses the **local model backend** when the runtime reports ZeroGPU, so the selected GPU is used; the optional `@spaces.GPU` wrapper is used when `spaces` is installed. Do not choose a paid hardware flavor if your goal is free hosting.
4. Choose a model backend using Space **Variables** / **Secrets**:

### Option A — no locally downloaded weights (remote inference)

Set Space Variables:

```text
EIM_BACKEND=hf
EIM_MODEL=Mungert/Qwen3-Coder-30B-A3B-Instruct-GGUF
```

Add `HF_TOKEN` under Space **Secrets** with permission to use the required Hugging Face Inference Provider. Remote inference availability/pricing depends on the provider and model; a token by itself does not guarantee free inference.

### Option B — load model weights in the Space

Set Space Variables:

```text
EIM_BACKEND=local
EIM_MODEL=Mungert/Qwen3-Coder-30B-A3B-Instruct-GGUF
EIM_GPU_SECONDS=60
```

This downloads model weights and can be slow or exceed CPU RAM/time limits on CPU Basic. ZeroGPU is the preferred option when available. The code preloads a local model when the runtime identifies itself as ZeroGPU; check Space logs before assuming GPU inference is active.

### Optional OpenAI-compatible endpoint

Set `EIM_BACKEND=openai`, `EIM_API_BASE`, `EIM_MODEL`, and, when required, `EIM_API_KEY` as a **Secret**. Do not commit tokens or keys into source files.

### Save corrections and repair memory across Space restarts

Hugging Face Space local disks are not a durable user-memory store. To sync memories to a private Dataset repository, create a private Dataset and add these Space Secrets/Variables:

```text
HF_TOKEN=<write-enabled token for that private Dataset>
EIM_MEMORY_REPO=your-account/your-private-dataset
```

`EIM_MEMORY_REPO` enables sync of EIM experience/repair memory and the correction-memory JSONL by default. Set `EIM_CORRECTIONS_REPO` to a different private Dataset if you want correction memory stored separately. The code syncs best-effort; failures are kept out of chat flow and surfaced only in its internal status. On a repository's first use, check Space logs and repository contents to verify that the secret has write access. Without remote sync or external persistent storage, corrections can disappear when an ephemeral Space restarts.

Do **not** enable any of these without understanding the risk: `EIM_ALLOW_NETWORK=1`, `EIM_ALLOW_RISKY_CODE=1`, `EIM_ALLOW_OUTSIDE_WORKSPACE=1`. Source checks and subprocess limits are defense-in-depth, not an OS security boundary; never execute untrusted code on a host with sensitive data. For a public multi-user service, use a disposable container/VM per job with filesystem isolation and outbound networking blocked.

## Local install and run

Use Python 3.10+ (the Space configuration above uses Python 3.12):

```bash
python -m pip install -r requirements.txt
python chat_app.py
```

The app needs a working model backend to answer new questions. Offline self-tests do not require model weights or a token.

## Tests (offline, no model inference)

Run from this directory:

```bash
python -m py_compile app.py eim_plus.py eval_eim.py chat_app.py correction_memory.py terminal_verify.py test_terminal_verify.py test_correction_memory.py test_eval_benchmarks.py
python app.py --selftest
python eim_plus.py --selftest
python chat_app.py --selftest
python eval_eim.py --selftest
python test_terminal_verify.py
python test_correction_memory.py
python test_eval_benchmarks.py
python app.py --backend-check
```

The tests include real subprocess executions and test negative cases (wrong output, exceptions, non-zero exits, timeouts and safety-policy blocks). Some tests intentionally use a scripted/fake language-model adapter to test engine behavior deterministically; those are explicitly offline engine tests, **not** evidence of the live model's coding accuracy. `test_eval_benchmarks.py` checks the benchmark reference implementations and their assertions in-process; it does not measure the model.

## Real before/after model benchmark

Use the **same deployed backend/model, prompt conditions, test set, iteration/candidate settings, and hardware class**; don't label an offline scripted test as a model benchmark. With a working model backend:

```bash
python eval_eim.py --label baseline --iterations 4 --candidates 3
# apply one engine version/change, then run exactly the same settings:
python eval_eim.py --label improved --iterations 4 --candidates 3
python eval_eim.py --compare baseline improved
```

The evaluator reports solve rate, uncertainty interval, runtime, candidate attempts, restarts, newly solved cases, and regressions. A local run with no model credentials/weights cannot support a claim that the model itself improved. For valid comparisons, use separate clean environments/runs and make sure both labels were produced by the model—not by `_ScriptedLM` self-tests.