Expanded_Repetition / README.md
Expanded-Repetition's picture
Upload 2 files
0d566ba verified
|
Raw History Blame Contribute Delete
6.68 kB

A newer version of the Gradio SDK is available: 6.30.0

Upgrade
metadata
title: EIM+ Real Verification Engineer
emoji: 🧪
colorFrom: blue
colorTo: indigo
sdk: gradio
sdk_version: 4.44.0
python_version: 3.12.12
app_file: chat_app.py
pinned: false

EIM+ Real Verification Engineer

A Gradio coding assistant with an iterative repair engine, actual child-process test execution, persistent correction/repair memory (optional private Hugging Face Dataset sync), and a repeatable benchmark harness.

Files

  • app.py — core engine: model providers, constrained subprocess execution, scoring, and experience memory.
  • eim_plus.py — candidate generation, independent test verification, holdout checks, weak-test/mutation checks, repair logs, retries, and result cache controls.
  • chat_app.py — bilingual chat UI and file-aware assistant.
  • terminal_verify.py — stand-alone real terminal runner with stdout/stderr/exit/timeout evidence, explicit opt-ins for higher-risk operations, and JSON reports.
  • correction_memory.py — explicit user-correction memory in Arabic/English, plus optional Hugging Face Dataset sync.
  • eval_eim.py — 25-task model benchmark; stores labeled runs and compares rates, confidence intervals, attempts, restarts, regressions, and time.

Deploy on Hugging Face Spaces

  1. Create a Gradio Space and push all files from this folder to its repository.
  2. Select ZeroGPU hardware in the Space settings. On a free personal account, Hugging Face currently requires a verified email and an account older than 30 days; eligible accounts can host up to two ZeroGPU Spaces. Current official docs list a 5 GPU-minute daily quota for Free accounts, shared at account level; this is not unlimited compute. Details can change: ZeroGPU official documentation.
  3. This app automatically chooses the local model backend when the runtime reports ZeroGPU, so the selected GPU is used; the optional @spaces.GPU wrapper is used when spaces is installed. Do not choose a paid hardware flavor if your goal is free hosting.
  4. Choose a model backend using Space Variables / Secrets:

Option A — no locally downloaded weights (remote inference)

Set Space Variables:

EIM_BACKEND=hf
EIM_MODEL=Mungert/Qwen3-Coder-30B-A3B-Instruct-GGUF

Add HF_TOKEN under Space Secrets with permission to use the required Hugging Face Inference Provider. Remote inference availability/pricing depends on the provider and model; a token by itself does not guarantee free inference.

Option B — load model weights in the Space

Set Space Variables:

EIM_BACKEND=local
EIM_MODEL=Mungert/Qwen3-Coder-30B-A3B-Instruct-GGUF
EIM_GPU_SECONDS=60

This downloads model weights and can be slow or exceed CPU RAM/time limits on CPU Basic. ZeroGPU is the preferred option when available. The code preloads a local model when the runtime identifies itself as ZeroGPU; check Space logs before assuming GPU inference is active.

Optional OpenAI-compatible endpoint

Set EIM_BACKEND=openai, EIM_API_BASE, EIM_MODEL, and, when required, EIM_API_KEY as a Secret. Do not commit tokens or keys into source files.

Save corrections and repair memory across Space restarts

Hugging Face Space local disks are not a durable user-memory store. To sync memories to a private Dataset repository, create a private Dataset and add these Space Secrets/Variables:

HF_TOKEN=<write-enabled token for that private Dataset>
EIM_MEMORY_REPO=your-account/your-private-dataset

EIM_MEMORY_REPO enables sync of EIM experience/repair memory and the correction-memory JSONL by default. Set EIM_CORRECTIONS_REPO to a different private Dataset if you want correction memory stored separately. The code syncs best-effort; failures are kept out of chat flow and surfaced only in its internal status. On a repository's first use, check Space logs and repository contents to verify that the secret has write access. Without remote sync or external persistent storage, corrections can disappear when an ephemeral Space restarts.

Do not enable any of these without understanding the risk: EIM_ALLOW_NETWORK=1, EIM_ALLOW_RISKY_CODE=1, EIM_ALLOW_OUTSIDE_WORKSPACE=1. Source checks and subprocess limits are defense-in-depth, not an OS security boundary; never execute untrusted code on a host with sensitive data. For a public multi-user service, use a disposable container/VM per job with filesystem isolation and outbound networking blocked.

Local install and run

Use Python 3.10+ (the Space configuration above uses Python 3.12):

python -m pip install -r requirements.txt
python chat_app.py

The app needs a working model backend to answer new questions. Offline self-tests do not require model weights or a token.

Tests (offline, no model inference)

Run from this directory:

python -m py_compile app.py eim_plus.py eval_eim.py chat_app.py correction_memory.py terminal_verify.py test_terminal_verify.py test_correction_memory.py test_eval_benchmarks.py
python app.py --selftest
python eim_plus.py --selftest
python chat_app.py --selftest
python eval_eim.py --selftest
python test_terminal_verify.py
python test_correction_memory.py
python test_eval_benchmarks.py
python app.py --backend-check

The tests include real subprocess executions and test negative cases (wrong output, exceptions, non-zero exits, timeouts and safety-policy blocks). Some tests intentionally use a scripted/fake language-model adapter to test engine behavior deterministically; those are explicitly offline engine tests, not evidence of the live model's coding accuracy. test_eval_benchmarks.py checks the benchmark reference implementations and their assertions in-process; it does not measure the model.

Real before/after model benchmark

Use the same deployed backend/model, prompt conditions, test set, iteration/candidate settings, and hardware class; don't label an offline scripted test as a model benchmark. With a working model backend:

python eval_eim.py --label baseline --iterations 4 --candidates 3
# apply one engine version/change, then run exactly the same settings:
python eval_eim.py --label improved --iterations 4 --candidates 3
python eval_eim.py --compare baseline improved

The evaluator reports solve rate, uncertainty interval, runtime, candidate attempts, restarts, newly solved cases, and regressions. A local run with no model credentials/weights cannot support a claim that the model itself improved. For valid comparisons, use separate clean environments/runs and make sure both labels were produced by the model—not by _ScriptedLM self-tests.