| # HANDOFF β what you need to do next |
|
|
| This is the ordered checklist for taking the mindxtrain repo from "code is |
| done" to "demo is live." Each step is concrete; check it off when finished. |
|
|
| The repo state at handoff: |
|
|
| - Single canonical package at `mindxtrain/` (12 subpackages, ~100 modules). |
| - All stub `NotImplementedError` paths replaced with real Python (lazy imports |
| for heavyweight deps). |
| - 112/112 tests pass on a CPU-only laptop (`uv sync` + `uv run pytest -q`). |
| - Optional dep groups in `pyproject.toml`: `ml`, `eval`, `data`, `serve`, |
| `chain`, `obs`. Install only what you need. |
| - 12 YAML training recipes wired through the CLI. |
| - Coach UI (`/coach/`) serves all 12 recipes without GPU. |
|
|
| --- |
|
|
| ## 1. Local setup (no GPU; 10 minutes) |
|
|
| ```bash |
| cd /home/hacker/Desktop/mindXtrain |
| cp .env.example .env # then edit .env to fill in HF_TOKEN, etc. |
| uv sync # base install |
| uv run pytest -q # β 112 passed |
| uv run mindxtrain --help # all 9 verbs listed |
| ``` |
|
|
| **What goes in `.env`** (rest of the file is sane defaults): |
|
|
| | Var | Where to get it | |
| |---|---| |
| | `HF_TOKEN` | https://huggingface.co/settings/tokens (write scope) | |
| | `HF_HUB_USERNAME` | your HF handle | |
| | `LIGHTHOUSE_API_KEY` | https://files.lighthouse.storage/dashboard/apikey | |
| | `MINDXTRAIN_OPENAI_API_KEY` | optional; only if you want to use openai_compat backend | |
| |
| > **Optional (on-chain anchors):** `MINDXTRAIN_REGISTRY_ADDR` (ERC-8004 contract), |
| > `MINDXTRAIN_FACILITATOR_URL` (x402 facilitator). The publish path skips |
| > these gracefully if unset. |
| |
| ## 2. Provision the MI300X droplet (sign-up + 30 min) |
| |
| > **Fast path (Coach UI):** if you've populated `GITHUB_TOKEN`, |
| > `AMD_DEV_CLOUD_TOKEN`, and `AMD_DEV_CLOUD_SSH_KEY_ID` in `.env`, you can skip |
| > the manual SSH dance entirely: |
| > |
| > 1. `uv run uvicorn mindxtrain.operator.app:app --port 8080` |
| > 2. Open <http://localhost:8080/coach/>, scroll to step 6 ("Deploy"). |
| > 3. Click β **Push to GitHub** β β‘ **Provision MI300X droplet**. The droplet |
| > boots, cloud-init clones the repo from the SHA you just pushed, pulls the |
| > container, and runs `mindxtrain bench` automatically. All output streams |
| > live in the browser via SSE. |
| > |
| > Equivalent CLI: `mindxtrain github push && mindxtrain droplet provision`. |
| > |
| > The manual sequence below is preserved for scripted / CI use and as a |
| > fallback when the Coach UI isn't available. |
|
|
| ```bash |
| # Sign up at https://devcloud.amd.com β request a single MI300X. |
| # Wait for the droplet (typically same-day). |
| # SSH in: |
| ssh ubuntu@<droplet-ip> |
| |
| # Install podman if missing: |
| sudo apt-get update && sudo apt-get install -y podman podman-compose |
| |
| # Pull the canonical training container: |
| podman pull docker.io/rocm/primus:v26.2 |
| |
| # Snapshot the digest into the repo so others can reproduce: |
| podman inspect --format '{{index .RepoDigests 0}}' rocm/primus:v26.2 \ |
| | tee -a ops/containerfiles/digest.lock |
| |
| # Verify the GPU is visible: |
| podman run --rm --device=/dev/kfd --device=/dev/dri rocm/primus:v26.2 \ |
| rocminfo | head -50 |
| # β should show gfx942, 192 GB HBM3 |
| ``` |
|
|
| > **Cost watch:** $1.99/hr Γ planned hours. Budget ~$30 for the full demo |
| > pipeline (~15 GPU-hours). Leave the droplet **stopped** when not actively |
| > training. |
|
|
| ## 3. Install heavyweight deps inside the container |
|
|
| ```bash |
| # On the MI300X: |
| git clone <your-repo-url> /workspace/mindxtrain |
| cd /workspace/mindxtrain |
| podman run -it --rm \ |
| --device=/dev/kfd --device=/dev/dri \ |
| -v /workspace/mindxtrain:/workspace/mindxtrain \ |
| -w /workspace/mindxtrain \ |
| rocm/primus:v26.2 bash |
| |
| # Inside the container: |
| pip install -e ".[ml,eval,data,obs]" |
| # (skip `serve` and `chain` until you need them β they pull large wheels) |
| ``` |
|
|
| ## 4. Run the autotune probe (real, ~60 s) |
|
|
| ```bash |
| mindxtrain bench --gpu 0 --out plan.json |
| cat plan.json | jq '.attention_backend, .gemm_heuristic, .rccl_config' |
| # β "ck", "hipblaslt_default", "1gpu_noop" |
| ``` |
|
|
| Snapshot `plan.json` into the repo so the run is reproducible: |
|
|
| ```bash |
| cp plan.json ops/k8s/plan-mi300x.json |
| git add ops/k8s/plan-mi300x.json |
| git commit -m "snapshot autotune plan from mi300x" |
| ``` |
|
|
| ## 5. Train + eval + quantize (~ 2 hours total for the demo recipe) |
|
|
| ```bash |
| # Pick a recipe: instella_3b_lora is the AMD-on-AMD demo path (~30 min). |
| # Or qwen3_8b_sft_lora for the Qwen side prize (~75 min). |
| mindxtrain init --template instella_3b_lora --out run.yaml |
| |
| # Optional: edit run.yaml for your project name, dataset, output path. |
| $EDITOR run.yaml |
| |
| # Dataset prep (pulls + dedupes + tokenizes + packs): |
| mindxtrain dataset prep run.yaml --out ./out/dataset |
| |
| # Training: |
| mindxtrain train run.yaml --plan plan.json |
| # β ./out/runs/<run_name>/checkpoint/ |
| |
| # Evaluation (MMLU subset): |
| mindxtrain eval run.yaml |
| # β ./out/runs/<run_name>/eval/lm_eval.json |
| |
| # Quantize to FP8: |
| mindxtrain quantize run.yaml |
| # β ./out/runs/<run_name>/quantized/ |
| ``` |
|
|
| If `mindxtrain train` fails with `accelerate not found`: you forgot |
| `pip install -e ".[ml]"` inside the container (step 3). |
|
|
| ## 6. Build the manifest + verify |
|
|
| ```bash |
| # Generate the provenance manifest by hashing every artifact: |
| uv run python -c " |
| from pathlib import Path |
| from mindxtrain.config.loader import load_config |
| from mindxtrain.provenance.manifest import emit_receipt, ProvenanceHashes |
| cfg = load_config('run.yaml') |
| run = Path('./out/runs') / cfg.meta.run_name |
| m = emit_receipt( |
| cfg, |
| cfg.meta.run_name, |
| config_yaml_path=Path('run.yaml'), |
| dataset_manifest_path=run / 'dataset_manifest.json', |
| checkpoint_dir=run / 'checkpoint', |
| eval_json_path=run / 'eval/lm_eval.json', |
| ) |
| out = run / 'manifest.json' |
| out.write_text(m.model_dump_json(indent=2)) |
| print(out) |
| " |
| |
| # Verify it round-trips: |
| mindxtrain receipt ./out/runs/<run_name>/manifest.json --config run.yaml |
| # β all BLAKE3 fields = true (config, checkpoint, autotune_plan; dataset/eval if present) |
| ``` |
|
|
| > **Auto-emitted receipts (operator + CPU lane).** Runs launched through the |
| > operator β Coach UI or `POST /v1/training/jobs` β now write `manifest.json` |
| > automatically at completion via `provenance.manifest.emit_receipt_for_run`, |
| > alongside `config.snapshot.yaml` and `autotune_plan.json` in the run dir. The |
| > receipt **binds the frozen AutotunePlan hash to the checkpoint hash** β this is |
| > the AOT artifact that makes a run bitwise-verifiable (cf. Verde/RepOps). The |
| > Coach "Verifiable receipt" card re-checks it live; `mindxtrain receipt` does the |
| > same from a shell. The manual `emit_receipt` above remains the full GPU path |
| > (dataset + eval JSON included). On MI300X, also snapshot the AOTriton / |
| > hipBLASLt tuning caches next to `autotune_plan.json` so the compiled artifact β |
| > not just the plan β is reproducible across machines. |
|
|
| ## 7. Publish (HF Hub + Lighthouse + mindX register) |
|
|
| ```bash |
| # Push to HF (uses HF_TOKEN; private=False for the demo): |
| mindxtrain publish run.yaml --manifest ./out/runs/<run_name>/manifest.json |
| # β updates manifest.json in-place with hf_repo_id + lighthouse_cid |
| ``` |
|
|
| If `LIGHTHOUSE_API_KEY` is unset, the pin step skips gracefully and the |
| manifest gets a `cid://stub-β¦` placeholder. |
|
|
| ## 8. Deploy contracts (optional) |
|
|
| The demo can ship without on-chain anchors. Do these once, when ready: |
|
|
| ```bash |
| cd contracts |
| forge install |
| forge test # local Foundry tests pass |
| forge script script/Deploy.s.sol \ |
| --rpc-url $MINDXTRAIN_BASE_RPC_URL \ |
| --private-key $DEPLOYER_KEY \ |
| --broadcast |
| # β records contract address; paste into .env as MINDXTRAIN_REGISTRY_ADDR |
| ``` |
|
|
| Once `MINDXTRAIN_REGISTRY_ADDR` is set, `mindxtrain.provenance.erc8004.broadcast_attestation` |
| can anchor the manifest BLAKE3 on-chain. |
|
|
| ## 9. Serve the model + wire the production URL |
|
|
| The production URL is `https://mindx.pythai.net` β the Coach UI is at `/coach/` |
| and the public training-jobs API is at `/v1/training/jobs`. |
|
|
| ```bash |
| # Inside the rocm/vllm-dev container: |
| podman-compose -f ops/compose/compose_dev.yaml up -d |
| # β vLLM-ROCm at :8000, mindxtrain operator FastAPI at :8080 |
| |
| # Verify locally: |
| curl http://localhost:8080/coach/api/health |
| # β {"coach_version":"0.1.0", "recipes_available":>=14, ...} |
| |
| # Public training-jobs API smoke (bearer auth via MINDXTRAIN_API_KEY): |
| curl -X POST http://localhost:8080/v1/training/jobs \ |
| -H "Authorization: Bearer $MINDXTRAIN_API_KEY" \ |
| -H "Content-Type: application/json" \ |
| -d '{"recipe":"mindx_fallback_qwen3_1_5b_cpu_smoke"}' |
| # β {"job_id":"...", "status":"running", "backend":"trl_cpu", ...} |
| |
| # Reverse-proxy mindx.pythai.net β MI300X:8080 (Caddy/Cloudflare). |
| ``` |
|
|
| Once the proxy is live, `curl https://mindx.pythai.net/coach/api/health` |
| returns 200 from the public internet. |
|
|
| ## 10. Publish & demo |
|
|
| ```bash |
| # Push code: |
| git push origin main |
| |
| # End-to-end demo walk-through: |
| # 1. mindxtrain init β show CLI verbs |
| # 2. mindxtrain bench β 60-second autotune (the differentiator) |
| # 3. mindxtrain train β timelapse of training |
| # 4. mindxtrain quantize β FP8 weights |
| # 5. curl /v1/chat/completions β live inference |
| # 6. mindxtrain receipt β BLAKE3 reverify |
| # 7. Open /coach/ β click through the UI |
| # 8. Open /coach/dcoach β Imprint & Prove (CPU recall proof) |
| ``` |
|
|
| ## 11. Quality gates (run before every push) |
|
|
| ```bash |
| uv run ruff check . |
| uv run mypy mindxtrain/config mindxtrain/provenance |
| uv run pytest -q # β 112 passed |
| ``` |
|
|
| All three must pass before pushing to `main`. CI runs the same gates on the |
| `main` branch. |
|
|
| --- |
|
|
| ## What's still TODO |
|
|
| These paths are wired but require runtime/contracts/services to actually |
| flow end-to-end: |
|
|
| - **x402 metering** (`mindxtrain.provenance.x402`) β wired to httpx, needs |
| a deployed facilitator URL. |
| - **ERC-8004 broadcast** (`mindxtrain.provenance.erc8004.broadcast_attestation`) |
| β needs deployed attestation registry + signer key. |
| - **BANKON ENS** allocation (`mindxtrain.provenance.algorand.allocate_ens_subname`) |
| β needs the BANKON allocation service deployed. |
| - **AgenticPlace listing** (`mindxtrain.deploy.api_client.list_on_agenticplace`) |
| β needs `agenticplace.pythai.net` live. |
| - **mindX agent register** (`mindxtrain.deploy.api_client.register_with_mindx`) |
| β needs `mindx.pythai.net/v1/agents` live. |
|
|
| The framework itself ships as production-ready Apache-2.0; the integrations |
| above are paid/external services you stand up at your own pace. |
|
|
| --- |
|
|
| ## Quick reference |
|
|
| | What | Where | |
| |---|---| |
| | All CLI verbs | `mindxtrain --help` | |
| | All recipes | `mindxtrain init --list` | |
| | Coach UI | http://localhost:8080/coach/ | |
| | Per-module status | `docs/actualization_status.md` | |
| | Architecture | `docs/architecture.md` | |
| | Autotune detail | `docs/autotune.md` | |
| | Coach detail | `docs/coach.md` | |
| | CLI reference | `docs/cli.md` | |
| | YAML schema | `docs/yaml_schema.md` | |
| | dcoach proof loop | `docs/dcoach.md` | |
| | Frozen blueprints | `docs/blueprints/{mindXtrain,mindXtrain2}.md` | |
|
|