Faaz
Add project context doc and WebSight batch uploader
07de2d7
|
Raw
History Blame
25.1 kB

MINDI 1.5 Vision-Coder β€” Complete Project Context

Last updated: April 16, 2026 Purpose: This file contains ALL context needed to continue development with any AI assistant. It covers architecture decisions, errors encountered, fixes applied, training state, and exact next steps.


1. PROJECT OVERVIEW

MINDI 1.5 Vision-Coder is a multimodal AI model that generates frontend code (HTML/CSS/JS, Next.js, Tailwind) from UI screenshots and text prompts. It combines:

  • Qwen/Qwen2.5-Coder-7B-Instruct β€” 7.62B param base LLM (Apache 2.0)
  • CLIP ViT-L/14 β€” Frozen vision encoder for UI screenshot understanding
  • LoRA adapters β€” Efficient fine-tuning (r=64, alpha=128)
  • Vision-Language Fusion β€” Prepend visual tokens to text embeddings
  • 22 MINDI Special Tokens β€” Structured agentic reasoning (think, code, critique, fix, etc.)
  • 3-Phase Training Strategy β€” Progressive training on MI300X 192GB

Repos:

  • GitHub: https://github.com/Faaz345/MINDI-1.5-Vision-Coder.git (branch: master)
  • HuggingFace Model: Mindigenous/MINDI-1.5-Vision-Coder (private, push as master:main)
  • HuggingFace Dataset: Mindigenous/MINDI-1.5-training-data (private)
  • HF Token: Set as HF_TOKEN environment variable (stored separately, not in repo)

2. DIRECTORY STRUCTURE

MINDI-1.5-Vision-Coder/
β”œβ”€β”€ src/
β”‚   β”œβ”€β”€ model/
β”‚   β”‚   β”œβ”€β”€ architecture.py       # Qwen2.5-Coder + LoRA wrapper (NOT nn.Module)
β”‚   β”‚   β”œβ”€β”€ mindi_model.py        # MINDI15 main class (nn.Module)
β”‚   β”‚   β”œβ”€β”€ vision_encoder.py     # CLIP ViT-L/14 (frozen) + trainable projection
β”‚   β”‚   β”œβ”€β”€ fusion_layer.py       # VisionLanguageFusion with text_gate
β”‚   β”‚   └── __init__.py
β”‚   β”œβ”€β”€ training/
β”‚   β”‚   β”œβ”€β”€ mindi_trainer.py      # MINDITrainer: 3-phase loop, streaming data
β”‚   β”‚   β”œβ”€β”€ data_pipeline.py      # Data processing pipeline
β”‚   β”‚   └── __init__.py
β”‚   β”œβ”€β”€ agents/                   # Agentic pipeline (orchestrator, error fixer, UI critic)
β”‚   β”œβ”€β”€ inference/                # Generation pipeline
β”‚   β”œβ”€β”€ evaluation/               # Evaluation framework
β”‚   β”œβ”€β”€ search/                   # Tavily search agent
β”‚   β”œβ”€β”€ sandbox/                  # E2B/Docker code execution
β”‚   β”œβ”€β”€ tokenizer/                # MINDI tokenizer wrapper
β”‚   └── utils/                    # Config & env loaders
β”œβ”€β”€ scripts/
β”‚   β”œβ”€β”€ train.py                  # Master training launcher (--dry_run, --phase, --resume)
β”‚   β”œβ”€β”€ download_websight.py      # Download WebSight v0.2 from HF
β”‚   β”œβ”€β”€ upload_websight_images.py # Upload images to HF in batches (10K/dir limit)
β”‚   β”œβ”€β”€ gpu_diagnostic.py         # 6-stage GPU test for MI300X
β”‚   └── ... (data processing scripts)
β”œβ”€β”€ configs/
β”‚   β”œβ”€β”€ training_config.yaml      # Training hyperparameters
β”‚   β”œβ”€β”€ model_config.yaml         # Model architecture config
β”‚   β”œβ”€β”€ data_config.yaml          # Data sources and processing
β”‚   └── search_config.yaml        # Tavily search settings
β”œβ”€β”€ data/
β”‚   β”œβ”€β”€ processed/                # Text training data (train.jsonl, val.jsonl, test.jsonl)
β”‚   β”œβ”€β”€ websight/                 # Vision data (52,500 images in subdirs + JSONL)
β”‚   β”‚   β”œβ”€β”€ train.jsonl           # 50,000 vision-code pairs
β”‚   β”‚   β”œβ”€β”€ val.jsonl             # 2,500 vision-code pairs
β”‚   β”‚   └── images/
β”‚   β”‚       β”œβ”€β”€ 00/               # ws_0000000.jpg - ws_0009999.jpg (10K each)
β”‚   β”‚       β”œβ”€β”€ 01/
β”‚   β”‚       β”œβ”€β”€ 02/
β”‚   β”‚       β”œβ”€β”€ 03/
β”‚   β”‚       β”œβ”€β”€ 04/
β”‚   β”‚       └── 05/               # ws_0050000.jpg - ws_0052499.jpg (2,500)
β”‚   β”œβ”€β”€ tokenizer/
β”‚   β”‚   β”œβ”€β”€ mindi_tokenizer/      # Custom tokenizer (vocab 151,685)
β”‚   β”‚   └── base_tokenizer/       # Original Qwen tokenizer
β”‚   └── raw/                      # Raw downloaded data sources
β”œβ”€β”€ api/                          # FastAPI endpoints
β”œβ”€β”€ checkpoints/                  # Model checkpoints
β”œβ”€β”€ logs/                         # Training logs
β”œβ”€β”€ requirements.txt              # Full requirements
β”œβ”€β”€ requirements-training.txt     # Lean MI300X Docker requirements
β”œβ”€β”€ setup_mi300x.sh               # MI300X Docker setup script
β”œβ”€β”€ .gitattributes                # LFS tracking for large tokenizer files
└── .gitignore

3. ARCHITECTURE DETAILS

3.1 Model Components

Component Class File Params Trainable
Base LLM MINDIArchitecture architecture.py 7.62B No (frozen)
LoRA via PEFT architecture.py 161.5M Yes
CLIP Vision VisionEncoder vision_encoder.py 304M 4.2M (projection only)
Fusion VisionLanguageFusion fusion_layer.py 16.8M Yes
Total MINDI15 mindi_model.py 8.1B 182.5M (2.25%)

3.2 CRITICAL Architecture Notes

  1. MINDIArchitecture is NOT an nn.Module β€” it's a plain Python wrapper class. The actual trainable PeftModel is accessed via self.architecture.get_model() and registered as self.llm in MINDI15.__init__().

  2. self.llm = self.architecture.get_model() β€” This line in mindi_model.py registers the PeftModel as a proper submodule so model.parameters() can find LoRA params. Without this, the optimizer gets zero trainable parameters.

  3. Vision encoder uses float32 projection β€” CLIP backbone is frozen, only self.projection (Linear 1024β†’4096) trains. The projection operates in float32 for stability even though the rest is bf16.

  4. Fusion layer has text_gate β€” A learnable scalar parameter (init=0) that creates a residual path for text-only inputs. This ensures gradients flow to the fusion layer during Phase 2 even when processing text-only batches (which have no vision tokens and would otherwise be pure passthrough with no gradient).

3.3 Forward Pass Flow

Image β†’ CLIP (frozen) β†’ 256 patches (1024) β†’ projection (4096) β†’ visual_tokens
Text β†’ tokenizer β†’ input_ids β†’ LLM embedding layer β†’ text_embeds

With image:   fusion = [gated_visual_tokens; text_embeds]  (prepend)
Without image: fusion = text_embeds + sigmoid(text_gate) * (transformed - text_embeds)

fusion β†’ LLM layers (with LoRA) β†’ logits β†’ loss (cross-entropy, labels=-100 for padding)

3.4 LoRA Configuration

LoraConfig(
    r=64,
    lora_alpha=128,
    lora_dropout=0.05,
    target_modules=["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"],
    bias="none",
    task_type=TaskType.CAUSAL_LM,
)

3.5 MINDI Special Tokens (22 total, 11 pairs)

<|think_start|> / <|think_end|>         β€” Internal reasoning
<|code_start|> / <|code_end|>           β€” Generated code blocks
<|file_start|> / <|file_end|>           β€” File references
<|critique_start|> / <|critique_end|>   β€” Self-critique
<|suggest_start|> / <|suggest_end|>     β€” Suggestions
<|search_start|> / <|search_end|>       β€” Search context
<|error_start|> / <|error_end|>         β€” Error messages
<|fix_start|> / <|fix_end|>             β€” Fix attempts
<|vision_start|> / <|vision_end|>       β€” Vision input markers
<|sandbox_start|> / <|sandbox_end|>     β€” Sandbox execution
<|context_start|> / <|context_end|>     β€” Context block

4. TRAINING PIPELINE

4.1 Three-Phase Training Strategy

Phase Name Steps LR Batch Components Data Purpose
1 phase1_lora 5,000 2e-4 16 LoRA only Text-only code Teach coding patterns
2 phase2_vision_bridge 2,500 1e-5 8 Vision+Fusion WebSight images Align visual tokens
3 phase3_all 2,500 5e-5 12 All trainable Mixed text+vision Joint fine-tuning

Total: 10,000 steps

4.2 Training Data

Text data (Phase 1 + Phase 3):

  • data/processed/train.jsonl β€” 1,304,486 examples, 4.18 GB
  • data/processed/val.jsonl β€” 72,471 examples
  • Sources: CodeAlpaca, CodeFeedback, EvolCode, MagicCoder, StarCoder (5 langs), Synthetic Next.js

Vision data (Phase 2 + Phase 3):

  • data/websight/train.jsonl β€” 50,000 image+code pairs, 114 MB JSONL
  • data/websight/val.jsonl β€” 2,500 image+code pairs, 5.7 MB JSONL
  • data/websight/images/ β€” 52,500 JPG screenshots in 6 subdirectories (11.6 GB)
  • Source: HuggingFaceM4/WebSight v0.2 (UI screenshot β†’ HTML/CSS pairs)

WebSight JSONL format:

{
  "id": "websight_0000001",
  "type": "vision_code",
  "source": "websight_v0.2",
  "image_path": "data/websight/images/00/ws_0000001.jpg",
  "messages": [
    {"role": "system", "content": "You are MINDI 1.5 Vision-Coder..."},
    {"role": "user", "content": "<|vision_start|><|vision_end|>\nGenerate the HTML/CSS code for this UI screenshot."},
    {"role": "assistant", "content": "<|think_start|>...<|think_end|>\n<|code_start|>\n...HTML/CSS...\n<|code_end|>"}
  ],
  "metadata": {"dataset": "websight", "version": "v0.2"}
}

IMPORTANT: Images are organized in subdirectories of ≀10,000 files each because HuggingFace has a 10K files/directory limit. The JSONL image_path fields reference the subdirectory structure (e.g., data/websight/images/00/ws_0000001.jpg).

4.3 Data Loading

  • StreamingJSONLDataset (in mindi_trainer.py) β€” Streams from disk line-by-line, tokenizes on-the-fly
  • Shuffle buffer of 10,000 examples (reservoir-style)
  • Image loading via _load_image() β€” loads PIL images from relative paths
  • Custom collate function β€” stacks tensors, keeps images as a list
  • Phase routing β€” Phase 1 uses text data, Phase 2 uses WebSight, Phase 3 uses text (with inline images if present)

4.4 Key Training Features

  • bf16 precision β€” Required for MI300X stability (NOT fp16)
  • Gradient checkpointing β€” Enabled even with 192GB VRAM
  • torch.compile() β€” Optional, works on ROCm
  • Cosine LR with warmup β€” Per-phase schedules
  • Gradient accumulation β€” Configurable per phase (default: 4)
  • Emergency checkpoint β€” Saved on Ctrl+C
  • Crash checkpoint β€” Saved on unhandled exceptions

5. TRAINING HISTORY & RESULTS

5.1 Phase 1 Dry Run β€” SUCCESS βœ…

Date: April 15, 2026 (on DigitalOcean MI300X) Command: python3 scripts/train.py --dry_run --no_wandb Result: Loss dropped from 1.94 β†’ 0.85 in 10 steps, completed in 12.1 minutes VRAM usage: ~14.3 GB

5.2 Phase 2 β€” First Attempt FAILED ❌

Error: element 0 of tensors does not require grad and does not have a grad_fn Root cause: Phase 2 trains vision+fusion with LoRA frozen. Text-only data means fusion is pure passthrough (no gradient path). The fusion layer was getting zero gradients because without vision tokens, the text-only path was return text_embeds, attention_mask β€” a pure passthrough with no learnable operation. Fix: Added text_gate learnable residual parameter to VisionLanguageFusion. Text-only path changed to: text_embeds + sigmoid(text_gate) * (transformed - text_embeds). Also built the WebSight vision data pipeline to provide actual image+code pairs for Phase 2.

5.3 Full 3-Phase Dry Run β€” NOT YET COMPLETED

The MI300X GPU kept hanging/wedging (see Section 6). Phase 2 and 3 with the new WebSight data pipeline have NOT been tested yet.


6. ERRORS & FIXES β€” COMPLETE HISTORY

6.1 GPU Hang #1 β€” HSA_OVERRIDE_GFX_VERSION

Symptom: GPU completely unresponsive. torch.cuda.get_device_name(0) returns blank, any CUDA operation hangs. Root cause: HSA_OVERRIDE_GFX_VERSION=11.0.0 was set in the Docker container. This conflicts with ROCm 7.0's native MI300X/gfx942 support. Fix: Do NOT set HSA_OVERRIDE_GFX_VERSION. ROCm 7.0 natively supports gfx942. Remove it from all scripts/env. Commit: 4a33f96 Remove HSA_OVERRIDE_GFX_VERSION

6.2 No Trainable Parameters in Optimizer

Symptom: RuntimeError: No trainable parameters in phase 'phase1_lora' Root cause: MINDIArchitecture is a plain Python class (not nn.Module). When MINDI15 calls model.parameters(), it doesn't find the LoRA parameters because the PeftModel isn't registered as a submodule. Fix: Added self.llm = self.architecture.get_model() in MINDI15.__init__() to register the PeftModel as a proper nn.Module submodule. Updated forward() and generate() to use self.llm instead of self.architecture.get_model(). Commit: cdc806e Fix: register LLM as nn.Module submodule so optimizer finds LoRA params

6.3 extra_special_tokens Format Error

Symptom: TypeError when loading tokenizer β€” transformers 4.55 expects extra_special_tokens as a dict, not a list. Fix: Changed data/tokenizer/mindi_tokenizer/tokenizer_config.json: converted extra_special_tokens from list format to {"token_name": {"content": "..."}} dict format. Commit: 02eef51 Fix extra_special_tokens: list to dict for transformers 4.55

6.4 Phase 2 Gradient Flow Crash

Symptom: element 0 of tensors does not require grad and does not have a grad_fn during Phase 2 Root cause: Text-only data β†’ no vision tokens β†’ fusion is pure passthrough β†’ no gradient path to fusion parameters. Fix: (1) Added text_gate learnable residual gate in VisionLanguageFusion for text-only gradient flow. (2) Built WebSight vision data pipeline with actual image+code pairs. Commit: 4e9835e Fix Phase 2: fusion layer processes text-only via learnable residual gate

6.5 Git LFS Issues

Symptom: tokenizer.json files >10MB causing push failures to HuggingFace. Fix: Configured .gitattributes for LFS tracking. Ran git lfs migrate import to rewrite history. Force-pushed to both GitHub and HF. Commit: 161c946 Track large tokenizer files with Git LFS

6.6 HuggingFace Auth for MI300X Clone

Symptom: git clone from HF failed with auth error in Docker container. Fix: Use token as both username and password: https://hf_TOKEN:hf_TOKEN@huggingface.co/Mindigenous/MINDI-1.5-Vision-Coder.git Also needed: apt-get install -y git-lfs && git lfs install

6.7 GPU Hang #2 β€” Driver Wedge After Heavy I/O

Symptom: After interrupted HF upload + training attempt, GPU shows 100% utilization with 0% VRAM in rocm-smi. Even torch.randn(device='cuda') hangs. Docker restart insufficient. Kernel log: amdgpu: GPU reset begin! β†’ device wedged, but recovered through reset β†’ But GPU% stays at 100%. Fix:

  1. docker stop rocm
  2. echo 1 > /sys/bus/pci/devices/0000:83:00.0/reset (PCI address from lspci | grep AMD)
  3. If GPU% still 100%: modprobe -r amdgpu && modprobe amdgpu
  4. Verify rocm-smi shows GPU% = 0% before restarting Docker Status: Droplet was deleted. Will need to handle this on fresh droplet if it recurs.

6.8 HuggingFace Upload Limits

Symptom: 413 Payload Too Large (25K files/commit) and 400 Bad Request (10K files/directory) Fix: Reorganized 52,500 images into 6 subdirectories of ≀10K files (00/ through 05/). Upload in separate commits per subdirectory. Updated JSONL image_path fields to include subdirectory. Script: scripts/upload_websight_images.py


7. MI300X DEPLOYMENT

7.1 Infrastructure

  • Provider: DigitalOcean GPU Droplet
  • GPU: AMD Instinct MI300X (192GB HBM3 VRAM)
  • Cost: $1.99/hr
  • Docker container: Named rocm, accessed via docker exec -it rocm /bin/bash
  • ROCm/HIP: 7.0.51831-a3e329ad8
  • PyTorch: 2.9.0.dev20250821+rocm7.0.0
  • Python: 3.10

7.2 Critical Environment Variables

export HF_TOKEN=<your-hf-token>     # Get from HF settings page
export HF_HUB_DISABLE_PROGRESS_BARS=1
export PYTORCH_ROCM_ARCH=gfx942
# DO NOT SET: HSA_OVERRIDE_GFX_VERSION (causes GPU hang on ROCm 7.0)

7.3 Fresh Droplet Setup Procedure

# 1. SSH into droplet
ssh root@<DROPLET_IP>

# 2. Start Docker
docker start rocm
docker exec -it rocm /bin/bash

# 3. Set environment (inside Docker)
export HF_TOKEN=<your-hf-token>     # Get from HF settings page
export HF_HUB_DISABLE_PROGRESS_BARS=1
export PYTORCH_ROCM_ARCH=gfx942

# 4. Quick GPU test
python3 -c "import torch; print('GPU:', torch.cuda.get_device_name(0)); x=torch.randn(100,device='cuda'); print('OK:', x.sum().item())"

# 5. Install git-lfs
apt-get update && apt-get install -y git-lfs
git lfs install

# 6. Clone code repo
cd /workspace
git clone https://$HF_TOKEN:$HF_TOKEN@huggingface.co/Mindigenous/MINDI-1.5-Vision-Coder.git
cd MINDI-1.5-Vision-Coder

# 7. Install requirements
pip install -r requirements-training.txt

# 8. Download training data from HF dataset repo
python3 -c "
from huggingface_hub import snapshot_download
import os
# HF_TOKEN must be set in environment
snapshot_download(
    repo_id='Mindigenous/MINDI-1.5-training-data',
    repo_type='dataset',
    local_dir='data',
    token=os.environ['HF_TOKEN'],
)
print('Data download complete!')
"

# 9. Verify data
ls -la data/processed/
ls -la data/websight/
ls data/websight/images/ | head

# 10. Run GPU diagnostic
python3 scripts/gpu_diagnostic.py

# 11. Dry run
python3 scripts/train.py --dry_run --no_wandb

# 12. Full training
python3 scripts/train.py --no_wandb

7.4 GPU Hang Recovery (if it happens again)

# From HOST (not inside Docker):
docker stop rocm
echo 1 > /sys/bus/pci/devices/0000:83:00.0/reset   # PCI address may differ
rocm-smi  # Verify GPU% = 0%
# If still 100%:
modprobe -r amdgpu && modprobe amdgpu
rocm-smi  # Should show 0% now
docker start rocm

8. HF DATASET REPO STRUCTURE

Repo: Mindigenous/MINDI-1.5-training-data (private, type: dataset)

β”œβ”€β”€ .gitattributes
β”œβ”€β”€ README.md
β”œβ”€β”€ processed/
β”‚   β”œβ”€β”€ train.jsonl              # 1.3M text examples
β”‚   β”œβ”€β”€ val.jsonl
β”‚   β”œβ”€β”€ test.jsonl
β”‚   β”œβ”€β”€ filter_report.json
β”‚   β”œβ”€β”€ mindi_filtered.jsonl
β”‚   └── split_meta.json
β”œβ”€β”€ raw/                         # Original data sources (11 files)
β”œβ”€β”€ tokenizer/
β”‚   β”œβ”€β”€ base_tokenizer/
β”‚   └── mindi_tokenizer/
└── websight/
    β”œβ”€β”€ train.jsonl              # 50K vision-code JSONL
    β”œβ”€β”€ val.jsonl                # 2.5K vision-code JSONL
    └── images/
        β”œβ”€β”€ 00/                  # 10,000 JPGs
        β”œβ”€β”€ 01/                  # 10,000 JPGs
        β”œβ”€β”€ 02/                  # 10,000 JPGs
        β”œβ”€β”€ 03/                  # 10,000 JPGs
        β”œβ”€β”€ 04/                  # 10,000 JPGs (uploading as of April 16)
        └── 05/                  # 2,500 JPGs  (uploading as of April 16)

NOTE: As of April 16, 2026, subdirectories 00-03 are uploaded. 04 and 05 are being uploaded via scripts/upload_websight_images.py. If upload was interrupted, re-run the script β€” it skips already-uploaded subdirs.


9. GIT HISTORY (CHRONOLOGICAL)

553fbf7 feat: initial project scaffold for MINDI 1.5 Vision-Coder
11e0d89 Day 1 Complete: Tokenizer setup β€” 22 MINDI special tokens (vocab 151,685)
59c6c97 Day 2 COMPLETE: 1.48M examples processed, 6GB dataset, WebSight done
2ff5c54 Day 3 COMPLETE: Full model architecture (7 files)
1c36b28 Fix train.py: mem -> memory on line 225
f04f58b Fix setup_mi300x.sh step 2 + add project context summary
35fd5fc Fix setup_mi300x.sh for Docker container on MI300X droplet
5fb9ec3 Add GPU diagnostic script, fix architecture loading with sync
161c946 Track large tokenizer files with Git LFS
4a33f96 Remove HSA_OVERRIDE_GFX_VERSION - ROCm 7.0 native MI300X support
24b5fb1 Add requirements-training.txt for MI300X Docker
02eef51 Fix extra_special_tokens: list to dict for transformers 4.55
cdc806e Fix: register LLM as nn.Module submodule so optimizer finds LoRA params
4e9835e Fix Phase 2: fusion layer text_gate for gradient flow
672896a Add WebSight vision data pipeline: download, image-aware loader, phase routing

10. WHAT WORKS (VERIFIED) βœ…

  1. Tokenizer β€” 151,685 vocab with 22 MINDI special tokens, loads correctly
  2. Model initialization β€” MINDI15 loads all 4 components, 182.5M trainable params
  3. GPU diagnostic — All 6 tests pass (bf16 matmul, 1GB alloc, CPU→CUDA transfer, forward pass)
  4. Phase 1 dry run β€” Loss 1.94 β†’ 0.85 in 10 steps βœ…
  5. WebSight download β€” 52,500 images (11.6 GB) downloaded and organized
  6. Data format β€” JSONL with image_path references, streaming dataset works
  7. Git LFS β€” Large tokenizer files tracked correctly
  8. Code pushed β€” All code on GitHub master + HF model repo main

11. WHAT REMAINS (TODO) ❌

  1. Complete WebSight upload to HF β€” Subdirs 04 and 05 still uploading (re-run scripts/upload_websight_images.py if interrupted)
  2. Full 3-phase dry run β€” Phase 2 (WebSight) and Phase 3 (mixed) NOT yet tested with the vision pipeline
  3. Full production training β€” 10,000 steps total (Phase 1: 5K, Phase 2: 2.5K, Phase 3: 2.5K)
  4. Inference testing β€” Generate code from screenshots after training
  5. Commit upload_websight_images.py and context.md β€” These new files need to be pushed

12. KNOWN ISSUES & GOTCHAS

DO NOT:

  • Set HSA_OVERRIDE_GFX_VERSION=11.0.0 β€” kills GPU on ROCm 7.0
  • Use fp16 on MI300X β€” use bf16 for stability
  • Try to upload >10K files to a single HF directory β€” split into subdirs
  • Try to commit >25K files in a single HF commit β€” batch commits
  • Use the global Python (base env) on Windows β€” use venv (global torch DLL is broken)

WATCH OUT FOR:

  • GPU hanging after heavy I/O β€” check rocm-smi shows 0% GPU before training
  • Data paths β€” WebSight images use relative paths from project root in JSONL
  • MINDIArchitecture is NOT nn.Module β€” always use self.llm inside MINDI15
  • The text_gate in fusion starts at 0 (sigmoid=0.5) β€” this is intentional
  • On MI300X, Docker container named rocm β€” always docker exec -it rocm /bin/bash

13. COMMANDS REFERENCE

Local (Windows, PowerShell, in venv):

# Activate venv
& ".\venv\Scripts\Activate.ps1"

# Download WebSight
$env:HF_TOKEN="<your-hf-token>"
python scripts/download_websight.py --num_train 50000 --num_val 2500

# Upload WebSight images to HF (handles subdirs, retry, skip)
python scripts/upload_websight_images.py

# Push code to GitHub + HF
git push origin master
git push hf master:main

MI300X (Linux, Docker, inside container):

# Dry run (10 steps per phase)
python3 scripts/train.py --dry_run --no_wandb

# Full training
python3 scripts/train.py --no_wandb

# Single phase
python3 scripts/train.py --phase 1 --no_wandb
python3 scripts/train.py --phase 2 --no_wandb
python3 scripts/train.py --phase 3 --no_wandb

# Resume from checkpoint
python3 scripts/train.py --resume checkpoints/training/phase1_lora_step5000 --no_wandb

# GPU diagnostic
python3 scripts/gpu_diagnostic.py

14. NEXT SESSION CHECKLIST

When continuing with a new AI assistant:

  1. Open this directory in your IDE
  2. Read this file first to get full context
  3. Check WebSight upload status:
    python -c "import os; from huggingface_hub import HfApi; api=HfApi(token=os.environ['HF_TOKEN']); files=[f for f in api.list_repo_files('Mindigenous/MINDI-1.5-training-data', repo_type='dataset') if 'websight/images' in f]; print(f'{len(files)} images in HF repo')"
    
  4. If <52,500: re-run python scripts/upload_websight_images.py
  5. Push any uncommitted files:
    git add scripts/upload_websight_images.py context.md
    git commit -m "Add WebSight batch uploader and project context"
    git push origin master
    git push hf master:main
    
  6. Spin up fresh MI300X droplet on DigitalOcean
  7. Follow Section 7.3 for setup procedure
  8. Run dry run first to verify all 3 phases work
  9. Then full training β€” python3 scripts/train.py --no_wandb

15. DATA FILE LOCATIONS ON HF DATASET REPO

When cloning data on MI300X using snapshot_download, files will land at:

HF Repo Path Local Path (relative to project root)
processed/train.jsonl data/processed/train.jsonl
processed/val.jsonl data/processed/val.jsonl
websight/train.jsonl data/websight/train.jsonl
websight/val.jsonl data/websight/val.jsonl
websight/images/00/*.jpg data/websight/images/00/*.jpg
tokenizer/mindi_tokenizer/* data/tokenizer/mindi_tokenizer/*

The snapshot_download(local_dir='data') call places everything correctly because the HF repo structure mirrors the local data/ directory.


This context file was created on April 16, 2026 during Claude Opus 4.6 session to ensure project continuity.