Production-ready: VRAM-sized batches, frontend API URL fix, repo hygiene 7ac44e4 Running Luturtun Claude Opus 4.8 commited on Jun 16
Add temporary /api/debug/gpu to read real VRAM for batch sizing 0e61f52 Luturtun Claude Opus 4.8 commited on Jun 16
Symmetric encoder batching: ENC_BATCH cap + chunking + GPU probe 0204ef7 Luturtun Claude Opus 4.8 commited on Jun 16
Probe: report actual prompt_tokens to confirm worst-case length e602665 Luturtun Claude Opus 4.8 commited on Jun 16
Probe: binary-search exact max batch (not just powers of two) 26157e5 Luturtun Claude Opus 4.8 commited on Jun 16
Encoder: use batch word_ids (drop per-sentence re-tokenize); per-model DEC_BATCH + GPU probe 7dc0517 Luturtun Claude Opus 4.8 commited on Jun 16
Fix encoder LoRA scaling (read alpha from ckpt); chunk decoder batch 246db50 Luturtun Claude Opus 4.8 commited on Jun 16
Fix ZeroGPU startup: load PEFT adapter on CPU; torch_dtype->dtype 567d349 Luturtun Claude Opus 4.8 commited on Jun 16
Replace models: new encoder + gemma-3-1b/qwen3 decoders, drop llama3 34cc2e8 Luturtun commited on Jun 16
Next.js and React implementation. UI improvements. Cloze test removed from the website. 161900c Luturtun commited on May 9