Competitive roadmap — accurate status vs competition
Law: (S=K(T_1+T_2+T_3)) pin D1D38A (unchanged; never fitted to BLEU).
Measured: 2026-07-21 SOTA push (m6_sota_push_report.json / v0.2.0 metrics).
1. Two competitions (do not mix them)
| Arena | Competitor definition of “win” | PFLT status |
|---|---|---|
| A. Inventory / form→gloss | Coverage + sense under offline densify | Winning our catalog (~99.99% on 113 langs) |
| B. Sentence MT | sacreBLEU / chrF / COMET on held-out sentences | Chat: competitive; News: trailing DeepL-class |
Google/DeepL/NLLB are judged almost entirely on B.
We already win parts of A (especially classical / offline / FSOT). Beating them “accurately” means closing B without lying about A.
2. Scoreboard (measured)
A — Catalog
| Metric | Us | Google-ish | DeepL-ish | NLLB-ish |
|---|---|---|---|---|
| Lang count | 113 | ~249 | ~30–100 | ~200 |
| Form→gloss on our catalog | ~99.99% | N/A (not their product metric) | N/A | N/A |
| Classical / dead / hieroglyph | Strong | Weak | Weak | Weak |
| FSOT law pin | Yes | No | No | No |
Verdict A: Competitive / leading on our product definition. Not “more languages than Google.”
B1 — Chat / easy parallel (Tatoeba-style open-set)
| System path | sacreBLEU (mean) | Notes |
|---|---|---|
| Neural best-of (opus/mul/NLLB) | 50.19 | Fair open-set |
| Hybrid oracle densify|neural | 53.58 | Product upper bound |
| Staged internal bar | 45 | Passed |
| Strong open MT chat (rough) | ~45–65+ | We sit mid/high |
Per-lang neural highs: it 68 · es/pt 61 · de/ru ~57–58 · hi 55
Gaps: ja 37 · zh 32 · la neural 13 (densify wins classical)
Verdict B1: Competitive on chat open-set. Not SOTA vs every commercial pair, but past mid bar.
B2 — News (WMT14 de→en test, n=3003) — the hard public bar
| System | sacreBLEU | Gap to 40 mid | Gap to 48 stretch |
|---|---|---|---|
| Densify-only | ~0.4 | — | — |
| opus-mt-de-en beams=5 | 33.88 | −6.1 | −14.1 |
| NLLB-600M beams=5 | 33.37 | −6.6 | −14.6 |
| DeepL-class mid (staged) | ~40 | 0 | — |
| Strong commercial stretch | ~45–55 | — | 0 |
Verdict B2: Near mid open-MT, not DeepL-class yet. This is the main gap if “beat the competition” means news fluency.
3. Stage ladder (honest)
S0 Dict ████████████ DONE
S1 Phrase ████████████ DONE
S2 Chat strong ████████████ DONE (neural ~50)
S3 Mid news ████████░░░░ HERE (~34 / need ~40)
S4 DeepL-class ████░░░░░░░░ (need ~45–55)
4. What is left to be competitive / beat them (priority order)
| # | Lever | Closes | Effort | Expected gain |
|---|---|---|---|---|
| 1 | Ship hybrid router (densify short/classical; neural long/news/CJK) | Product accuracy vs single-path | Low | Realizes measured ~53.6 chat hybrid; better UX |
| 2 | WMT student finetune (opus or NLLB on WMT train; law fixed) | News −6 gap | Med–High | Often +2–6+ sacre if done right |
| 3 | Stronger student (NLLB-1.3B / more beams / length / ensembling) | News + chat CJK | Med | +0.5–3 sacre (varies) |
| 4 | CJK order always neural SPM path (no EN-dep densify for ja/zh) | ja/zh chat | Low | Lift ja/zh from ~32–37 toward EU band |
| 5 | FLORES same-file public bar | Comparable claim vs NLLB papers | Blocked until Hub access | Credibility, not training |
| 6 | Breadth more Kaikki langs toward 150–200 | Catalog vs Google count | Med | Count game only |
| 7 | COMET / human spot checks | Beyond BLEU honesty | Med | Avoid BLEU-only overclaim |
What will not beat DeepL alone
- More form→gloss densify on chat templates (product ceiling ≠ news SOTA)
- Refitting the FSOT law scalar to BLEU (forbidden / false)
- Claiming product densify BLEU-4 ~83 as open-set MT
5. Definition of “we beat the competition” (accurate)
Pick a claim level:
| Claim level | Criteria | Distance now |
|---|---|---|
| L1 Product unique | Offline FSOT + classical + catalog depth | Met |
| L2 Chat competitive | Mean chat sacre ≥45 open-set multi-lang | Met (~50) |
| L3 News mid-parity | WMT14 de-en sacre ≥40 | ~6 pts short |
| L4 News strong | WMT14 de-en sacre ≥45–48 | ~11–14 pts short |
| L5 Broad DeepL-class | Many pairs FLORES/WMT mid-40s+ | Not started (FLORES gated; multi-pair news unrun) |
Accurate one-liner:
We beat / own unique offline+classical+law. We match mid open-MT on chat. We do not yet beat commercial systems on news full-sentence fluency — need roughly +6 sacre for mid-parity and +14 for stretch SOTA on de→en.
6. Immediate execution plan (this push)
- Implement product hybrid router + measure non-oracle (feature rules, not ref peeking)
- WMT decode push: higher beams, dual-system sentence-level ensemble upper bound
- Write measured gaps into
reports/COMPETITIVE_PUSH.md - Keep law pin D1D38A fixed
Optional next session: WMT finetune loop under densify law wrap.
7. Measured this push (2026-07-22 competitive run)
| Track | Result | Notes |
|---|---|---|
| Product hybrid chat sacreBLEU | 40.64 | No ref peeking; densify-heavy (2800) + neural CJK (400) |
| WMT opus-mt beams=8 | 33.79 | Slightly under prior beams=5 33.88 |
| WMT NLLB-600M beams=8 | 33.49 | |
| WMT oracle dual ensemble | 37.61 | Best of opus/nllb per sentence (upper bound) |
| Gap ensemble → mid 40 | 2.39 | Was ~6.1 for single student |
| Gap ensemble → stretch 48 | 10.39 | Still needs finetune / larger model |
Insight: Beams alone do not close news. Dual-student ensemble recovers ~+3.7 sacre toward bar 40 without finetune. Remaining ~2.4 points need quality model upgrade or WMT finetune.
8. Beat levers run (2026-07-22)
| Lever | Result | Clears bar? |
|---|---|---|
| L2 Neural-first hybrid chat | 48.74 sacre | Yes (≥45) |
| L1 Product NLL ensemble WMT | 34.54 | No (gap 5.46 to 40) |
| L1 Oracle ensemble upper | 37.13 | No (gap 2.87) |
| L3 Aggressive FT | 32.2 | No (regressed) |
| L3b Safe FT freeze-enc | 33.86 | No (flat) |
| Base opus-mt-de-en | 33.88 | No |
Still required to beat news mid-parity (40): better ensemble selection (~2.6 pts headroom to oracle), multi-epoch WMT FT with validation early-stop, or larger NLLB.
9. News push v3 (2026-07-22)
| System | sacreBLEU |
|---|---|
| Product cross-NLL ensemble | 34.11 |
| Base opus | 33.88 |
| FT v3 test / best val | 33.41 / 35.76 |
| Oracle upper | 37.13 |
| Gap product → 40 | 5.89 |
Chat neural-first hybrid remains 48.74 (mid bar met).