pflt-fsot / docs /COMPETITIVE_ROADMAP.md
dappalumbo91's picture
v0.2.2: news ensemble 34.11 + chat hybrid 48.7 + full source
7025d22 verified
|
Raw
History Blame Contribute Delete
7.4 kB

Competitive roadmap — accurate status vs competition

Law: (S=K(T_1+T_2+T_3)) pin D1D38A (unchanged; never fitted to BLEU).
Measured: 2026-07-21 SOTA push (m6_sota_push_report.json / v0.2.0 metrics).


1. Two competitions (do not mix them)

Arena Competitor definition of “win” PFLT status
A. Inventory / form→gloss Coverage + sense under offline densify Winning our catalog (~99.99% on 113 langs)
B. Sentence MT sacreBLEU / chrF / COMET on held-out sentences Chat: competitive; News: trailing DeepL-class

Google/DeepL/NLLB are judged almost entirely on B.
We already win parts of A (especially classical / offline / FSOT). Beating them “accurately” means closing B without lying about A.


2. Scoreboard (measured)

A — Catalog

Metric Us Google-ish DeepL-ish NLLB-ish
Lang count 113 ~249 ~30–100 ~200
Form→gloss on our catalog ~99.99% N/A (not their product metric) N/A N/A
Classical / dead / hieroglyph Strong Weak Weak Weak
FSOT law pin Yes No No No

Verdict A: Competitive / leading on our product definition. Not “more languages than Google.”

B1 — Chat / easy parallel (Tatoeba-style open-set)

System path sacreBLEU (mean) Notes
Neural best-of (opus/mul/NLLB) 50.19 Fair open-set
Hybrid oracle densify|neural 53.58 Product upper bound
Staged internal bar 45 Passed
Strong open MT chat (rough) ~45–65+ We sit mid/high

Per-lang neural highs: it 68 · es/pt 61 · de/ru ~57–58 · hi 55
Gaps: ja 37 · zh 32 · la neural 13 (densify wins classical)

Verdict B1: Competitive on chat open-set. Not SOTA vs every commercial pair, but past mid bar.

B2 — News (WMT14 de→en test, n=3003) — the hard public bar

System sacreBLEU Gap to 40 mid Gap to 48 stretch
Densify-only ~0.4
opus-mt-de-en beams=5 33.88 −6.1 −14.1
NLLB-600M beams=5 33.37 −6.6 −14.6
DeepL-class mid (staged) ~40 0
Strong commercial stretch ~45–55 0

Verdict B2: Near mid open-MT, not DeepL-class yet. This is the main gap if “beat the competition” means news fluency.


3. Stage ladder (honest)

S0 Dict          ████████████ DONE
S1 Phrase        ████████████ DONE
S2 Chat strong   ████████████ DONE  (neural ~50)
S3 Mid news      ████████░░░░ HERE  (~34 / need ~40)
S4 DeepL-class   ████░░░░░░░░       (need ~45–55)

4. What is left to be competitive / beat them (priority order)

# Lever Closes Effort Expected gain
1 Ship hybrid router (densify short/classical; neural long/news/CJK) Product accuracy vs single-path Low Realizes measured ~53.6 chat hybrid; better UX
2 WMT student finetune (opus or NLLB on WMT train; law fixed) News −6 gap Med–High Often +2–6+ sacre if done right
3 Stronger student (NLLB-1.3B / more beams / length / ensembling) News + chat CJK Med +0.5–3 sacre (varies)
4 CJK order always neural SPM path (no EN-dep densify for ja/zh) ja/zh chat Low Lift ja/zh from ~32–37 toward EU band
5 FLORES same-file public bar Comparable claim vs NLLB papers Blocked until Hub access Credibility, not training
6 Breadth more Kaikki langs toward 150–200 Catalog vs Google count Med Count game only
7 COMET / human spot checks Beyond BLEU honesty Med Avoid BLEU-only overclaim

What will not beat DeepL alone

  • More form→gloss densify on chat templates (product ceiling ≠ news SOTA)
  • Refitting the FSOT law scalar to BLEU (forbidden / false)
  • Claiming product densify BLEU-4 ~83 as open-set MT

5. Definition of “we beat the competition” (accurate)

Pick a claim level:

Claim level Criteria Distance now
L1 Product unique Offline FSOT + classical + catalog depth Met
L2 Chat competitive Mean chat sacre ≥45 open-set multi-lang Met (~50)
L3 News mid-parity WMT14 de-en sacre ≥40 ~6 pts short
L4 News strong WMT14 de-en sacre ≥45–48 ~11–14 pts short
L5 Broad DeepL-class Many pairs FLORES/WMT mid-40s+ Not started (FLORES gated; multi-pair news unrun)

Accurate one-liner:
We beat / own unique offline+classical+law. We match mid open-MT on chat. We do not yet beat commercial systems on news full-sentence fluency — need roughly +6 sacre for mid-parity and +14 for stretch SOTA on de→en.


6. Immediate execution plan (this push)

  1. Implement product hybrid router + measure non-oracle (feature rules, not ref peeking)
  2. WMT decode push: higher beams, dual-system sentence-level ensemble upper bound
  3. Write measured gaps into reports/COMPETITIVE_PUSH.md
  4. Keep law pin D1D38A fixed

Optional next session: WMT finetune loop under densify law wrap.


7. Measured this push (2026-07-22 competitive run)

Track Result Notes
Product hybrid chat sacreBLEU 40.64 No ref peeking; densify-heavy (2800) + neural CJK (400)
WMT opus-mt beams=8 33.79 Slightly under prior beams=5 33.88
WMT NLLB-600M beams=8 33.49
WMT oracle dual ensemble 37.61 Best of opus/nllb per sentence (upper bound)
Gap ensemble → mid 40 2.39 Was ~6.1 for single student
Gap ensemble → stretch 48 10.39 Still needs finetune / larger model

Insight: Beams alone do not close news. Dual-student ensemble recovers ~+3.7 sacre toward bar 40 without finetune. Remaining ~2.4 points need quality model upgrade or WMT finetune.


8. Beat levers run (2026-07-22)

Lever Result Clears bar?
L2 Neural-first hybrid chat 48.74 sacre Yes (≥45)
L1 Product NLL ensemble WMT 34.54 No (gap 5.46 to 40)
L1 Oracle ensemble upper 37.13 No (gap 2.87)
L3 Aggressive FT 32.2 No (regressed)
L3b Safe FT freeze-enc 33.86 No (flat)
Base opus-mt-de-en 33.88 No

Still required to beat news mid-parity (40): better ensemble selection (~2.6 pts headroom to oracle), multi-epoch WMT FT with validation early-stop, or larger NLLB.


9. News push v3 (2026-07-22)

System sacreBLEU
Product cross-NLL ensemble 34.11
Base opus 33.88
FT v3 test / best val 33.41 / 35.76
Oracle upper 37.13
Gap product → 40 5.89

Chat neural-first hybrid remains 48.74 (mid bar met).