Spaces:
Running on Zero
A newer version of the Gradio SDK is available: 6.22.0
title: SphinxEyes — Hieroglyph Decoder
emoji: 🔆
colorFrom: yellow
colorTo: indigo
sdk: gradio
sdk_version: 5.49.1
app_file: app.py
pinned: false
license: agpl-3.0
short_description: Detect, order and read Middle Egyptian hieroglyphs
SphinxEyes
End-to-end reading of Middle Egyptian hieroglyphs from a photograph:
photo -> YOLO11l detection (150 Gardiner classes)
-> quadrat clustering + reading order
-> Viterbi correction (dictionary trie + bigram LM + confusion priors)
-> cartouche matching (Needleman-Wunsch vs 41 Beckerath-validated kings)
-> [optional] GPT-4o transliteration + English gloss
Everything up to the last step runs locally in the Space and is free and unlimited. Only the GPT stage is metered.
Setup
Hardware: set the Space to ZeroGPU. The detector runs a raw PyTorch
forward inside @spaces.GPU; on CPU it still works but takes ~3–4 s.
Secret: add OPENAI_API_KEY under Settings → Secrets to enable
transliteration. Without it the app runs normally and the toggle reports that
the stage is unavailable — it never errors out.
Optional environment overrides:
| Variable | Default | Meaning |
|---|---|---|
OPENAI_MODEL |
gpt-4o |
Transliteration model |
TRANSLIT_MAX_SIGNS |
20 |
Signs per LLM chunk |
TRANSLIT_TEMPERATURE |
0.1 |
Sampling temperature |
SPHINX_PER_IP_DAILY |
3 |
Transliterations per visitor per UTC day |
SPHINX_GLOBAL_DAILY |
100 |
Space-wide daily ceiling (key-spend guard) |
SPHINX_GPU_SECONDS |
60 |
ZeroGPU duration budget per forward |
SPHINX_WEIGHTS |
src/artifacts/best_modelv10.pt |
Detector checkpoint |
SPHINX_ONNX |
(unset) | Local dev only — use an ONNX file on CPU instead of torch |
Using it
Layout is required and cannot be auto-detected reliably — the geometric
voter misvotes on real walls, and a wrong choice scrambles the reading order.
Pick columns for vertical text, rows for horizontal.
Reading direction is geometrically undecidable. Right-to-left is the
default; signs normally face the start of the line, so if the figures look
left, choose ltr.
The archaeological-context panel is optional but materially improves the GPT
reading — a temple wall favours royal formulae, pyramid texts favour Old
Kingdom spellings, and a known reign anchors the titulary. Fields left as
unknown are explicitly flagged to the model as unsupplied so it does not
invent them.
Rate limiting
The 3/day cap is best-effort cost control, not security: state is
in-memory (a Space restart or sleep resets it) and identity is the hashed
client IP, which is shared behind NAT and rotatable. SPHINX_GLOBAL_DAILY is
the backstop that actually bounds API spend. A failed or unavailable LLM call
does not consume quota.
Model
best_modelv10.pt — YOLO11l, 150 classes (148 Gardiner + cartouche +
unknown), imgsz 1024, trained with Ultralytics 8.4.102.
Class ordering is Aa1=0, a1=1 … Aa15=149 and is identical to v9, including
the three transmuted indices (30 m4, 61 f34, 70 o29).
Known gap: the substitution priors ship as
confusion_matrix_v9_normalized.csv — no v10 confusion matrix exists yet.
It feeds only the Viterbi substitution cost, and since v10 shares v9's exact
class ordering the priors stay meaningful, but they are one generation stale.
There are also no published v10 validation metrics.
Why a raw forward, not model.predict()
The spatial layer needs the top-3 class candidates per detection to feed
the Viterbi corrector. predict() returns top-1 after its own NMS, which
would gut the correction layer. So src/infer_torch.py runs
DetectionModel.forward to get the raw [1, 154, 21504] tensor — the same
contract as an ONNX export with nms=False — and hands it to the unchanged
spatial_logic.postprocess_onnx.
Layout
app.py Gradio UI + two-stage orchestration
src/ flat, NO __init__.py (see below)
pipeline.py SphinxPipeline; detection backend is injected
infer_torch.py ZeroGPU backend (v10 .pt)
llm.py GPT-4o transliteration (env-var config)
ratelimit.py per-IP + global daily cap
spatial_logic.py detection postprocess, NMS, quadrats, reading order
sphinx_corrector.py Viterbi + bigram LM + linguistic priors
sphinx_trie.py lexicon trie (must stay importable as `sphinx_trie`)
cartouche_matcher.py Needleman-Wunsch vs royal_names.json
layout_detector.py fallback layout voter
enhance_img.py CLAHE/denoise presets (off by default)
royal_names.json determinatives.json initial_particles.json
artifacts/ weights, class map, trie, corpus, confusion matrix
src/ has no __init__.py on purpose: sphinx_trie_v4.pkl has the module
path sphinx_trie baked into the pickle, so importing it as src.sphinx_trie
makes pickle.load raise ModuleNotFoundError. app.py puts src/ on
sys.path and imports flat.
Image enhancement is off by default and should stay that way: A/B testing showed it deletes real signs and hallucinates others on clean photographs.
Local development
onnxruntime is not in requirements.txt, but the ONNX backend lets you
verify the pipeline without installing torch:
pip install onnxruntime
python src/pipeline.py /path/to/Unas1c.jpg columns # SPHINX_ONNX or repo default
SPHINX_ONNX=/path/to/best_model_v9.onnx python app.py
Both backends letterbox identically and end in postprocess_onnx, so a result
that holds under ONNX holds under torch.
License
The Space is AGPL-3.0: it loads a YOLO checkpoint via Ultralytics, which is AGPL-3.0 licensed, and that obligation extends to this deployed service.