reverse-etymology / docs /DESIGN.md
Mongoosetross's picture
Deploy Reverse Etymology Atlas
9729dcb verified
|
Raw
History Blame Contribute Delete
17.6 kB

Reverse Etymology Atlas

Purpose

Ordinary etymology asks "where did this word come from?" Reverse etymology asks the opposite: given a base word, what was derived from it?

This tool turns the lukeslp/etymology-atlas graph (4.17 million relationships across 19,401 languages) into an interactive, mobile-capable explorer. A user picks an etymon such as Latin mater or Proto-Indo-European meh2ter and immediately sees the descendant cloud: inherited daughters, later borrowings, compounds, clippings, and (optionally) cognates.

The product is a local web application: a FastAPI service over an in-memory compressed graph, plus a canvas-rendered frontend. It is designed to feel like a living atlas, not a spreadsheet with circles.

Data sources

The Hugging Face dataset is the system of record. Five parquet tables are ingested, then enriched with two CLDF catalogs so that empty fields in the published atlas (notably lang*_family and most ISO codes) do not cripple filtering.

Source Role
etymologies.parquet (4,173,462 rows) Directed word graph
languages.parquet (19,401 rows) Glottolog skeleton: coordinates, macroarea, vitality, phoneme counts
cognate_sets.parquet (4,981 IE-CoR sets) Expert cognate groups, shown on node detail
phonemes.parquet (105,484 rows) PHOIBLE inventories by glottocode
linguistic_features.parquet (76,475 rows) WALS typological values
Glottolog CLDF languages.csv Family names, ISO 639-3, coordinates, language vs family vs dialect
WALS CLDF languages.csv + codes.csv Join WALS codes to glottocodes and human-readable feature values

License of the atlas and Wiktionary-derived etymology-db content is CC BY-SA 3.0. The UI cites Steuber 2026 and the upstream sources.

Graph semantics

Edge direction

Each etymology row is:

(term1, lang1)  relationship_type, confidence, source  (term2, lang2)

Empirical inspection of the atlas matches Wiktionary / etymology-db convention:

  • term1 is the derived or attested form.
  • term2 is the etymon, stem, or related form.

Examples:

  • English portmanteau borrowed Middle French portemanteau
  • English mother inherited Middle English moder, Old English modor, Proto-Germanic moder, PIE meh2ter

Reverse etymology therefore walks edges from term2 toward term1: "everything that lists this node as a source."

Forward etymology (etymons of the current node) is still stored and shown in the inspector, but it is not the primary layout.

Node identity

A node is the pair (term, lang) after trimming whitespace and dropping empty terms (a few dozen nulls exist in the parquet file). lang is the atlas language key: a Wiktionary language name (latin, proto-indo-european) for etymology-db rows, or an ISO 639-3 code (ben, hin) for Lexibank IE-CoR rows.

Display names prefer, in order: atlas lang*_name when non-empty, Glottolog Name, then title-cased lang.

Relation taxonomy

The atlas already normalizes Wiktionary's ~31 templates into 11 types:

Type Count Default in reverse walk Notes
inherited 327,261 on Unbroken descent. Highest-signal descendants.
derived 325,776 on Derivation that is not specifically inherited/borrowed.
borrowed 231,814 on Loanwords.
calque 10,318 on Loan translations.
clipping 2,482 on Spoken shortening.
back_formation 2,173 on Affix removal.
abbreviation 404 on Written shortening.
compound 304,470 on term1 contains term2 as a part. Useful for affix/root exploration; noisier for strict lexical descent — uncheck if a root gets bushy.
blend 15,796 on Portmanteau parts.
other 2,162,894 on Catch-all (prefixes, mentions, mixed Wiktionary templates).
cognate 790,074 on Lateral, not descent. Includes Lexibank pairwise cognates; can expand into cliques — uncheck for stricter trees.

Global edge filters (allowed types + minimum confidence) apply before traversal. They define which graph is being searched, not which nodes survive afterwards.

Confidence in the atlas is discrete: 0.70 (other), 0.75, 0.85, 0.90, 0.92 (inherited), 0.95 (derived), 1.00 (Lexibank). Default floor is 0.70 (everything). A slider exposes the steps.

Cycles, multi-parents, and depth

The graph is not a tree. A word can list several etymons; walking reverse edges can revisit nodes through different parents. Traversal is a bounded BFS:

  1. Start at the selected root.
  2. Follow reverse edges whose type/confidence pass the global edge filters.
  3. Ignore edges back to an already-seen node (cycle cut).
  4. Record the BFS parent so the visible structure is a tree even when the underlying graph is a DAG with extra links.
  5. Stop at max_depth (default 3, max 8) or max_visit (50,000 nodes).

Extra parents that were not chosen as the BFS parent are attached as also_from on the node payload so the inspector can show them without breaking layout.

Filtering model (the core requirement)

Filters are not a single global predicate applied to every visible node. That would be wrong for reverse etymology. A user who asks "English words that came from Latin caput by way of French" wants:

  • Leaves (terminal descendants, or nodes at max depth treated as terminals for this query) to be English.
  • Some intermediate hop on the unique BFS path to be French.
  • The Latin root itself to remain visible even though it is not English.

Three independent layers are applied in order.

Layer A — graph construction (edges)

Applies to every hop, including those that will later be pruned.

  • Allowed relationship_type set.
  • min_confidence.
  • Optional max_depth.

Layer B — leaf predicates

When no leaf constraints are set, candidate leaves are structural sinks under the construction graph (or nodes at max_depth). When leaf constraints are set, every visited non-root node that matches those constraints is a candidate — including productive words that still have further descendants (Spanish verde under Latin viridis must survive a Spanish leaf filter even though Tagalog berde was borrowed from it). A candidate is kept only if it matches all specified leaf constraints:

  • Term contains / regex (case-insensitive, Unicode-aware).
  • Language key or display-name set (multi-select).
  • Language family (Glottolog Family_ID resolved to a name such as Indo-European).
  • Macroarea (Africa, Eurasia, Papunesia, Australia, North America, South America).
  • Vitality / Glottolog endangerment when known.
  • Geographic: has coordinates; optional lat/lon bounding box.
  • WALS: one or more (feature_id, allowed_values) clauses, AND-combined. Example: 81A in {SOV, SVO}.
  • Phonology: inventory contains these IPA segments; inventory lacks these segments; has tone; segment class present (click, implosive approximated via IPA).
  • Source dataset: etymology-db, lexibank, or both (inherited from incident edges).
  • Minimum descendant fan-out of the leaf itself (rarely used; more often a search ranking feature).

Unknown metadata: by default a positive constraint fails closed (a language with no WALS 81A value does not count as SOV). A per-layer toggle keep_unknown lets sparse typological data through.

Layer C — path / intermediate predicates

Let path(leaf) = [root, hop1, hop2, ..., leaf].

Intermediates are path[1:-1] (root and leaf excluded). If the path has length 1 (the root is also a leaf), intermediate filters do not apply unless the user opts into apply_path_to_root.

Each intermediate constraint set is evaluated with an explicit quantifier:

Quantifier Meaning
any (default) At least one intermediate matches. "Passed through French."
all Every intermediate matches. "Never left Romance."
none No intermediate matches. "Did not pass through English."
exactly Exactly k intermediates match.

Path constraints may include the same trait vocabulary as leaves (language, family, macroarea, WALS, phonemes, term pattern) plus two edge-aware clauses evaluated on the path's edges rather than nodes:

  • path_relations_any: at least one hop has this relation (e.g. a borrowing occurred somewhere).
  • path_relations_none: no hop has this relation (e.g. no compounds).

Pruning

After BFS:

  1. Evaluate every candidate leaf against Layer B.
  2. Evaluate path(leaf) against Layer C.
  3. Keep a leaf only if both pass.
  4. Rebuild the visible tree as the union of surviving root-to-leaf paths.
  5. Drop nodes that are not on any surviving path (except the root, always kept).
  6. Cluster remaining high-degree sibling lists for display (see Layout).

This is the only definition of "filter correctly ... for the leaf words and intermediate nodes along the path." Applying leaf language = English to every node would hide the French and Latin ancestors the user needs in order to understand the path.

Worked example

Root: Latin caput.

  • Layer A: relations {inherited, borrowed, derived}, confidence >= 0.9, depth 4.
  • Layer B: leaf language = English.
  • Layer C: any intermediate language = French.

Kept: English chief, chef, capital when their BFS paths include a French (or Middle French) hop. Dropped: English head if it arrived only through Germanic, even though it is English; dropped French chef because it is not an English leaf.

Search and ranking

Autocomplete queries SQLite FTS5 on term with a Unicode61 tokenizer, plus a normalized shadow column (nfc, lowercased, leading * stripped, h2/h₂ folded) so PIE reconstructions remain findable.

Ranking for prefix q:

  1. Exact term match.
  2. Prefix match.
  3. Substring / FTS rank.
  4. Tie-break: child_count descending (productive etymons float up), then language name.

Each hit returns {term, lang, display_lang, family, child_count, sample_children}. Homographs are separate rows (Latin mater vs English mater).

A curated example list is shipped for first-run exploration: PIE meh2ter, Latin mater, Latin caput, Latin habere, English computer, Arabic kitab, Sanskrit cakra, etc.

Visualization

Views

  1. Descent tree (default). Left-to-right tidy tree: etymon on the left, descendants growing right. This matches the reverse-etymology metaphor.
  2. Radial. Root at center, depth as radius. Better on tablets for bushy trees.
  3. Map. Surviving leaves (and optionally intermediates) plotted by language coordinates. Marker size = number of surviving words in that language; color = family. Clicking a marker lists those words and can isolate the language in the tree.
  4. Table. Virtualized, sortable: term, language, family, relation from parent, confidence, depth, full path as a breadcrumb. CSV export.
  5. Stats. Relation mix, family mix, depth histogram, borrowing vs inheritance ratio, macroarea breakdown for the current filtered tree.

Canvas renderer

SVG is too slow above a few hundred labeled nodes. The tree and radial views use a single high-DPI canvas:

  • Layout via d3-hierarchy (tree / cluster) computed on a web worker when the node count exceeds 800.
  • Camera: pan, wheel zoom, pinch zoom, double-tap zoom, keyboard arrows.
  • Hit testing: uniform grid hashed by screen cell, rebuilt on camera change.
  • Level of detail: labels fade in by zoom and by node importance (child_count, depth, selection).
  • Minimap in the corner on desktop; hidden on narrow screens unless toggled.
  • Selection: click a node to open the inspector; hover previews the BFS path back to the root in a brighter stroke.
  • Re-root: "descend from here" makes the node the new query without leaving the UI.

Clustering (speed and legibility)

Unfiltered reverse walks of productive affixes are huge (ly English has 13,422 direct children). After filtering, if a node still has more than cluster_threshold (default 28) children:

  1. Group by language.
  2. If a single language still exceeds the threshold, group by initial letter.
  3. Emit a cluster stub {kind: "cluster", label, count, relation_mix} that expands on click.

The server never ships more than max_payload_nodes (2,500) in one response. Expansion requests pass expand keys so only the requested cluster is unpacked, still subject to the same filters.

Color and encoding

User-selectable color mode:

  • Relation from parent (default): inherited teal, borrowed copper, derived moss, calque violet, clipping/abbreviation slate, compound/blend ochre, other ink-wash gray, cognate muted magenta.
  • Language family (hashed qualitative palette, stable per family name).
  • Confidence (sequential).
  • Depth (sequential).
  • Macroarea.

Stroke width encodes child_count on the visible subtree. A legend is always present.

Inspector

Slide-over on desktop, bottom sheet on mobile:

  • Term, language, family, coordinates, vitality.
  • Incident reverse children (counts by relation).
  • Forward etymons ("this word comes from...").
  • IE-CoR cognate sets containing the form, when any.
  • WALS highlights (word order, case, tone).
  • Phoneme inventory summary (counts by class, tone yes/no, searchable list).
  • Wiktionary deep link for etymology-db nodes.
  • Actions: re-root, isolate this language, copy path, exclude this language from leaves / from path.

Mobile

"Viewable on mobile" is a first-class layout, not a shrunk desktop.

  • Responsive shell: below 840px the filter panel and inspector become bottom sheets; the canvas is full-bleed under a compact chrome (search + view switcher + filter chip count).
  • Touch: one-finger pan, pinch zoom, tap select, two-finger rotate disabled (it fights tree reading).
  • Filter chips summarize active leaf vs path constraints so a phone user can see the query without opening the sheet.
  • PWA manifest + service worker cache the app shell (not the 4M-edge index). Add-to-home-screen works when served over HTTPS or localhost.
  • The API binds 0.0.0.0 so a phone on the same network can open http://<lan-ip>:8000. A "Open on phone" dialog shows that origin as a QR code.
  • Shareable URLs encode the full query (root, layers A/B/C, view, color mode). Opening the URL on a phone restores the same tree.

Architecture

parquet + CLDF CSVs
        |
        v
 pipeline/build_index.py
        |
        +--> data/index/atlas.sqlite   (nodes, FTS, languages, WALS, phonemes, cognates)
        +--> data/index/graph.npz      (CSR: offsets, targets, rel_code, conf_code, src_code)
        +--> data/index/meta.json      (vocabularies, color maps, examples)
        |
        v
 FastAPI  --loads CSR into RAM once--  React canvas SPA

Why this split:

  • SQLite FTS is the right tool for autocomplete and inspector joins.
  • Pure SQL recursive CTEs cannot apply Layer C quantifiers quickly enough on 13k-way affixes. An in-memory CSR BFS in Python over int32 arrays does this in tens of milliseconds.
  • The SPA never downloads the full graph. It only receives the pruned, clustered payload for the current query.

Estimated RAM after load: well under 500 MB (CSR ~30 MB, interned trait arrays ~20 MB, SQLite mmap, Python overhead).

API

All JSON. CORS open. No auth.

Method Path Purpose
GET /api/health Index stats, load time
GET /api/suggest?q=&limit= Autocomplete
GET /api/meta Vocabularies for filter widgets
POST /api/tree Filtered descendant tree
GET /api/node?term=&lang= Inspector payload
GET /api/cognates?term=&lang= IE-CoR sets
POST /api/export Same tree as CSV or JSON download
GET /api/examples Curated roots

POST /api/tree body is the serialized filter document (see Implementation plan). Response:

{
  "root": {"id": 0, "term": "mater", "lang": "latin", "...": ""},
  "nodes": [{"id": 1, "parent": 0, "kind": "word"|"cluster", "...": ""}],
  "stats": {"visited": 1800, "kept_leaves": 42, "elapsed_ms": 18},
  "truncated": false
}

Performance budget

Action Target
Process start including mmap load < 3 s on a laptop
Autocomplete < 30 ms
Typical tree (depth 3, < 2k visit) < 80 ms
Worst productive affix, filtered < 400 ms
Canvas pan/zoom 60 fps at 2,500 nodes
Mobile first interaction usable within one round trip

Techniques: CSR int32, BFS with typed arrays, cluster stubs, canvas LOD, SQLite covering indexes, gzip JSON, immutable filter objects so React does not relayout on every keystroke (search input is debounced 120 ms).

Accessibility and i18n

  • Canvas is supplemented by the table view (keyboard reachable, screen-reader friendly).
  • Focus states on all controls; / focuses search; f toggles filters; 1-4 switch views.
  • Color is never the only encoding (relation also appears as a short label and in the inspector).
  • Terms are rendered as stored (IPA, Greek, reconstructions). UI chrome is English.

Non-goals (this version)

  • Editing or crowdsourcing etymologies.
  • Full phylogenetic tree inference.
  • Offline copies of Wiktionary prose.
  • Account/sync in the cloud.

Visual language

The UI should read as a night atlas: deep ink navy, warm parchment text, copper and sea-teal relation colors, serif display for headwords, sans for chrome. No generic dashboard cards, no rainbow force-directed hairball as the default.