You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

TaxEmbed — a hyperbolic embedding of the complete NCBI tree of cellular life

TaxEmbed places all 1,102,163 named cellular taxa of the NCBI Taxonomy (new_taxdump, release 2026-06-09) into a single 100-dimensional Poincaré ball. It is a standalone, reusable representation of organismal relatedness: a differentiable, fixed-dimensional metric of how organisms relate, learned from one identical training recipe that holds from a few thousand taxa up to the full 1.1 million.

In the geometry:

  • radius (embedding norm ‖x‖) tracks taxonomic depth — the root (cellular organisms) sits near the centre, tips near the rim (Pearson r ≈ 0.95 vs. root distance);
  • direction encodes learned lineage — taxa in the same clade point the same way (family-level kNN purity 0.907, angular separation up to 7.6× at family rank).

Files

File Size Contents
cellular_embedding.safetensors 441 MB The embedding matrix. Tensor key embedding, shape (1102163, 100), float32.
taxid_to_index.tsv 16 MB Two columns taxididx. Row i of the matrix is the taxon whose idx == i.
LICENSE Apache-2.0.

All release metadata (dimension, taxon count, taxdump date, geometry) is embedded in the safetensors header (__metadata__).

Quick start

import numpy as np, pandas as pd
from safetensors.numpy import load_file

emb = load_file("cellular_embedding.safetensors")["embedding"]   # (1102163, 100) float32
idx = pd.read_csv("taxid_to_index.tsv", sep="\t").set_index("taxid")["idx"]

def poincare_distance(u, v):
    """Geodesic distance in the Poincaré ball."""
    sq = np.sum((u - v) ** 2)
    return np.arccosh(1 + 2 * sq / ((1 - u @ u) * (1 - v @ v)))

# Homo sapiens (9606) vs Mus musculus (10090)
d = poincare_distance(emb[idx[9606]], emb[idx[10090]])
print(f"Poincaré distance human–mouse: {d:.3f}")

# radius encodes depth: deeper taxa sit closer to the rim
print("‖x‖ human:", np.linalg.norm(emb[idx[9606]]))   # ~ deep taxon, near the rim

To look up a taxon by name, join taxid_to_index.tsv against the NCBI names.dmp from the same new_taxdump release (2026-06-09).

Model details

  • Geometry: Poincaré ball, curvature −1, d = 100.
  • Coverage: 1,102,163 named cellular taxa — Eukaryota 878,524; Bacteria 216,507; Archaea 7,131 (+ the root node). This mirrors NCBI's own sampling (~80% Eukaryota, ~0.6% Archaea); it is not a balanced census of the domains.
  • Training signal: the depth-weighted transitive closure of the taxonomy (every ancestor–descendant pair), topology only — no branch lengths and no molecular data.
  • Recipe (five load-bearing components): Euclidean tangent parametrization, a softmax/NLL loss over Poincaré distances, a depth curriculum, a radial nudge toward depth-based target radii, and a scale-aware effective batch. This single recipe is what carries the embedding to 1.1M taxa, the scale at which the prior (Nickel–Kiela) approach collapses.

Intended use

A general substrate for organismal relatedness that machine-learning models can consume directly: a differentiable phylogenetic prior, a relatedness metric for retrieval and taxonomy quality-control, or a factor to fuse with protein/other embeddings in one vector space. As one demonstrated application, the paper shows that taxonomy and protein function occupy linearly separable subspaces of a protein language model, so protein embeddings can be linearly debiased against phylogeny using this geometry.

Limitations

  • Transductive. The embedding is fixed to the 2026-06-09 taxdump. A taxon added later has no coordinate; out-of-sample projection is future work.
  • Topology only. It captures internal taxonomy consistency, not branch-length or molecular truth.
  • Sampling skew. ~80% Eukaryota / ~0.6% Archaea, inherited from NCBI; this bounds any per-domain claim (the Archaeal arm in particular is thin).
  • "Complete" means the named cellular taxa retained by the reproducible clean pass, not every NCBI node. Viruses are excluded by design.

License

Apache-2.0. Copyright 2025–2026 the TaxPointCare authors (@jcoludar and contributors). The implementation began as a fork of the Nickel & Kiela (2017) Poincaré-embeddings reference code and has since been fully reimplemented.

Citation

If you use this embedding, please cite:

Koludarov, I. & Rost, B. A hyperbolic embedding of the complete NCBI tree of cellular life (1.1 million named taxa) from a single recipe (2026).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support