protea-sparse-knn
Protein function prediction by nearest-neighbour Gene Ontology transfer, performed in the space of a learned sparse encoder rather than in the raw embedding space.
This repository holds the frozen artefacts a container needs to run the method: the trained encoder and the reference bank. It is a LAFA submission bundle, released so the evaluation can be reproduced.
The idea
A protein language model places proteins near each other when their sequences look alike. Function does not follow sequence that closely, and the gap is what limits nearest-neighbour transfer.
This encoder is a single linear map from the 768-dimensional Ankh-base embedding to a 2048-dimensional dictionary, keeping the 128 largest components. It is trained so that the cosine between two proteins' codes matches the information-content-weighted Lin semantic similarity of their propagated GO closures, with hard negatives mined from embedding-near neighbours. In other words, it learns a geometry in which distance means functional similarity.
Held out on 4,000 proteins absent from the training pool, over 120,000 pairs, correlation against Lin semantic similarity:
| representation | Spearman |
|---|---|
| raw Ankh-base embedding | +0.168 |
| encoder, dense | +0.530 |
| encoder, top-128 (what ships here) | +0.467 |
The gain of the shipped code over the raw embedding is +0.300. Fitting inside the earlier release instead moves it to +0.293, so the result does not depend on which release it was fitted in. The figures are differences rather than ratios on purpose: the baseline is small enough that the same two runs read as +179 and +162 percent, a spread that is an artefact of the denominator.
The encoder also removes most of the anisotropy of the raw space: mean cosine between random pairs falls from 0.551 to 0.086.
Contents
| file | what it is |
|---|---|
encoder.pt |
the trained encoder, 6.3 MB |
MANIFEST.json |
what it was fitted on |
VALIDATION.json |
the held-out measurement above |
codes_idx.npy |
dictionary positions of the kept components, uint16 |
codes_val.npy |
their values, float16 |
accessions.npy |
row order, mapping a code back to a protein |
donors.tsv.gz |
the annotations transferred by the neighbours |
BANK.json |
what the bank was built from |
EMBEDDING_RECIPE.json |
the backbone recipe the bank was built with |
EMBEDDING_RECIPE.json is not documentation. A consumer must embed its
queries with that exact recipe, or the queries and the bank end up in
different geometries and accuracy degrades without anything failing.
Training and cutoff
Everything is fitted once and frozen. Nothing is refitted at inference.
| item | value |
|---|---|
| backbone | ElnaggarLab/ankh-base, last layer, mean pooled, L2-normalised, truncated at 2048 |
| annotations | GOA release 227, published 2025-09-04 |
| ontology | releases/2025-07-22 |
| bank rows | 575,503 canonical proteins, over 487,237 distinct sequences |
| donor annotations | 5,880,402 over 557,071 proteins |
| encoder training pool | 60,000 proteins |
| dictionary | 2048, top 128 kept |
| seed | 42 |
The ontology is the last snapshot at or before the annotation release. The training driver refuses to run if that ordering is violated, because a later ontology would put terms that did not exist at the cutoff on the label side of the objective.
The earlier release, GOA 226 of 2025-05-03, is reserved as the window for choosing hyperparameters and is not used to fit what ships here.
Running it
The container that consumes this bundle lives in the PROTEA repository
under apps/lafa_sparse_knn, with its bind-mount contract in
METHOD_CARD.md. Retrieval is exact brute force on the processor, and
the bank is held sparse, so the whole method fits in about 1.5 GB plus
the backbone.
Two things a consumer should know
The bank is keyed by protein, not by sequence. 575,503 canonical accessions resolve to 487,237 distinct sequences, so proteins sharing a sequence appear as separate rows carrying their own annotations. They are distinct donors, which is the intent, but they are also exact ties in retrieval: a query matching a sequence held by five accessions fills five neighbour slots with the same vector.
Self-retrieval is the parity check. Embedding a protein that is in
the bank returns that protein first, at cosine 0.99 or above (measured on
20 proteins: 19 return themselves and the twentieth returns an accession
sharing its exact sequence). If a consumer sees lower, its embedding
recipe has drifted from EMBEDDING_RECIPE.json and the results should
not be trusted.
Scope
This is the retrieval core on its own. It carries no reranker, no alignment features, no taxonomic prior and no per-aspect model. It is an intermediate release of a system under active development, published so the containerised inference path can be exercised end to end.
Citation
Francisco Miguel Perez Canales. PROTEA: protein functional embedding-based annotation.