protea-sparse-knn / BANK.json
XaxiPiruli's picture
Correct the distinct-sequence count: 487,237, not the embedding row total
916ee45 verified
Raw
History Blame Contribute Delete
637 Bytes
{
"backbone": "ElnaggarLab/ankh-base",
"dict_dim": 2048,
"top_k": 128,
"bank_rows": 575503,
"distinct_sequences": 487237,
"donor_annotations": 5880402,
"annotation_release": "227",
"annotation_published": "2025-09-04",
"ontology": "releases/2025-07-22",
"encoder_trained_at": "2026-08-10T17:46:26Z",
"built_at": "2026-08-10T18:06:53Z",
"note": "bank_rows are canonical protein accessions; proteins sharing a sequence are separate rows with their own annotations. distinct_sequences counts the sequences behind them, so bank_rows minus distinct_sequences is the number of rows that duplicate another row's code."
}