YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
license: apache-2.0 library_name: transformers tags: [single-cell, scRNA-seq, biology, genomics, lucaone, lucacell, cell-foundation-model] pipeline_tag: feature-extraction
LucaCell-v1.0
A cell foundation model: gene representations (LucaOne nucleotide-sequence embeddings) + binned expression values β a 40-layer Transformer encoder.
| Backbone | 40 layers / 40 heads / hidden 2560 / FFN 10240 / RoPE / pre-LN |
| Parameters | β 3.15 B (bf16 β 6.3 GB) |
| Gene representation | LucaGroup/LucaOne-step60M, float32, nucleotide sequences right-truncated to 10240, mean(last_hidden[:, 1:-1, :]) β 2560-dim |
| Cell length | 1200 genes + [E_CLS] + [E_SEP] = 1202 |
| Expression vocabulary | 6 special tokens + 50 bins = 56 |
| Pre-training task | express_token_mask (MLM, mlm_probability = 0.5, 80/10/10) |
| Prediction targets | [E_NON] (gene not expressed) + all expression bins |
| Disabled components | express_sorted embedding, gene_type embedding, nucleotide token encoder |
| Checkpoint dtype | bf16 |
Expression-value encoding rules (must be followed)
| Gene state | What to write | id |
|---|---|---|
| Not expressed / zero count | "[E_NON]" |
5 |
| Expressed (50 bins, 1-indexed) | "1" β¦ "50" |
6 β¦ 55 |
β οΈ Do NOT use "0" to indicate "not expressed" or the lowest bin β it is not in the vocabulary and will be mapped to [E_UNK]. Such positions will not participate in MLM and will not contribute to the loss, and the model will consume an almost untrained embedding β all without raising an error.
- The
qcut_binsandfeature_idscolumns must correspond one-to-one; non-expressed genes go into thezero_freature_idscolumn (the collator automatically sets their expression values to[E_NON]). - Before you start, run:
python tools/check_bin_labels.py --model_dir <repo> --data_file <your.csv>. The reported "fraction of labels not found in the vocabulary" must be 0. - The tokenizer defaults to
unknown_bin_policy="error", which raises an exception on unknown labels (can be changed to"warn").
dtype rules
| Scenario | Model | Weight loading | Compute precision | Output dtype | Key arguments |
|---|---|---|---|---|---|
| Inference / embedding extraction | LucaCell | float32 |
bfloat16 |
float32 (auto-cast) |
--use_bf16 |
| LucaOne | float32 |
float32 |
float32 |
β | |
| Continued pre-training | LucaCell | float32 (master weights) |
bfloat16 (AMP) |
β | torch_dtype=torch.float32 + --bf16 |
| LucaOne | float32 |
β | float32 |
Frozen offline, no gradient updates |
Installation
pip install torch==2.5.1 torchvision torchaudio --extra-index-url https://download.pytorch.org/whl/cu121
pip install -r requirements.txt
Meta data
gene_positions.txt: the sorted index of all seqs.express_vocab.txt: the vocab of the gene expression bins.
Data format (gene sequence FASTA file + cell CSV file)
binned_logcpm.test.fa: FASTA file used for sequence embedding.binned_logcpm.test.csv: the cell data, one record per cell, including: sample_id, sample_type, seq_ids (list), expression_bins.
1) Pre-compute gene embeddings with LucaOne (only once)
python lucaone_embedding.py \
--gene_fasta ./test_data/binned_logcpm.test.fa \
--emb_dir ../emb/seq/mean_vector \
--max_len 10240 \
--pooling_type mean \
--overwrite \
--gpu_id 0
2) Cell embedding
python lucacell_embedding.py \
--model_name_or_path LucaGroup/LucaCell-v1.0-step6.4M \
--input_file ./test_data/binned_logcpm.test.csv \
--max_gene_len 1202 \
--save_path ../emb/cell/ \
--gene_emb_dirs ../emb/seq/mean_vector \
--embedding_type matrix \
--overwrite \
--use_bf16 \
--add_special_tokens \
--attn_impl sdpa \
--gpu_id 0
- Downloads last month
- 52