YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

I need to translate the README into English while preserving the YAML frontmatter and keeping any code blocks as-is.

license: apache-2.0 library_name: transformers tags: [single-cell, scRNA-seq, biology, genomics, lucaone, lucacell, cell-foundation-model] pipeline_tag: feature-extraction

LucaCell-v1.0

A cell foundation model: gene representations (LucaOne nucleotide-sequence embeddings) + binned expression values β†’ a 40-layer Transformer encoder.

Backbone 40 layers / 40 heads / hidden 2560 / FFN 10240 / RoPE / pre-LN
Parameters β‰ˆ 3.15 B (bf16 β‰ˆ 6.3 GB)
Gene representation LucaGroup/LucaOne-step60M, float32, nucleotide sequences right-truncated to 10240, mean(last_hidden[:, 1:-1, :]) β†’ 2560-dim
Cell length 1200 genes + [E_CLS] + [E_SEP] = 1202
Expression vocabulary 6 special tokens + 50 bins = 56
Pre-training task express_token_mask (MLM, mlm_probability = 0.5, 80/10/10)
Prediction targets [E_NON] (gene not expressed) + all expression bins
Disabled components express_sorted embedding, gene_type embedding, nucleotide token encoder
Checkpoint dtype bf16

Expression-value encoding rules (must be followed)

Gene state What to write id
Not expressed / zero count "[E_NON]" 5
Expressed (50 bins, 1-indexed) "1" … "50" 6 … 55

⚠️ Do NOT use "0" to indicate "not expressed" or the lowest bin β€” it is not in the vocabulary and will be mapped to [E_UNK]. Such positions will not participate in MLM and will not contribute to the loss, and the model will consume an almost untrained embedding β€” all without raising an error.

  • The qcut_bins and feature_ids columns must correspond one-to-one; non-expressed genes go into the zero_freature_ids column (the collator automatically sets their expression values to [E_NON]).
  • Before you start, run: python tools/check_bin_labels.py --model_dir <repo> --data_file <your.csv>. The reported "fraction of labels not found in the vocabulary" must be 0.
  • The tokenizer defaults to unknown_bin_policy="error", which raises an exception on unknown labels (can be changed to "warn").

dtype rules

Scenario Model Weight loading Compute precision Output dtype Key arguments
Inference / embedding extraction LucaCell float32 bfloat16 float32 (auto-cast) --use_bf16
LucaOne float32 float32 float32 β€”
Continued pre-training LucaCell float32 (master weights) bfloat16 (AMP) β€” torch_dtype=torch.float32 + --bf16
LucaOne float32 β€” float32 Frozen offline, no gradient updates

Installation

pip install torch==2.5.1 torchvision torchaudio --extra-index-url https://download.pytorch.org/whl/cu121
pip install -r requirements.txt

Meta data

gene_positions.txt: the sorted index of all seqs.
express_vocab.txt: the vocab of the gene expression bins.

Data format (gene sequence FASTA file + cell CSV file)

binned_logcpm.test.fa: FASTA file used for sequence embedding.
binned_logcpm.test.csv: the cell data, one record per cell, including: sample_id, sample_type, seq_ids (list), expression_bins.

1) Pre-compute gene embeddings with LucaOne (only once)

python lucaone_embedding.py \
      --gene_fasta ./test_data/binned_logcpm.test.fa \
      --emb_dir ../emb/seq/mean_vector \
      --max_len 10240 \
      --pooling_type mean \
      --overwrite \
      --gpu_id 0 

2) Cell embedding

python lucacell_embedding.py \
     --model_name_or_path LucaGroup/LucaCell-v1.0-step6.4M \
     --input_file ./test_data/binned_logcpm.test.csv \
     --max_gene_len 1202 \
     --save_path ../emb/cell/ \
     --gene_emb_dirs ../emb/seq/mean_vector \
     --embedding_type matrix  \
     --overwrite \
     --use_bf16 \
     --add_special_tokens \
     --attn_impl sdpa \
     --gpu_id 0
Downloads last month
52
Safetensors
Model size
3B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Collection including LucaGroup/LucaCell-v1.0-step6.4M