A newer version of the Gradio SDK is available: 6.20.0
title: Ouroboros
emoji: π§¬
colorFrom: blue
colorTo: purple
sdk: gradio
python_version: '3.10'
app_file: app.py
pinned: false
Ouroboros: Bidirectional Protein-Text Generation
Dual T5 model for protein sequence β text description bidirectional generation.
- Protein β Text: Given an amino acid sequence, generate a functional description.
- Text β Protein: Given a text description, generate candidate protein sequences.
Architecture
Model: ouroboros_model.py
Two independent T5-large models connected by cross-modal FC projection layers.
Protein Modality Text Modality
ββββββββββββββββββββββββ ββββββββββββββββββββββββ
β p_t5_model (T5-large)β β t_t5_model (T5-large)β
β d_model=1024, 24L β β d_model=1024, 24L β
β vocab=128 (ProtT5) β β vocab=32128 (T5-base)β
β ββ β β ββ β
β p_t5_encoder β β t_t5_encoder β
β β β β β β
β fc_p2t (1024β1024) βββββββ fc_t2p (1024β1024) β
β β β β β β
β t_t5_decoder β β p_t5_decoder β
ββββββββββββββββββββββββ ββββββββββββββββββββββββ
- share_fc:
fc_t2p.weight = transpose(fc_p2t.weight), no bias - text_share_param / protein_share_param: Encoder-decoder SA weights tied within each T5
- use_mean_pooling: False (full sequence cross-attention)
- FP32, ~1.23B params (after weight sharing)
- Checkpoint:
dual_model.pt(2.45 GB) β single full state_dict
Weight Sharing (share_weight_t5)
Replicated from dual_training_t5/utils/share_weight.py:126-159:
- SA q/k/v/o weights: encoder.block[i] β decoder.block[i]
- SA layer_norm: encoder.block[i] β decoder.block[i]
- FFN DenseReluDense: encoder.block[i] β decoder.block[0] (
block[0]is intentional β mirrors training behavior) - No bias sharing, no FFN layer_norm sharing
Tokenizer Handling
| Protein | Text | |
|---|---|---|
| Tokenizer | Rostlab/prot_t5_xl_uniref50 |
google-t5/t5-base |
| vocab_size (model) | 128 (hardcoded) | 32128 (hardcoded) |
| Vocab from tokenizer | 28 (transformers >=4.40) | 32000 (transformers >=4.40) |
Critical: Model vocab_sizes are hardcoded, not read from tokenizer. Newer transformers report 28/32000 but checkpoint was trained with 128/32128. Loading from tokenizer causes RuntimeError: size mismatch.
Inference Pipeline
p2t: protein_seq β [" ".join(chars)] β protein_tokenizer β p_t5_encoder β fc_p2t β t_t5_model.generate() β text_tokenizer.decode()
t2p: text_desc β text_tokenizer β t_t5_encoder β fc_t2p β p_t5_model.generate() β protein_tokenizer.decode() β .replace(" ", "")
Generation params: temperature=1.0, top_k=40, top_p=0.9, num_beams=4, max_length=512
Checkpoint Loading
state_dict = torch.load(path, map_location='cpu', weights_only=True)
# Strip 'module.' prefix from DDP training
new_state_dict = {k.replace('module.', ''): v for k, v in state_dict.items()}
model.load_state_dict(new_state_dict, strict=False)
strict=False because shared weights (encoder-decoder tied params, share_fc transposed params) may cause missing/unexpected key warnings.
Gradio UI (app.py)
Color Scheme
| Element | Protein (red/pink) | Text (blue) |
|---|---|---|
| Accent | #E74C3C |
#2471A3 |
| Border | #E6B0AA |
#A9CCE3 |
| Background | #FDEDEC |
#EBF5FB |
Tab 1: Protein β Text
- Monospace input with red border/focus
- Output with blue left-accent border
- Advanced: temperature, top_k, top_p, num_beams, greedy toggle
Tab 2: Text β Protein
Three input modes:
Structured Prompt Builder (default): 29 annotation fields matching UniProt categories
- 7 common fields visible (similarity, function, catalytic activity, subunit, cofactor, pathway, subcellular location)
- 22 more fields in collapsible accordion
- "Build Prompt" button: assembles all fields per ORDER_LIST sequence
- "Random Example" button: loads from
100sample.json
Direct Input: Collapsible accordion, paste raw text prompt
Assembled Prompt textbox: always editable, serve as final input to generation
Prompt Assembly Rule
ORDER_LIST traversal β if field has content β append to prompt (no separator)
Training format: similarity + function + catalytic activity + path way + ...
100sample.json Data
JSON array of UniProt samples, each with:
uniprot_id,sequence(amino acid string)- Annotation fields matching ORDER_LIST keys (each a single-string list)
prompt: assembled concatenation of all annotation fieldsfunc_sim: short version (similarity + function only)
HF Spaces Deployment
Current Setup
- Space:
REXWind/ouroboros(Gradio SDK, Python 3.10, CPU basic) - Model Repo:
REXWind/ouroboros-weights(storesdual_model.pt, ~2.5 GB via Git LFS) - Auto-download:
app.pycallshuggingface_hub.hf_hub_download()on startup
Requirements (requirements.txt)
torch>=2.0.0
transformers>=4.36.0
sentencepiece>=0.1.99
huggingface_hub>=0.20.0
Do NOT pin gradio β HF Spaces injects its own gradio version during build. Pinning causes dependency conflicts.
YAML Front Matter (README.md)
sdk: gradio
python_version: "3.10"
app_file: app.py
- Python 3.10 required (3.13 has pydub/audioop incompatibility)
- Gradio SDK version auto-managed by HF
Key Files
| File | Purpose |
|---|---|
app.py |
Gradio web interface (model loading, UI layout, event bindings) |
ouroboros_model.py |
Self-contained model class (no dependency on training utils/) |
t5config_large.json |
T5-large architecture config |
100sample.json |
Sample UniProt annotations for Random Example button |
requirements.txt |
Minimal Python dependencies (gradio managed by HF) |
Source Project Paths
All paths on gpfs shared storage:
Training code: /mnt/shared-storage-user/dnacoding/protein-text-dual-learn/dual_t5_workspace/dual_training_t5/
Checkpoint: /mnt/shared-storage-user/dnacoding/protein-text-dual-learn/dual_t5_workspace/model_weights/dual_t5_large_newds_v5/best_recovery/dual_model.pt
Tokenizer: /mnt/shared-storage-user/dnacoding/protein-text-dual-learn/dual_t5_workspace/hf_weights/prot_t5_xl_uniref50/
Tokenizer: /mnt/shared-storage-user/dnacoding/protein-text-dual-learn/dual_t5_workspace/hf_weights/t5-base/
Demo source: /mnt/shared-storage-user/dnacoding/protein-text-dual-learn/dual_t5_workspace/Ouroboros_hf_demo/
HF Space clone: /mnt/shared-storage-user/dnacoding/protein-text-dual-learn/dual_t5_workspace/hf_space_ouroboros/
Known Issues & Fixes
| Issue | Cause | Fix |
|---|---|---|
size mismatch: [32128] vs [32000] |
New transformers reports vocab=32000 | Hardcode configs_t.vocab_size = 32128 |
ImportError: cannot import HfFolder |
gradio 5 + huggingface_hub >=1.0 | Don't pin gradio version in requirements |
TypeError: unhashable type: 'dict' |
gradio 4.44 + jinja2 >=3.1.5 | Don't pin gradio; use Python 3.10 |
Cannot install gradio<5.0.0 and gradio==6.15.2 |
HF installs gradio 6 | Don't specify gradio in requirements |
Repository storage limit (Max: 1 GB) |
Space only; model 2.5GB | Put checkpoint in separate Model Repo |
| Protein vocab=28 | New transformers prot_t5 vocab | configs_p.vocab_size = 128 (hardcoded) |