DrugSpace-8B
DrugSpace-8B is the final standalone DrugSpace embedding model for English
drug descriptions. It combines PubTator-based biomedical domain adaptation
with contrastive alignment on OpenFDA drug-label text.
The Stage II LoRA adapter has already been merged into the Stage I model. This repository therefore contains a complete model and does not require a separate base model or PEFT adapter at inference time.
Model lineage
| Stage | Model or data |
|---|---|
| Foundation model | meta-llama/Meta-Llama-3.1-8B-Instruct |
| Stage I | MNTP adaptation on PubTator biomedical text: cczzzyyy/DrugSpace-mntp-8B |
| Stage II | OpenFDA contrastive-alignment adapter: cczzzyyy/DrugSpace-openfda-lora |
| Final release | Stage II adapter merged into the Stage I model |
Usage
Install LLM2Vec and its dependencies:
pip install llm2vec transformers accelerate
Load the merged model directly:
import torch
import torch.nn.functional as F
from llm2vec import LLM2Vec
device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.bfloat16 if torch.cuda.is_available() else torch.float32
model = LLM2Vec.from_pretrained(
base_model_name_or_path="cczzzyyy/DrugSpace-8B",
enable_bidirectional=True,
pooling_mode="mean",
max_length=512,
doc_max_length=400,
skip_instruction=True,
device_map=device,
torch_dtype=dtype,
)
texts = [
"Metformin is an oral biguanide medication used to improve glycemic control in type 2 diabetes.",
"Another complete English description of a drug.",
]
embeddings = model.encode(
texts,
batch_size=2,
show_progress_bar=True,
convert_to_tensor=True,
)
similarity = F.cosine_similarity(
embeddings[0].unsqueeze(0),
embeddings[1].unsqueeze(0),
).item()
print(embeddings.shape)
print(similarity)
The released configuration uses mean pooling, a maximum sequence length of
512, a document maximum length of 400, and skip_instruction=true.
Model details
- Model type: bidirectional Llama-based LLM2Vec embedding model
- Parameters: 8B
- Weight dtype: bfloat16
- Task: drug-description embedding, semantic similarity, and retrieval
- Language: English
- Output: one fixed-size embedding per input description
Training data
Stage I: PubTator
The Stage I model was adapted to biomedical literature using an MNTP objective on PubTator text.
Stage II: OpenFDA
Stage II used OpenFDA-derived single-ingredient labeling records. Text views were drawn from the following label sections when available:
- indications and usage
- clinical pharmacology
- mechanism of action
- pharmacodynamics
- pharmacokinetics
- contraindications
- warnings and cautions
- drug interactions
- adverse reactions
The processed source contained 1,867 records. Of these, 1,857 records with at least two distinct non-empty text views were eligible for contrastive training.
Stage II training procedure
- Fixed training set constructed once with seed 42 and reused for all 20 epochs
- 3,714 constructed triplets; 3,712 used per epoch after dropping the final incomplete batch
- Global batch size 64
- Learning rate 1e-4
- 300 warm-up steps
- Maximum sequence length 512
- bfloat16 training
- Final training step used as the released checkpoint; no alignment validation split or early stopping
Merge details
The Stage II adapter was merged into cczzzyyy/DrugSpace-mntp-8B with PEFT
merge_and_unload(safe_merge=True) and saved as bfloat16 safetensors. The
saved standalone checkpoint was reloaded successfully and checked against the
unmerged base-plus-adapter pipeline.
Intended use
This model is intended for research involving English drug descriptions, including embedding generation, semantic similarity, clustering, and retrieval. Inputs should be complete, clinically meaningful descriptions and should use a consistent level of detail and writing style when results are compared.
Limitations
- The model is intended for representation learning, not text generation.
- Training data and label-section availability may introduce coverage and documentation biases.
- Performance may degrade for non-English text, very short names, incomplete descriptions, or text outside the biomedical drug domain.
- Similarity in the embedding space does not establish therapeutic equivalence, safety, efficacy, or causal relationships.
- The model is for research use and must not be used as a substitute for professional medical judgment.
Evaluation variant
For experiments that require the pre-2020 Stage I MNTP base, use
cczzzyyy/DrugSpace-eval-8B.
License
Use of this model is subject to the applicable Llama 3.1 license and acceptable-use terms.
- Downloads last month
- 5
Model tree for cczzzyyy/DrugSpace-8B
Base model
meta-llama/Llama-3.1-8B