DrugSpace-8B

DrugSpace-8B is the final standalone DrugSpace embedding model for English drug descriptions. It combines PubTator-based biomedical domain adaptation with contrastive alignment on OpenFDA drug-label text.

The Stage II LoRA adapter has already been merged into the Stage I model. This repository therefore contains a complete model and does not require a separate base model or PEFT adapter at inference time.

Model lineage

Stage Model or data
Foundation model meta-llama/Meta-Llama-3.1-8B-Instruct
Stage I MNTP adaptation on PubTator biomedical text: cczzzyyy/DrugSpace-mntp-8B
Stage II OpenFDA contrastive-alignment adapter: cczzzyyy/DrugSpace-openfda-lora
Final release Stage II adapter merged into the Stage I model

Usage

Install LLM2Vec and its dependencies:

pip install llm2vec transformers accelerate

Load the merged model directly:

import torch
import torch.nn.functional as F
from llm2vec import LLM2Vec

device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.bfloat16 if torch.cuda.is_available() else torch.float32

model = LLM2Vec.from_pretrained(
    base_model_name_or_path="cczzzyyy/DrugSpace-8B",
    enable_bidirectional=True,
    pooling_mode="mean",
    max_length=512,
    doc_max_length=400,
    skip_instruction=True,
    device_map=device,
    torch_dtype=dtype,
)

texts = [
    "Metformin is an oral biguanide medication used to improve glycemic control in type 2 diabetes.",
    "Another complete English description of a drug.",
]

embeddings = model.encode(
    texts,
    batch_size=2,
    show_progress_bar=True,
    convert_to_tensor=True,
)

similarity = F.cosine_similarity(
    embeddings[0].unsqueeze(0),
    embeddings[1].unsqueeze(0),
).item()

print(embeddings.shape)
print(similarity)

The released configuration uses mean pooling, a maximum sequence length of 512, a document maximum length of 400, and skip_instruction=true.

Model details

  • Model type: bidirectional Llama-based LLM2Vec embedding model
  • Parameters: 8B
  • Weight dtype: bfloat16
  • Task: drug-description embedding, semantic similarity, and retrieval
  • Language: English
  • Output: one fixed-size embedding per input description

Training data

Stage I: PubTator

The Stage I model was adapted to biomedical literature using an MNTP objective on PubTator text.

Stage II: OpenFDA

Stage II used OpenFDA-derived single-ingredient labeling records. Text views were drawn from the following label sections when available:

  • indications and usage
  • clinical pharmacology
  • mechanism of action
  • pharmacodynamics
  • pharmacokinetics
  • contraindications
  • warnings and cautions
  • drug interactions
  • adverse reactions

The processed source contained 1,867 records. Of these, 1,857 records with at least two distinct non-empty text views were eligible for contrastive training.

Stage II training procedure

  • Fixed training set constructed once with seed 42 and reused for all 20 epochs
  • 3,714 constructed triplets; 3,712 used per epoch after dropping the final incomplete batch
  • Global batch size 64
  • Learning rate 1e-4
  • 300 warm-up steps
  • Maximum sequence length 512
  • bfloat16 training
  • Final training step used as the released checkpoint; no alignment validation split or early stopping

Merge details

The Stage II adapter was merged into cczzzyyy/DrugSpace-mntp-8B with PEFT merge_and_unload(safe_merge=True) and saved as bfloat16 safetensors. The saved standalone checkpoint was reloaded successfully and checked against the unmerged base-plus-adapter pipeline.

Intended use

This model is intended for research involving English drug descriptions, including embedding generation, semantic similarity, clustering, and retrieval. Inputs should be complete, clinically meaningful descriptions and should use a consistent level of detail and writing style when results are compared.

Limitations

  • The model is intended for representation learning, not text generation.
  • Training data and label-section availability may introduce coverage and documentation biases.
  • Performance may degrade for non-English text, very short names, incomplete descriptions, or text outside the biomedical drug domain.
  • Similarity in the embedding space does not establish therapeutic equivalence, safety, efficacy, or causal relationships.
  • The model is for research use and must not be used as a substitute for professional medical judgment.

Evaluation variant

For experiments that require the pre-2020 Stage I MNTP base, use cczzzyyy/DrugSpace-eval-8B.

License

Use of this model is subject to the applicable Llama 3.1 license and acceptable-use terms.

Downloads last month
5
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for cczzzyyy/DrugSpace-8B

Finetuned
(1)
this model