OLMo-7B Full Fine-Tune — Chemistry SMILES CPT

Model Description

This model is a full-parameter fine-tuned version of Codemaster67/Olmo-7b-spe trained on chemistry SMILES strings from the Codemaster67/Causal_lm_chemistry_1M_rows dataset.

The base model's tokenizer was pre-extended with ~300 SPE (SMILES Pair Encoding) chemistry tokens plus <|start_of_smiles|> / <|end_of_smiles|> special tokens, and its embedding & LM-head layers were resized with mean-initialised vectors for the new tokens.

Training Details

Parameter Value
Method Full Fine-Tune (all weights updated)
Parallelism FSDP (Fully Sharded Data Parallel)
Epochs 1
Learning Rate 5e-06
Batch Size (per device) 16
Gradient Accumulation 1
Max Sequence Length 512
Warmup Ratio 0.1
Weight Decay 0.01
Scheduler Cosine
Precision bf16
Augmentation OFF
Training Samples 250000
Eval Samples 25000

Evaluation Results

Metric Value
Final Eval Loss 0.9727568626403809
Final Eval Perplexity 2.645226943673604
Training Loss 1.1177

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("harindhar10/olmo_chem_lora_cpt_LoRA_500k", trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained("harindhar10/olmo_chem_lora_cpt_LoRA_500k", trust_remote_code=True)

smiles_input = "<|start_of_smiles|>CC(=O)Oc1ccccc1C(=O)O<|end_of_smiles|>"
inputs = tokenizer(smiles_input, return_tensors="pt")
outputs = model.generate(**inputs, max_new_tokens=128)
print(tokenizer.decode(outputs[0], skip_special_tokens=False))

Intended Use

Chemistry-domain language modelling, SMILES generation and completion, and downstream molecular property prediction via fine-tuning.

Limitations

  • Trained primarily on SMILES strings; natural-language instruction-following ability may degrade compared to the base OLMo checkpoint.
  • Augmentation was disabled for this run.
Downloads last month
50
Safetensors
Model size
7B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Codemaster67/olmo_chem_250k

Finetuned
(2)
this model

Dataset used to train Codemaster67/olmo_chem_250k