README / README.md
spood spoody
Upload README.md with huggingface_hub
681710f verified
|
Raw
History Blame Contribute Delete
2.39 kB
---
title: Opus Research
---
# Opus Research
Independent research on small language models β€” pretraining from scratch,
post-training, and the practical engineering around getting models to run on
constrained hardware.
We publish weights, datasets, and methodology, including the parts that did not
work.
## Models
**[opus-1.5](https://huggingface.co/opus-research/opus-1.5)** β€” 0.88B parameter
language model pretrained from scratch on 2Γ— RTX 4090. Trained in 42 hours;
energy use documented in the model card.
**[opus-2.0](https://huggingface.co/opus-research/opus-2.0)** β€” successor model,
targeting a larger parameter count. **Training is currently paused: we do not
have the compute budget to continue.** The architecture and data pipeline are
ready; the run is blocked on GPU access rather than on research.
**[gemma-2-2b-thinking](https://huggingface.co/opus-research/gemma-2-2b-thinking)**
β€” experiment in adding explicit chain-of-thought to a base model that lacks it.
LoRA on Gemma 2 2B, trained on our own reasoning dataset. Training loss 3.56 β†’
0.69 over 566 steps. Exploratory; not benchmarked.
**[bernard-gpt-oss-20b-lora](https://huggingface.co/opus-research/bernard-gpt-oss-20b-lora)**
β€” style-transfer LoRA for gpt-oss-20b, a 21B mixture-of-experts model. Notable
mainly for the training notes: a four-run ablation showing that the checkpoint
with the *worst* validation loss was the correct one to ship, plus a set of
gpt-oss-specific engineering gotchas we could not find documented elsewhere.
## Datasets
**[opus-thinking-10k](https://huggingface.co/datasets/opus-research/opus-thinking-10k)**
β€” ~9,000 chain-of-thought examples, used to train `gemma-2-2b-thinking`.
**[opus-thinking](https://huggingface.co/datasets/opus-research/opus-thinking)**
β€” earlier iteration of the same dataset.
## What we work on
- Pretraining small models end-to-end on consumer hardware
- Post-training: LoRA, style transfer, capability injection
- Dataset construction and filtering methodology
- Quantization and CPU inference for models that would otherwise need a
datacenter
Our constraint is compute, not ideas. Everything above was produced on consumer
GPUs, and the ceiling that imposes is the main thing limiting what we publish
next.
## Contact
Open to collaboration, compute partnerships, and grant programs. Reach us
through the organization page.