Spaces:
Configuration error
Configuration error
Update README.md
Browse files
README.md
CHANGED
|
@@ -4,71 +4,45 @@ title: Opus Research
|
|
| 4 |
|
| 5 |
# Opus Research
|
| 6 |
|
| 7 |
-
Independent research on small language models
|
| 8 |
-
|
| 9 |
-
constrained hardware.
|
| 10 |
|
| 11 |
-
|
| 12 |
-
work.
|
| 13 |
|
| 14 |
-
|
|
|
|
|
|
|
|
|
|
| 15 |
|
| 16 |
-
|
| 17 |
-
language model pretrained from scratch on 2Γ RTX 4090. Trained in 42 hours;
|
| 18 |
-
energy use documented in the model card.
|
| 19 |
|
| 20 |
-
**[
|
| 21 |
-
|
| 22 |
-
|
| 23 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
| 24 |
|
| 25 |
-
|
| 26 |
-
β experiment in adding explicit chain-of-thought to a base model that lacks it.
|
| 27 |
-
LoRA on Gemma 2 2B, trained on our own reasoning dataset. Training loss 3.56 β
|
| 28 |
-
0.69 over 566 steps. Exploratory; not benchmarked.
|
| 29 |
|
| 30 |
-
**[
|
| 31 |
-
β
|
| 32 |
-
|
| 33 |
-
|
| 34 |
-
|
| 35 |
-
|
| 36 |
-
**[bernard-qwen3-14b-thinking](https://huggingface.co/opus-research/bernard-qwen3-14b-thinking)**
|
| 37 |
-
β follow-up asking whether fine-tuning can change a model's *internal reasoning*
|
| 38 |
-
rather than only its output voice. The reasoning block is part of the training
|
| 39 |
-
target, with a same-length neutral-reasoning control run so that any effect is
|
| 40 |
-
attributable. It works, and the informative result is what that costs: replies
|
| 41 |
-
faithfully execute whatever the reasoning decided, so training traces that framed
|
| 42 |
-
an input as having "nothing to work with" taught the model to reason its way out
|
| 43 |
-
of the persona entirely. Also documents two things a standard eval will not tell
|
| 44 |
-
you β that a persona fine-tune can be *worse* than its base model at following
|
| 45 |
-
the system prompt, and that adapters overfit to a personal corpus memorise the
|
| 46 |
-
names in it at 12Γ the base rate, invisibly to any benchmark that does not ask
|
| 47 |
-
for them.
|
| 48 |
|
| 49 |
## Datasets
|
| 50 |
|
| 51 |
-
**[opus-thinking-10k](https://huggingface.co/datasets/opus-research/opus-thinking-10k)**
|
| 52 |
-
β ~
|
| 53 |
-
|
| 54 |
-
**[opus-thinking](https://huggingface.co/datasets/opus-research/opus-thinking)**
|
| 55 |
-
β earlier iteration of the same dataset.
|
| 56 |
-
|
| 57 |
-
## What we work on
|
| 58 |
-
|
| 59 |
-
- Pretraining small models end-to-end on consumer hardware
|
| 60 |
-
- Post-training: LoRA, style transfer, capability injection, and supervising a
|
| 61 |
-
model's reasoning rather than only its output
|
| 62 |
-
- Dataset construction and filtering methodology, including privacy auditing of
|
| 63 |
-
corpora mined from personal data
|
| 64 |
-
- Quantization and CPU inference for models that would otherwise need a
|
| 65 |
-
datacenter
|
| 66 |
|
| 67 |
-
|
| 68 |
-
GPUs, and the ceiling that imposes is the main thing limiting what we publish
|
| 69 |
-
next.
|
| 70 |
|
| 71 |
-
|
|
|
|
|
|
|
| 72 |
|
| 73 |
-
|
| 74 |
-
through the organization page.
|
|
|
|
| 4 |
|
| 5 |
# Opus Research
|
| 6 |
|
| 7 |
+
Independent research on small language models, done entirely on consumer
|
| 8 |
+
hardware. We publish weights, methodology, and the experiments that failed.
|
|
|
|
| 9 |
|
| 10 |
+
## Pretraining
|
|
|
|
| 11 |
|
| 12 |
+
- **[opus-1.5](https://huggingface.co/opus-research/opus-1.5)** β 0.88B model
|
| 13 |
+
pretrained from scratch in 42 hours on 2Γ RTX 4090.
|
| 14 |
+
- **[opus-2.0](https://huggingface.co/opus-research/opus-2.0)** β successor,
|
| 15 |
+
currently paused on compute, not on research.
|
| 16 |
|
| 17 |
+
## Fine-tunes
|
|
|
|
|
|
|
| 18 |
|
| 19 |
+
- **[bernard-qwen3-14b-thinking](https://huggingface.co/opus-research/bernard-qwen3-14b-thinking)**
|
| 20 |
+
β can fine-tuning change a model's *reasoning*, not just its voice? Yes, and
|
| 21 |
+
the card documents what that costs.
|
| 22 |
+
- **[bernard-gpt-oss-20b-lora](https://huggingface.co/opus-research/bernard-gpt-oss-20b-lora)**
|
| 23 |
+
β style-transfer LoRA where the worst validation loss was the right
|
| 24 |
+
checkpoint to ship.
|
| 25 |
+
- **[gemma-2-2b-thinking](https://huggingface.co/opus-research/gemma-2-2b-thinking)**
|
| 26 |
+
β adding chain-of-thought to a base model that lacks it.
|
| 27 |
|
| 28 |
+
## Classifiers
|
|
|
|
|
|
|
|
|
|
| 29 |
|
| 30 |
+
- **[opus-moderation-1](https://huggingface.co/opus-research/opus-moderation-1)**
|
| 31 |
+
β toxicity + jailbreak detection in one 149M model. The card shows three
|
| 32 |
+
failed recipes and the identity bias each one had.
|
| 33 |
+
- **[opus-emotion-1](https://huggingface.co/opus-research/opus-emotion-1)** β
|
| 34 |
+
28 emotions, trained in 14 minutes on an unsupported AMD GPU.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 35 |
|
| 36 |
## Datasets
|
| 37 |
|
| 38 |
+
- **[opus-thinking-10k](https://huggingface.co/datasets/opus-research/opus-thinking-10k)**
|
| 39 |
+
β ~9k chain-of-thought examples.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 40 |
|
| 41 |
+
## Team
|
|
|
|
|
|
|
| 42 |
|
| 43 |
+
- **[spoodyzz](https://huggingface.co/spoodyzz)** (Michael)
|
| 44 |
+
- **[cgcristi0](https://huggingface.co/cgcristi0)** (CGCristi)
|
| 45 |
+
- **[targeteater](https://huggingface.co/targeteater)** (target)
|
| 46 |
|
| 47 |
+
Our constraint is compute, not ideas. Open to collaboration and grant
|
| 48 |
+
programs β reach us through the organization page.
|