amethyst-1-mini / README.md
VertexAIco's picture
Upload Amethyst 1 Mini
8babf6a verified
|
Raw
History Blame Contribute Delete
4.8 kB
---
license: gemma
base_model: google/gemma-3-4b-it
language:
- en
library_name: transformers
pipeline_tag: text-generation
tags:
- gemma
- gemma3
- lora
- mlx
- conversational
- distillation
---
# Amethyst 1 Mini
**Amethyst 1 Mini** is a general-purpose chat and instruction-following model, fine-tuned from **Gemma 3 4B IT** using LoRA on a small, high-quality distilled instruction dataset. It is the first model in the Amethyst family β€” an early, small-scale release focused on validating an end-to-end distillation β†’ fine-tune β†’ evaluate pipeline on consumer hardware.
## Model Details
| | |
|---|---|
| **Developed by** | Independent research project |
| **Base model** | [google/gemma-3-4b-it](https://huggingface.co/google/gemma-3-4b-it) |
| **Fine-tuning base checkpoint** | [mlx-community/gemma-3-4b-it-qat-4bit](https://huggingface.co/mlx-community/gemma-3-4b-it-qat-4bit) |
| **Architecture** | Gemma 3, 4B parameters (dense, decoder-only transformer) |
| **Fine-tuning method** | LoRA (rank 8, scale 20.0), fused into the base weights and dequantized to fp16 for this release |
| **Fine-tuning framework** | [MLX](https://github.com/ml-explore/mlx) / `mlx-lm`, on Apple Silicon |
| **Trained modules** | `q_proj`, `k_proj`, `v_proj`, `o_proj`, `gate_proj`, `up_proj`, `down_proj` across 16 layers |
| **Language** | English |
| **License** | [Gemma Terms of Use](https://ai.google.dev/gemma/terms) |
## Training Data
Amethyst 1 Mini was fine-tuned on **1,122 instruction/response pairs** (1,082 train / 40 validation), synthetically generated via knowledge distillation from **`nvidia/nemotron-3-super-120b-a12b`** (Nemotron-3-Super, a 120B-parameter MoE model, ~12B active) through the OpenRouter API.
The dataset spans a deliberately broad set of general-chat categories to encourage well-rounded conversational ability rather than narrow task performance:
- Explanation & misconceptions
- Reasoning & math reasoning
- Code generation
- Extraction & structured output
- Planning
- Roleplay & creative writing
- Translation
- Sentiment classification
- Brainstorming
Each example was generated from a unique prompt, distilled from the teacher model, then filtered/cleaned before training.
## Training Procedure
- **Method:** Supervised fine-tuning via LoRA (rank 8, dropout 0.0, scale 20.0)
- **Optimizer:** Adam, learning rate 1e-5 (constant schedule)
- **Sequence length:** 1024 tokens
- **Gradient checkpointing:** enabled
- **Training steps:** 3,246 iterations total, with validation every 200 steps
- **Checkpoint selection:** the released weights use the **iteration 2,600 checkpoint**, selected for lowest validation loss (1.513) β€” later checkpoints began overfitting on this small dataset (validation loss rose to ~1.95 by the final iterations)
This release merges the selected LoRA adapter into the base model and dequantizes the result to fp16, so it can be loaded directly with `transformers` without any MLX or quantization dependencies.
## Intended Use
Amethyst 1 Mini is intended as a lightweight, general-purpose conversational assistant β€” for experimentation, research into small-scale distillation pipelines, and hobbyist deployment. It is **not** intended for high-stakes, safety-critical, or production use.
## Limitations
- Trained on a small (1,122-example) synthetic dataset β€” behavior can be inconsistent outside the categories represented in training.
- Inherits the general limitations and knowledge cutoff of its base model, Gemma 3 4B IT.
- Distilled from a single teacher model without human review of every example; synthetic-data artifacts (teacher biases, occasional factual errors) may be present.
- This is an early, first-generation checkpoint in an ongoing series β€” later Amethyst releases are expected to use larger, cleaner datasets.
## Usage
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "VertexAIco/amethyst-1-mini"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")
messages = [{"role": "user", "content": "Explain how vaccines train the immune system, in simple terms."}]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(inputs, max_new_tokens=512)
print(tokenizer.decode(out[0][inputs.shape[-1]:], skip_special_tokens=True))
```
## Citation
If you reference this model, please cite it as:
```
@misc{amethyst1mini,
title = {Amethyst 1 Mini},
author = {Independent research project},
year = {2026},
note = {LoRA fine-tune of Gemma 3 4B IT, distilled from Nemotron-3-Super-120B-A12B}
}
```
This model is built on Gemma and subject to the [Gemma Terms of Use](https://ai.google.dev/gemma/terms).