--- license: gemma base_model: google/gemma-3-4b-it language: - en library_name: transformers pipeline_tag: text-generation tags: - gemma - gemma3 - lora - mlx - conversational - distillation --- # Amethyst 1 Mini **Amethyst 1 Mini** is a general-purpose chat and instruction-following model, fine-tuned from **Gemma 3 4B IT** using LoRA on a small, high-quality distilled instruction dataset. It is the first model in the Amethyst family — an early, small-scale release focused on validating an end-to-end distillation → fine-tune → evaluate pipeline on consumer hardware. ## Model Details | | | |---|---| | **Developed by** | Independent research project | | **Base model** | [google/gemma-3-4b-it](https://huggingface.co/google/gemma-3-4b-it) | | **Fine-tuning base checkpoint** | [mlx-community/gemma-3-4b-it-qat-4bit](https://huggingface.co/mlx-community/gemma-3-4b-it-qat-4bit) | | **Architecture** | Gemma 3, 4B parameters (dense, decoder-only transformer) | | **Fine-tuning method** | LoRA (rank 8, scale 20.0), fused into the base weights and dequantized to fp16 for this release | | **Fine-tuning framework** | [MLX](https://github.com/ml-explore/mlx) / `mlx-lm`, on Apple Silicon | | **Trained modules** | `q_proj`, `k_proj`, `v_proj`, `o_proj`, `gate_proj`, `up_proj`, `down_proj` across 16 layers | | **Language** | English | | **License** | [Gemma Terms of Use](https://ai.google.dev/gemma/terms) | ## Training Data Amethyst 1 Mini was fine-tuned on **1,122 instruction/response pairs** (1,082 train / 40 validation), synthetically generated via knowledge distillation from **`nvidia/nemotron-3-super-120b-a12b`** (Nemotron-3-Super, a 120B-parameter MoE model, ~12B active) through the OpenRouter API. The dataset spans a deliberately broad set of general-chat categories to encourage well-rounded conversational ability rather than narrow task performance: - Explanation & misconceptions - Reasoning & math reasoning - Code generation - Extraction & structured output - Planning - Roleplay & creative writing - Translation - Sentiment classification - Brainstorming Each example was generated from a unique prompt, distilled from the teacher model, then filtered/cleaned before training. ## Training Procedure - **Method:** Supervised fine-tuning via LoRA (rank 8, dropout 0.0, scale 20.0) - **Optimizer:** Adam, learning rate 1e-5 (constant schedule) - **Sequence length:** 1024 tokens - **Gradient checkpointing:** enabled - **Training steps:** 3,246 iterations total, with validation every 200 steps - **Checkpoint selection:** the released weights use the **iteration 2,600 checkpoint**, selected for lowest validation loss (1.513) — later checkpoints began overfitting on this small dataset (validation loss rose to ~1.95 by the final iterations) This release merges the selected LoRA adapter into the base model and dequantizes the result to fp16, so it can be loaded directly with `transformers` without any MLX or quantization dependencies. ## Intended Use Amethyst 1 Mini is intended as a lightweight, general-purpose conversational assistant — for experimentation, research into small-scale distillation pipelines, and hobbyist deployment. It is **not** intended for high-stakes, safety-critical, or production use. ## Limitations - Trained on a small (1,122-example) synthetic dataset — behavior can be inconsistent outside the categories represented in training. - Inherits the general limitations and knowledge cutoff of its base model, Gemma 3 4B IT. - Distilled from a single teacher model without human review of every example; synthetic-data artifacts (teacher biases, occasional factual errors) may be present. - This is an early, first-generation checkpoint in an ongoing series — later Amethyst releases are expected to use larger, cleaner datasets. ## Usage ```python from transformers import AutoModelForCausalLM, AutoTokenizer model_id = "VertexAIco/amethyst-1-mini" tokenizer = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto") messages = [{"role": "user", "content": "Explain how vaccines train the immune system, in simple terms."}] inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device) out = model.generate(inputs, max_new_tokens=512) print(tokenizer.decode(out[0][inputs.shape[-1]:], skip_special_tokens=True)) ``` ## Citation If you reference this model, please cite it as: ``` @misc{amethyst1mini, title = {Amethyst 1 Mini}, author = {Independent research project}, year = {2026}, note = {LoRA fine-tune of Gemma 3 4B IT, distilled from Nemotron-3-Super-120B-A12B} } ``` This model is built on Gemma and subject to the [Gemma Terms of Use](https://ai.google.dev/gemma/terms).