Instructions to use VertexAIco/amethyst-1-mini with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use VertexAIco/amethyst-1-mini with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="VertexAIco/amethyst-1-mini") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("VertexAIco/amethyst-1-mini") model = AutoModelForMultimodalLM.from_pretrained("VertexAIco/amethyst-1-mini", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - MLX
How to use VertexAIco/amethyst-1-mini with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("VertexAIco/amethyst-1-mini") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- vLLM
How to use VertexAIco/amethyst-1-mini with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "VertexAIco/amethyst-1-mini" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VertexAIco/amethyst-1-mini", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/VertexAIco/amethyst-1-mini
- SGLang
How to use VertexAIco/amethyst-1-mini with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "VertexAIco/amethyst-1-mini" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VertexAIco/amethyst-1-mini", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "VertexAIco/amethyst-1-mini" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VertexAIco/amethyst-1-mini", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - MLX LM
How to use VertexAIco/amethyst-1-mini with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "VertexAIco/amethyst-1-mini"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "VertexAIco/amethyst-1-mini" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VertexAIco/amethyst-1-mini", "messages": [ {"role": "user", "content": "Hello"} ] }' - Docker Model Runner
How to use VertexAIco/amethyst-1-mini with Docker Model Runner:
docker model run hf.co/VertexAIco/amethyst-1-mini
license: gemma
base_model: google/gemma-3-4b-it
language:
- en
library_name: transformers
pipeline_tag: text-generation
tags:
- gemma
- gemma3
- lora
- mlx
- conversational
- distillation
Amethyst 1 Mini
Amethyst 1 Mini is a general-purpose chat and instruction-following model, fine-tuned from Gemma 3 4B IT using LoRA on a small, high-quality distilled instruction dataset. It is the first model in the Amethyst family — an early, small-scale release focused on validating an end-to-end distillation → fine-tune → evaluate pipeline on consumer hardware.
Model Details
| Developed by | Independent research project |
| Base model | google/gemma-3-4b-it |
| Fine-tuning base checkpoint | mlx-community/gemma-3-4b-it-qat-4bit |
| Architecture | Gemma 3, 4B parameters (dense, decoder-only transformer) |
| Fine-tuning method | LoRA (rank 8, scale 20.0), fused into the base weights and dequantized to fp16 for this release |
| Fine-tuning framework | MLX / mlx-lm, on Apple Silicon |
| Trained modules | q_proj, k_proj, v_proj, o_proj, gate_proj, up_proj, down_proj across 16 layers |
| Language | English |
| License | Gemma Terms of Use |
Training Data
Amethyst 1 Mini was fine-tuned on 1,122 instruction/response pairs (1,082 train / 40 validation), synthetically generated via knowledge distillation from nvidia/nemotron-3-super-120b-a12b (Nemotron-3-Super, a 120B-parameter MoE model, ~12B active) through the OpenRouter API.
The dataset spans a deliberately broad set of general-chat categories to encourage well-rounded conversational ability rather than narrow task performance:
- Explanation & misconceptions
- Reasoning & math reasoning
- Code generation
- Extraction & structured output
- Planning
- Roleplay & creative writing
- Translation
- Sentiment classification
- Brainstorming
Each example was generated from a unique prompt, distilled from the teacher model, then filtered/cleaned before training.
Training Procedure
- Method: Supervised fine-tuning via LoRA (rank 8, dropout 0.0, scale 20.0)
- Optimizer: Adam, learning rate 1e-5 (constant schedule)
- Sequence length: 1024 tokens
- Gradient checkpointing: enabled
- Training steps: 3,246 iterations total, with validation every 200 steps
- Checkpoint selection: the released weights use the iteration 2,600 checkpoint, selected for lowest validation loss (1.513) — later checkpoints began overfitting on this small dataset (validation loss rose to ~1.95 by the final iterations)
This release merges the selected LoRA adapter into the base model and dequantizes the result to fp16, so it can be loaded directly with transformers without any MLX or quantization dependencies.
Intended Use
Amethyst 1 Mini is intended as a lightweight, general-purpose conversational assistant — for experimentation, research into small-scale distillation pipelines, and hobbyist deployment. It is not intended for high-stakes, safety-critical, or production use.
Limitations
- Trained on a small (1,122-example) synthetic dataset — behavior can be inconsistent outside the categories represented in training.
- Inherits the general limitations and knowledge cutoff of its base model, Gemma 3 4B IT.
- Distilled from a single teacher model without human review of every example; synthetic-data artifacts (teacher biases, occasional factual errors) may be present.
- This is an early, first-generation checkpoint in an ongoing series — later Amethyst releases are expected to use larger, cleaner datasets.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "VertexAIco/amethyst-1-mini"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")
messages = [{"role": "user", "content": "Explain how vaccines train the immune system, in simple terms."}]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(inputs, max_new_tokens=512)
print(tokenizer.decode(out[0][inputs.shape[-1]:], skip_special_tokens=True))
Citation
If you reference this model, please cite it as:
@misc{amethyst1mini,
title = {Amethyst 1 Mini},
author = {Independent research project},
year = {2026},
note = {LoRA fine-tune of Gemma 3 4B IT, distilled from Nemotron-3-Super-120B-A12B}
}
This model is built on Gemma and subject to the Gemma Terms of Use.