Text Generation
Transformers
Safetensors
MLX
English
gemma3
image-text-to-text
gemma
lora
conversational
distillation
text-generation-inference
Instructions to use VertexAIco/amethyst-1-mini with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use VertexAIco/amethyst-1-mini with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="VertexAIco/amethyst-1-mini") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("VertexAIco/amethyst-1-mini") model = AutoModelForMultimodalLM.from_pretrained("VertexAIco/amethyst-1-mini", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - MLX
How to use VertexAIco/amethyst-1-mini with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("VertexAIco/amethyst-1-mini") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- vLLM
How to use VertexAIco/amethyst-1-mini with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "VertexAIco/amethyst-1-mini" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VertexAIco/amethyst-1-mini", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/VertexAIco/amethyst-1-mini
- SGLang
How to use VertexAIco/amethyst-1-mini with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "VertexAIco/amethyst-1-mini" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VertexAIco/amethyst-1-mini", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "VertexAIco/amethyst-1-mini" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VertexAIco/amethyst-1-mini", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - MLX LM
How to use VertexAIco/amethyst-1-mini with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "VertexAIco/amethyst-1-mini"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "VertexAIco/amethyst-1-mini" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "VertexAIco/amethyst-1-mini", "messages": [ {"role": "user", "content": "Hello"} ] }' - Docker Model Runner
How to use VertexAIco/amethyst-1-mini with Docker Model Runner:
docker model run hf.co/VertexAIco/amethyst-1-mini
| license: gemma | |
| base_model: google/gemma-3-4b-it | |
| language: | |
| - en | |
| library_name: transformers | |
| pipeline_tag: text-generation | |
| tags: | |
| - gemma | |
| - gemma3 | |
| - lora | |
| - mlx | |
| - conversational | |
| - distillation | |
| # Amethyst 1 Mini | |
| **Amethyst 1 Mini** is a general-purpose chat and instruction-following model, fine-tuned from **Gemma 3 4B IT** using LoRA on a small, high-quality distilled instruction dataset. It is the first model in the Amethyst family β an early, small-scale release focused on validating an end-to-end distillation β fine-tune β evaluate pipeline on consumer hardware. | |
| ## Model Details | |
| | | | | |
| |---|---| | |
| | **Developed by** | Independent research project | | |
| | **Base model** | [google/gemma-3-4b-it](https://huggingface.co/google/gemma-3-4b-it) | | |
| | **Fine-tuning base checkpoint** | [mlx-community/gemma-3-4b-it-qat-4bit](https://huggingface.co/mlx-community/gemma-3-4b-it-qat-4bit) | | |
| | **Architecture** | Gemma 3, 4B parameters (dense, decoder-only transformer) | | |
| | **Fine-tuning method** | LoRA (rank 8, scale 20.0), fused into the base weights and dequantized to fp16 for this release | | |
| | **Fine-tuning framework** | [MLX](https://github.com/ml-explore/mlx) / `mlx-lm`, on Apple Silicon | | |
| | **Trained modules** | `q_proj`, `k_proj`, `v_proj`, `o_proj`, `gate_proj`, `up_proj`, `down_proj` across 16 layers | | |
| | **Language** | English | | |
| | **License** | [Gemma Terms of Use](https://ai.google.dev/gemma/terms) | | |
| ## Training Data | |
| Amethyst 1 Mini was fine-tuned on **1,122 instruction/response pairs** (1,082 train / 40 validation), synthetically generated via knowledge distillation from **`nvidia/nemotron-3-super-120b-a12b`** (Nemotron-3-Super, a 120B-parameter MoE model, ~12B active) through the OpenRouter API. | |
| The dataset spans a deliberately broad set of general-chat categories to encourage well-rounded conversational ability rather than narrow task performance: | |
| - Explanation & misconceptions | |
| - Reasoning & math reasoning | |
| - Code generation | |
| - Extraction & structured output | |
| - Planning | |
| - Roleplay & creative writing | |
| - Translation | |
| - Sentiment classification | |
| - Brainstorming | |
| Each example was generated from a unique prompt, distilled from the teacher model, then filtered/cleaned before training. | |
| ## Training Procedure | |
| - **Method:** Supervised fine-tuning via LoRA (rank 8, dropout 0.0, scale 20.0) | |
| - **Optimizer:** Adam, learning rate 1e-5 (constant schedule) | |
| - **Sequence length:** 1024 tokens | |
| - **Gradient checkpointing:** enabled | |
| - **Training steps:** 3,246 iterations total, with validation every 200 steps | |
| - **Checkpoint selection:** the released weights use the **iteration 2,600 checkpoint**, selected for lowest validation loss (1.513) β later checkpoints began overfitting on this small dataset (validation loss rose to ~1.95 by the final iterations) | |
| This release merges the selected LoRA adapter into the base model and dequantizes the result to fp16, so it can be loaded directly with `transformers` without any MLX or quantization dependencies. | |
| ## Intended Use | |
| Amethyst 1 Mini is intended as a lightweight, general-purpose conversational assistant β for experimentation, research into small-scale distillation pipelines, and hobbyist deployment. It is **not** intended for high-stakes, safety-critical, or production use. | |
| ## Limitations | |
| - Trained on a small (1,122-example) synthetic dataset β behavior can be inconsistent outside the categories represented in training. | |
| - Inherits the general limitations and knowledge cutoff of its base model, Gemma 3 4B IT. | |
| - Distilled from a single teacher model without human review of every example; synthetic-data artifacts (teacher biases, occasional factual errors) may be present. | |
| - This is an early, first-generation checkpoint in an ongoing series β later Amethyst releases are expected to use larger, cleaner datasets. | |
| ## Usage | |
| ```python | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| model_id = "VertexAIco/amethyst-1-mini" | |
| tokenizer = AutoTokenizer.from_pretrained(model_id) | |
| model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto") | |
| messages = [{"role": "user", "content": "Explain how vaccines train the immune system, in simple terms."}] | |
| inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device) | |
| out = model.generate(inputs, max_new_tokens=512) | |
| print(tokenizer.decode(out[0][inputs.shape[-1]:], skip_special_tokens=True)) | |
| ``` | |
| ## Citation | |
| If you reference this model, please cite it as: | |
| ``` | |
| @misc{amethyst1mini, | |
| title = {Amethyst 1 Mini}, | |
| author = {Independent research project}, | |
| year = {2026}, | |
| note = {LoRA fine-tune of Gemma 3 4B IT, distilled from Nemotron-3-Super-120B-A12B} | |
| } | |
| ``` | |
| This model is built on Gemma and subject to the [Gemma Terms of Use](https://ai.google.dev/gemma/terms). | |