Solar-Open2-250B-MLX-4bit

Built with Solar. This is an MLX 4-bit affine quantization of upstage/Solar-Open2-250B, converted for Apple Silicon / MLX workflows.

Details

  • Source model: upstage/Solar-Open2-250B
  • Quantization: 4-bit affine, group size 64
  • Local size: 131G
  • Weight shards: 29
  • Architecture: Solar Open 2 hybrid-attention MoE, 250B total / ~15B active parameters
  • Context: source model advertises 1M-token context; practical MLX context depends on memory and runtime settings

Important runtime notes

Solar Open2 is not yet a stock mlx-lm architecture in many installs. This repo includes solar_open2.py; launch with --trust-remote-code when serving or loading from Hugging Face.

mlx_lm.server \
  --model Vontra/Solar-Open2-250B-MLX-4bit \
  --host 0.0.0.0 \
  --port 8021 \
  --trust-remote-code \
  --temp 0.2 \
  --top-p 0.9 \
  --max-tokens 32768

You may see a transformers warning that mentions loading model_type=solar_open2 into a blank model type. With the included custom MLX loader this warning is expected; the important check is that the model actually loads.

The tokenizer template uses Solar/Whale-style tool markers such as <|tool_call:start|> and <|tool_arg:start|>. For OpenAI-compatible tool calling, your serving runtime must parse those markers into structured tool_calls. Plain text generation does not need this parser.

Use with MLX

This repo includes a small solar_open2.py MLX loader because upstream mlx-lm does not yet ship native Solar Open 2 support.

pip install -U mlx-lm
from mlx_lm import load, generate

model, tokenizer = load("Vontra/Solar-Open2-250B-MLX-4bit")
prompt = "Write a short Python function that validates an IPv4 CIDR string."
print(generate(model, tokenizer, prompt=prompt, max_tokens=256, verbose=True))

Notes

This is an independent community conversion under the Vontra organization. It is not an official Upstage release.

License

The source model is released under the Upstage Solar License. A copy is included in LICENSE. Please review the upstream model card and license before use or redistribution.

Downloads last month
-
Safetensors
Model size
250B params
Tensor type
BF16
·
U32
·
F32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Vontra/Solar-Open2-250B-MLX-4bit

Quantized
(8)
this model