Vontra
Clean public model snapshot
c444e5d
|
Raw
History Blame Contribute Delete
2.5 kB
---
language:
- en
- ko
- ja
library_name: mlx
pipeline_tag: text-generation
license: other
license_name: upstage-solar-license
license_link: LICENSE
base_model: upstage/Solar-Open2-250B
tags:
- mlx
- solar
- solar-open2
- moe
- text-generation
- quantized
- 4bit
---
# Solar-Open2-250B-MLX-4bit
Built with Solar. This is an MLX 4-bit affine quantization of [upstage/Solar-Open2-250B](https://huggingface.co/upstage/Solar-Open2-250B), converted for Apple Silicon / MLX workflows.
## Details
- Source model: `upstage/Solar-Open2-250B`
- Quantization: 4-bit affine, group size 64
- Local size: 131G
- Weight shards: 29
- Architecture: Solar Open 2 hybrid-attention MoE, 250B total / ~15B active parameters
- Context: source model advertises 1M-token context; practical MLX context depends on memory and runtime settings
## Important runtime notes
Solar Open2 is not yet a stock `mlx-lm` architecture in many installs. This repo includes `solar_open2.py`; launch with `--trust-remote-code` when serving or loading from Hugging Face.
```bash
mlx_lm.server \
--model Vontra/Solar-Open2-250B-MLX-4bit \
--host 0.0.0.0 \
--port 8021 \
--trust-remote-code \
--temp 0.2 \
--top-p 0.9 \
--max-tokens 32768
```
You may see a `transformers` warning that mentions loading `model_type=solar_open2` into a blank model type. With the included custom MLX loader this warning is expected; the important check is that the model actually loads.
The tokenizer template uses Solar/Whale-style tool markers such as `<|tool_call:start|>` and `<|tool_arg:start|>`. For OpenAI-compatible tool calling, your serving runtime must parse those markers into structured `tool_calls`. Plain text generation does not need this parser.
## Use with MLX
This repo includes a small `solar_open2.py` MLX loader because upstream `mlx-lm` does not yet ship native Solar Open 2 support.
```bash
pip install -U mlx-lm
```
```python
from mlx_lm import load, generate
model, tokenizer = load("Vontra/Solar-Open2-250B-MLX-4bit")
prompt = "Write a short Python function that validates an IPv4 CIDR string."
print(generate(model, tokenizer, prompt=prompt, max_tokens=256, verbose=True))
```
## Notes
This is an independent community conversion under the Vontra organization. It is not an official Upstage release.
## License
The source model is released under the Upstage Solar License. A copy is included in `LICENSE`. Please review the upstream model card and license before use or redistribution.