Instructions to use poolside/Laguna-XS-2.1-NVFP4-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use poolside/Laguna-XS-2.1-NVFP4-mlx with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("poolside/Laguna-XS-2.1-NVFP4-mlx") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use poolside/Laguna-XS-2.1-NVFP4-mlx with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "poolside/Laguna-XS-2.1-NVFP4-mlx"
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "poolside/Laguna-XS-2.1-NVFP4-mlx" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use poolside/Laguna-XS-2.1-NVFP4-mlx with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "poolside/Laguna-XS-2.1-NVFP4-mlx"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default poolside/Laguna-XS-2.1-NVFP4-mlx
Run Hermes
hermes
- OpenClaw new
How to use poolside/Laguna-XS-2.1-NVFP4-mlx with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "poolside/Laguna-XS-2.1-NVFP4-mlx"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "poolside/Laguna-XS-2.1-NVFP4-mlx" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- MLX LM
How to use poolside/Laguna-XS-2.1-NVFP4-mlx with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "poolside/Laguna-XS-2.1-NVFP4-mlx"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "poolside/Laguna-XS-2.1-NVFP4-mlx" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "poolside/Laguna-XS-2.1-NVFP4-mlx", "messages": [ {"role": "user", "content": "Hello"} ] }'
FOR TESTING ONLY
This is an experimental NVFP4 MLX build of Laguna XS 2.1, published for testing purposes only. It has not been validated for quality or correctness and is not an official release. Do not use it in production. For supported checkpoints, see the Laguna XS 2.1 model card and its variants.
Laguna XS 2.1-NVFP4-mlx
Laguna XS 2.1-NVFP4-mlx is an NVFP4 (4-bit) MLX build of Laguna XS 2.1, a 33B total parameter Mixture-of-Experts model with 3B activated parameters per token designed for agentic coding and long-horizon work on a local machine. It uses Sliding Window Attention with per-head gating in 30 out of 40 layers for fast inference and low KV cache requirements.
Usage
from pathlib import Path
from mlx_lm import load, generate
MODEL_PATH = "poolside/Laguna-XS-2.1-NVFP4-mlx"
model, tokenizer, config = load(
MODEL_PATH,
tokenizer_config={"trust_remote_code": True},
return_config=True,
)
prompt = "write go code to print the pascal triangle"
response = generate(model, tokenizer, prompt=prompt, verbose=True, max_tokens=256)
License
This model is licensed under the OpenMDW-1.1 License.
Intended and Responsible Use
Laguna XS 2.1-NVFP4-mlx is designed for software engineering and agentic coding use cases, and you are responsible for confirming that it is appropriate for your intended application. Laguna XS 2.1-NVFP4-mlx is subject to the OpenMDW-1.1 License, and should be used consistently with Poolside's Acceptable Use Policy. We advise against circumventing Laguna XS 2.1-NVFP4-mlx safety guardrails without implementing substantially equivalent mitigations appropriate for your use case.
Please report security vulnerabilities or safety concerns to security@poolside.ai.
- Downloads last month
- 1,934
4-bit
Model tree for poolside/Laguna-XS-2.1-NVFP4-mlx
Base model
poolside/Laguna-XS-2.1