Instructions to use Remek/basal-1.5-mini-FP8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Remek/basal-1.5-mini-FP8 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Remek/basal-1.5-mini-FP8") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Remek/basal-1.5-mini-FP8") model = AutoModelForCausalLM.from_pretrained("Remek/basal-1.5-mini-FP8", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Remek/basal-1.5-mini-FP8 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Remek/basal-1.5-mini-FP8" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Remek/basal-1.5-mini-FP8", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Remek/basal-1.5-mini-FP8
- SGLang
How to use Remek/basal-1.5-mini-FP8 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Remek/basal-1.5-mini-FP8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Remek/basal-1.5-mini-FP8", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Remek/basal-1.5-mini-FP8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Remek/basal-1.5-mini-FP8", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Remek/basal-1.5-mini-FP8 with Docker Model Runner:
docker model run hf.co/Remek/basal-1.5-mini-FP8
basal-1.5-mini-FP8
FP8 (NVIDIA Model Optimizer) checkpoint of basal-1.5-mini (1.5B) for vLLM on NVIDIA
GPUs. Same model, same prompt format and the same typed answers (choice, noul, score) with a probability for every
option; see the main card for the description, benchmarks, limitations and license.
Agreement with the bf16 model (fp32 reference): 0.979; accuracy 0.903 (−0.013; bf16 0.916) (1,000 development decisions, both option orders, vLLM 0.30.0 on one RTX PRO 6000 Blackwell; reference: the bf16 weights in fp32). Speed: 8.8 ms per decision at batch 1, 337.2 decisions/s batched, on one RTX PRO 6000 Blackwell (vLLM 0.30.0). Size: 1.7 GB.
What it is
Post-training quantisation with NVIDIA Model Optimizer (FP8_DEFAULT_CFG), calibrated on
basal-1.5 calibration decisions: FP8 (E4M3) weights and activations in the decoder layers. The token embeddings and the output head stay in bf16, because the decision
is read from the output head's logits. It is a standard ModelOpt Hugging Face export.
Hardware: NVIDIA Hopper (H100, H200), Ada (RTX 40xx, L40S) and Blackwell (B200, B300, RTX 50xx, RTX PRO 6000, DGX Spark).
Run
vLLM brings its own torch, so use a separate environment:
uv pip install "basal[vllm] @ https://github.com/rkinas/basal/archive/refs/tags/v1.5.0.tar.gz"
basal-serve --model Remek/basal-1.5-mini-FP8 --mode vllm --port 8000
--mode vllm starts with a 4,096-token context and refuses longer states; raise it with --max-len (e.g. 32768 for long
documents). The request and response are the same as for the bf16 model. The basal engine's own modes do not load this
checkpoint; for on-the-fly FP8 in the basal engine, use the bf16 repository with --mode fp8. Evidence spans need the
basal engine.
Calibration
CALIBRATION.json carries the temperatures and confidence thresholds of the bf16 model. The thresholds are not validated
for this checkpoint: refit them on your own labelled requests before automating decisions with them.
License
Apache-2.0, like basal-1.5-mini.
- Downloads last month
- 23
Model tree for Remek/basal-1.5-mini-FP8
Base model
speakleash/Bielik-1.5B-v3