Instructions to use CentIo/Cent1-1B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use CentIo/Cent1-1B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="CentIo/Cent1-1B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("CentIo/Cent1-1B") model = AutoModelForCausalLM.from_pretrained("CentIo/Cent1-1B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use CentIo/Cent1-1B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "CentIo/Cent1-1B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "CentIo/Cent1-1B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/CentIo/Cent1-1B
- SGLang
How to use CentIo/Cent1-1B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "CentIo/Cent1-1B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "CentIo/Cent1-1B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "CentIo/Cent1-1B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "CentIo/Cent1-1B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use CentIo/Cent1-1B with Docker Model Runner:
docker model run hf.co/CentIo/Cent1-1B
Cent1-1B: Financial Tool-Calling Model for Agentic Workflows
Model Description
Cent1-1B is a compact language model purpose-built for financial tool calling and numeric reasoning over financial statements. Given a user request and a set of available tool schemas, it selects the correct function and emits a well-formed call with correctly typed arguments — or, when no tool is appropriate, answers conversationally instead of forcing a call.
It targets the analysis workflows that dominate finance work: pulling figures out of statements, computing ratios and changes, and routing requests to the right backend function.
At a ~3.4GB fp16 footprint it runs on a single consumer GPU or a free-tier T4.
- Developed by: CentIo
- Model type: Causal language model for tool calling and financial reasoning
- Architecture: Qwen3 dense transformer, 1.72B parameters (embeddings tied), 28 layers, hidden size 2048, 40,960-token context window
- Fine-tuned from model:
Qwen/Qwen3-1.7B - License: Apache 2.0
Get Started
Transformers
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_id = "CentIo/Cent1-1B"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id, torch_dtype=torch.float16, device_map="auto"
)
tools = [{
"name": "calculate_net_margin",
"description": "Compute net margin from revenue and net income.",
"parameters": {
"type": "object",
"properties": {
"revenue": {"type": "number"},
"net_income": {"type": "number"},
},
"required": ["revenue", "net_income"],
},
}]
messages = [
{"role": "system", "content": f"You are a precise financial tool caller. Available tools:\n{tools}"},
{"role": "user", "content": "What's the net margin on 480 revenue and 62 net income?"},
]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
enable_thinking=False, # the mode this model was tuned in -- see the note below
return_tensors="pt",
return_dict=True,
).to(model.device)
out = model.generate(**inputs, max_new_tokens=256, do_sample=False)
print(tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
Pass
enable_thinking=False. Qwen3's template enables its native thinking channel by default, and Cent1-1B was fine-tuned with that channel off (the model emits the tool call directly, without a precedingthinkingblock). Omitting the flag puts the model on a prompt format it was not tuned for, which degrades both call formatting and routing accuracy.
The model emits calls as:
<tool_call>{"name": "calculate_net_margin", "arguments": {"revenue": 480, "net_income": 62}}</tool_call>
vLLM / SGLang
vllm serve CentIo/Cent1-1B --dtype float16 --max-model-len 8192
Results
Cent1-1B was evaluated on BFCL v3 (simple) for function calling, an MMLU finance/business/economics subset (1,461 items across accounting, econometrics, macro- and microeconomics, business ethics, management and marketing) for domain knowledge, and FinQA for numeric reasoning over real filings. Every model was run on the identical suite with greedy decoding, so the rows are directly comparable.
Table 1: Function calling, financial knowledge and financial numeric reasoning.
| Model | Params | BFCL v3 simple ↑ | MMLU finance ↑ | FinQA ↑ |
|---|---|---|---|---|
| Cent1-1B (ours) | 1.7B | 0.7225 | 0.5756 | 0.0220 |
| Qwen3-1.7B (base) | 1.7B | 0.6750 | 0.5476 | 0.0549 |
| Qwen3-0.6B | 0.6B | 0.5475 | 0.4449 | 0.1209 |
| Qwen3-4B | 4B | 0.7375 | 0.7276 | 0.3407 |
Table 2: Where the function-calling gain comes from. BFCL v3 simple decomposes into choosing the right function and filling its arguments correctly — the two fail independently, so they are reported separately.
| Model | Emitted a call | Function name correct | Arguments correct | Exact match |
|---|---|---|---|---|
| Cent1-1B (ours) | 0.9950 | 0.9875 | 0.7250 | 0.7225 |
| Qwen3-1.7B (base) | 1.0000 | 0.9975 | 0.6775 | 0.6750 |
| Qwen3-0.6B | 0.9975 | 0.9950 | 0.5500 | 0.5475 |
| Qwen3-4B | 0.9950 | 0.9925 | 0.7400 | 0.7375 |
Three things to read from this, stated plainly:
- The gain is in argument binding, not tool selection. Cent1-1B fills arguments correctly +4.75 points more often than its base (0.7250 vs 0.6775), while its function-name accuracy is marginally lower (0.9875 vs 0.9975). The base model is already near-ceiling at picking the function; what it gets wrong is the arguments. That is the specific thing this model was tuned for, and it is the specific thing that improved.
- It lands in the same band as a model 2.4× its size. 0.7225 against Qwen3-4B's 0.7375 — a 1.5-point gap for a model that fits in ~3.4 GB of fp16.
- FinQA regresses, and honestly so. 0.0220 vs the base's 0.0549 (2 vs 5 correct out of 91). This is a real weakness, not a rounding artifact: the model produces well-formed, concise numeric answers — the right format — but the values are wrong, and often only slightly wrong. Every model at or below 1.7B is near the floor here; only the 4B clears 0.34. Treat Cent1-1B as a router and formatter, not a calculator, and verify figures before acting on them.
A correction worth stating. An earlier draft of this card predicted the base model would score ~0 on BFCL because "an untuned Qwen3 emits no
<tool_call>tags". Measurement falsified that. When the prompt spells out the call format and supplies the function schemas — as this harness does — the base model emits a parseable call on 100% of items and scores 0.6750. The base's0.0000figure that motivated the prediction came from a different, stricter internal suite that requires the exact fine-tuned output format with no format instruction in the prompt. The two numbers are not in conflict, but only the BFCL one belongs in this table.
Harness note. The BFCL column uses a faithful subset of the official harness — function name plus argument binding against the published acceptable-value lists — not the official leaderboard scorer, which additionally performs AST normalisation and optional live execution. Scores are comparable across the rows above; they are not directly comparable to published BFCL leaderboard numbers. Argument matching is strict: an argument that the ground truth accepts (including as the empty string) still counts as wrong if the model omits the key entirely, so Table 2's argument column is a conservative lower bound.
Reproducibility. All four rows come from a single kernel run on identical inputs with greedy decoding; the suite definitions and scorers are unit-tested on CPU before any GPU is spent (65 checks, including a full-score reconstruction from the published BFCL ground truth).
Limitations
- 1.7B capacity. Strong at single-call routing and short numeric chains; multi-step tool planning and long-horizon agentic workflows are out of scope.
- Schema-bound. Tool selection is conditioned on the schemas supplied in the prompt. Ambiguous or overlapping tool descriptions degrade routing quality.
- Numeric precision is the weak axis — measured, not hedged. FinQA is 0.0220, below the 1.7B base model's 0.0549 and below even Qwen3-0.6B. The model reliably produces a clean, correctly formatted number; the number is frequently wrong. Use it to route and to format, and compute the figures elsewhere.
- Not financial advice. Outputs are generated text, not professional financial guidance.
Citation
@misc{centio2026cent1,
title={Cent1-1B: Financial Tool-Calling Model for Agentic Workflows},
author={CentIo},
year={2026},
url={https://huggingface.co/CentIo/Cent1-1B},
}
@misc{qwen3,
title={Qwen3 Technical Report},
year={2025}
}
- Downloads last month
- 288