NimbusLM

NimbusLM

1. Introduction

NimbusLM is the newest open-weight release from Toola Labs, trained from a 7B base with an extended curriculum of verified reasoning traces and tool-use dialogues. This release concentrates on long-horizon reasoning: the chat template carries a scratchpad section by default, and function-calling data was trained with balanced positive and negative examples. Across public and internal evaluations the model closes most of the gap to frontier systems, especially on quantitative and coding suites.

Two practical upgrades ship together with the weights. First, hallucination on factual lookups is measurably lower: on an internal 4,000-question probe, the unsupported-answer rate dropped from 6.1% to 2.4%. Second, the tokenizer inherits a compact 92k vocabulary with clean tool-call delimiters, which keeps prompts short when agents interleave system commands.

2. Evaluation Results

Comprehensive Benchmark Results

Benchmark BaselineA BaselineB BaselineA-v2 NimbusLM
Core Reasoning Tasks Math Reasoning 0.502 0.519 0.527 0.550
Logical Reasoning 0.841 0.859 0.869 0.884
Common Sense 0.730 0.741 0.752 0.775
Language Understanding Reading Comprehension 0.697 0.712 0.726 0.745
Question Answering 0.574 0.589 0.601 0.616
Text Classification 0.798 0.812 0.823 0.841
Sentiment Analysis 0.764 0.779 0.790 0.807
Generation Tasks Code Generation 0.689 0.702 0.715 0.734
Creative Writing 0.622 0.640 0.653 0.671
Dialogue Generation 0.578 0.596 0.607 0.625
Summarization 0.728 0.740 0.758 0.775
Specialized Capabilities Translation 0.758 0.772 0.784 0.803
Knowledge Retrieval 0.601 0.617 0.630 0.648
Instruction Following 0.718 0.733 0.744 0.765
Safety Evaluation 0.712 0.698 0.726 0.755

Overall Performance Summary

NimbusLM posts consistent gains over every baseline in the suite, with the largest margins concentrated in the reasoning and generation groups.

3. Chat Website & API Platform

A hosted playground and an OpenAI-compatible endpoint are available on the Toola Labs website; bring your own key or sign up for the free tier there.

4. How to Run Locally

The weights load with the standard transformers API, exactly like the base model the series was built on. Two notes for this release:

  1. The scratchpad section is toggled by a chat template flag, no code patch needed.
  2. Tool-call traces should be passed as plain JSON strings inside the scratchpad.

Temperature

A sampling temperature of 0.7 works well for general chat; drop to 0.3 for extractive tasks.

5. License

Weights and code are released under the Apache-2.0 License, including commercial use and distillation.

6. Contact

Bug reports and reproduction questions go to the public issue tracker or to nimbus@toola.ai.

Downloads last month
23
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support