NexusLM

NexusLM

1. Introduction

NexusLM is the latest generation language model trained with a curriculum-based data mixture and a revised post-training recipe. During development we tracked fifteen benchmark suites after every 100 training steps, and the release candidate was selected as the checkpoint with the strongest overall evaluation accuracy rather than simply the last one. The selected model balances deep reasoning ability with robustness on open-ended generation tasks.

One notable finding from our training runs: pushing past the selected checkpoint did not help. In later checkpoints, several reasoning benchmarks actually regress while a handful of easy classification tasks keep inching up, which is why checkpoint selection based on the full benchmark suite matters for this release.

2. Evaluation Results

Comprehensive Benchmark Results

Benchmark Baseline-A Baseline-B Baseline-A-v2 NexusLM
Core Reasoning Tasks Math Reasoning 0.512 0.534 0.541 0.900
Logical Reasoning 0.778 0.792 0.805 0.900
Common Sense 0.702 0.694 0.717 0.705
Language Understanding Reading Comprehension 0.665 0.681 0.688 0.722
Question Answering 0.576 0.598 0.604 0.850
Text Classification 0.801 0.812 0.823 0.578
Sentiment Analysis 0.775 0.779 0.788 0.777
Generation Tasks Code Generation 0.622 0.641 0.652 0.860
Creative Writing 0.585 0.577 0.599 0.557
Dialogue Generation 0.618 0.633 0.641 0.740
Summarization 0.741 0.752 0.758 0.728
Specialized Capabilities Translation 0.784 0.796 0.799 0.779
Knowledge Retrieval 0.648 0.662 0.671 0.620
Instruction Following 0.729 0.744 0.748 0.850
Safety Evaluation 0.714 0.699 0.721 0.820

Overall Performance Summary

NexusLM shows well-rounded performance across all evaluated benchmark categories, with the strongest gains concentrated in reasoning-heavy and code-related tasks.

3. Chat Website & API Platform

We host an interactive demo and an API endpoint for NexusLM. See our project homepage for access details and rate limits.

4. How to Run Locally

The code for loading and serving NexusLM lives in our main GitHub repository. Check there for installation steps and quantization recipes.

Compared to the previous generation, the usage recommendations for NexusLM changed as follows:

  1. A system prompt is now supported and encouraged.
  2. No special tokens are needed at the start of the output to trigger a particular reasoning mode.

The architecture of NexusLM-Small matches the base model and shares its tokenizer configuration, so it runs with the same code paths.

System Prompt

We recommend the following system prompt with a concrete date.

You are NexusLM, a helpful AI assistant.
Today is {current date}.

For example,

You are NexusLM, a helpful AI assistant.
Today is October 5, 2026, Monday.

Temperature

We recommend setting the temperature parameter $T_{model}$ to 0.7.

Prompts for File Uploading and Web Search

For file uploading, build prompts from the following template, where {file_name}, {file_content} and {question} are arguments.

file_template = \
"""[file name]: {file_name}
[file content begin]
{file_content}
[file content end]
{question}"""

For web-search-augmented generation, the recommended template places the search results ahead of the user question and instructs the model to cite the sources it actually used, as in the example shipped in our repository.

5. License

This code repository is licensed under the Apache License 2.0. Use of the NexusLM model weights is also governed by the Apache License 2.0. Commercial deployment and distillation are both permitted.

6. Contact

Questions are welcome via GitHub issues on our repository, or reach the team at contact@nexuslm.dev.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support