NovaReasoner

NovaReasoner

1. Introduction

NovaReasoner is the newest release from our lab, rebuilt on a larger pre-training corpus and a refined post-training pipeline. This release pushes deeper on multi-step reasoning, code understanding, and factual grounding while keeping the footprint of the previous generation. Across our internal benchmark suite the model closes most of the gap to much larger open-weight systems.

The jump is most visible on hard reasoning: on the AIME 2025 set NovaReasoner moves from 68% to 84.5% pass@1, and the median chain-of-thought length grows from 11K to 21K tokens per problem, evidence that the model "thinks longer" before committing to an answer rather than pattern-matching.

We also tightened hallucination guardrails and expanded reliable tool/function-calling support in this version.

2. Evaluation Results

Comprehensive Benchmark Results

Benchmark Baseline-A Baseline-B Baseline-A-v2 NovaReasoner
Core Reasoning Tasks Math Reasoning 0.495 0.518 0.503 0.537
Logical Reasoning 0.772 0.785 0.796 0.801
Common Sense 0.698 0.684 0.707 0.727
Language Understanding Reading Comprehension 0.652 0.666 0.671 0.689
Question Answering 0.563 0.580 0.582 0.600
Text Classification 0.786 0.794 0.803 0.820
Sentiment Analysis 0.760 0.764 0.773 0.786
Generation Tasks Code Generation 0.598 0.614 0.623 0.636
Creative Writing 0.571 0.562 0.584 0.595
Dialogue Generation 0.604 0.618 0.622 0.634
Summarization 0.728 0.738 0.743 0.759
Specialized Capabilities Translation 0.764 0.781 0.783 0.800
Knowledge Retrieval 0.634 0.651 0.653 0.670
Instruction Following 0.715 0.731 0.733 0.750
Safety Evaluation 0.700 0.683 0.707 0.732

Overall Performance Summary

NovaReasoner posts consistent gains over both Baseline-A and Baseline-B, with the largest margins in the reasoning and generation groups.

3. Chat Website & API Platform

A hosted chat surface and REST API for NovaReasoner are available on our official site.

4. How to Run Locally

See our source repository for instructions on serving NovaReasoner locally.

Relative to the prior release, usage notes changed as follows:

  1. A system prompt is now supported.
  2. You no longer need to prepend a special token to force the model into a reasoning mode.

NovaReasoner-Small shares the architecture of its base model and reuses the NovaReasoner tokenizer; run it exactly like the base model.

System Prompt

We suggest the following system prompt, with the current date substituted in.

You are NovaReasoner, a helpful AI assistant.
Today is {current date}.

For example,

You are NovaReasoner, a helpful AI assistant.
Today is June 2, 2026, Monday.

Temperature

We recommend setting the temperature parameter $T_{model}$ to 0.7.

Prompts for File Uploading and Web Search

For file uploading, follow the template below, where {file_name}, {file_content} and {question} are placeholders.

file_template = \
"""[file name]: {file_name}
[file content begin]
{file_content}
[file content end]
{question}"""

For web-search-augmented generation, use the prompt template below, where {search_results}, {cur_date}, and {question} are placeholders.

search_answer_en_template = \
'''# The following contents are the search results related to the user's message:
{search_results}
In the search results I provide to you, each result is formatted as [webpage X begin]...[webpage X end], where X represents the numerical index of each article. Please cite the context at the end of the relevant sentence when appropriate. Use the citation format [citation:X] in the corresponding part of your answer. If a sentence is derived from multiple contexts, list all relevant citation numbers, such as [citation:3][citation:5]. Be sure not to cluster all citations at the end; instead, include them in the corresponding parts of the answer.
When responding, please keep the following points in mind:
- Today is {cur_date}.
- Not all content in the search results is closely related to the user's question. You need to evaluate and filter the search results based on the question.
- For listing-type questions (e.g., listing all flight information), try to limit the answer to 10 key points and inform the user that they can refer to the search sources for complete information. Prioritize providing the most complete and relevant items in the list. Avoid mentioning content not provided in the search results unless necessary.
- For creative tasks (e.g., writing an essay), ensure that references are cited within the body of the text, such as [citation:3][citation:5], rather than only at the end of the text. You need to interpret and summarize the user's requirements, choose an appropriate format, fully utilize the search results, extract key information, and generate an answer that is insightful, creative, and professional. Extend the length of your response as much as possible, addressing each point in detail and from multiple perspectives, ensuring the content is rich and thorough.
- If the response is lengthy, structure it well and summarize it in paragraphs. If a point-by-point format is needed, try to limit it to 5 points and merge related content.
- For objective Q&A, if the answer is very brief, you may add one or two related sentences to enrich the content.
- Choose an appropriate and visually appealing format for your response based on the user's requirements and the content of the answer, ensuring strong readability.
- Your answer should synthesize information from multiple relevant webpages and avoid repeatedly citing the same webpage.
- Unless the user requests otherwise, your response should be in the same language as the user's question.
# The user's message is:
{question}'''

5. License

This code repository is licensed under the Apache-2.0 License. Use of the NovaReasoner model weights is also governed by the Apache-2.0 License. The model family permits commercial use and distillation.

6. Contact

Open an issue on our GitHub repository, or reach us at contact@novareasoner.ai.


Downloads last month
12
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support