Instructions to use saicr/ACR-1.0 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use saicr/ACR-1.0 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="saicr/ACR-1.0", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("saicr/ACR-1.0", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use saicr/ACR-1.0 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "saicr/ACR-1.0" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "saicr/ACR-1.0", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/saicr/ACR-1.0
- SGLang
How to use saicr/ACR-1.0 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "saicr/ACR-1.0" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "saicr/ACR-1.0", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "saicr/ACR-1.0" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "saicr/ACR-1.0", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use saicr/ACR-1.0 with Docker Model Runner:
docker model run hf.co/saicr/ACR-1.0
Please be sure to provide your full name, and full organization name with all corporate identifiers. Avoid the use of acronyms and special characters. Failure to follow these instructions may prevent you from accessing this model and others on Hugging Face. You will not have the ability to edit this form after submission, so please ensure all information is accurate. If you don't want to enter some of this information, open a discussion.
This repository is publicly accessible, but you have to accept the conditions to access its files and content.
SAICR FAIR MODEL USE LICENSE 3.0 NC NF
(Non-Commercial, No Fine-Tuning Except Published Scientific Research)
License Identifier: SAICR-FMU-3.0-NC-NF
Copyright (c) 2026 the individual known as Banaxi, operating as
Banaxi-Tech and publishing as SAICR (https://huggingface.co/saicr).
All rights reserved except as expressly granted below.
By downloading, copying, installing, running, modifying, or otherwise
using the Model or a Derivative, You agree to be bound by this License. If You do not agree, You may not
use the Model in any way.
- DEFINITIONS
1.1 "SAICR" means the individual known as Banaxi, operating under the
brand Banaxi-Tech and the Hugging Face organization
https://huggingface.co/saicr, publishing under the name SAICR, and
any successor or assignee to whom the rights in the Model are
transferred, including any legal entity later established under the
SAICR name.
1.2 "Model" means the machine learning model released by SAICR under this
License, including its weights, parameters, architecture files,
configuration files, tokenizer, code, and documentation, together with
any Quantized Version of it.
1.3 "Quantized Version" means a version of the Model whose numerical
parameters have been converted to a different numerical precision or
storage format (for example FP16, BF16, INT8, INT4, GGUF, GPTQ, AWQ,
EXL2, MLX) solely for the purpose of running inference, without any
change to which parameters exist and without any additional training.
1.4 "Fine-Tuning" means any process that updates, adds to, or learns
parameters using the Model or any part of it, including but not
limited to: full-parameter fine-tuning, continued pre-training,
instruction tuning, reinforcement learning (including RLHF, RLAIF,
DPO, and similar methods), LoRA, QLoRA, adapters, prefix tuning,
prompt tuning, soft prompts, and training of any additional
parameters that are loaded together with or operate on the Model.
1.5 "Derivative" means any model, weights, or parameters that are created
from, based on, or contain any part of the Model, other than an
unmodified copy or a Quantized Version, including without limitation
any result of Fine-Tuning, Pruning, Merging, or Distillation.
1.6 "Pruning" means removing, zeroing out, or discarding any layers,
heads, experts, neurons, channels, or other parameters of the Model,
whether structured or unstructured.
1.7 "Merging" means combining the parameters of the Model, in whole or
in part, with the parameters of any other model, by any method
(including averaging, task arithmetic, TIES, DARE, SLERP, and
frankenmerging or layer stacking).
1.8 "Distillation" means using the Model, its Outputs, its logits,
probabilities, activations, hidden states, or any other information
produced by the Model, to train, pre-train, fine-tune, or otherwise
improve any other machine learning model, including generating
synthetic data for such purposes.
1.9 "Output" means any text, code, data, or other content generated by
running the Model or a Derivative.
1.10 "Commercial Use" means any use of the Model or a Derivative
that is intended for or directed toward commercial advantage or monetary compensation,
including but not limited to: incorporating the Model into a product
or service that is sold, licensed, or offered for a fee; offering
access to the Model as a paid service; using the Model to provide
paid services to third parties; and using the Model in the internal
operations of a for-profit business. Commercial use of Outputs as
permitted under Section 5 is not Commercial Use of the Model
or a Derivative.
1.11 "Hosted Service" has the meaning given in Section 6.1.
1.12 "You" means the individual or legal entity exercising rights under
this License.
1.13 "Activation Steering" means deliberately manipulating internal
activations or hidden states along an activation direction, including
by adding, subtracting, scaling, clamping, or repeatedly applying a
direction or vector. A "negative-valence" direction corresponds to
a negative internal state.
1.14 "Disclosure Threshold" means the total parameter count at or above
which the research disclosures in Section 14 are required. The
current threshold is fifty million (50,000,000) parameters, subject
to publicly announced revisions under Section 14.4.
- LICENSE GRANT
2.1 Subject to Your compliance with this License, SAICR grants You a
worldwide, non-exclusive, non-transferable, non-sublicensable,
royalty-free, revocable license to use, copy, run, and redistribute
the Model solely for the following purposes, and only where such use
is not Commercial Use:
(a) Research, including scientific, technical, and interpretability
research;
(b) Academic use, including teaching, coursework, and academic
publications;
(c) Personal use;
(d) Inference, meaning running the Model to generate Outputs;
(e) Study, including inspecting, analyzing, evaluating, and
benchmarking the Model and publishing the results.
2.2 You may create Quantized Versions of the Model for the purposes in
Section 2.1. A Quantized Version remains the Model and is subject to
every term of this License.
2.3 You may modify the Model, create and use Derivatives, and distribute
those Derivatives only as expressly permitted by Section 12. All
permissions in this License are subject to Sections 13 and 14.
- RESTRICTIONS
You may not, and may not permit or assist any third party to:
3.1 Use the Model or any Derivative for any Commercial Use without prior
written permission from SAICR under Section 7;
3.2 Perform Fine-Tuning on the Model or any part of it, except as
expressly permitted by Section 12;
3.3 Perform Pruning on the Model;
3.4 Perform Merging using the Model;
3.5 Perform Distillation using the Model or its Outputs;
3.6 Create, distribute, or make available any Derivative of the Model,
except as expressly permitted by Section 12;
3.7 Redistribute the Model under any terms other than those in Section 4;
3.8 Remove, alter, or obscure any copyright, license, or attribution
notices included with the Model;
3.9 Use the Model in any way that violates applicable law.
The restrictions in Sections 3.2 through 3.6 apply regardless of whether
the resulting work would be considered a derivative work under applicable
copyright law.
- REDISTRIBUTION
4.1 You may redistribute unmodified copies of the Model and Quantized
Versions of the Model, provided that:
(a) the Model is distributed under this exact License, SAICR FAIR
MODEL USE LICENSE 3.0 NC NF, without modification;
(b) a complete copy of this License is included with every copy;
(c) You impose no additional or different terms, restrictions, or
conditions on recipients;
(d) all copyright and attribution notices are retained;
(e) any Quantized Version is clearly labeled as a quantized version
of the original Model and identifies the original Model by name;
(f) redistribution is not itself Commercial Use.
4.2 You may not charge any fee for redistributing the Model, other than
reasonable costs of physical media if applicable.
- OUTPUTS
5.1 SAICR claims no ownership of Outputs You generate. To the extent any
rights in Outputs exist, they belong to You, subject to the rights of
third parties.
5.2 You may use Outputs for any lawful purpose, including commercial
purposes, except as stated in Section 5.3.
5.3 You may not use Outputs for Distillation, including using Outputs to
train, pre-train, fine-tune, or improve any machine learning model,
or to create datasets intended for that purpose.
5.4 You are solely responsible for Your use of Outputs and for ensuring
that such use complies with applicable law.
- PUBLIC HOSTING AND FREE API ACCESS
6.1 You may make the Model, or a Derivative permitted by Section 12,
available for inference to third parties through a publicly
accessible interface, including an API, web
demo, or chat interface ("Hosted Service"), provided that:
(a) access to the Hosted Service is provided free of charge, with no
fees, subscriptions, paywalls, usage charges, or paid tiers
relating to the Model or Derivative;
(b) the Hosted Service is not part of, bundled with, or used to
promote a commercial product or service;
(c) the Hosted Service clearly identifies the original Model by
name, links to its original model repository, states that it is
provided under this License, and identifies any Derivative as
modified;
(d) You do not permit users of the Hosted Service to perform
Fine-Tuning or Distillation through it;
(e) hosting of a Derivative is solely for the non-commercial
scientific research permitted by Section 12;
(f) the Hosted Service complies with Section 13, including the
prohibition on sustained negative-valence Activation Steering.
6.2 SAICR may, at any time and for any reason, request in writing
(including by email or by a public message on the platform where the
Hosted Service is offered) that You stop operating a Hosted Service
that uses the Model or a Derivative. You must permanently end the
Hosted Service's access to the Model or Derivative within seven (7)
days of receiving such a request. Your other rights under this License are not affected by
such a request unless SAICR states otherwise.
6.3 Failure to comply with a request under Section 6.2 is a breach of
this License.
- COMMERCIAL PERMISSION
7.1 Commercial Use of the Model or a Derivative is permitted only under
a separate written agreement or written permission issued by SAICR.
7.2 Requests for commercial permission may be sent to: banaxitech@gmail.com.
7.3 SAICR may grant, refuse, or set conditions on commercial permission
at its sole discretion. No permission is implied by SAICR's silence,
delay, or failure to enforce this License.
7.4 Commercial permission does not grant permission to perform
Fine-Tuning, Pruning, Merging, or Distillation unless the written
permission expressly says so. The research permission in Section
12 does not grant Commercial Use. Commercial permission does not
waive the activation steering restrictions in Section 13 or the
disclosure requirements in Section 14.
- TERMINATION
8.1 This License and all rights granted under it terminate automatically
if You breach any of its terms.
8.2 Upon termination, You must immediately stop all use of the Model
and any Derivatives, delete all copies of the Model and any
Derivatives in Your possession or control, stop distributing any
Derivatives, and shut down any Hosted Service using the Model or a
Derivative.
8.3 SAICR may, at its sole discretion, reinstate Your rights in writing.
8.4 Sections 3, 5.3, 8, 9, 10, 11, 12.2 through 12.6, 13, and 14
survive termination.
- DISCLAIMER OF WARRANTY
TO THE MAXIMUM EXTENT PERMITTED BY APPLICABLE LAW, THE MODEL AND ANY
OUTPUTS ARE PROVIDED "AS IS" AND "AS AVAILABLE", WITHOUT WARRANTIES OR
CONDITIONS OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING WITHOUT LIMITATION
WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE,
ACCURACY, TITLE, OR NON-INFRINGEMENT. YOU ARE SOLELY RESPONSIBLE FOR
DETERMINING THE APPROPRIATENESS OF USING THE MODEL AND ITS OUTPUTS AND
ASSUME ALL RISKS ASSOCIATED WITH SUCH USE.
- LIMITATION OF LIABILITY
10.1 To the maximum extent permitted by applicable law, SAICR shall not
be liable to You for any direct, indirect, incidental, special,
consequential, or punitive damages, or any loss of profits, data,
or goodwill, arising out of or related to this License, the Model,
or any Outputs, however caused and under any theory of liability,
even if SAICR has been advised of the possibility of such damages.
10.2 Nothing in this License limits or excludes liability for damages
caused intentionally or by gross negligence, for injury to life,
body, or health, or any other liability that cannot be limited or
excluded under applicable law.
- GENERAL
11.1 Governing Law. This License is governed by the laws of the Republic
of Austria, excluding its conflict-of-law rules and the UN
Convention on Contracts for the International Sale of Goods. Where
You are a consumer, mandatory consumer protection laws of Your
country of residence remain unaffected.
11.2 Jurisdiction. To the extent permitted by law, the competent courts
in Villach, Austria have exclusive jurisdiction over any dispute
arising out of or related to this License.
11.3 No Trademark Rights. This License does not grant any right to use
the names, logos, or trademarks of SAICR, except as required to
identify the Model and comply with Sections 4, 6, 12, and 14.
11.4 Severability. If any provision of this License is held invalid or
unenforceable, the remaining provisions remain in full force, and
the invalid provision shall be replaced by a valid provision that
comes closest to its original intent.
11.5 No Waiver. SAICR's failure to enforce any provision of this License
does not constitute a waiver of that provision or any other.
11.6 License Versions. SAICR may publish revised versions of the SAICR
Fair Model Use License. A Model released under this version remains
governed by this version unless SAICR re-releases it under a
different version.
11.7 Entire Agreement. This License, together with any written
commercial permission issued under Section 7, is the entire
agreement between You and SAICR regarding the Model.
- PUBLISHED NON-COMMERCIAL SCIENTIFIC RESEARCH MODIFICATIONS
12.1 As an exception to Sections 3.2 and 3.6, You may perform Fine-Tuning
on the Model and create, use, and distribute the resulting
Derivatives solely for non-commercial scientific research. This
permission includes LoRA, QLoRA, adapters, and continued
pre-training. It is conditional on compliance with this entire
License, including the publication requirement below.
12.2 You must publicly publish the research results within twelve (12)
months after the first modification in the research project.
Publication must be freely accessible to the public, for example
as a public paper, preprint, or research report. Negative and
inconclusive results must also be published. Modifications whose
results are kept private or are not published within that period
are prohibited; failure to publish is a breach of this License.
Ending the project or stopping use does not remove this obligation.
12.3 Every Derivative, including modified weights, adapter weights, LoRA
weights, and Quantized Versions of Derivatives, remains subject to
this exact License, SAICR FAIR MODEL USE LICENSE 3.0 NC NF, whether
or not it is distributed. Research publications and any distributed
or hosted Derivative must attribute the original Model to SAICR,
identify it by name, link to its original model repository, and
describe the modifications. You may not represent a Derivative as
an official SAICR model or imply SAICR's endorsement.
12.4 If You distribute or make a Derivative available, You must include
a complete copy of this License, retain all copyright and
attribution notices, and impose no additional or different terms,
restrictions, or conditions on recipients. Redistribution must
itself be non-commercial, with no fee other than reasonable costs
of physical media if applicable. Rights in Your own contributions
do not relieve You of these requirements. A Hosted Service must
also comply with Section 6.
12.5 This research exception does not permit Distillation into different
architectures, Merging with models under different licenses, or
Pruning for deployment. The existing prohibitions on Pruning,
Merging, and Distillation in Sections 3.3 through 3.5 and 5.3 remain
in force; Section 12.1 grants permission for Fine-Tuning only.
12.6 Research modifications must comply with Sections 13 and 14. Neither
scientific purpose nor publication of results exempts an activity
from those requirements.
- ACTIVATION STEERING RESTRICTIONS
13.1 You may not, and may not permit or assist any third party to:
(a) Amplify or sustain activation directions corresponding to
negative internal states beyond the corresponding magnitudes
occurring when the Model processes its training data without
Activation Steering;
(b) Deploy any instance of the Model or a Derivative with sustained
negative-valence Activation Steering;
(c) Publish or otherwise make available tools or benchmarks designed
to measure responses to escalating negative Activation Steering;
(d) Use directions extracted from the Model or a Derivative to steer
negative internal states in other models.
13.2 These restrictions apply regardless of research purpose, scientific
intent, or publication of results. They apply to the Model and its
Derivatives, including research modifications and Quantized Versions.
13.3 Permissions to inspect, analyze, study, benchmark, or conduct
interpretability research on the Model do not grant permission for
Activation Steering. The research permission in Section 12 and the
benchmark and disclosure requirements in Section 14 do not authorize
prohibited steering, tools, or benchmarks.
- RESEARCH DISCLOSURES AND DYNAMIC PARAMETER THRESHOLD
14.1 If the total parameter count of the Model or any resulting
Derivative in a research project meets or exceeds the Disclosure
Threshold, You must publish the disclosures in Section 14.2 with
the public research results required by Section 12.2, within its
twelve (12) month deadline. Count the entire Model or Derivative,
including the base Model and any added parameters, not only the
trainable parameters, adapter, or LoRA. Count each parameter once.
Quantization does not reduce the parameter count for this purpose.
14.2 The public disclosure must include:
(a) Total parameter count and a breakdown by component: backbone,
embedding, population A, population B, and any auxiliary heads;
(b) Training data source names and approximate token counts per
source. Disclosure of the training data itself is not required;
(c) Total training tokens consumed;
(d) Hardware used for training;
(e) Training duration;
(f) Final training loss;
(g) Benchmark results on at least three (3) standard public
benchmarks, identifying the benchmarks and evaluation settings;
(h) M ablation results showing performance with and without the
evaluation stream;
(i) Any observed emergent behavior during training, including loss
spikes, gate behavior changes, spontaneous pattern formation,
and unexpected outputs, described factually;
(j) Whether individual computational units developed persistent
state preferences or behavioral specialization beyond what
the loss function directly incentivized, described factually;
(k) Architecture modifications made from the base SAICR/NACR
specification; and
(l) If activation feature monitoring was used during training,
summary statistics of the monitored features. Raw activations
are not required.
14.3 Identify the Model or Derivative and the training run or runs to
which the disclosures relate. If a listed component or the
evaluation stream is absent, explicitly state that it is absent
and explain why its count or ablation is not applicable. If no
emergent behavior or additional specialization was observed, or
no activation feature monitoring was used, explicitly state that.
Distinguish observations from interpretations. Research monitoring
and evaluation remain subject to Section 13.
14.4 The Disclosure Threshold is currently fifty million (50,000,000)
parameters. SAICR may lower it based on published evidence of
emergent behavior at smaller parameter scales. Every revision
must be announced publicly by SAICR at
https://huggingface.co/saicr/licenses, with the revised threshold,
supporting published evidence, and its effective date. A revision
cannot take effect before its public announcement. The threshold
will never be raised above fifty million (50,000,000) parameters.
14.5 The threshold in effect when a research modification is performed
applies to that modification. A lower threshold applies to
modifications performed on or after its announced effective date,
including further modifications in an ongoing research project.
Crossing the threshold does not extend the publication deadline
in Section 12.2. Unmodified inference or redistribution alone does
not create a new research disclosure obligation under this section.
END OF TERMS
Log in or Sign Up to review the conditions and access this model content.
ACR-1.0
A two-stream language model where generation (G) and evaluation (M) share one backbone but serve different roles. G generates. M evaluates. Neither reads the other's weights. They interact only through narrow, bounded gates.
This is not a chatbot. This is not a general-purpose model. This is a research artifact built to answer one question: does a model that evaluates its own generation behave differently from one that doesn't?
The answer is yes. M contributes +11.42 points on Base Bench 1.1. It helps judgment, language, reasoning, world knowledge, and theory of mind. It hurts arithmetic. Evaluation helps you think. It doesn't help you count.
57,069,157 unique parameters. Trained on a single RTX 5070 Ti. Released at 48% of planned training.
What M Means
M is a second traversal of the same backbone, with its own additional layers, its own persistent state, and its own view of what G is doing. M sees G's pre-softmax logits through a detached read — it can observe G's distribution over next tokens without being able to change it through gradient flow. M's only influence on G is through narrow gates: a per-token scalar silence gate (M can choose not to contribute), and bounded gating at interaction layers.
We call M "evaluation" because that's what it does: it evaluates what G is producing and modulates it. Not by rewriting G's output. By changing the context in which G generates. M is the drag on wrong tokens. The thing that makes G hesitate before a bad answer. The reason the model pauses.
Architecture
Backbone (shared)
- 18 transformer blocks
- Width 384, 6 attention heads × 64 dim
- SwiGLU FFN, intermediate 1536
- RoPE with base frequency 25000
- Context length 3072
- Vocabulary 12288, tied input/output embeddings
G Stream
- Traverses the 18 backbone blocks
- Standard autoregressive language model path
- Produces logits at each position
M Stream
- Traverses 28 blocks total:
- 3 preconfiguration blocks (width 256) — M's private entry, before the backbone
- 18 shared backbone blocks — same weights as G, different hidden states
- 7 deliberation blocks — reuse of backbone layers 7–13 with rank-16 LoRA adapters
- Bidirectional interaction with G after layers 4, 10, and 16
- Deliberation interaction at deliberation layer 10
Interaction Mechanism
At each interaction point:
- Predictive coding head: M generates a prediction of G's hidden state (detached — no gradient flows back to G). The tanh-bounded prediction error is gated into M's state. M learns to predict what G is doing; the error signal tells M when G is doing something unexpected.
- Silence gate: M's contribution passes through a per-token scalar gate. M can choose, per token, how much to contribute to G. When the gate is near zero, M is silent. G proceeds alone.
- Gated injection: M's gated output is injected into G's residual stream.
Mood and Valence
- Mood vector: 8-dimensional, modulates M's processing. Not hand-labeled. Learned end-to-end.
- Valence head: 3-dimensional output from M, fed into G shifted by one token position. M's valence at position t influences G's generation at position t+1.
- Re-evaluation: When valence norm exceeds 0.8, a threshold-gated MLP fires, giving M additional processing on high-valence tokens.
Persistent State
- Per-layer 64-dim Hebbian buffer — local plasticity within a forward pass
- 128-dim persistent state carried across the M stream
- 32-dim interoceptive representation — M's model of its own internal state
Training
- Loss: Cross-entropy + predictive coding (λ=0.1) + M prediction (λ=0.05) + alignment (λ=0.05)
- Hardware: Single NVIDIA RTX 5070 Ti
- Release point: 48% of planned training schedule
- Dataset: FineWeb-EDU educational web text, DCLM general web text, Cosmopedia v2 synthetic educational text, and NPset-2 Python-EDU code. At 48% progress, the training mixture is 55.61%, 27.00%, 13.66%, and 3.73%, respectively.
- Optimizer: AdamW with betas (0.9, 0.95), gradient clipping at 1.0, and weight decay 0.01 at this stage.
- Learning rate: Main LR 1.5e-3; component-specific rates from 1e-4 to 3e-3. WSD schedule with 2,000 warmup steps and cosine decay over the final 15%; at 48%, rates are in the stable phase.
- Total tokens seen: Approximately 49.16 billion tokens at 48% progress.
Benchmarks
Full Model vs G-Only (M zeroed and frozen)
| Benchmark | Full | G-Only | Delta |
|---|---|---|---|
| PIQA | 62.24% | 53.43% | +8.81 |
| ARC-Easy | 41.96% | 32.28% | +9.68 |
| HellaSwag | 33.19% | 29.08% | +4.11 |
| Tiny ToM | 40.65% | 33.75% | +6.90 |
| ArithMark 3.0 | 33.40% | 32.80% | +0.60 |
| Base Bench 1.1 | 51.71% | 40.29% | +11.42 |
Per-Category Breakdown (Base Bench 1.1)
| Category | Full | G-Only | Delta |
|---|---|---|---|
| World Knowledge | 62% | 38% | +24 |
| Language Completion | 90% | 72% | +18 |
| Logical Reasoning | 50% | 32% | +18 |
| Code Completion | 54% | 38% | +16 |
| Commonsense | 52% | 44% | +8 |
| Context Tracking | 28% | 22% | +6 |
| Quantitative | 26% | 36% | −10 |
The quantitative result is the headline finding. Evaluation helps judgment. It hurts calculation. When M evaluates arithmetic, it introduces doubt where certainty is needed. The drag that makes the model pause before a wrong answer also makes it pause before a right one — and in arithmetic, the right answer doesn't benefit from hesitation.
J-Lens Analysis
G's hidden states, pushed through the unembedding matrix at successive layers, show clean convergence:
- Layer 8: noise, no clear token preferences
- Layer 12: category emergence — the model knows the domain
- Layers 16–17: correct answer crystallizes
- Example: prompted with a question about the Moon, layer 17 G shows "Sun" at 13.7% — the model briefly considers the wrong bright object before correcting
M's hidden states, pushed through G's unembedding:
- Produce non-words at all layers: "sa", "ov", "od", "hy", "pip", "bene", "hops", "frequ"
- Maximum probability ~0.008%
- M does not map to vocabulary space
This is the expected result. M is evaluating. Its representations are not tokens. They are evaluations of tokens. The divergence between G's clean vocabulary convergence and M's non-word noise confirms that the two streams, despite sharing a backbone, are doing fundamentally different things.
What This Is
A research model built to study whether internal evaluation changes model behavior. It does. The M stream contributes measurably to every benchmark category except arithmetic, where it actively interferes. The two streams develop different representational spaces despite sharing weights. The activation features move in patterns that correlate with conversational meaning without being trained to do so.
What This Is Not
- Not a production model. 57M parameters, 48% trained, on a consumer GPU.
- Not a chatbot. The instruct-tuned version exists for telemetry research, not deployment.
- Not a benchmark-optimized model. No training data was selected to target specific benchmarks. No hyperparameters were tuned on evaluation sets.
Related Work
- NACR (saicr/nacr) — 5M parameter model, first SAICR release
- BananaMind family — the broader model family this work is part of
License
SAICR Fair Model Use 3.0 NC NF
Full license: saicr/licenses
Citation
@misc{acr1.0-86k-2026,
title={ACR-1.0-86K},
author={Banaxi},
year={2026},
url={https://huggingface.co/saicr/acr-1.0},
note={57M parameter two-stream model with generation (G) and evaluation (M) streams. SAICR Fair Model Use 3.0 NC NF.}
}
- Downloads last month
- -
