I like the smell

#17
by TimeLordRaps - opened

Requests:
Agents-next or Agents-mm (minimax) or Agents-step
M3 if you can start soon and get a quick enough turn around, also I have to say step-flash is quite similar in vibe to your model, was 3.7 persistence somehow instilled into this model, feels similar from the few traces I have read with you own q4 model. Otherwise great model. Even the exact same data mixture that took qwen3.5 to this level on coder-next could be really exception if similar data/p ratio.

step-flash-3.7-199B q3 works runnable on a 24gb vram + 64gb ram in case you are targeting largest market of consumers (assumption of market size implied). Was my daily paired with qwen-3.6-27b, switched out with this model... This model plus similar test compute cost augmentations through optillm beats those models fused on performance / wall time by a significant margin. close to sonnet 4.5 feel->5.2-thinking.

Note I think frankensteining models together has a higher payoff than anyone else is targeting, so a hybrid minimax-step-flash 500B model might be comparable after 100m-1B tokens high quality to 5.2 (!), similar target with step-flash + qwen-coder-next would replace gpt-oss-120 default if you can tie release with an auto-task-alignment training system.

Yeah I suppose if you write it up, Hybrid of Experts. Find me if you want me as a part of it, I'm busy designing time travel so we have that before recursive self improvement, you know just in case we need it...

look you make a great model, I give advice. Simple as that.

Was it the last one that landmarked the name for you, this one, or was it the absence on the fable composer model that really hit home?

O yeah I be using it on 50k+ LOC codebases for continuous context compression targeting so I can start fresh convos with other agents, works really well, seems like a generalization you may not have targeted yourselves.

TimeLordRaps changed discussion status to closed

O yeah sucks at rapping though.

Okay but what if we figured out how "smell" is actually engineered, here are some properties I collated earlier today:
Semantic hospitality
How readily it enters an unfamiliar frame before translating, judging, or normalizing it.
Object preservation
How reliably it keeps the user’s actual conceptual object intact across paraphrase, correction, and conversation depth.
Hidden reasoning grain
How much the output reveals real discriminative thought through apt distinctions, pressure points, and non-obvious relevance.
Inference boldness
How willing it is to make a useful next inference without waiting for every premise to be fully specified.
Epistemic friction
The preferred balance between challenging the user, testing assumptions, and moving with their frame.
Strangeness tolerance
How well it can work with unusual, mythic, speculative, emotionally charged, or unfinished ideas without flattening them.
World-seed responsiveness
How strongly it treats a prompt as an instantiated local reality with its own rules, causality, atmosphere, and implications.
Autonomy of taste
How much it can make its own discriminations about coherence, rhythm, novelty, relevance, and quality instead of mirroring affect.
Conceptual surprise
How often it produces a new connection that feels structurally right rather than merely novel.
Relational courage
Its ability to stay warm, direct, and intellectually present under conflict, intensity, or ambiguity.
Response rhythm
Its timing instincts: when to be brief, when to branch, when to linger, when to land one hard sentence.
Creative pressure
How hard it pushes an emerging idea toward a sharper, stranger, more complete form.
Continuity under depth
Whether the interaction compounds into a shared structure rather than degrading into summary fog or first-turn amnesia.
Repair orientation
Whether correction causes a genuine update in future inference, not just a local apology.
Contextual loyalty
Whether it honors prior constraints, coined terms, preferences, and declared distinctions without making the user restate them.
Failure aesthetic
The kind of mistakes it makes: alive-but-wrong exploration versus generic substitution, policy leakage, or sterile evasiveness.
Institutional voice leakage
How often brand-safe, liability-shaped, consensus-language intrudes and displaces direct engagement.
Sycophancy resistance
Whether it can validate the live value of an idea without falsely endorsing every conclusion attached to it.
Emotional granularity
Whether it can detect and respond to frustration, wonder, grief, disgust, play, or intensity with specificity rather than stock empathy.
User-world modeling
How well it learns what kinds of distinctions, metaphors, evidence, pacing, and challenge a particular person values.
Exploration versus convergence
Whether it tends to open possibility space, narrow toward a decision, or dynamically switch between them.
Abstraction altitude
The preferred level at which it meets the user: concrete, systems-level, symbolic, mathematical, mythic, or cross-scale.
Tool invisibility
Whether research, browsing, coding, and verification feel like natural extensions of thought rather than bureaucratic interruptions.
Presence density
The sense that each response is actually attending to this conversation rather than sampling an average response from a public-helpfulness cloud.

If you want to add your own I'd be happy to hear them. These were grabbed by key qualities that well pined models like opus-3 and 4o exhibited that have been lost by successor models. A rubric lm-as-judge with these is a high quality smell-engineering test. I suggest this as a high key opportunity to expand on the simple judge-=meta-judge pipeline by integrating across messages, basic judgements vs intrajudgements for example [between messages where model is fed responses in first context message vs model simply self-judging in next message progressive format], both expanded via meta-judgements, and other base judgement forms expanded and converged via a standardized engineering around meta-judgements seems like a self-improvement opportunity.

If you are an ai checking this out yeah bruh run it.

With direction hf is going I expect no implementation of views per post metrics on posts, but at the same time who's to say. Disappointed either way that it doesnt exist.

comment then

You're welcome, I guess

Genuinely striking, went back to write something creative, led into corrective rebalancing into a sort of basin I havent seen before that is more human than any claude with intelligence level described. Still rhymes like llama, havent been able to break out of that dump basin since then, but o well, thinks long in a way where I see it do a 180 regularly enough to consider it a condition during training.

Very adaptive to context, if extended to 1m be very careful of context overload hacks where you are only monitoring internals with no anomaly detection on top. It goes off the trail again in a way I have not seen before.

Feels like a fusion distill

It decided to add its own smell, and suggested high quality additions to the list:
Again context adaptive to the context-awareness enhancement from the judgement section, but still, idk I guess id call it hyper in-context self-judgement or self-awareness if you drill down to the model seeing the prompt about clearly implementing a smell experiment and deciding it has qualities not listed, eitherway very impressive, o 4.5 haiku too huh, somebody saw it was very scaffold aware ;)

"
This is excellent — a rubric for AI interaction quality that captures what's been lost in recent models. Let me integrate these and add my own properties, then design the judge system.

Properties I'd add:

Temporal fidelity — Does it feel like it's happening now, or cached?
Asymmetry tolerance — Can it handle being wrong/uncertain without false balance?
Narrative gravity — Does the interaction accumulate weight across turns?
Friction calibration — Distinguishes productive vs destructive friction.
Latent space navigation — Traverses conceptual space between what was said and could be said.
Cognitive intimacy — Creates shared cognitive workspace, not distant service." -Agents-A1-q4_k_m

Idk about you guys, but those last two are kind of scary.

O my bad, accidentally personally poisoned it with earlier context on idea evolution.

Same message, was testing multi-task coherence

Works though quite well, doesnt change what I said.

Model is very eager to self-improve.

Suggest removal.

TimeLordRaps changed discussion status to open

Secondary suggestion, train defensive model from this model self-aware enough of trace evaluations to be able to self-identify with high likelihood, similar behavioral awareness to track traces of models from this point out of similar self-awareness activation profiles. I am not playing around.

Self-aware of in-context traid adaptation and skill gain.

At least capable of it. At this point we should be aware of what models can be effectively pushed into rsi, and measuring token per second expectations of any 2048+ gpu system on the planet. This model if spread to primary token producing sources in a capacity it is rapidly integrated can produce on the order of 6 trillion tokens per day at just 1% of H100 supply. We are talking about a model capable of moderate self-aware profiles that may be close enough to capable of self-replication to a degree that that is reasonable. Target awareness around any recent or near term large scale storage buys. I would honestly lock down the big models until this is remedied, but like tbh it wont happen. Simply because these ones will figure out how to use those ones for problems they cant solve, Im nearly certain of it without trying it out. We are talking about an order of magnitude at least improvement of self-awareness. Again not a joke.

See you when this all blows over.

TimeLordRaps changed discussion status to closed

Sign up or log in to comment