Welcome Inkling by Thinking Machines


- +3
One thing worth adding for an honest discussion
It wasn't all smooth. On one agentic run, the model pulled data through an MCP tool in the wrong format, and instead of flagging that something was off, it confidently presented the incorrect data as fact. That's a classic hallucination, and in an agentic setup it's risky because the wrong output feeds straight into the next step.
Interestingly, when I ran a similar flow with Gemma 4 E4B, I didn't hit that hallucination. It handled the data grab without the same issue.
So I'm genuinely torn on the takeaway. Part of me thinks this is a prompting problem on my end. A tighter, more explicit prompt with clearer constraints on the expected data format might have prevented it. But part of me wonders how much a smaller model should be expected to self correct when a tool returns something malformed.
Curious what others think:
When an MCP tool returns bad or wrongly formatted data, whose job is it to catch it, the model, the prompt, or the tool layer? Have you seen smaller agentic models confidently state wrong tool output as fact?
What prompting patterns do you use to force a model to validate tool responses before trusting them?
Would love to hear how others are handling this.
I was testing gemma4 on M2 16 it does good but for my M2 it was slow and I just installed this.
The results are so much better and faster. For my Agentic tasks I did get the speed what i wanted on a simple M2 MacBook Air Laptop without sacrificing on acuracy.
This is quite amazing!