Dipankar Sarkar PRO
AI & ML interests
Recent Activity
Organizations
naghamo/mist-vlm-judges
It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them
PatchHolmes: Agentic Patch Retrieval via Listwise Selection
CheatBench: Measuring Reward Gaming in AI Agents
See it, Say it, Sorted: Mechanistic Diagnosis and Parameter-Space Mitigation of Emergent Misalignment in LLMs
SkillSeek: Revisiting Agent Skill Retrieval at Marketplace Scale
Dating the Model: Hidden Dates in System Prompts Affect LLM Evaluation
WUSH-KV: KV Cache Quantization with Data-Adaptive Transforms
Thinking Outside the Box: Can Language Models Rely on External Guidance Selectively?
PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents
A2Z GameSpec-Bench: How Faithfully Can Coding Agents Generate Games from Game Design Specifications?
OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software
A small demo built on imajev-4b: a closet stylist ๐
Request you to star it here so we can make it better - https://github.com/mohit67890/imajev.
Tap a piece and it reads the photo (red 75%, checked 99%), then the app picks bottoms, shoes and a bag from your own closet in the colours you like. Change your colours and the outfit changes.
Under the hood it's one request with one photo and 4 typed questions. Every option gets a probability, so the app applies its rules (one pattern per outfit) and ranks what's left. About 1.1 s per outfit on a Mac (MLX). Every % in the video is the model's real answer.
Also, thanks to @zenmagnets for the FP8 version for Blackwell GPUs: same calibration, 0.1 pt less accuracy on all 23,900 DecisionBench rows, ~1.6ร faster ๐
zenmagnets/Imajev-4B-FP8-SM120
๐ง Weights: mohit67890/imajev-4b
๐ Demo: mohit67890/imajev
๐ Leaderboard: https://benchmarkheaven.com/image-jev-bench
report.
OpenAI scrapped GPT-6.1 Astra (per the WSJ: more deception, acting without
the user's permission). Google gave Gemini 4 Argon only to vetted cyber
defenders, per its own post. Anthropic's August risk report describes a
staged internal rollout of "Model 2" and says, about internal use:
"we do not have strict technical safeguards on internal deployment"
For most models. Early snapshots of future public releases included.
What I did about the same gap in our own stack, with receipts:
- A public incident dataset, 97 entries, each with a source. Three were added
today from that report (sections 2.18 and 2.8; the 2.8 items are quoted
there from the system card, which I did not open).
- A gate that refuses to run an action without a signed verdict. The verdict
comes from a separate service: its own unix user, a key the agent's process
cannot read, RS256, bound to the exact command, 60 seconds, single use. STOP
never runs. CONFIRM needs a human.
- Checked live today: rm -rf came back STOP. A replayed verdict was refused.
A verdict issued for one command was refused for another.
What is not done: it is not wired into any agent yet. And with the current
seed table the service never issues PASS, because anything it has no data on
lands exactly on the CONFIRM threshold. Conservative on purpose, but it means
no action is auto-approved today.
Astra's reasons are secondhand (WSJ via a third-party writeup). Argon's
claims are Google's own, not independently measured.
Day zero is not when an exploit finds the bug. It is when the bug is already
inside your own action. Capability is shipping faster than the thing that
catches it.
Dataset: huggingface.co/datasets/SoulInPsyAbstract/sipa-os-governance
pollix/stuntd-support-triage is three small heads on the Laya encoder that triage a support ticket in one request: category, urgency and needs_human. About 50 MB each, all three answers come back at a p50 of 71ms through the daemon.
On 1,000 tickets they never saw, each head answers on its own when it's sure: category 99.9%, needs_human 92%, urgency 76%. A ticket only skips the big model when all three are sure, that's 72.7% of them, and all three are right on 97.1% of those.
It's the support demo from the repo, so the tickets are generated and the teacher is a rule. The point is to show what a head looks like and how fast it is, then you train the same thing on your own traffic with your own LLM as the teacher.
hf download pollix/stuntd-support-triage --local-dir support-headsModel: pollix/stuntd-support-triage
Everything in one place: pollix/stuntd-6abe0a33303828e10c72ab41
Code: https://github.com/bladedevoff/stuntd
The 1,000 eval tickets are unseen states, not unseen tickets.
I ran examples/support/generate.py at 51186ae1 with client.py's own draw (seed 10001, training states skipped).
Every label is a function of three things: the template, the impact sentence and the tail sentence. Channel, plan, product, when, amount and number are slot fill no rule reads.
997 of the 1,000 eval tickets have a template + impact + tail that is already in the 3,000 training rows. Median 17 training copies each. Across train and eval, no such combination ever carries two different labels, on any of the three questions.
So a lookup table on those three fields answers 997/1000, all correct.
That makes 97.1% a measure of something else: how often a field no rule reads moves a head. Your category meta.json points the same way. Two of its five confident errors are the "What is included in the {amount} EUR tier?" template, labelled billing, answered other.
One teacher detail before anyone copies the rule: 21 of the 368 needs_human=true eval tickets are stuck payouts. They trip the "money back" cue through "our sellers are waiting for their money back". Nobody asked for a refund.
Two cheap tests would say what happens on someone's own traffic:
- hold out one template per category, train on the rest
- flip only channel and plan on each eval ticket, count changed answers
Which one do you expect the 72.7% to survive?