ShoreCrab-128M

ShoreCrab-128M · A60
Image + question + candidate answers → choice probabilities

Notion · Technical Report

Code · Read the technical report

ShoreCrab-128M has 128.1M parameters. A60 is the version name. The model calculates probabilities for 2–16 answer choices in one pass. It also calculates a probability for “none of these,” shown as None.

The model accepts English and Korean. This non-commercial research release contains the model card, demonstrations, examples, and the A60 Safetensors bundle in bundle/. Use the weight terms in LICENSE.md and the notices. Code uses Apache-2.0; A60 weights are for non-commercial research under separate terms.

I am a student. I run a company alone. I welcome investment for this research and future work.
저는 혼자 일하며 회사를 운영하는 학생입니다. 다양한 연구와 이 연구의 후속 연구를 위해 투자를 환영합니다.
Contact: guritopabaem@gmail.com

Download and run

The weights are available for non-commercial research under LICENSE.md. The inference code uses Apache-2.0. This is a research preview; the abstention release criterion did not pass.

git clone https://github.com/Topabaem05/Crab-cookbook.git
cd Crab-cookbook
python -m pip install -e '.[serve]' huggingface_hub
hf download Haverbex/ShoreCrab-128M --include 'bundle/*' 'LICENSE.md' 'notices/*' --local-dir model
crab-serve --bundle model/bundle --device cpu

In another terminal, run either example from the cookbook directory:

python examples/return_probabilities.py assets/kitchen.png
python examples/select_answer.py assets/kitchen.png

Use A60: probabilities or a selected answer

A60 uses an image, a question, and 2–16 answer choices. Both output modes use the same probabilities, including None. The first mode gives all probabilities. The second mode selects the answer with the highest probability.

If None has the highest probability, the result is None. A60 does not generate new answer text.

English cookbook · Pinned engine installation

Tests cover macOS CPU execution in FP32. llama.cpp uses the custom llama-shorecrab extension. vLLM uses a registered pooling model with LLM.encode. These tests do not cover GPU execution, quantization, or the standard llama-server.

Run the examples from the cookbook directory. Use bundle/ from this package after reviewing LICENSE.md.

Python / HTTP

To get all probabilities, run this example:

from shorecrab import CrabClient
from shorecrab.outputs import probability_view

client = CrabClient("http://127.0.0.1:8090")
choices = ["white", "red", "blue", "black"]
result = client.score("assets/kitchen.png", "What color is the mug?", choices)
print(probability_view(choices, result))

To select the answer with the highest probability, run this example:

from shorecrab import CrabClient

client = CrabClient("http://127.0.0.1:8090")
print(client.choose("assets/kitchen.png", "What color is the mug?",
                    ["white", "red", "blue", "black"]))

If you already have the probabilities, use this function to select an answer:

from shorecrab.outputs import select_answer
print(select_answer(choices, result))

vLLM

Use the CPU versions specified in Native engines. Convert the inference bundle with that guide.

The worker calls LLM.encode(..., pooling_task="plugin"). This interface does not use the chat-completions endpoint.

# Mode 1: all candidate probabilities
OMP_NUM_THREADS=2 python -m shorecrab.native.vllm_cli \
  --model /path/to/a60-vllm --image assets/kitchen.png \
  --question 'What color is the mug?' --choices white red blue black \
  --output-mode probabilities
# Mode 2: the most likely candidate or None
OMP_NUM_THREADS=2 python -m shorecrab.native.vllm_cli \
  --model /path/to/a60-vllm --image assets/kitchen.png \
  --question 'What color is the mug?' --choices white red blue black \
  --output-mode select

llama.cpp

Build the specified llama-shorecrab extension with the Native engines guide. Convert the inference bundle with that guide. Install gguf-py in the PyTorch 2.10 environment from the cookbook. Run the commands in that environment.

The CLI removes temporary input and output files after each call, including failed calls.

# Mode 1: all candidate probabilities
OMP_NUM_THREADS=2 python -m shorecrab.native.llama_cpp_cli \
  --executable build/a60/llama-shorecrab --model /path/to/a60-f32.gguf \
  --tokenizer /path/to/a60-inference-bundle/tokenizer \
  --image assets/kitchen.png --question 'What color is the mug?' \
  --choices white red blue black --output-mode probabilities
# Mode 2: the most likely candidate or None
OMP_NUM_THREADS=2 python -m shorecrab.native.llama_cpp_cli \
  --executable build/a60/llama-shorecrab --model /path/to/a60-f32.gguf \
  --tokenizer /path/to/a60-inference-bundle/tokenizer \
  --image assets/kitchen.png --question 'What color is the mug?' \
  --choices white red blue black --output-mode select

Python prepares the input. The executable runs A60 through llama_shorecrab_eval in libllama. It does not use an external inference service. The standard llama.cpp loader cannot load the A60 architecture.

The probabilities for the answer choices and None add up to one. The selection mode keeps None when None has the highest probability. These outputs are raw model probabilities. They do not give a calibrated confidence level.

Game control

Freedoom: aim and eliminate

The excerpt shows 9 kills in 8 seconds, with no cut. We selected the excerpt after the complete run. The complete run contains 14 kills in 24.3 seconds. The player loses at the end.

A60 calculates visual features from RGB pixels. A separate game head has 28,708 parameters and uses these features. A fixed controller turns, searches, or fires. This game head is separate from the general A60 reader.

Two held-out game seeds give 13 and 14 kills. The original A60 reader gives 3 and 1 kills. A controller that always fires gives 2 and 1 kills. A random controller gives 3 and 1 kills.

The records contain the complete runs and failures: seed 41 · seed 43.

Paddle Catch: a second pixel game

Paddle Catch is an original CLI game. The paddle must catch balls that fall. The controller catches 10/10 and 10/10 balls on the two locked game seeds. A fixed-center controller catches 0/10 and 2/10 balls.

The video contains the complete run for the first seed. We selected this seed before evaluation. A separate head has 16,408 parameters. It reads A60 visual features and calculates scores for eight horizontal regions.

The base weights stay fixed. Game coordinates supply offline labels and report data. The controller does not use these coordinates for perception or decisions during a run.

Both games use raw scores from separate task heads. The scores do not give calibrated confidence levels. Playback follows simulation time, not the time for inference. Both games keep the original A60 base checkpoint.

Relative depth

Actual A60 relative-depth decisions

The original A60 reader gives 8/12 correct answers in one realistic kitchen simulation. The example has six object pairs. Each pair has two orders for the answer choices. The image shows all four incorrect answers.

Scene geometry gives reference distances from the camera to each object center. A60 selects an object from the answer choices. Its output does not contain a dense depth map or metric distances. This small simulation is not a depth benchmark.

Approaching traffic: which object is closest?

A60 uses RGB frames and four answer choices: red car, blue car, traffic cone, and yellow barrier. The base checkpoint stays fixed. 7/24 decisions match the reference distance order. Both orders of the answer choices give 14/48 matches.

The video shows every decision, including errors. Code creates this simulation. A60 gives incorrect selections of the nearest object.

Parking: which obstacle is closer?

The camera moves sideways. This movement changes which obstacle is nearer: the traffic cone or the yellow barrier. 8/16 decisions match the reference. The result is 16/32 for both orders of the answer choices.

The model selects the barrier at every time step. It fails to follow the change in the nearest obstacle.

Both videos show actual K+1 probabilities at 2 Hz and separate reference distances. K is the number of answer choices. The probabilities do not give calibrated confidence levels. The RGB inputs contain no annotations or geometry.

Playback follows simulation time. The renderer supplies the distances in meters. These distances are not model outputs. The results include distance differences below 0.25 m.

These scenes do not establish performance for driving or collision avoidance. English example code and native commands.

Separate native-engine retest

Tests used three new traffic frames, both orders of answer choices, both output modes, and both engines. All 24 real CLI calls passed the comparison with the original reader. 2 tests of invalid input also passed.

The probability error was 0 for vLLM. The maximum error for llama.cpp was 3.70e-6. The specified tolerance was 2e-5. These tests cover macOS CPU execution in FP32.

These results show agreement between implementations, not scene accuracy or speed. The tests use the custom libllama extension and the registered vLLM pooling model.

Method

ShoreCrab candidate-scoring method

A resampler uses the question to compress the image representation. A shared reader uses the visual tokens and question tokens to calculate a score for each answer choice. The answerability head adds “none of these.”

Inputs that fit the 512-pixel canvas use one view. Larger inputs add local tiles. The illustration shows the original reader. The game examples use the separate task heads specified above.

Measured results

Development-set highlights

ShoreCrab development benchmark highlights: 70.0% accuracy and 48.0% both-of-pair accuracy

Same 1,000-question, 500-pair development subset. A60 is a GQA specialist; the comparison models are zero-shot. The gauges use a fixed 0–100% scale. Timing uses a separate recorded run. These measurements do not establish a general VLM ranking.

System Accuracy Both of pair
ShoreCrab-128M 70.0% 48.0%
SmolVLM-256M 47.1% 17.6%
CLM-8B + Qwen3-VL-2B captions 45.6% 14.4%
CLM-8B · text only 39.4% 0.0%
I-A60 measurement Result Scope
GQA-derived development accuracy 70.77% 2,600 questions, seed 42
Blank-image control 44.88% Same development set
Three-seed mean ± sample SD 70.50% ± 0.68 pp Final-stage seed runs
Korean answer choices 67.38% Same images with translated choices
Apple M5 Pro GPU p50 / p95 25.6 / 65.9 ms FP32, batch 1, 60 warm requests
Apple M5 Pro CPU p50 / p95 109.3 / 145.5 ms FP32, batch 1, 60 warm requests

These measurements use development data. The latency includes image decoding and all steps through the decision. Each request starts with empty text caches. Larger images with tiles can take more time.

The original sealed A60 release test has no evaluation result. The shared image comparison below is a separate cohort. We calculated the calibration parameters. However, the 2% risk policy accepts 0 examples and fails the release criterion. These results do not establish correct abstention on new inputs or a general VLM ranking.

Shared image comparison

ShoreCrab-128M versus Qwen3.5-0.8B on the shared 512-question image cohort

Clef-flash has no image benchmark score: its model card reports text benchmarks only (see the text-only comparison below), so it is not part of this image comparison.

Model Mode Strict accuracy Format-normalized accuracy* None recall*
ShoreCrab-128M Candidate probabilities 59.96% 59.96% (307/512) 0.00% (0/128)
Qwen3.5-0.8B Generated MC letter 4.30% 84.18% (431/512) 0.00% (0/128)
SmolVLM-256M Generated MC letter 2.15% 58.79% (301/512) 0.00% (0/128)

*Format normalization is a post-hoc, label-blind sensitivity analysis. The original strict scores remain visible. All 512 images, questions, answer choices, and their order are fixed across models. None recall uses 128 separate answer-absent questions. No model was tuned from these results.

The strict rule accepts only a single option letter or NONE. It rejects 481/512 Qwen outputs and 498/512 SmolVLM outputs, often for forms such as A. no or Answer: B. The low strict scores mainly measure format compliance, not general visual ability.

The supplemental extractor accepts an unambiguous letter, an optional Answer: prefix, letter plus matching choice text, or exact unique choice text. It rejects conflicting or ambiguous answers and extra explanations. The extractor never reads the correct label. It applies the same rule to all retained outputs; no generation was repeated and no row was removed. Qwen has zero remaining format errors; SmolVLM has 19. ShoreCrab's candidate probabilities need no text parser, so its score is unchanged.

This custom GQA-derived task is not the official GQA leaderboard. Unknown upstream pretraining overlap remains a limitation. A60 ran on CPU and the generation baselines on RTX 4090, all in FP32. Generation was greedy with at most 32 new tokens. Native model preprocessing was used. These results compare task accuracy, not speed or calibrated confidence.

Sources: Qwen3.5-0.8B and SmolVLM-256M-Instruct.

Text-only Decision Index comparison

Benchmark A60 Clef-flash (reported) Answered / all requests
BFCL 0.00 98.8 0 / 1,694
ToolRet 0.00 66.4 0 / 685
API-Bank 0.00 93.1 0 / 508
BANKING77 0.00 90.9 0 / 3,080
CLINC150+OOS 0.00 66.8 0 / 5,500
RouterBench 0.00 79.9 0 / 10,000
Home appliance simulator 0.00 97.7 0 / 88
SGD/SGD-X 0.00 34.2 0 / 2,500
ContractNLI 0.00 84.3 0 / 123
ANLI 0.20 59.1 13 / 3,200
BPoMP 20.71 95.4 2,559 / 5,000
Humicroedit 49.20 75.1 2,628 / 2,628
POP909-CL 0.00 1.6 0 / 2,000
cfcolor 0.00 65.8 0 / 5,000
MMLU 16.84 91.8 9,748 / 14,033
GPQA Diamond 3.06 51.0 33 / 196
ARC-Easy 24.70 99.5 2,290 / 2,376
ARC-Challenge 23.21 98.3 1,105 / 1,172
WinoGrande 50.43 97.5 1,267 / 1,267
HellaSwag 7.69 98.6 3,235 / 10,042
GSM8K 0.00 67.3 0 / 2,638
ChessBench 0.00 23.0 0 / 5,000
MuSR 0.00 86.0 0 / 752
SATA-Bench 0.00 36.7 0 / 1,650
BRIGHT 0.00 39.3 0 / 220
Amazon ESCI 0.47 57.4 159 / 5,000
ACOS 0.00 25.9 0 / 1,565
FinEntity 0.00 97.1 0 / 979
VAST 0.00 49.6 0 / 3,006
NLI4CT 0.00 78.6 0 / 5,500
CRUXEval 0.00 86.1 0 / 570
CLadder 0.00 97.7 0 / 5,000
ForecastBench · Brier, lower is better N/A 10.6 0 / 10,139
Habermas Machine 0.00 71.8 0 / 1,676
MMLU-Pro 6.03 65.3 7,013 / 12,032
BBH fixed-option tasks 13.82 68.9 1,671 / 5,507
RAGTruth response-level hallucination 0.00 35.6 0 / 2,700
HoVer claim verification 0.00 61.2 0 / 4,000
When2Call MCQ 0.33 65.6 109 / 3,652
New Yorker caption matching 0.00 66.1 0 / 528
PhishNChips phishing decisions 0.00 75.0 0 / 2,000

Scores use each benchmark’s native metric ×100. Zero supported requests indicate an input-capacity failure. Different metrics are not averaged.

The comparison fails the relative ±20% criterion. Forty defined comparisons are outside the band. ForecastBench Brier has no defined result because A60 cannot answer any request in that benchmark.

A60 answered 31,830 / 145,206 requests (21.92%). The other 113,376 requests were outside its input limits. There are no runtime errors or remaining requests.

28 benchmarks have zero supported requests. Their scores use the full number of requests as the denominator. Their zero scores show input capacity failures. They do not mean that the model processed every request and gave incorrect answers.

The original A60 checkpoint receives a fixed black image for this text-only evaluation. The input keeps the complete state and all answer choices. The limits are 96 question tokens, 48 tokens per answer choice, and 2–16 answer choices.

The evaluation does not shorten inputs or remove answer choices. We did not adjust the model from these benchmark results. A60 processed every request in Humicroedit and WinoGrande.

Humicroedit results are 49.20 vs 75.1. WinoGrande results are 50.43 vs 97.5. Each pair shows our A60 measurement first and the published Clef-flash result second. The values use metric ×100.

We did not rerun Clef-flash. The comparison uses its published model-card scores. A60 uses the specified public Decision Index 0.2.1 kit. We verified source hashes and fixed the groups, exclusions, and score calculations.

The exact internal Clef corpus hash is not public. We cannot establish identical private execution conditions. These results do not measure image understanding or a general VLM ranking. The comparison does not average different metrics.

We do not compare A60 CPU FP32 latency with Clef GPU latency. We cannot calculate paired confidence intervals without the individual Clef predictions.

Read the technical report.

Everyday and robotics examples

The kitchen replay shows 3/3 correct answers to fixed questions on one camera frame. The Panda simulation shows 0/3 correct answers, including failed abstention. Code controls the robot motion and grasp with kinematics. This example does not show success on a physical robot.

Both examples use the original A60 reader. The records include the failures.

Evidence and limitations

  • Base checkpoint SHA256: 427fbf2559e470543e19578b3f9a831c6cfed62782cbf91b4e04cad928aee3bb.
  • Game heads are separate adapters. The game results use task heads with the compatible A60 base. They do not show the base reader alone.
  • Small game and scene samples show behavior. They do not establish general performance. Relative-depth references use object-center distances from simulations.
  • The Korean translations have no review by a native speaker. This repository does not distribute GQA images.
  • Source and asset notices · technical report.

License

The inference code, example code, and original traffic simulation media use Apache-2.0. This license does not change the terms for model weights or earlier third-party assets. A60 weight terms stay separate. This repository does not establish permission to distribute the weights under Apache-2.0 alone.

License scope · Apache-2.0 text · Cookbook notices.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using Haverbex/ShoreCrab-128M 1