Cubert Hyperspectral

Cuvis.AI docs Cuvis.AI on GitHub

Foreign-object detection on walnuts: trained Cuvis.AI pipelines

Six ready-to-run Cuvis.AI pipelines that find foreign objects among walnut kernels and shell fragments on a moving belt, recorded with a Cubert Ultris XMR hyperspectral camera. Each pipeline fuses an unsupervised anomaly detector on a near-infrared false-colour projection with a supervised shell segmentation, so the operator sees one mask: shells in one level, foreign objects in another. They are the pipelines shown at the Cubert stand in October 2026, with the gate values of that stand.

What it detects

Foreign objects on a belt of walnut kernels and shell fragments: stones, aluminium, rubber, plastic imitation shells, loose stems, dark specks and other material the detector has not seen in clean product. The anomaly detector was fitted on clean product only, so it has no class list; anything that does not look like kernel, shell or belt in the near-infrared projection scores high. The segmentation part outlines the real shell fragments (two RF-DETR models on an RGB and a CIR projection, fused), so shells and foreign objects come out as two levels of one composite mask (Composite.mask: 85 = shell, 255 = foreign object on top; Composite.scores: 0.5 and 1.0). The single outputs stay selectable: gate.scores (the foreign-object heatmap), FOMask.decisions, ShellMask.decisions, ShellHeatmap.scores.

The numbers below hold for the stand setup they were measured on: two production LED lights at 20 degrees, the camera 33 cm above the belt, 50 to 80 ms integration time, a white reference taken at the camera's own integration time, a slow turntable or belt. Other lighting shifts the scores; recalibrate before use.

Pipelines

Folder Model Frame gate Pixel threshold When to use
walnut_best_refit/ Refit banks, robust mask 1.1512 1.2178 the production stand with its two lights, where the refit banks find more foreign objects with fewer false marks (565 of 575 on the 1 October labels, 0.15 false marks per frame)
walnut_best_original/ Original banks, robust mask 1.35 1.35 setups and lighting other than the October production stand: the original banks hold up better there (330 of 536 on the sessions before October, against 232 for the refit)
walnut_best_refit_cut/ Refit banks, robust mask, pixel cut 1.1512 1.2178 a display where the mark should sit on the object and not spill over the belt: the cut removes 81 to 84 % of the marked area off foreign objects and shells
walnut_best_original_cut/ Original banks, robust mask, pixel cut 1.35 1.35 lighting other than the production stand, with a display where the mark should sit on the object: 880 marked px per frame off foreign objects and shells instead of 5,539
walnut_best_refit_shellaware/ Refit banks, robust mask, shell-aware gate 1.1512 1.2178 real shell fragments next to a foreign object light up (the anomaly score spills across touching objects)
walnut_best_original_shellaware/ Original banks, robust mask, shell-aware gate 1.35 1.35 real shell fragments next to a foreign object light up, under lighting other than the production stand

Each folder holds the pipeline yaml, its fitted weights (.pt, same stem) and a README with the knobs and a quick start. The frame gate and the pixel threshold are the values tuned on the recordings this card describes: set them for your own product before use (the calibration section and the per-pipeline READMEs say how).

Pipeline diagram

Pipeline diagram

Requirements

  • Fitted on Cuvis.AI 0.17.2 (the cuvis_ai_version stamp inside each .pt); running needs Cuvis.AI 0.17 or later with cuvis-ai-core 0.17.4 or later, the floor of the plugins below, installed from the git-tag manifests in plugins/ (copy them into a configs/plugins/ folder next to your pipelines, or point --plugins-dir at plugins/):
  • A CUDA GPU for the pipelines as shipped (a laptop RTX 4070 runs these pipelines at 136 to 189 ms per frame with the torch backend as shipped and at 56 to 64 ms with the TensorRT fp16 build the stand ran, measured on six reference cubes; a Jetson AGX Thor takes 120 to 140 ms with TensorRT; CPU works but is slow).
  • Third-party weights the nodes fetch on first use: SteerViT (steervit_dinov2_base.pth) (MIT, https://huggingface.co/JonaRuthardt/SteerViT), DINOv2 ViT-B/14 (timm vit_base_patch14_dinov2.lvd142m) (Apache-2.0, https://github.com/facebookresearch/dinov2), RoBERTa-large (MIT, https://huggingface.co/roberta-large).
  • The first construction of a pipeline needs network access for the third-party downloads above, or a Hugging Face cache that already holds them; an offline runtime (HF_HUB_OFFLINE=1) fails until they are mirrored under cubert-gmbh.
  • cuvis-ai-dataloader[cu3s] with the Cuvis SDK to read .cu3s recordings; CuvisNEXT feeds the cube directly.

How to use

With the registry (one command per pipeline, all files pinned and sha256-checked):

uv run download-model download walnut_best_refit

Then run it on a recording from the snapshot folder the command prints (the yaml refers to the shared weight files under weights/ relative to that folder):

cd <snapshot folder>
uv run restore-pipeline --pipeline-path walnut_best_refit/walnut_best_refit.yaml \
  --weights-path walnut_best_refit/walnut_best_refit.pt --plugins-dir plugins \
  --data-module cu3s --data-arg cu3s_file_path=<recording>.cu3s --data-arg processing_mode=Reflectance

Without Cuvis.AI's registry: python fetch.py in a checkout of this repo downloads every file listed in manifest.json and verifies it; hf download cubert-gmbh/Industrial_Foreign_Object_Detection_Walnuts --local-dir ./Industrial_Foreign_Object_Detection_Walnuts does the same without the check.

In CuvisNEXT: pipeline = the yaml, weights = the .pt with the same name, started from the snapshot folder so relative paths resolve.

Calibration at the stand

Every pipeline has two gate values in its gate node, and they are stand values. Before the first use at a new setup, and every morning, run the pipeline on clean product only for 16 to 20 frames (kernels and shells moving, no foreign object, no hands) with the gate logging on (log_scores: true, one line per frame: frame_score, threshold, passed, pmax). Then set

  • gate.threshold = 1.10 x the highest clean frame_score,
  • gate.mask_threshold = the mask margin x the highest clean pmax (1.0 for the original pipelines, 1.10 for the refit pipelines),

leaving out a frame whose score lies more than six robust standard deviations above the window's median (a blank camera frame opens every gate). The two pipelines of one model share the gate, so one clean turn calibrates both; the shell-aware pipelines see a damped map and need a clean turn of their own. Reload the pipeline afterwards: the thresholds are read when it loads.

Dataset

Fitted and measured on recordings of the walnut demonstration stand at Cubert, 18 August to 5 October 2026: kernels and shell fragments on a dark belt and on a turntable, with foreign objects placed by hand (stones, aluminium, rubber, plastic shells, stems), plus clean turns for calibration. The labelled test frames (575 foreign objects on 1 October, 318 on the 30 September turntable, 536 on the sessions of 18 August to 22 September) are the ones the evaluation below counts. The recordings and their COCO labels are published separately as the dataset repository of the same name.

Training details

Anomaly detector: a frozen SteerViT (DINOv2 ViT-B/14 trunk with a text-conditioned head) extracts features from a CIR projection (870, 640 and 550 nm, stretched jointly between the 2nd and 98th percentile) at two scales, the whole frame and a 2 x 2 tiling; one PatchCore memory bank per scale (coreset 7,200 of at most 80,000 patches, seed 0, top 0.1 % of the patch distances per frame) scores each pixel; the two maps are range normalised on 34 validation frames (1st to 99th percentile) and averaged. The original banks were fitted on 200 clean training frames of the first stand sessions; the refit banks on the same 200 frames plus 567 clean frames of the 1 October stand session under the production lights. The frame gate alarms above threshold and masks every pixel above mask_threshold; a mark stays only with at least 100 pixels of which at least 16 lie on an object (spectral angle to the frame's own belt above 6 degrees).

Shell segmentation: two RF-DETR large segmentation models fine-tuned on the stand recordings, one on an RGB projection and one on a CIR projection (v2, EMA weights, 24 September 2026), run on the whole frame at 504 px, fused by the mean of their scores and thresholded at 0.62. The shell mask shrinks the foreign-object mask where the two overlap (foreign object on top).

Evaluation

Measured offline with the stand calibration of each session. A foreign object counts as found with at least 50 marked pixels on it (a quarter of it under 200 px); false marks are marks touching no foreign object, per frame.

session refit refit + cut original original + cut
1 October labels, foreign objects found (of 575) 565 565 552 552
1 October labels, false marks per frame 0.15 0.16 0.18 0.19
1 October labels, marked pixels off foreign objects and shells per frame 6,127 1,170 5,539 880
30 September moving turntable, found (of 318) 240 240 303 303
18 August to 22 September sessions, found (of 536) 232 232 330 329
18 August to 22 September sessions, false marks per frame 0.47 0.52 0.78 0.87
2 October production belt, marks on empty belt per frame (wrong / right white reference) 0.09 / 0.05 0.10 / 0.07 0.04 / 0.05 0.05 / 0.14

The refit model is the better one at the production stand (more foreign objects, fewer false marks) and the weaker one on the older setups, so the original model stays for other lighting. The cut removes 81 to 84 % of the marked area off foreign objects and shells. The shell-aware pipelines were checked live on 5 October: they drop the marks on real shells next to a plastic imitation shell and keep the imitation, a stem and dark specks.

Limits

  • False marks on kernels and shells count as objects and stay (about 1.4 per frame on the moving turntable); hands are objects too.
  • A foreign object with the belt's own spectrum cannot be shown, and the alarm still counts belt specks; only the shown mask is filtered.
  • The scores drift with the lighting and the white reference: a reference taken at the wrong integration time produced belt and edge marks on 2 October. Calibrate at the stand, not from these files.
  • The cut's safety rule (pixels above 1.3 x the mask threshold always stay) was chosen on one production recording; check it on the next one before relying on it.
  • The shell-aware pipelines dampen a foreign object lying on a real shell by the same factor (it needs 1.25 x the thresholds).

License and credits

The fitted states, fine-tuned weights and pipeline yamls in this repository are released under the Apache-2.0 license (LICENSE). The third-party models they build on are credited in NOTICE.md with their licenses.

Contact

Trained and packaged by the AI team at Cubert GmbH: cuvis.ai@cubert-gmbh.de. For pilots on your own product line, reach out by email or through https://www.cubert-hyperspectral.com.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support