pianoCV

A piano segmentation model: given a camera frame of a piano or electronic keyboard, it marks every pixel that belongs to a white key, a black key, or the gap between two white keys, so each key comes out as its own region with its own outline.

It ships with two smaller companions, a detector that finds roughly where the keyboard is and a matcher that reads the key edges along it, and the three together find every key live in a web browser. They are the models behind https://github.com/matheusfillipe/pianoCV, which fits a real keyboard's layout to what they see and draws every key in the camera's own perspective.

They are deliberately small and fast. Each one runs through onnxruntime-web, and the key segmenter runs on the GPU through WebGPU in about 30 ms per frame.

file what it finds size
keybed_seg2.onnx roughly where the keyboard is in the frame 6.6 MB
keyseg.onnx every pixel that is a white key, a black key, or the gap between two white keys 6.6 MB
keymatch.onnx where the white-key gaps and the black-key edges are along the keyboard 0.3 MB

All three take RGB input normalised on ImageNet statistics, mean [0.485, 0.456, 0.406] and std [0.229, 0.224, 0.225], as planar float32.

keybed_seg2.onnx

input image, [1, 3, 288, 288], the whole frame squashed to a square
output mask, [1, 1, 144, 144], how likely each patch is to be keyboard, 0 to 1

A MobileNetV3-Small backbone pretrained on ImageNet with a small U-Net decoder. It gives a rough region, which the pipeline only uses to decide where to look.

keyseg.onnx

input crop, [1, 3, 224, 1024], a crop around the keyboard, rotated so the keys run left to right with the player's edge at the bottom
output classes, [1, 4, 224, 1024], per-pixel probabilities for background, white key, black key and white-key gap

The crop only rotates and scales the frame, so keys keep the shape the camera gives them. The same MobileNetV3-Small encoder with a decoder that climbs back to full resolution, since the gap between two white keys is a pixel or two wide. A black key's label covers its whole visible outline, raised top and sides included. On held out synthetic renders it overlaps the true key pixels with an IoU of 0.86 for black keys and 0.65 for white keys.

keymatch.onnx

input strip, [1, 3, 64, 768], the keybed rectified flat, far edge at the top
output heatmaps, [1, 3, 768], probability along the strip of a white-key gap, a black key's left edge and a black key's right edge

A small 1D network over a rectified strip. The pipeline reads it on two strips of different depth, and how far each gap moves between them gives the slant of every key line, which is how it straightens a skewed outline. On held out synthetic renders it finds edges with a precision of 0.85 and places them within 0.75 px on average.

How they were built

All three are trained on synthetic renders. A three.js generator builds a procedural keyboard with real key proportions and varies its size, case, lighting, background and camera, and saves every key's exact outline with each frame. The keybed detector is also fine tuned on hand labelled frames from real recordings.

Limitations, honestly

  • Mostly synthetic training data. The models transfer to real cameras, but the far end of a keyboard at a steep angle is still where they are weakest: small, distant black keys can merge. The pipeline covers that by keeping a key's fitted shape when the segmenter cannot separate it.
  • Very low camera angles, below about 20 degrees above the keys, mostly fail.
  • They only find the keys. They do not detect hands, read notes, or identify the instrument.

Licence

Apache 2.0.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support