pose-detection / docs /spec.md
Nishant Prasad
Upload folder using huggingface_hub
9da727c verified
|
Raw History Blame Contribute Delete
3.29 kB
# Model Specification
## YOLO26x-Pose
| Property | Value |
| Model | YOLO26x-Pose |
| Task | Human Pose Estimation |
| Framework | Ultralytics |
| Input Resolution | 960 × 960 |
| Dataset | COCO Keypoints |
| Classes | 1 |
| Class | person |
| Keypoints | 17 |
| Output features | 51 per person (17 keypoints × 3) + bbox `(x1, y1, x2, y2)` |
| Checkpoint | `models/yolo26x-pose.pt` |
## Model Configuration
- Input format: RGB image
- Input size: 960 × 960
- Detection class: person
- Pose keypoints: 17
- Pose task: human keypoint estimation
- Feature vector: 51 floats (indices 0–50 keypoint x/y/confidence); bounding box `(x1, y1, x2, y2)` returned separately as `box_xyxy`
## Output Feature Vector (51) + bbox
| Index range | Count | Block | Contents |
| --- | ---: | --- | --- |
| 0–50 | 51 | Keypoint features | 17 COCO keypoints × (x, y, confidence) |
Bounding box: `box_xyxy = [x1, y1, x2, y2]` (pixel coordinates, original frame). Downstream consumers assemble bbox + feature vector into a pandas DataFrame.
Keypoint feature order (each triplet x, y, confidence): nose, left_eye, right_eye, left_ear, right_ear, left_shoulder, right_shoulder, left_elbow, right_elbow, left_wrist, right_wrist, left_hip, right_hip, left_knee, right_knee, left_ankle, right_ankle.
## Reference Benchmark
| Metric | Score |
| mAP50-95 | 71.6% |
| mAP50 | 91.6% |
These are reference benchmark values for the model and are not presented as an independently reproduced local evaluation.
# Development specification
## Scope
YOLO26x-Pose is a human pose-estimation model for detecting people and predicting 17 human body keypoints from an input image. Each detected person is expanded into a fixed 51-feature vector (17 keypoints × 3) plus bounding box `(x1, y1, x2, y2)` for downstream analytics.
The packaged v1 artifact contains the upstream pretrained YOLO26x-Pose checkpoint from Ultralytics. The model performs person detection and pose estimation in a single model pipeline. Downstream applications can use the predicted bounding boxes, confidence scores, 17 keypoints, and the 51-feature vector for pose analysis.
## Architecture decisions
The model uses the Ultralytics YOLO26 pose architecture and is loaded through the Ultralytics framework.
The packaged checkpoint is configured for:
- Task: Human Pose Estimation
- Model: YOLO26x-Pose
- Input resolution: 960 × 960
- Number of classes: 1
- Class: `person`
- Number of keypoints: 17
The model accepts an RGB image and produces person detections together with human pose keypoints.
The packaged repository keeps the upstream model checkpoint and supporting inference, training-provenance, and evaluation files together. Application-specific tracking, quality filtering, identity association, or alert logic is outside the model itself.
## Starting checkpoint and training
The starting and final checkpoint for v1 is the upstream pretrained YOLO26x-Pose artifact from Ultralytics.
Select AI did not train or fine-tune the checkpoint.
`scripts/train.py` records this provenance and intentionally does not launch a training job or download a training dataset.
The published checkpoint is:
```text
models/yolo26x-pose.pt