Object Detection
ultralytics
ONNX
yolo
instance-segmentation
image-classification
pose-estimation
obb
tracking
semantic-segmentation
yolo26
Instructions to use prithivMLmods/YOLO26-ONNX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- ultralytics
How to use prithivMLmods/YOLO26-ONNX with ultralytics:
# Couldn't find a valid YOLO version tag. # Replace XX with the correct version. from ultralytics import YOLOvXX model = YOLOvXX.from_pretrained("prithivMLmods/YOLO26-ONNX") source = 'http://images.cocodataset.org/val2017/000000039769.jpg' model.predict(source=source, save=True) - Notebooks
- Google Colab
- Kaggle
| library_name: ultralytics | |
| license: agpl-3.0 | |
| base_model: | |
| - Ultralytics/YOLO26 | |
| language: | |
| - en | |
| - zh | |
| - ja | |
| - ru | |
| - de | |
| - fr | |
| - es | |
| - pt | |
| - tr | |
| - vi | |
| - ar | |
| pipeline_tag: object-detection | |
| tags: | |
| - ultralytics | |
| - yolo | |
| - object-detection | |
| - instance-segmentation | |
| - image-classification | |
| - pose-estimation | |
| - obb | |
| - tracking | |
| - semantic-segmentation | |
| - yolo26 | |
| # YOLO26-ONNX | |
| This repository contains ONNX exports of every [Ultralytics YOLO26](https://huggingface.co/Ultralytics/YOLO26) model variant. YOLO26 is a unified, real-time vision model family covering object detection, instance segmentation, oriented bounding box detection, pose estimation, and image classification, with native end-to-end NMS-free inference and no Distribution Focal Loss head. | |
| Each variant here was converted from the original PyTorch weights using `model.export(format="onnx")` and placed in its own folder. The exports are meant to be used directly with ONNX Runtime, with no `ultralytics` package required at inference time. | |
| ## Model Variants | |
| Five scales (n, s, m, l, x) across five tasks, 25 variants in total. | |
| | Size | Detect | Segment | OBB | Pose | Classify | | |
| |------|--------|---------|-----|------|----------| | |
| | n | yolo26n | yolo26n-seg | yolo26n-obb | yolo26n-pose | yolo26n-cls | | |
| | s | yolo26s | yolo26s-seg | yolo26s-obb | yolo26s-pose | yolo26s-cls | | |
| | m | yolo26m | yolo26m-seg | yolo26m-obb | yolo26m-pose | yolo26m-cls | | |
| | l | yolo26l | yolo26l-seg | yolo26l-obb | yolo26l-pose | yolo26l-cls | | |
| | x | yolo26x | yolo26x-seg | yolo26x-obb | yolo26x-pose | yolo26x-cls | | |
| ## Repository Structure | |
| ``` | |
| YOLO26-ONNX/ | |
| ├── README.md | |
| ├── yolo26n/yolo26n.onnx | |
| ├── yolo26s/yolo26s.onnx | |
| ├── yolo26m/yolo26m.onnx | |
| ├── yolo26l/yolo26l.onnx | |
| ├── yolo26x/yolo26x.onnx | |
| ├── yolo26n-seg/yolo26n-seg.onnx | |
| ├── yolo26s-seg/yolo26s-seg.onnx | |
| ├── yolo26m-seg/yolo26m-seg.onnx | |
| ├── yolo26l-seg/yolo26l-seg.onnx | |
| ├── yolo26x-seg/yolo26x-seg.onnx | |
| ├── yolo26n-obb/yolo26n-obb.onnx | |
| ├── yolo26s-obb/yolo26s-obb.onnx | |
| ├── yolo26m-obb/yolo26m-obb.onnx | |
| ├── yolo26l-obb/yolo26l-obb.onnx | |
| ├── yolo26x-obb/yolo26x-obb.onnx | |
| ├── yolo26n-pose/yolo26n-pose.onnx | |
| ├── yolo26s-pose/yolo26s-pose.onnx | |
| ├── yolo26m-pose/yolo26m-pose.onnx | |
| ├── yolo26l-pose/yolo26l-pose.onnx | |
| ├── yolo26x-pose/yolo26x-pose.onnx | |
| ├── yolo26n-cls/yolo26n-cls.onnx | |
| ├── yolo26s-cls/yolo26s-cls.onnx | |
| ├── yolo26m-cls/yolo26m-cls.onnx | |
| ├── yolo26l-cls/yolo26l-cls.onnx | |
| ├── yolo26x-cls/yolo26x-cls.onnx | |
| └── yolo26_onnx_utils.py | |
| ``` | |
| ## Run with ONNX | |
| ``` | |
| !pip install \ | |
| gradio \ | |
| ultralytics \ | |
| huggingface_hub \ | |
| onnx>=1.16.0 \ | |
| onnxslim>=0.1.31 \ | |
| onnxruntime>=1.18.0 \ | |
| opencv-python-headless>=4.9.0 \ | |
| numpy | |
| ``` | |
| ```python | |
| import json | |
| import tempfile | |
| import time | |
| from pathlib import Path | |
| import cv2 | |
| import gradio as gr | |
| import numpy as np | |
| import onnxruntime as ort | |
| import PIL.Image as Image | |
| from huggingface_hub import hf_hub_download | |
| # ============================================================================ | |
| # Config | |
| # ============================================================================ | |
| MODEL_REPO = "prithivMLmods/YOLO26-ONNX" | |
| SIZES = ["n", "s", "m", "l", "x"] | |
| TASK_SUFFIX = {"detect": "", "seg": "-seg", "obb": "-obb", "pose": "-pose", "cls": "-cls"} | |
| TASK_ORDER = ["detect", "seg", "obb", "pose", "cls"] | |
| TASK_IMGSZ = {"detect": 640, "seg": 640, "obb": 640, "pose": 640, "cls": 224} | |
| IMAGE_SIZE_CHOICES = [320, 640, 1024] | |
| def variant_name(size, task): | |
| return f"yolo26{size}{TASK_SUFFIX[task]}" | |
| MODEL_CHOICES = [variant_name(s, t) for t in TASK_ORDER for s in SIZES] | |
| COCO80 = [ | |
| "person", "bicycle", "car", "motorcycle", "airplane", "bus", "train", "truck", "boat", | |
| "traffic light", "fire hydrant", "stop sign", "parking meter", "bench", "bird", "cat", | |
| "dog", "horse", "sheep", "cow", "elephant", "bear", "zebra", "giraffe", "backpack", | |
| "umbrella", "handbag", "tie", "suitcase", "frisbee", "skis", "snowboard", "sports ball", | |
| "kite", "baseball bat", "baseball glove", "skateboard", "surfboard", "tennis racket", | |
| "bottle", "wine glass", "cup", "fork", "knife", "spoon", "bowl", "banana", "apple", | |
| "sandwich", "orange", "broccoli", "carrot", "hot dog", "pizza", "donut", "cake", "chair", | |
| "couch", "potted plant", "bed", "dining table", "toilet", "tv", "laptop", "mouse", | |
| "remote", "keyboard", "cell phone", "microwave", "oven", "toaster", "sink", "refrigerator", | |
| "book", "clock", "vase", "scissors", "teddy bear", "hair drier", "toothbrush", | |
| ] | |
| # COCO 17-keypoint skeleton (0-indexed) + keypoint colors | |
| POSE_SKELETON = [ | |
| (15, 13), (13, 11), (16, 14), (14, 12), (11, 12), (5, 11), (6, 12), (5, 6), (5, 7), | |
| (6, 8), (7, 9), (8, 10), (1, 2), (0, 1), (0, 2), (1, 3), (2, 4), (3, 5), (4, 6), | |
| ] | |
| def class_color(idx): | |
| """Deterministic BGR color per class index.""" | |
| hue = (idx * 47) % 180 | |
| hsv = np.uint8([[[hue, 200, 255]]]) | |
| bgr = cv2.cvtColor(hsv, cv2.COLOR_HSV2BGR)[0][0] | |
| return tuple(int(c) for c in bgr) | |
| # ============================================================================ | |
| # Model cache + preload | |
| # ============================================================================ | |
| MODEL_CACHE = {} # name -> dict(session, task, imgsz, input_name, output_names) | |
| _IMAGENET_LABELS = None | |
| def task_for_name(name): | |
| for task in ("cls", "seg", "obb", "pose"): # check longest/most-specific suffixes first | |
| if name.endswith(TASK_SUFFIX[task]) and TASK_SUFFIX[task]: | |
| return task | |
| return "detect" | |
| def load_imagenet_labels(): | |
| """Lazily fetch ImageNet-1k id->label names (used by yolo26*-cls variants).""" | |
| global _IMAGENET_LABELS | |
| if _IMAGENET_LABELS is not None: | |
| return _IMAGENET_LABELS | |
| try: | |
| path = hf_hub_download(repo_id="huggingface/label-files", filename="imagenet-1k-id2label.json", repo_type="dataset") | |
| with open(path) as f: | |
| id2label = json.load(f) | |
| _IMAGENET_LABELS = [id2label[str(i)] for i in range(len(id2label))] | |
| except Exception as e: | |
| print(f"[YOLO26-ONNX] could not load ImageNet labels ({e}); classification will show raw indices.") | |
| _IMAGENET_LABELS = [] | |
| return _IMAGENET_LABELS | |
| def preload_all_models(): | |
| providers = ["CUDAExecutionProvider", "CPUExecutionProvider"] | |
| print(f"[YOLO26-ONNX] preloading {len(MODEL_CHOICES)} variants from {MODEL_REPO} ...") | |
| for i, name in enumerate(MODEL_CHOICES): | |
| task = task_for_name(name) | |
| t0 = time.time() | |
| try: | |
| onnx_path = hf_hub_download(repo_id=MODEL_REPO, filename=f"{name}/{name}.onnx") | |
| session = ort.InferenceSession(onnx_path, providers=providers) | |
| in_shape = session.get_inputs()[0].shape | |
| h = in_shape[2] if isinstance(in_shape[2], int) else TASK_IMGSZ[task] | |
| w = in_shape[3] if isinstance(in_shape[3], int) else TASK_IMGSZ[task] | |
| MODEL_CACHE[name] = { | |
| "session": session, | |
| "task": task, | |
| "imgsz": (h, w), | |
| "input_name": session.get_inputs()[0].name, | |
| "output_names": [o.name for o in session.get_outputs()], | |
| } | |
| print(f"[YOLO26-ONNX] [{i + 1}/{len(MODEL_CHOICES)}] loaded {name} ({task}) in {time.time() - t0:.1f}s") | |
| except Exception as e: | |
| print(f"[YOLO26-ONNX] [{i + 1}/{len(MODEL_CHOICES)}] FAILED to load {name}: {e}") | |
| if any(task_for_name(n) == "cls" for n in MODEL_CACHE): | |
| load_imagenet_labels() | |
| print(f"[YOLO26-ONNX] preload complete: {len(MODEL_CACHE)}/{len(MODEL_CHOICES)} models ready.") | |
| preload_all_models() | |
| # ============================================================================ | |
| # Preprocessing | |
| # ============================================================================ | |
| def letterbox(image, new_shape=640, color=(114, 114, 114)): | |
| if isinstance(new_shape, int): | |
| new_shape = (new_shape, new_shape) | |
| h0, w0 = image.shape[:2] | |
| r = min(new_shape[0] / h0, new_shape[1] / w0) | |
| new_unpad = (int(round(w0 * r)), int(round(h0 * r))) | |
| dw = (new_shape[1] - new_unpad[0]) / 2 | |
| dh = (new_shape[0] - new_unpad[1]) / 2 | |
| if (w0, h0) != new_unpad: | |
| image = cv2.resize(image, new_unpad, interpolation=cv2.INTER_LINEAR) | |
| top, bottom = int(round(dh - 0.1)), int(round(dh + 0.1)) | |
| left, right = int(round(dw - 0.1)), int(round(dw + 0.1)) | |
| image = cv2.copyMakeBorder(image, top, bottom, left, right, cv2.BORDER_CONSTANT, value=color) | |
| return image, r, (left, top) | |
| def preprocess_spatial(image_bgr, imgsz): | |
| padded, ratio, pad = letterbox(image_bgr, imgsz) | |
| img = padded[:, :, ::-1].transpose(2, 0, 1) | |
| img = np.ascontiguousarray(img, dtype=np.float32) / 255.0 | |
| return img[None], ratio, pad, image_bgr.shape[:2] | |
| def preprocess_cls(image_bgr, imgsz): | |
| """Classification preprocessing: plain resize (no letterbox padding).""" | |
| h, w = imgsz | |
| resized = cv2.resize(image_bgr, (w, h), interpolation=cv2.INTER_LINEAR) | |
| img = resized[:, :, ::-1].transpose(2, 0, 1) | |
| img = np.ascontiguousarray(img, dtype=np.float32) / 255.0 | |
| return img[None] | |
| def scale_xy_back(xy, ratio, pad, orig_shape): | |
| out = xy.copy() | |
| out[..., 0] -= pad[0] | |
| out[..., 1] -= pad[1] | |
| out /= ratio | |
| h0, w0 = orig_shape | |
| out[..., 0] = out[..., 0].clip(0, w0) | |
| out[..., 1] = out[..., 1].clip(0, h0) | |
| return out | |
| def scale_boxes_back(boxes_xyxy, ratio, pad, orig_shape): | |
| boxes = boxes_xyxy.copy() | |
| boxes[:, [0, 2]] -= pad[0] | |
| boxes[:, [1, 3]] -= pad[1] | |
| boxes /= ratio | |
| h0, w0 = orig_shape | |
| boxes[:, [0, 2]] = boxes[:, [0, 2]].clip(0, w0) | |
| boxes[:, [1, 3]] = boxes[:, [1, 3]].clip(0, h0) | |
| return boxes | |
| def extra_nms(boxes, scores, classes, iou_thres): | |
| """Optional belt-and-suspenders NMS per class on top of YOLO26's native | |
| end-to-end (NMS-free) output. Usually a no-op since the model already | |
| returns filtered detections, but keeps the IoU slider meaningful.""" | |
| if len(boxes) == 0: | |
| return np.arange(0) | |
| keep_all = [] | |
| for c in np.unique(classes): | |
| idx = np.where(classes == c)[0] | |
| b = boxes[idx] | |
| s = scores[idx] | |
| xywh = np.stack([b[:, 0], b[:, 1], b[:, 2] - b[:, 0], b[:, 3] - b[:, 1]], axis=1) | |
| keep = cv2.dnn.NMSBoxes(xywh.tolist(), s.tolist(), score_threshold=0.0, nms_threshold=iou_thres) | |
| if len(keep) > 0: | |
| keep = np.array(keep).reshape(-1) | |
| keep_all.extend(idx[keep].tolist()) | |
| return np.array(sorted(keep_all), dtype=int) | |
| # ============================================================================ | |
| # Inference + decode | |
| # ============================================================================ | |
| def run_inference(model_name, image_bgr, conf_thres, iou_thres, imgsz_override=None): | |
| entry = MODEL_CACHE[model_name] | |
| session, task = entry["session"], entry["task"] | |
| imgsz = (int(imgsz_override), int(imgsz_override)) if imgsz_override else entry["imgsz"] | |
| if task == "cls": | |
| blob = preprocess_cls(image_bgr, imgsz) | |
| outputs = session.run(entry["output_names"], {entry["input_name"]: blob}) | |
| logits = outputs[0][0] | |
| exp = np.exp(logits - logits.max()) | |
| probs = exp / exp.sum() | |
| top5_idx = probs.argsort()[::-1][:5] | |
| labels = load_imagenet_labels() | |
| top5_labels = [labels[i] if labels and i < len(labels) else str(i) for i in top5_idx] | |
| return {"task": "cls", "top5_idx": top5_idx, "top5_conf": probs[top5_idx], "top5_labels": top5_labels} | |
| blob, ratio, pad, orig_shape = preprocess_spatial(image_bgr, imgsz) | |
| outputs = session.run(entry["output_names"], {entry["input_name"]: blob}) | |
| det = outputs[0] | |
| if det.ndim == 3: | |
| det = det[0] | |
| keep = det[:, 4] >= conf_thres | |
| det = det[keep] | |
| result = {"task": task, "boxes": np.zeros((0, 4)), "scores": np.zeros((0,)), "classes": np.zeros((0,), dtype=int)} | |
| if det.shape[0] == 0: | |
| return result | |
| boxes = scale_boxes_back(det[:, :4].astype(np.float32), ratio, pad, orig_shape) | |
| scores = det[:, 4] | |
| classes = det[:, 5].astype(int) | |
| order = extra_nms(boxes, scores, classes, iou_thres) | |
| if len(order) == 0: | |
| return result | |
| boxes, scores, classes, det = boxes[order], scores[order], classes[order], det[order] | |
| result.update(boxes=boxes, scores=scores, classes=classes) | |
| if task == "obb" and det.shape[1] >= 7: | |
| result["angles"] = det[:, 6] | |
| elif task == "pose" and det.shape[1] > 6: | |
| n_kpt = (det.shape[1] - 6) // 3 | |
| kpts = det[:, 6:6 + n_kpt * 3].reshape(-1, n_kpt, 3).copy() | |
| kpts[..., :2] = scale_xy_back(kpts[..., :2], ratio, pad, orig_shape) | |
| result["keypoints"] = kpts | |
| elif task == "seg" and len(outputs) > 1: | |
| proto = outputs[1][0] # (32, mh, mw) | |
| n_mask = proto.shape[0] | |
| mask_coeffs = det[:, 6:6 + n_mask] | |
| result["masks"] = decode_masks(mask_coeffs, proto, pad, orig_shape, imgsz) | |
| return result | |
| def decode_masks(mask_coeffs, proto, pad, orig_shape, imgsz): | |
| n_mask, mh, mw = proto.shape | |
| proto_flat = proto.reshape(n_mask, -1) | |
| masks = 1 / (1 + np.exp(-(mask_coeffs @ proto_flat))) | |
| masks = masks.reshape(-1, mh, mw) | |
| ih, iw = imgsz | |
| h0, w0 = orig_shape | |
| left, top = pad | |
| out = np.zeros((masks.shape[0], h0, w0), dtype=np.uint8) | |
| for i, m in enumerate(masks): | |
| m_full = cv2.resize(m, (iw, ih), interpolation=cv2.INTER_LINEAR) | |
| m_cropped = m_full[top: ih - top, left: iw - left] | |
| if m_cropped.size == 0: | |
| m_cropped = m_full | |
| m_resized = cv2.resize(m_cropped, (w0, h0), interpolation=cv2.INTER_LINEAR) | |
| out[i] = (m_resized > 0.5).astype(np.uint8) | |
| return out | |
| def obb_corners(box, angle): | |
| """box=[x1,y1,x2,y2] axis-aligned extent, angle in radians -> 4 rotated corner points. | |
| NOTE: verify angle sign/units against your specific export in Netron before | |
| relying on this for anything beyond visualization.""" | |
| x1, y1, x2, y2 = box | |
| cx, cy = (x1 + x2) / 2, (y1 + y2) / 2 | |
| w, h = x2 - x1, y2 - y1 | |
| cos_a, sin_a = np.cos(angle), np.sin(angle) | |
| pts = np.array([[-w / 2, -h / 2], [w / 2, -h / 2], [w / 2, h / 2], [-w / 2, h / 2]]) | |
| rot = np.array([[cos_a, -sin_a], [sin_a, cos_a]]) | |
| pts = pts @ rot.T + np.array([cx, cy]) | |
| return pts.astype(int) | |
| # ============================================================================ | |
| # Drawing | |
| # ============================================================================ | |
| def draw_result(image_bgr, result, show_labels=True, show_conf=True, class_names=None): | |
| img = image_bgr.copy() | |
| task = result["task"] | |
| if task == "cls": | |
| for i, (label, conf) in enumerate(zip(result["top5_labels"], result["top5_conf"])): | |
| text = f"{label}: {conf * 100:.1f}%" if show_conf else label | |
| cv2.putText(img, text, (10, 30 + 28 * i), cv2.FONT_HERSHEY_SIMPLEX, 0.7, (0, 0, 0), 4, cv2.LINE_AA) | |
| cv2.putText(img, text, (10, 30 + 28 * i), cv2.FONT_HERSHEY_SIMPLEX, 0.7, (255, 255, 255), 1, cv2.LINE_AA) | |
| return img | |
| names = class_names or (COCO80 if task != "pose" else ["person"]) | |
| boxes, scores, classes = result["boxes"], result["scores"], result["classes"] | |
| for i in range(len(boxes)): | |
| cls_id = int(classes[i]) | |
| color = class_color(cls_id) | |
| label_txt = names[cls_id] if cls_id < len(names) else str(cls_id) | |
| if show_conf: | |
| label_txt = f"{label_txt} {scores[i]:.2f}" | |
| if task == "obb" and "angles" in result: | |
| pts = obb_corners(boxes[i], result["angles"][i]) | |
| cv2.polylines(img, [pts], True, color, 2) | |
| anchor = tuple(pts[0]) | |
| else: | |
| x1, y1, x2, y2 = boxes[i].astype(int) | |
| cv2.rectangle(img, (x1, y1), (x2, y2), color, 2) | |
| anchor = (x1, y1) | |
| if show_labels: | |
| (tw, th), _ = cv2.getTextSize(label_txt, cv2.FONT_HERSHEY_SIMPLEX, 0.55, 1) | |
| ax, ay = anchor | |
| cv2.rectangle(img, (ax, max(0, ay - th - 6)), (ax + tw + 4, ay), color, -1) | |
| cv2.putText(img, label_txt, (ax + 2, max(12, ay - 4)), cv2.FONT_HERSHEY_SIMPLEX, 0.55, (255, 255, 255), 1, cv2.LINE_AA) | |
| if task == "seg" and "masks" in result and i < len(result["masks"]): | |
| mask = result["masks"][i].astype(bool) | |
| overlay = img.copy() | |
| overlay[mask] = color | |
| img = cv2.addWeighted(overlay, 0.4, img, 0.6, 0) | |
| if task == "pose" and "keypoints" in result: | |
| kpts = result["keypoints"][i] | |
| for x, y, v in kpts: | |
| if v > 0.3: | |
| cv2.circle(img, (int(x), int(y)), 3, (0, 255, 0), -1) | |
| for a, b in POSE_SKELETON: | |
| if a < len(kpts) and b < len(kpts) and kpts[a][2] > 0.3 and kpts[b][2] > 0.3: | |
| pa, pb = kpts[a][:2].astype(int), kpts[b][:2].astype(int) | |
| cv2.line(img, tuple(pa), tuple(pb), (0, 200, 255), 2) | |
| return img | |
| # ============================================================================ | |
| # Gradio callbacks | |
| # ============================================================================ | |
| def predict_image(img, conf_threshold, iou_threshold, model_name, show_labels, show_conf, imgsz): | |
| if img is None or model_name not in MODEL_CACHE: | |
| return None | |
| image_bgr = cv2.cvtColor(np.array(img.convert("RGB")), cv2.COLOR_RGB2BGR) | |
| result = run_inference(model_name, image_bgr, conf_threshold, iou_threshold, imgsz) | |
| annotated = draw_result(image_bgr, result, show_labels, show_conf) | |
| return Image.fromarray(cv2.cvtColor(annotated, cv2.COLOR_BGR2RGB)) | |
| def predict_video(video_path, conf_threshold, iou_threshold, model_name, show_labels, show_conf, imgsz): | |
| if video_path is None or model_name not in MODEL_CACHE: | |
| return None | |
| cap = cv2.VideoCapture(video_path) | |
| if not cap.isOpened(): | |
| return None | |
| fps = int(cap.get(cv2.CAP_PROP_FPS)) or 25 | |
| width = int(cap.get(cv2.CAP_PROP_FRAME_WIDTH)) | |
| height = int(cap.get(cv2.CAP_PROP_FRAME_HEIGHT)) | |
| temp_output = tempfile.NamedTemporaryFile(suffix=".mp4", delete=False) | |
| output_path = temp_output.name | |
| temp_output.close() | |
| fourcc = cv2.VideoWriter_fourcc(*"mp4v") | |
| out = cv2.VideoWriter(output_path, fourcc, fps, (width, height)) | |
| while True: | |
| ret, frame = cap.read() | |
| if not ret: | |
| break | |
| result = run_inference(model_name, frame, conf_threshold, iou_threshold, imgsz) | |
| annotated = draw_result(frame, result, show_labels, show_conf) | |
| out.write(annotated) | |
| cap.release() | |
| out.release() | |
| return output_path | |
| def predict_webcam(frame, conf_threshold, iou_threshold, model_name, show_labels, show_conf, imgsz): | |
| if frame is None or model_name not in MODEL_CACHE: | |
| return None | |
| if isinstance(frame, np.ndarray): | |
| frame_bgr = cv2.cvtColor(frame, cv2.COLOR_RGB2BGR) | |
| result = run_inference(model_name, frame_bgr, conf_threshold, iou_threshold, imgsz) | |
| annotated = draw_result(frame_bgr, result, show_labels, show_conf) | |
| return cv2.cvtColor(annotated, cv2.COLOR_BGR2RGB) | |
| return None | |
| # ============================================================================ | |
| # UI | |
| # ============================================================================ | |
| with gr.Blocks(title="YOLO26-ONNX") as demo: | |
| gr.Markdown( | |
| f""" | |
| # YOLO26-ONNX | |
| Real-time detection, segmentation, OBB, pose, and classification, running | |
| pure ONNX Runtime, no ultralytics install required at inference time. | |
| All {len(MODEL_CHOICES)} model variants are pulled from | |
| [{MODEL_REPO}](https://huggingface.co/{MODEL_REPO}) and preloaded at startup | |
| ({len(MODEL_CACHE)}/{len(MODEL_CHOICES)} loaded successfully). | |
| """ | |
| ) | |
| with gr.Tabs(): | |
| with gr.TabItem("Image"): | |
| with gr.Row(): | |
| with gr.Column(): | |
| img_input = gr.Image(type="pil", label="Upload Image") | |
| img_conf = gr.Slider(minimum=0, maximum=1, value=0.25, label="Confidence threshold") | |
| img_iou = gr.Slider(minimum=0, maximum=1, value=0.7, label="IoU threshold (extra NMS pass)") | |
| img_model = gr.Radio(choices=MODEL_CHOICES, label="Model", value="yolo26n") | |
| img_labels = gr.Checkbox(value=True, label="Show Labels") | |
| img_conf_show = gr.Checkbox(value=True, label="Show Confidence") | |
| img_size = gr.Radio(choices=IMAGE_SIZE_CHOICES, label="Image Size", value=640) | |
| img_btn = gr.Button("Detect", variant="primary") | |
| with gr.Column(): | |
| img_output = gr.Image(type="pil", label="Result") | |
| img_btn.click( | |
| predict_image, | |
| inputs=[img_input, img_conf, img_iou, img_model, img_labels, img_conf_show, img_size], | |
| outputs=img_output, | |
| ) | |
| gr.Examples( | |
| examples=[ | |
| ["https://ultralytics.com/images/bus.jpg", 0.25, 0.7, "yolo26n", True, True, 640], | |
| ["https://ultralytics.com/images/zidane.jpg", 0.25, 0.7, "yolo26n-seg", True, True, 640], | |
| ["https://ultralytics.com/images/boats.jpg", 0.25, 0.7, "yolo26n-obb", True, True, 1024], | |
| ["https://ultralytics.com/images/zidane.jpg", 0.25, 0.7, "yolo26n-pose", True, True, 640], | |
| ["https://ultralytics.com/images/bus.jpg", 0.25, 0.7, "yolo26n-cls", True, True, 224], | |
| ], | |
| inputs=[img_input, img_conf, img_iou, img_model, img_labels, img_conf_show, img_size], | |
| ) | |
| with gr.TabItem("Video"): | |
| with gr.Row(): | |
| with gr.Column(): | |
| vid_input = gr.Video(label="Upload Video") | |
| vid_conf = gr.Slider(minimum=0, maximum=1, value=0.25, label="Confidence threshold") | |
| vid_iou = gr.Slider(minimum=0, maximum=1, value=0.7, label="IoU threshold (extra NMS pass)") | |
| vid_model = gr.Radio(choices=MODEL_CHOICES, label="Model", value="yolo26n") | |
| vid_labels = gr.Checkbox(value=True, label="Show Labels") | |
| vid_conf_show = gr.Checkbox(value=True, label="Show Confidence") | |
| vid_size = gr.Radio(choices=IMAGE_SIZE_CHOICES, label="Image Size", value=640) | |
| vid_btn = gr.Button("Process Video", variant="primary") | |
| with gr.Column(): | |
| vid_output = gr.Video(label="Result") | |
| vid_btn.click( | |
| predict_video, | |
| inputs=[vid_input, vid_conf, vid_iou, vid_model, vid_labels, vid_conf_show, vid_size], | |
| outputs=vid_output, | |
| ) | |
| with gr.TabItem("Webcam"): | |
| gr.Markdown("### Real-time Webcam Inference") | |
| with gr.Row(): | |
| with gr.Column(): | |
| webcam_conf = gr.Slider(minimum=0, maximum=1, value=0.25, label="Confidence threshold") | |
| webcam_iou = gr.Slider(minimum=0, maximum=1, value=0.7, label="IoU threshold (extra NMS pass)") | |
| webcam_model = gr.Radio(choices=MODEL_CHOICES, label="Model", value="yolo26n") | |
| webcam_labels = gr.Checkbox(value=True, label="Show Labels") | |
| webcam_conf_show = gr.Checkbox(value=True, label="Show Confidence") | |
| webcam_size = gr.Radio(choices=IMAGE_SIZE_CHOICES, label="Image Size", value=640) | |
| with gr.Column(): | |
| webcam_input = gr.Image(sources=["webcam"], type="numpy", label="Webcam (streaming)", streaming=True) | |
| webcam_output = gr.Image(type="numpy", label="Detection Result") | |
| webcam_input.stream( | |
| predict_webcam, | |
| inputs=[webcam_input, webcam_conf, webcam_iou, webcam_model, webcam_labels, webcam_conf_show, webcam_size], | |
| outputs=webcam_output, | |
| ) | |
| if __name__ == "__main__": | |
| demo.launch(ssr_mode=False) | |
| ``` | |
| ## Output Format Notes | |
| - Detection heads are exported end-to-end (NMS-free): the primary output is shaped `(1, 300, 6)` as `[x1, y1, x2, y2, confidence, class_id]` in letterboxed pixel coordinates, already filtered without a separate NMS step. | |
| - Segmentation, pose, and OBB variants append extra columns to that same row: mask coefficients plus a separate prototype tensor for segment, keypoints for pose, and a rotation angle for OBB. | |
| - Classification variants output a single `(1, num_classes)` logits vector; softmax is applied client-side. | |
| - The OBB angle convention (sign and units) has not been independently verified against every export; check it in Netron before relying on it beyond visualization. | |
| - Semantic segmentation (`-sem`) variants are not included in this repository, only the five task heads shown above. | |
| ## License | |
| Ultralytics YOLO26 code, training, and export paths are distributed under AGPL-3.0, with a separate Enterprise License available from Ultralytics for closed-source use. These ONNX exports inherit that license from the base model. |