Video Classification
Transformers
Safetensors
egosieve
robotics
egocentric-video
video-quality
dataset-curation
physical-ai
custom_code
Eval Results (legacy)
Instructions to use itspublu/EgoSieve-S with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use itspublu/EgoSieve-S with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("video-classification", model="itspublu/EgoSieve-S", trust_remote_code=True)# Load model directly from transformers import AutoModelForVideoClassification model = AutoModelForVideoClassification.from_pretrained("itspublu/EgoSieve-S", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
| license: apache-2.0 | |
| library_name: transformers | |
| pipeline_tag: video-classification | |
| base_model: facebook/dinov2-small | |
| datasets: | |
| - itspublu/EgoSieve-Eval | |
| tags: | |
| - robotics | |
| - egocentric-video | |
| - video-quality | |
| - dataset-curation | |
| - physical-ai | |
| model-index: | |
| - name: EgoSieve-S | |
| results: | |
| - task: | |
| type: video-classification | |
| name: Egocentric video readiness | |
| dataset: | |
| type: itspublu/EgoSieve-Eval | |
| name: EgoSieve-Eval | |
| split: test | |
| metrics: | |
| - type: f1 | |
| name: Readiness macro F1 | |
| value: 0.67394871 | |
| - type: roc_auc | |
| name: Issue macro AUROC | |
| value: 0.74073329 | |
| - type: average_precision | |
| name: Issue macro average precision | |
| value: 0.72219606 | |
| - type: f1 | |
| name: Boundary micro F1 | |
| value: 0.10303030 | |
| # EgoSieve-S | |
| EgoSieve-S ranks manipulation-ready spans in first-person video. It produces | |
| three readiness logits (`KEEP`, `REVIEW`, `REJECT`), seven observable issue | |
| scores, diagnostic start/end boundary proposals, and a normalized retrieval | |
| embedding. It is a dataset-curation model, not a robot policy. | |
| ## Usage | |
| ```python | |
| from transformers import AutoModelForVideoClassification, AutoProcessor | |
| processor = AutoProcessor.from_pretrained( | |
| "itspublu/EgoSieve-S", revision="v0.1.0", trust_remote_code=True | |
| ) | |
| model = AutoModelForVideoClassification.from_pretrained( | |
| "itspublu/EgoSieve-S", revision="v0.1.0", trust_remote_code=True | |
| ).eval() | |
| outputs = model(**processor(frames, return_tensors="pt")) | |
| ``` | |
| The timestamp-aware scanner and JSONL compiler are provided by the `egosieve` | |
| package. The checkpoint expects 12 center-sampled | |
| RGB frames per window; use its bundled processor. | |
| ## Training and evaluation | |
| Data represented in the held-out evaluation: HoloAssist (CDLA-Permissive-2.0), HoloAssist controlled corruptions (CDLA-Permissive-2.0). Splits are grouped by | |
| original capture unit. Readiness, calibration, and boundary results use | |
| 142 human-grounded readiness rows: | |
| 0 direct human and | |
| 142 human-derived. Boundary results use | |
| 97 human-grounded boundary rows. Issue results | |
| use 145 labeled rows: 0 | |
| human, 109 human-derived, and | |
| 36 programmatic | |
| controlled corruptions. Unlabeled task targets are masked. The release bundle | |
| includes raw held-out predictions with task-level provenance, split | |
| assignments, exact run configuration, and metric provenance. | |
| Human-derived rows are not direct, independent EgoSieve rubric judgments. Treat | |
| them as proxy evidence and consult the source dataset cards and their | |
| dataset-specific proxy details before comparing or interpreting these metrics. | |
| For v0.1, readiness and boundaries come from a fixed-grid occupancy rule over | |
| reviewed HoloAssist fine-action intervals. `low_hand_activity` is an occupancy | |
| proxy, and `acting_hand_not_visible` follows HoloAssist's acting-hand modifier; | |
| it does not assert that every hand is absent. The other five issue metrics | |
| measure injected-corruption versus unmodified-reference discrimination. Those | |
| references were not independently audited as natural issue negatives. | |
| - Readiness macro F1: 0.6739 | |
| - Issue macro AUROC: 0.7407 | |
| - Issue macro average precision: 0.7222 | |
| - Boundary micro F1 at 0.30s: 0.1030 | |
| - Readiness ECE: 0.1018 | |
| ## Intended use and limitations | |
| Use the model to rank raw egocentric windows, route uncertain spans for review, | |
| and create embeddings for near-duplicate search. Readiness remains dependent on | |
| the published rubric and capture domain. RGB cannot establish force, physical | |
| success, consent, safety, metric depth, or legal publishability. Boundary | |
| scores are proposals and are diagnostic-only in the v0.1 compiler. | |
| First-person recordings can contain faces, screens, homes, and bystanders. | |
| Apply a separate privacy and consent review before sharing any media. | |