| --- |
| library_name: pytorch |
| license: apache-2.0 |
| tags: |
| - driving |
| - multimodal |
| - computer-vision |
| - time-series |
| - maneuver-forecasting |
| - custom-handler |
| - research |
| metrics: |
| - f1 |
| - accuracy |
| model-index: |
| - name: YellowCab v0.1-causal |
| results: |
| - task: |
| type: image-classification |
| name: Multimodal maneuver forecasting |
| dataset: |
| name: YellowCab private forward-route holdout |
| type: private |
| split: test |
| metrics: |
| - type: f1 |
| name: Macro F1 |
| value: 0.3915542658 |
| - type: accuracy |
| name: Balanced accuracy |
| value: 0.3915633404 |
| - type: accuracy |
| name: Accuracy |
| value: 0.5755178031 |
| --- |
| |
| # YellowCab |
|
|
| YellowCab is a compact multimodal model for forecasting a taxi's observed |
| maneuver approximately 20 seconds ahead. It combines three chronological road |
| images with twelve past-and-current telemetry signals and returns calibrated |
| probabilities for: |
|
|
| - `continue` |
| - `slow` |
| - `stop` |
| - `turn_left` |
| - `turn_right` |
|
|
| The model has **5.29 million parameters**. It is intended for fleet-video |
| indexing, maneuver analytics, multimodal research, and prototyping—not vehicle |
| control. |
|
|
| ## What is technically useful about it? |
|
|
| YellowCab tests a practical hypothesis: recent visual context and vehicle motion |
| history are more useful together than either signal alone. |
|
|
| On the route-held-out silver benchmark, the causal temporal-fusion checkpoint |
| reaches **0.3916 macro F1**, compared with **0.3414** for the strongest tested |
| past-only telemetry baseline and **0.3050** for the tested current-frame |
| vision baseline. This supports uses such as: |
|
|
| - assigning likely maneuver tags to fleet footage; |
| - retrieving candidate stop and turn events for human review; |
| - prioritizing uncertain sequences for data collection or annotation; |
| - studying disagreement between visual and motion signals; and |
| - prototyping calibrated, abstaining multimodal classifiers. |
|
|
| These results do not establish a global rank. YellowCab uses a private, |
| task-specific benchmark and has not been evaluated under a directly comparable |
| public autonomous-driving leaderboard protocol. |
|
|
| ## Architecture |
|
|
| - Frozen ImageNet-pretrained EfficientNet-B0 visual encoder |
| - Three frames at nominal offsets `t-20 s`, `t-10 s`, and `t` |
| - Layer normalization followed by a one-layer unidirectional GRU |
| - Telemetry MLP over twelve normalized causal features |
| - Fusion classifier over five maneuver classes |
| - Temperature scaling and optional confidence-based abstention |
|
|
| The repository contains the exact architecture, configuration, integrity-checked |
| Safetensors checkpoint, and custom Hugging Face endpoint handler. |
|
|
| ## Evaluation |
|
|
| ### Protocol |
|
|
| - **74,680** examples across **36** routes |
| - **55,711** training examples from 26 routes |
| - **10,375** validation examples from 5 routes |
| - **8,594** test examples from the newest 5 completely held-out routes |
| - Complete forward-in-time route separation |
| - No frame from a validation or test route enters training |
| - Model inputs use only information available at or before the newest frame |
| - Labels are GPS-derived maneuver proxies approximately 20 seconds ahead |
| - Primary metric: macro F1 |
| - 95% confidence interval: 500 route-clustered bootstrap samples |
|
|
| ### Headline results |
|
|
| | Metric | YellowCab | |
| | --- | ---: | |
| | Macro F1 | **0.3916** | |
| | Route-bootstrap 95% CI | **0.3476–0.4281** | |
| | Balanced accuracy | **0.3916** | |
| | Accuracy | **0.5755** | |
| | Weighted F1 | **0.5807** | |
| | Log loss | **1.0841** | |
| | Brier score | **0.5563** | |
| | ECE, 15 bins | **0.0118** | |
|
|
| Raw accuracy is not the primary metric because `continue` is the majority |
| class. A majority classifier reaches 0.6232 accuracy while scoring only 0.1536 |
| macro F1. |
|
|
| ### Baselines |
|
|
| | Model | Macro F1 | Balanced accuracy | Accuracy | |
| | --- | ---: | ---: | ---: | |
| | **YellowCab: temporal vision + telemetry** | **0.3916** | 0.3916 | 0.5755 | |
| | Telemetry gradient boosting | 0.3414 | **0.4217** | 0.4147 | |
| | Current-frame vision linear model | 0.3050 | 0.2968 | 0.5752 | |
| | Past-motion rule | 0.2235 | 0.2238 | 0.4550 | |
| | Telemetry logistic regression | 0.1887 | 0.2872 | 0.1963 | |
| | Majority class | 0.1536 | 0.2000 | **0.6232** | |
|
|
| YellowCab improves macro F1 over every tested baseline, but the gradient-boosted |
| telemetry baseline has higher balanced accuracy. We report both rather than |
| claiming uniform dominance. |
|
|
| ### Per-class results |
|
|
| | Class | Precision | Recall | F1 | Support | |
| | --- | ---: | ---: | ---: | ---: | |
| | Continue | 0.7595 | 0.7394 | **0.7493** | 5,356 | |
| | Slow | 0.4031 | 0.3544 | **0.3772** | 951 | |
| | Stop | 0.3754 | 0.3423 | **0.3581** | 634 | |
| | Turn left | 0.2447 | 0.3157 | **0.2757** | 833 | |
| | Turn right | 0.1897 | 0.2061 | **0.1975** | 820 | |
|
|
| Turn performance is the clearest weakness. YellowCab should be used to produce |
| ranked candidates or analytics, not treated as a reliable maneuver oracle. |
|
|
| ### Confidence-based abstention |
|
|
| The default confidence threshold is 0.45. On this held-out set it retains |
| **69.5%** of examples, with **0.6583 selective accuracy** and **0.4327 selective |
| macro F1** among retained examples. Raising the threshold trades coverage for |
| higher accuracy: |
|
|
| | Threshold | Coverage | Selective accuracy | Selective macro F1 | |
| | --- | ---: | ---: | ---: | |
| | 0.00 | 100.0% | 0.5755 | 0.3916 | |
| | 0.35 | 91.0% | 0.6007 | 0.4061 | |
| | **0.45** | **69.5%** | **0.6583** | **0.4327** | |
| | 0.55 | 47.2% | 0.7324 | 0.4702 | |
| | 0.65 | 30.2% | 0.8092 | 0.5275 | |
|
|
| This makes the model more useful for candidate retrieval and review queues, but |
| these thresholds must be recalibrated after domain shift. Full results are in |
| [`selective_metrics.csv`](selective_metrics.csv). |
|
|
|  |
|
|
|  |
|
|
| The complete checkpoint-specific report is in |
| [`evaluation.json`](evaluation.json). Recomputable, privacy-sanitized held-out |
| probabilities are in [`eval_predictions.csv`](eval_predictions.csv), and the |
| metric script is in [`recompute_metrics.py`](recompute_metrics.py). Additional |
| methodology and limitations are documented in |
| [`EVALUATION.md`](EVALUATION.md). |
|
|
| ## Checkpoint identity |
|
|
| - Version: `v0.1-causal` |
| - Published checkpoint SHA-256: |
| `af1a98cd96bb68d642f7a0adeedfacac37b230917afc15f3c87230a99a15b8a2` |
| - Training manifest SHA-256: |
| `8b92ad563892776cc25148f39fec6bf207e1e442c618b58fb07af7e2b899364d` |
| - Split SHA-256: |
| `bfc5e960c8af32d0ea75a5e293bb95b3125111ace2cf8452eb42feb4bc2a1a5e` |
| - Training seed: `20260726` |
|
|
| The causal-manifest audit replaced unavailable current headings with the neutral |
| encoding `(sin=0, cos=1)` rather than deriving them from future motion. The |
| published checkpoint was trained and evaluated on that corrected manifest. |
|
|
| ## Input |
|
|
| Requests require exactly three base64-encoded JPEG or PNG frames ordered oldest |
| to newest. Optional frame timestamps must be strictly increasing. |
|
|
| All telemetry fields are required: |
|
|
| `current_speed_mps`, `past_speed_mps`, `acceleration_mps2`, |
| `past_turn_degrees`, `past_yaw_rate_deg_s`, `heading_sin`, `heading_cos`, |
| `gps_accuracy_m`, `hour_sin`, `hour_cos`, `dow_sin`, and `dow_cos`. |
|
|
| ```json |
| { |
| "inputs": { |
| "frames": ["<base64 t-20>", "<base64 t-10>", "<base64 now>"], |
| "frame_timestamps": [1721990000.0, 1721990010.2, 1721990020.1], |
| "telemetry": { |
| "current_speed_mps": 7.2, |
| "past_speed_mps": 8.1, |
| "acceleration_mps2": -0.045, |
| "past_turn_degrees": 3.1, |
| "past_yaw_rate_deg_s": 0.15, |
| "heading_sin": 0.5, |
| "heading_cos": 0.8660254038, |
| "gps_accuracy_m": 4.0, |
| "hour_sin": -0.7071067812, |
| "hour_cos": 0.7071067812, |
| "dow_sin": 0.7818314825, |
| "dow_cos": 0.6234898019 |
| }, |
| "abstention_threshold": 0.45 |
| } |
| } |
| ``` |
|
|
| The custom `EndpointHandler` validates the request and returns the predicted |
| label, ordered class probabilities, confidence, and abstention state. This |
| custom multimodal architecture is not compatible with the standard |
| image-classification `pipeline()` or Hub widget. |
|
|
| ## Local use |
|
|
| ```bash |
| python -m venv .venv |
| # Activate the environment for your shell. |
| python -m pip install --requirement requirements.txt |
| ``` |
|
|
| ```python |
| from handler import EndpointHandler |
| |
| predict = EndpointHandler(".") |
| result = predict(request_body) |
| ``` |
|
|
| ## Training data and privacy |
|
|
| The model was trained on privacy-redacted imagery collected from the operator's |
| taxi fleet under contributor agreements authorizing model training. No source |
| imagery, exact GPS records, route manifests, or personal identifiers are |
| distributed in this repository. |
|
|
| The published evaluation predictions use sequential example identifiers and |
| generic held-out route groups. Capture identifiers, timestamps, grid cells, and |
| locations have been removed. |
|
|
| ## Intended uses |
|
|
| - Offline fleet-video indexing and search |
| - Maneuver analytics and aggregate research |
| - Candidate-event retrieval followed by human review |
| - Multimodal classification and calibration research |
| - Prototyping and collection-planning experiments |
|
|
| ## Out-of-scope uses |
|
|
| Do not use YellowCab: |
|
|
| - as the sole input to steering, braking, throttle, or dispatch decisions; |
| - as a safety-certified driver-assistance component; |
| - to identify drivers, passengers, pedestrians, or locations; |
| - to infer intent, fault, liability, or legal compliance; or |
| - outside a target domain without measuring performance and calibration again. |
|
|
| ## Limitations |
|
|
| - Ground truth is GPS-derived silver labeling, not independent human annotation. |
| - The evaluation covers five held-out routes from one operating domain. |
| - Geography, weather, camera placement, traffic, and telemetry quality can shift |
| performance. |
| - Rare classes are substantially weaker than `continue`. |
| - Temperature scaling used the validation split rather than a separate |
| calibration split. |
| - No membership-inference, model-inversion, or differential-privacy assessment |
| has been completed. |
| - Benchmark results are author-reported and have not been independently |
| replicated on a public dataset. |
|
|
| ## License |
|
|
| Unless otherwise noted, the model weights and original code, configuration, and |
| documentation are licensed under the Apache License 2.0. Third-party components |
| remain subject to their respective terms; see |
| [`THIRD_PARTY_NOTICES.md`](THIRD_PARTY_NOTICES.md). |
|
|
| ## Citation |
|
|
| ```bibtex |
| @software{yellowcab_v0_1_causal_2026, |
| author = {{The General Data Corporation}}, |
| title = {YellowCab: Multimodal Taxi Maneuver Forecasting}, |
| year = {2026}, |
| version = {v0.1-causal}, |
| url = {https://huggingface.co/generaldata/YellowCab} |
| } |
| ``` |
|
|