File size: 10,588 Bytes
fcb9b70
 
 
 
 
 
52f55f8
fcb9b70
52f55f8
fcb9b70
 
52f55f8
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
fcb9b70
 
 
 
52f55f8
 
 
 
fcb9b70
 
 
 
 
 
 
52f55f8
 
 
fcb9b70
52f55f8
fcb9b70
52f55f8
 
fcb9b70
52f55f8
 
 
 
fcb9b70
52f55f8
 
 
 
 
fcb9b70
52f55f8
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
fcb9b70
 
 
 
52f55f8
fcb9b70
52f55f8
fcb9b70
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
52f55f8
 
 
 
fcb9b70
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
52f55f8
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
fcb9b70
 
 
52f55f8
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
---
library_name: pytorch
license: apache-2.0
tags:
  - driving
  - multimodal
  - computer-vision
  - time-series
  - maneuver-forecasting
  - custom-handler
  - research
metrics:
  - f1
  - accuracy
model-index:
  - name: YellowCab v0.1-causal
    results:
      - task:
          type: image-classification
          name: Multimodal maneuver forecasting
        dataset:
          name: YellowCab private forward-route holdout
          type: private
          split: test
        metrics:
          - type: f1
            name: Macro F1
            value: 0.3915542658
          - type: accuracy
            name: Balanced accuracy
            value: 0.3915633404
          - type: accuracy
            name: Accuracy
            value: 0.5755178031
---

# YellowCab

YellowCab is a compact multimodal model for forecasting a taxi's observed
maneuver approximately 20 seconds ahead. It combines three chronological road
images with twelve past-and-current telemetry signals and returns calibrated
probabilities for:

- `continue`
- `slow`
- `stop`
- `turn_left`
- `turn_right`

The model has **5.29 million parameters**. It is intended for fleet-video
indexing, maneuver analytics, multimodal research, and prototyping—not vehicle
control.

## What is technically useful about it?

YellowCab tests a practical hypothesis: recent visual context and vehicle motion
history are more useful together than either signal alone.

On the route-held-out silver benchmark, the causal temporal-fusion checkpoint
reaches **0.3916 macro F1**, compared with **0.3414** for the strongest tested
past-only telemetry baseline and **0.3050** for the tested current-frame
vision baseline. This supports uses such as:

- assigning likely maneuver tags to fleet footage;
- retrieving candidate stop and turn events for human review;
- prioritizing uncertain sequences for data collection or annotation;
- studying disagreement between visual and motion signals; and
- prototyping calibrated, abstaining multimodal classifiers.

These results do not establish a global rank. YellowCab uses a private,
task-specific benchmark and has not been evaluated under a directly comparable
public autonomous-driving leaderboard protocol.

## Architecture

- Frozen ImageNet-pretrained EfficientNet-B0 visual encoder
- Three frames at nominal offsets `t-20 s`, `t-10 s`, and `t`
- Layer normalization followed by a one-layer unidirectional GRU
- Telemetry MLP over twelve normalized causal features
- Fusion classifier over five maneuver classes
- Temperature scaling and optional confidence-based abstention

The repository contains the exact architecture, configuration, integrity-checked
Safetensors checkpoint, and custom Hugging Face endpoint handler.

## Evaluation

### Protocol

- **74,680** examples across **36** routes
- **55,711** training examples from 26 routes
- **10,375** validation examples from 5 routes
- **8,594** test examples from the newest 5 completely held-out routes
- Complete forward-in-time route separation
- No frame from a validation or test route enters training
- Model inputs use only information available at or before the newest frame
- Labels are GPS-derived maneuver proxies approximately 20 seconds ahead
- Primary metric: macro F1
- 95% confidence interval: 500 route-clustered bootstrap samples

### Headline results

| Metric | YellowCab |
| --- | ---: |
| Macro F1 | **0.3916** |
| Route-bootstrap 95% CI | **0.3476–0.4281** |
| Balanced accuracy | **0.3916** |
| Accuracy | **0.5755** |
| Weighted F1 | **0.5807** |
| Log loss | **1.0841** |
| Brier score | **0.5563** |
| ECE, 15 bins | **0.0118** |

Raw accuracy is not the primary metric because `continue` is the majority
class. A majority classifier reaches 0.6232 accuracy while scoring only 0.1536
macro F1.

### Baselines

| Model | Macro F1 | Balanced accuracy | Accuracy |
| --- | ---: | ---: | ---: |
| **YellowCab: temporal vision + telemetry** | **0.3916** | 0.3916 | 0.5755 |
| Telemetry gradient boosting | 0.3414 | **0.4217** | 0.4147 |
| Current-frame vision linear model | 0.3050 | 0.2968 | 0.5752 |
| Past-motion rule | 0.2235 | 0.2238 | 0.4550 |
| Telemetry logistic regression | 0.1887 | 0.2872 | 0.1963 |
| Majority class | 0.1536 | 0.2000 | **0.6232** |

YellowCab improves macro F1 over every tested baseline, but the gradient-boosted
telemetry baseline has higher balanced accuracy. We report both rather than
claiming uniform dominance.

### Per-class results

| Class | Precision | Recall | F1 | Support |
| --- | ---: | ---: | ---: | ---: |
| Continue | 0.7595 | 0.7394 | **0.7493** | 5,356 |
| Slow | 0.4031 | 0.3544 | **0.3772** | 951 |
| Stop | 0.3754 | 0.3423 | **0.3581** | 634 |
| Turn left | 0.2447 | 0.3157 | **0.2757** | 833 |
| Turn right | 0.1897 | 0.2061 | **0.1975** | 820 |

Turn performance is the clearest weakness. YellowCab should be used to produce
ranked candidates or analytics, not treated as a reliable maneuver oracle.

### Confidence-based abstention

The default confidence threshold is 0.45. On this held-out set it retains
**69.5%** of examples, with **0.6583 selective accuracy** and **0.4327 selective
macro F1** among retained examples. Raising the threshold trades coverage for
higher accuracy:

| Threshold | Coverage | Selective accuracy | Selective macro F1 |
| --- | ---: | ---: | ---: |
| 0.00 | 100.0% | 0.5755 | 0.3916 |
| 0.35 | 91.0% | 0.6007 | 0.4061 |
| **0.45** | **69.5%** | **0.6583** | **0.4327** |
| 0.55 | 47.2% | 0.7324 | 0.4702 |
| 0.65 | 30.2% | 0.8092 | 0.5275 |

This makes the model more useful for candidate retrieval and review queues, but
these thresholds must be recalibrated after domain shift. Full results are in
[`selective_metrics.csv`](selective_metrics.csv).

![Benchmark comparison](benchmark.png)

![Held-out confusion matrix](confusion_matrix.png)

The complete checkpoint-specific report is in
[`evaluation.json`](evaluation.json). Recomputable, privacy-sanitized held-out
probabilities are in [`eval_predictions.csv`](eval_predictions.csv), and the
metric script is in [`recompute_metrics.py`](recompute_metrics.py). Additional
methodology and limitations are documented in
[`EVALUATION.md`](EVALUATION.md).

## Checkpoint identity

- Version: `v0.1-causal`
- Published checkpoint SHA-256:
  `af1a98cd96bb68d642f7a0adeedfacac37b230917afc15f3c87230a99a15b8a2`
- Training manifest SHA-256:
  `8b92ad563892776cc25148f39fec6bf207e1e442c618b58fb07af7e2b899364d`
- Split SHA-256:
  `bfc5e960c8af32d0ea75a5e293bb95b3125111ace2cf8452eb42feb4bc2a1a5e`
- Training seed: `20260726`

The causal-manifest audit replaced unavailable current headings with the neutral
encoding `(sin=0, cos=1)` rather than deriving them from future motion. The
published checkpoint was trained and evaluated on that corrected manifest.

## Input

Requests require exactly three base64-encoded JPEG or PNG frames ordered oldest
to newest. Optional frame timestamps must be strictly increasing.

All telemetry fields are required:

`current_speed_mps`, `past_speed_mps`, `acceleration_mps2`,
`past_turn_degrees`, `past_yaw_rate_deg_s`, `heading_sin`, `heading_cos`,
`gps_accuracy_m`, `hour_sin`, `hour_cos`, `dow_sin`, and `dow_cos`.

```json
{
  "inputs": {
    "frames": ["<base64 t-20>", "<base64 t-10>", "<base64 now>"],
    "frame_timestamps": [1721990000.0, 1721990010.2, 1721990020.1],
    "telemetry": {
      "current_speed_mps": 7.2,
      "past_speed_mps": 8.1,
      "acceleration_mps2": -0.045,
      "past_turn_degrees": 3.1,
      "past_yaw_rate_deg_s": 0.15,
      "heading_sin": 0.5,
      "heading_cos": 0.8660254038,
      "gps_accuracy_m": 4.0,
      "hour_sin": -0.7071067812,
      "hour_cos": 0.7071067812,
      "dow_sin": 0.7818314825,
      "dow_cos": 0.6234898019
    },
    "abstention_threshold": 0.45
  }
}
```

The custom `EndpointHandler` validates the request and returns the predicted
label, ordered class probabilities, confidence, and abstention state. This
custom multimodal architecture is not compatible with the standard
image-classification `pipeline()` or Hub widget.

## Local use

```bash
python -m venv .venv
# Activate the environment for your shell.
python -m pip install --requirement requirements.txt
```

```python
from handler import EndpointHandler

predict = EndpointHandler(".")
result = predict(request_body)
```

## Training data and privacy

The model was trained on privacy-redacted imagery collected from the operator's
taxi fleet under contributor agreements authorizing model training. No source
imagery, exact GPS records, route manifests, or personal identifiers are
distributed in this repository.

The published evaluation predictions use sequential example identifiers and
generic held-out route groups. Capture identifiers, timestamps, grid cells, and
locations have been removed.

## Intended uses

- Offline fleet-video indexing and search
- Maneuver analytics and aggregate research
- Candidate-event retrieval followed by human review
- Multimodal classification and calibration research
- Prototyping and collection-planning experiments

## Out-of-scope uses

Do not use YellowCab:

- as the sole input to steering, braking, throttle, or dispatch decisions;
- as a safety-certified driver-assistance component;
- to identify drivers, passengers, pedestrians, or locations;
- to infer intent, fault, liability, or legal compliance; or
- outside a target domain without measuring performance and calibration again.

## Limitations

- Ground truth is GPS-derived silver labeling, not independent human annotation.
- The evaluation covers five held-out routes from one operating domain.
- Geography, weather, camera placement, traffic, and telemetry quality can shift
  performance.
- Rare classes are substantially weaker than `continue`.
- Temperature scaling used the validation split rather than a separate
  calibration split.
- No membership-inference, model-inversion, or differential-privacy assessment
  has been completed.
- Benchmark results are author-reported and have not been independently
  replicated on a public dataset.

## License

Unless otherwise noted, the model weights and original code, configuration, and
documentation are licensed under the Apache License 2.0. Third-party components
remain subject to their respective terms; see
[`THIRD_PARTY_NOTICES.md`](THIRD_PARTY_NOTICES.md).

## Citation

```bibtex
@software{yellowcab_v0_1_causal_2026,
  author = {{The General Data Corporation}},
  title = {YellowCab: Multimodal Taxi Maneuver Forecasting},
  year = {2026},
  version = {v0.1-causal},
  url = {https://huggingface.co/generaldata/YellowCab}
}
```