Tancho Observation Program v8
Tancho Observation Program v8 is a fine-tune of
nvidia/Cosmos3-Edge. One
checkpoint serves two prompted roles for disaster reconnaissance:
- the Reasoner selects operational questions that remain unanswered, and reports
satisfiedwhen the supplied image already answers every required question; - the Generator converts verified questions into semantic observation requirements.
The Generator does not emit flight coordinates. Trusted deterministic code binds opaque intent IDs, validates the model output, and compiles exact viewpoints and routes. Detection and area classification remain separate purpose-built components.
Experimental decision checkpoint
The Tancho Cosmos Decision Preview directory publishes a Jev-style candidate-selection fine-tune and its full export. The checkpoint scores two supplied observation programs and a reject option. It does not replace the v8 weights in this repository and is not flight-qualified. On 12 held-out real flood scenes from one related flight, each tested in two candidate orders, it agreed with AI-review labels on 11/24 decisions versus 12/24 for this v8 checkpoint using the same candidate mode. These labels are weak supervision, not an independently verified field benchmark. Its faster decision path compared with v8's long JSON generation comes primarily from candidate scoring in one forward pass, not from the fine-tuning itself.
| Modal L4 decision path, 12 real flood scenes x 2 orders | Median time | Relative to v8 JSON |
|---|---|---|
| v8 long JSON generation | 40.4724 s | 1x |
| v8 candidate selection | 0.5266 s | 76.8x faster |
| Decision Preview candidate selection | 0.5033 s | 80.4x faster |
These are warmed decision-stage times, excluding model load, image file I/O,
Reasoner, route compilation, and flight control. The JSON run used
use_cache=False; candidate selection changes the prompt and output contract.
The weight update alone did not produce the 80.4x speed difference. Jetson Orin
NX has not been measured. In all six candidate orders, this v8 checkpoint's
candidate-mode choice depended strongly on option position (on 24 synthetic road
images it chose the middle slot in 89 of 144 decisions); see the Decision
Preview card for the measurements.
Status
This is a research candidate, not a flight-qualified autonomy system. It passed its frozen synthetic final and completed 8 of 8 pre-fixed model-in-loop SITL scenes in a synthetic Gazebo world, one of them after a post-hoc rerun of a flight that hit a simulator telemetry fault. Independent real-image final evaluation and concurrent Jetson Orin NX measurement are still open Release 1 gates. Do not use the model output as an executable flight command.
Model and contracts
| Item | Value |
|---|---|
| Parent | nvidia/Cosmos3-Edge |
| Parent revision | 344d602b128d1bbdacb43b08d0a3626f46343e29 |
| Cosmos Framework revision | 96303bb0bdd14d9efa18f24d8ef98c7a0bfb8412 |
| Training | Modal, 1 x H100, 2,632 updates (4 epochs of 658 train records) |
| Trained parameters | autoregressive tower; vision encoder frozen |
| Reasoner output | tancho-observation-intent-1.0 |
| Generator output | tancho-observation-program-1.1 |
| Prompt contract | tancho-observation-program-prompts-1.1 |
| Release profiles | slope_failure_v1, access_clearance_v1 |
satisfied means every required question is either completed by trusted context (for
example a map relation) or answered by the supplied image. A mission ends only when the
Reasoner says satisfied and a deterministic capture check agrees.
Evaluation
The synthetic final contains 104 procedurally rendered cases (seeds 801-804): 72 Reasoner and 32 Generator cases, including answered close views framed like the compiler's -35 degree observation, far, too-small and cut-off views of the same target, targets shifted sideways, and wide targets with and without trees in front. It was generated only after the checkpoint was fixed and run once without output repair.
| Frozen evaluation | Exact | Schema valid | Paired answer changed |
|---|---|---|---|
| Synthetic final (v8) | 104/104 | 104/104 | 20/20 |
In every paired contrast the two prompts are byte-identical outside the image, so a changed answer can only come from the image. These are synthetic scenes; they do not establish performance on real disasters.
Correction to earlier releases. The v4 checkpoint previously published here
reported 32/32 and 16/16 on a synthetic final. In that data the context ID given to the
model named the image variant, so the answer could be read from the prompt, and paired
cases differed in that ID. Those numbers do not show image use. The data builders were
fixed before v7 and v8 were trained; an ID-text baseline that scored 52/56 on the
earlier validation set scores 24/56 on the corrected one. See evidence/label-leak-finding.json.
The synthetic final pack, its thresholds, and the 104 saved raw outputs are bundled under
evaluation/. python reproduce_scoring.py re-checks them against the recorded report
and re-scores them with the standard library only. It does not re-run the model.
Closed-loop SITL
Eight unseen scenes (seeds 413-420), fixed before the run with varied start positions and
the target shifted up to 4 m sideways, ran once each on Modal with no retries: far initial
image, Reasoner, Generator, deterministic compiler and RFL, a PX4 flight in Gazebo, a
re-captured image, and the completion Reasoner. Seven completed as run: in each the Generator requested one full-target overview, the
flight passed (tracking error at most 0.141 m, return error at most 0.041 m), and the
completion Reasoner returned satisfied, agreeing with the capture check. The eighth
(seed 416) stopped mid-flight on a simulator telemetry fault before the completion step.
After the results were known, its flight alone was rerun once with the recorded model
outputs (no model call repeated); the rerun passed and the completion Reasoner agreed.
The result is therefore 7 of 8 under the pre-fixed one-attempt rule and 8 of 8 including
that post-hoc rerun.
These are synthetic scenes built to match the training renders, not real-flight
evidence. See evidence/multi-scene-sitl.json.
On 2026-09-30 the same checkpoint was flown through the candidate-selection path: its candidate readout chose between two supplied observation programs instead of the Generator writing one. With the viewpoint compiler now aiming whole-boundary observations at the target centre (an opt-in scene contract), four preregistered procedural slope scenes (seeds 424-427) all reached a completed decision with settings frozen; the one-sided 95% lower bound on the success rate is about 47%. In a procedural road scene, a candidate readout head fitted on synthetic features chose the close two-viewpoint program, and both programs flew with every geometry check agreeing. These runs used procedural scenes; the slope candidates compile to the same viewpoint, and the road head is weak and does not reject. See the Decision Preview card for the run table.
Loading
Use the pinned Cosmos Framework revision. It registers the native
Cosmos3EdgeForConditionalGeneration class used by the export. Evaluation and SITL
used one CUDA GPU, bfloat16, SDPA, greedy generation, use_cache=False, EOS token 11,
and pad token 0. Although the inherited generation_config.json has sampling enabled,
Tancho evaluation and runtime calls explicitly set do_sample=False.
import torch
from transformers import AutoModelForImageTextToText
import cosmos_framework.model.generator.reasoner.cosmos3_edge
model = AutoModelForImageTextToText.from_pretrained(
"v13s/tancho",
torch_dtype=torch.bfloat16,
device_map="cuda:0",
attn_implementation="sdpa",
)
model.eval()
The processor is built from the pinned parent snapshot with
cosmos_framework.data.generator.processors.build_processor. Prompt construction and
strict output validation are included under tancho/; the exact evaluation entry
point is under training/. See REPRODUCE.md for the environment and integrity steps.
Package integrity
release-manifest.json records the size and SHA-256 of every v8 file except the
manifest itself and Hugging Face's generated .gitattributes. Run:
python verify_files.py
The 30-file training export is recorded in evidence/export-inventory.json. The
release omits only the export's local .cache/huggingface/trees/...json lookup cache.
The repository's earlier checkpoint/detect240/ release is retained as historical
research material; its results do not describe v8.
Training data and scope
The v8 training pack contains 762 records: 658 train and 104 validation; 538 Reasoner and 224 Generator records. The 34 real Sannoudani rows are Reasoner-only and reuse previously reviewed decisions, marked as deterministic legacy migration rather than new human review. All other labels come from deterministic procedural scene truth. Media, labels, and raw training data are not distributed in this model repository.
Limitations
- The final result is synthetic and covers two mission profiles.
- There is no independent real-image final evaluation for this checkpoint.
- Generator-path SITL covers 8 synthetic scenes of one layout, one completed only after a post-hoc flight rerun. Programs that compile to more than one viewpoint have since flown in the candidate-path runs; their capture check was corrected during those runs.
- Jetson Orin NX latency, memory, power, and concurrent-load behavior remain unmeasured.
- Exact routes depend on trusted geometry, policy, vehicle limits, and the deterministic compiler; the model alone cannot produce an approved flight plan.
- A detector or area classifier must supply target evidence.
License and attribution
The model weights and retained upstream artifacts are distributed under Open Model
Derivative Weights 1.1. Original Tancho code and documentation in this repository are
MIT licensed. See NOTICE.md, LICENSE-CODE, licenses/LICENSE.OpenMDW-1.1, and
SOURCES.md. Published by Vox in Tokyo, Japan, under the v13s namespace. Modal
supplied rented compute and did not contribute model design, data, or evaluation.
- Downloads last month
- 82
Model tree for v13s/tancho
Base model
nvidia/Cosmos3-Edge