TacSIm / README.md
pengsc's picture
Upload 8 files
102361a verified
|
Raw
History Blame Contribute Delete
2.57 kB
---
title: TacSIm Online Benchmark
emoji:
colorFrom: blue
colorTo: green
sdk: gradio
app_file: app.py
pinned: false
---
# TacSIm Full Test-Split Online Benchmark
This ZeroGPU Space evaluates one generated CSV for every trajectory in
the complete backend Test split.
## Evaluation workflow
```text
Backend:
Test/
├── 00001.csv
├── 00002.csv
└── ...
User upload:
submission.zip
├── 00001.csv
├── 00002.csv
└── ...
```
The evaluator:
1. securely extracts the submission ZIP;
2. matches CSV files by unique basename;
3. rejects missing or unexpected files;
4. evaluates every matched trajectory;
5. averages results across the entire Test split;
6. produces detailed and aggregate reports;
7. allows the evaluated result to be submitted to the leaderboard.
## Preferred hidden-test setup
Store the official Test split in a **private Hugging Face Dataset**
repository. Configure the Space variable:
```text
TEST_DATASET_REPO_ID=YOUR_USERNAME/tacsim-private-test
```
Optional variables:
```text
TEST_DATASET_SUBDIR=Test
TEST_DATASET_REVISION=main
```
Add a Space secret with read permission:
```text
TEST_DATASET_TOKEN=hf_...
```
The private dataset repository should contain:
```text
Test/
├── 00001.csv
├── 00002.csv
└── ...
```
## Local Test-folder setup
When `TEST_DATASET_REPO_ID` is not configured, the Space reads:
```text
Test/**/*.csv
```
from its own repository.
Do not use this mode in a public Space when the Test CSVs are supposed
to remain hidden.
## Leaderboard setup
Create a separate Dataset repository, then configure:
```text
LEADERBOARD_REPO_ID=YOUR_USERNAME/tacsim-leaderboard
```
Add the write token as a Space secret:
```text
LEADERBOARD_HF_TOKEN=hf_...
```
For compatibility, one shared `HF_TOKEN` may be used instead, provided
it has both private-test read permission and leaderboard write
permission.
## Official benchmark settings
```text
Benchmark version: v2.0-batch
Frame interval: 0.1 seconds
Horizons: 3s, 5s, 10s
Grids: 10×6, 15×10, 20×12, 30×20, 105×68
```
The TacSIm score for each file, horizon, and grid is:
```text
(Spatial Occupancy Similarity + Movement Vector Similarity) / 2
```
The leaderboard Overall score is the arithmetic mean over all test
files, horizons, and grids. Scores are displayed on a 0–100 scale.
## Files uploaded to the Space
```text
app.py
football_trajectory_similarity.py
requirements.txt
README.md
```
When using local Test-folder mode, also upload:
```text
Test/
└── all official test CSV files
```