--- title: Breakdown risk emoji: πŸ”§ colorFrom: red colorTo: gray sdk: docker app_port: 7860 short_description: Breakdown risk from the caller's turns pinned: false --- # Breakdown risk Demo of the breakdown-risk classifier: from the **caller's turns** of an inbound call, the model answers - `risk` β€” the vehicle is probably off the road, the call needs a tow truck or an urgent slot; - `no_risk` β€” the customer can still drive, a normal appointment will do. Calls are in French, so the transcripts, the examples and the dataset are French; the interface around them is English. ## The model [SetFit](https://github.com/huggingface/setfit): the sentence encoder `sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2` fine-tuned contrastively, plus a logistic regression on the embeddings. Published by `train.py` to [`bee2link/breakdown-risk-paraphrase-multilingual-MiniLM-L12-v2`](https://huggingface.co/bee2link/breakdown-risk-paraphrase-multilingual-MiniLM-L12-v2), trained on [`bee2link/breakdown-risk`](https://huggingface.co/datasets/bee2link/breakdown-risk). The **Evaluation** tab gives the current score: the test split keeps growing, so no number is copied in here. The demo's examples come from that same split β€” turns the model never saw during training. ## The three-turn window The model reads only the **last three caller turns**, and the demo reproduces that windowing: one turn per line, older ones are shown but ignored. Filtering comes before slicing. Taking the last three messages of a dialogue would count the agent's replies: in an alternating exchange only two caller turns would be left, and after two follow-up questions the complaint itself would fall out of the window β€” the classifier would see nothing but "no / no". ## Model access The model repo is private. The Space reads it through an `HF_TOKEN` secret (Settings β†’ Secrets), which needs read access to that repo. Secrets are not visible to visitors of the Space. ## Picking the model and the revision Repos in the organisation whose name starts with `breakdown-risk-` are listed at startup, and every commit of the chosen repo is offered as a revision β€” enough to compare two successive training runs without redeploying. `↻` reloads the list after a fresh `train.py` push. The list stays sorted by name, so a repo keeps its place as you scan it, but the initial selection is the most recently pushed repo, on `main`: at startup the demo therefore shows the latest training run, whichever backbone it used. The date of that push is shown under the two menus and follows the selected repo, which makes the default legible. Careful, `last_modified` applies to the whole repo β€” a README fix is enough to move the default, and the shown date says so. Every entry of the revision menu carries its timestamp, `main (latest)` included, which mirrors the tip commit's β€” so two runs pushed on the same day stay distinguishable. One format everywhere, `%Y-%m-%d %H:%M` in Paris time (`config.stamp`); the Hub reports UTC, and `tzdata` is a direct dependency so the conversion does not depend on the base image shipping a tz database. Both outputs of `train.py` are accepted: the presence of `model_head.pkl` selects the SetFit path, otherwise the model is loaded as an `AutoModelForSequenceClassification`. The **Evaluation** tab replays the test split and shows the confusion matrix, then the cases one by one, errors first. ## Clicking Classify never downloads Both action buttons start disabled, and the picked revision is fetched up front: on page load, on picking a model, and on picking a commit. While that runs, both buttons read *Downloading the model…* and are greyed out β€” they wait on the same weights, so they move together. A status line in the picker panel says the same thing, and Gradio's own progress animation sits on it via `show_progress_on`, so the feedback appears where the change was made rather than down in a tab. That line then reports `Ready Β· N ms per call`, which is the point of putting it there: the number cannot exist until the weights do, so until then the component is simply loading. It comes from two inference calls at warm time β€” the first absorbs torch's lazy init, the second is the one measured β€” so it matches what a click will actually cost instead of reporting the init as if it were latency. For the same reason the `ms` beside a prediction is timed after `load()`: inference alone, not a download that happened to be first. Picking again mid-download takes over. Gradio would otherwise drop the new pick (`trigger_mode` defaults to `"once"` everywhere except `.change`) or queue it behind the running download (`concurrency_limit` defaults to 1), so both are set explicitly. `warm` then decides which pick still owns the buttons: it records what the session is waiting on and, if that changed while it was loading, returns without touching anything and leaves the newer pick in charge. The abandoned download does finish in its thread β€” a blocking `hf_hub_download` cannot be interrupted β€” but it only fills the cache, and nothing stale reaches the screen. Warming runs once per pick. `model.change` updates the revision menu and *then* warms, so it reads the revision the switch just chose rather than the previous repo's commit; the revision menu warms on `.input`, which is user-only, so the programmatic update `model.change` just made does not warm a second time. A revision that fails to load hands the buttons back and reports why, instead of leaving a dead interface. ⇧ Enter still submits while a model is loading β€” you can paste a transcript during the download, and an early keypress simply waits on the same load. ## Nothing is ever stale `main` is resolved to a commit on every call, and the caches are keyed on the commit β€” not on the branch name. A `train.py --push` is therefore picked up without a restart: without this, `main` would forever mean the model loaded on the first click. The test split is resolved the same way, which matters: it went from 46 to 51 cases in two days. The reported score names both commits (`model @ sha Β· data sha`), so a number is always tied to what produced it. The list of repos and revisions reloads every `REFRESH_SECONDS` seconds, keeping the current selection, and `↻` forces the reload. ## Layout ``` main.py startup app/config.py constants, display names, timestamp format app/turns.py window over the last turns app/hub.py repos, revisions, downloads app/predictors.py SetFit or AutoModel, depending on the repo's files app/evaluation.py test split, accuracy, tables app/handlers.py functions the interface calls app/ui.py Gradio Blocks app/cmd_enter.js ⌘/Ctrl + Enter ``` ## Shortcuts ⇧ Enter fires `submit` natively on a multiline `Textbox`. ⌘/Ctrl + Enter does not exist in Gradio and goes through `app/cmd_enter.js`, loaded via `launch(js=...)` β€” the method described in the *Custom CSS and JS* guide. ## Dependencies Docker SDK rather than Gradio, so that versions come from a `uv.lock`: ``` uv sync # local environment, identical to the image's uv lock # after any change to pyproject.toml ``` `setfit` is not installed. A SetFit model is a `SentenceTransformer` followed by a pickled logistic regression, and `app/predictors.py` loads those two pieces β€” `gradio` 6 requires `transformers>=5`, which `setfit` 1.1 does not support (`ImportError: default_logdir`). Both paths give the same predictions, to within 6e-8. `torch` comes from the CPU index on Linux, which keeps ~2 GB of unused CUDA out of the image.