Spaces:
Sleeping
Sleeping
| title: Breakdown risk | |
| emoji: π§ | |
| colorFrom: red | |
| colorTo: gray | |
| sdk: docker | |
| app_port: 7860 | |
| short_description: Breakdown risk from the caller's turns | |
| pinned: false | |
| # Breakdown risk | |
| Demo of the breakdown-risk classifier: from the **caller's turns** of an inbound | |
| call, the model answers | |
| - `risk` β the vehicle is probably off the road, the call needs a tow truck or an | |
| urgent slot; | |
| - `no_risk` β the customer can still drive, a normal appointment will do. | |
| Calls are in French, so the transcripts, the examples and the dataset are French; | |
| the interface around them is English. | |
| ## The model | |
| [SetFit](https://github.com/huggingface/setfit): the sentence encoder | |
| `sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2` fine-tuned | |
| contrastively, plus a logistic regression on the embeddings. Published by | |
| `train.py` to | |
| [`bee2link/breakdown-risk-paraphrase-multilingual-MiniLM-L12-v2`](https://huggingface.co/bee2link/breakdown-risk-paraphrase-multilingual-MiniLM-L12-v2), | |
| trained on [`bee2link/breakdown-risk`](https://huggingface.co/datasets/bee2link/breakdown-risk). | |
| The **Evaluation** tab gives the current score: the test split keeps growing, so | |
| no number is copied in here. The demo's examples come from that same split β | |
| turns the model never saw during training. | |
| ## The three-turn window | |
| The model reads only the **last three caller turns**, and the demo reproduces | |
| that windowing: one turn per line, older ones are shown but ignored. | |
| Filtering comes before slicing. Taking the last three messages of a dialogue | |
| would count the agent's replies: in an alternating exchange only two caller turns | |
| would be left, and after two follow-up questions the complaint itself would fall | |
| out of the window β the classifier would see nothing but "no / no". | |
| ## Model access | |
| The model repo is private. The Space reads it through an `HF_TOKEN` secret | |
| (Settings β Secrets), which needs read access to that repo. Secrets are not | |
| visible to visitors of the Space. | |
| ## Picking the model and the revision | |
| Repos in the organisation whose name starts with `breakdown-risk-` are listed at | |
| startup, and every commit of the chosen repo is offered as a revision β enough to | |
| compare two successive training runs without redeploying. `β»` reloads the list | |
| after a fresh `train.py` push. | |
| The list stays sorted by name, so a repo keeps its place as you scan it, but the | |
| initial selection is the most recently pushed repo, on `main`: at startup the | |
| demo therefore shows the latest training run, whichever backbone it used. The | |
| date of that push is shown under the two menus and follows the selected repo, | |
| which makes the default legible. Careful, `last_modified` applies to the whole | |
| repo β a README fix is enough to move the default, and the shown date says so. | |
| Every entry of the revision menu carries its timestamp, `main (latest)` included, | |
| which mirrors the tip commit's β so two runs pushed on the same day stay | |
| distinguishable. One format everywhere, `%Y-%m-%d %H:%M` in Paris time | |
| (`config.stamp`); the Hub reports UTC, and `tzdata` is a direct dependency so the | |
| conversion does not depend on the base image shipping a tz database. | |
| Both outputs of `train.py` are accepted: the presence of `model_head.pkl` selects | |
| the SetFit path, otherwise the model is loaded as an | |
| `AutoModelForSequenceClassification`. | |
| The **Evaluation** tab replays the test split and shows the confusion matrix, | |
| then the cases one by one, errors first. | |
| ## Clicking Classify never downloads | |
| Both action buttons start disabled, and the picked revision is fetched up front: | |
| on page load, on picking a model, and on picking a commit. While that runs, both | |
| buttons read *Downloading the modelβ¦* and are greyed out β they wait on the same | |
| weights, so they move together. A status line in the picker panel says the same | |
| thing, and Gradio's own progress animation sits on it via `show_progress_on`, so | |
| the feedback appears where the change was made rather than down in a tab. | |
| That line then reports `Ready Β· N ms per call`, which is the point of putting it | |
| there: the number cannot exist until the weights do, so until then the component | |
| is simply loading. It comes from two inference calls at warm time β the first | |
| absorbs torch's lazy init, the second is the one measured β so it matches what a | |
| click will actually cost instead of reporting the init as if it were latency. For | |
| the same reason the `ms` beside a prediction is timed after `load()`: inference | |
| alone, not a download that happened to be first. | |
| Picking again mid-download takes over. Gradio would otherwise drop the new pick | |
| (`trigger_mode` defaults to `"once"` everywhere except `.change`) or queue it | |
| behind the running download (`concurrency_limit` defaults to 1), so both are set | |
| explicitly. `warm` then decides which pick still owns the buttons: it records what | |
| the session is waiting on and, if that changed while it was loading, returns | |
| without touching anything and leaves the newer pick in charge. The abandoned | |
| download does finish in its thread β a blocking `hf_hub_download` cannot be | |
| interrupted β but it only fills the cache, and nothing stale reaches the screen. | |
| Warming runs once per pick. `model.change` updates the revision menu and *then* | |
| warms, so it reads the revision the switch just chose rather than the previous | |
| repo's commit; the revision menu warms on `.input`, which is user-only, so the | |
| programmatic update `model.change` just made does not warm a second time. A | |
| revision that fails to load hands the buttons back and reports why, instead of | |
| leaving a dead interface. | |
| β§ Enter still submits while a model is loading β you can paste a transcript | |
| during the download, and an early keypress simply waits on the same load. | |
| ## Nothing is ever stale | |
| `main` is resolved to a commit on every call, and the caches are keyed on the | |
| commit β not on the branch name. A `train.py --push` is therefore picked up | |
| without a restart: without this, `main` would forever mean the model loaded on | |
| the first click. The test split is resolved the same way, which matters: it went | |
| from 46 to 51 cases in two days. | |
| The reported score names both commits (`model @ sha Β· data sha`), so a number is | |
| always tied to what produced it. | |
| The list of repos and revisions reloads every `REFRESH_SECONDS` seconds, keeping | |
| the current selection, and `β»` forces the reload. | |
| ## Layout | |
| ``` | |
| main.py startup | |
| app/config.py constants, display names, timestamp format | |
| app/turns.py window over the last turns | |
| app/hub.py repos, revisions, downloads | |
| app/predictors.py SetFit or AutoModel, depending on the repo's files | |
| app/evaluation.py test split, accuracy, tables | |
| app/handlers.py functions the interface calls | |
| app/ui.py Gradio Blocks | |
| app/cmd_enter.js β/Ctrl + Enter | |
| ``` | |
| ## Shortcuts | |
| β§ Enter fires `submit` natively on a multiline `Textbox`. β/Ctrl + Enter does not | |
| exist in Gradio and goes through `app/cmd_enter.js`, loaded via `launch(js=...)` | |
| β the method described in the *Custom CSS and JS* guide. | |
| ## Dependencies | |
| Docker SDK rather than Gradio, so that versions come from a `uv.lock`: | |
| ``` | |
| uv sync # local environment, identical to the image's | |
| uv lock # after any change to pyproject.toml | |
| ``` | |
| `setfit` is not installed. A SetFit model is a `SentenceTransformer` followed by a | |
| pickled logistic regression, and `app/predictors.py` loads those two pieces β | |
| `gradio` 6 requires `transformers>=5`, which `setfit` 1.1 does not support | |
| (`ImportError: default_logdir`). Both paths give the same predictions, to within | |
| 6e-8. | |
| `torch` comes from the CPU index on Linux, which keeps ~2 GB of unused CUDA out | |
| of the image. | |