Instructions to use Andesprit/dictaria-question-classifier with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers.js
How to use Andesprit/dictaria-question-classifier with Transformers.js:
// npm i @huggingface/transformers import { pipeline } from '@huggingface/transformers'; // Allocate pipeline const pipe = await pipeline('text-classification', 'Andesprit/dictaria-question-classifier');
Dictaria question classifier
Reads a live meeting transcript as it grows and says whether the text ENDS with a complete question (or request for information) that the listener should answer. Dictaria runs it on the user's device (browser or desktop app) with onnxruntime-web, at every change of the live text: no text leaves the device to be classified.
This revision judges the end of the text with the sentences before it as context. The first revision (commit 08d64712) judged one chunk of speech said before a pause; apps pinned to that commit keep working.
Labels
question, request, social, rhetorical, statement, fragment (in labels.json order). The label describes the last sentence of the input; earlier sentences are context only.
Dictaria treats question + request probability >= 0.5 as "answer this".
question: a real question someone else is expected to answer.request: asks for information without question form ("walk me through...", "cuéntame...").social: comprehension checks, tag questions, small talk, logistics, action requests.rhetorical: nobody is expected to answer, and that is clear when it ends.statement: asks nothing, including answers to earlier questions.fragment: speech still in progress, cut off mid-thought, or filler only.
Model
- Base:
paraphrase-multilingual-MiniLM-L12-v2(Apache 2.0) with a 6-way classification head, fine-tuned from the first revision. - Training data: 910 synthetic snippets of 2 to 5 consecutive meeting sentences (each sentence labeled, about half with an unfinished version) plus the first revision's 3,400 single chunks, in English, Spanish, French, German and Portuguese, labeled by an AI model. Each sentence end is seen with the sentences before it, and unfinished versions are labeled
fragment. - Exported to ONNX and quantized to int8 per channel (
model.onnx, 118 MB). - Input: at most 96 tokens (BOS + the last 94 tokens of the transcript up to the sentence being judged + EOS).
Results
180 independently written test snippets (never used for training). The model was chosen on half of them and is reported on the other half (363 sentence ends), scored through the browser code path with the int8 model:
| First revision | This revision | |
|---|---|---|
| Right verdict at a sentence end, with context | 81.0% | 95.3% |
| Questions caught | 80.8% | 90.4% |
| False alarms | 18.9% | 2.1% |
| Fires again on the sentence after a question | 45.5% | 0.0% |
| Unfinished sentence taken for a question | 29.6% | 1.6% |
| Same, text without punctuation or capitals: right verdict | 75.8% | 88.2% |
| First revision's 340 single chunks | 96.8% | 96.5% |
By language, right verdict with context: English 97.6%, Spanish 94.8%, French 94.8%, German 95.1%, Portuguese 93.7%.
Speed: 30 to 65 ms per check in a browser worker (wasm, one thread).
Limitations
- Most misses are real questions read as small talk or as rhetorical.
- It cannot know that a speaker is about to answer their own question.
- The int8 file gives slightly different probabilities in onnxruntime and onnxruntime-web on borderline texts; the numbers above are from onnxruntime-web.
- Trained on synthetic speech, not real transcripts; speech-recognition errors in real meetings may lower accuracy.
- Only the five languages above were tested.