Spaces:
Running
title: Conversation Data Extraction System
sdk: docker
emoji: 🗂️
colorFrom: indigo
colorTo: purple
app_port: 8000
Conversation Data Extraction System
A system that extracts structured information from chat conversations: named entities (NER), topics and sentiment, then pseudonymizes personal data (GDPR) and stores the structured result.
Project versions
This repository is developed incrementally as part of a diploma thesis. Each version is tagged so the progression is visible in the git history.
| Version | Tag | Description |
|---|---|---|
| V1 | v1.0 |
Single fine-tuned RoBERTa model, /predict endpoint, simple SQLite logging with regex anonymization. |
| V2 | v2.0 |
Full extraction pipeline: batch /ingest, NER, zero-shot topic classification, Presidio-based anonymization, dashboard with stored records and statistics. |
| V3 | v3.0 |
Dynamic model selection from the Hugging Face Hub (sentiment, NER and topic classification) at request time, plus XML export of stored records. |
| V4 | v4.0 |
Export of stored records in a choice of formats (XML, JSON, CSV) via a single endpoint, picked from a dropdown button on the dashboard. |
| V5 | v5.0 |
Per-step model choice (sentiment, NER, topics) switched from radio buttons to a <select> dropdown; a free-text field for a custom model appears only when "Custom model" is selected. |
| V6 | v6.0 |
Batched inference for /ingest: NER, topic classification and sentiment analysis each run once over the whole batch instead of once per message, significantly reducing total processing time for large batches (especially on CPU-only hosting). |
| V7 | v7.0 |
/ingest runs as a background job instead of one long blocking request; the dashboard polls job status and shows a progress bar and elapsed-time timer. |
See docs/class_diagram.md for the current architecture.
Pipeline
ingest → NER → topic classification → sentiment → anonymization → storage
| Step | Method / model |
|---|---|
| NER | dslim/bert-base-NER (inference) |
| Topics | facebook/bart-large-mnli (zero-shot) |
| Sentiment | fine-tuned RoBERTa (vojmahdal/roberta-sentiment-3labels) |
| Anonymization | Microsoft Presidio (spaCy NER + regex), regex fallback |
| Storage | SQLite (hash of original + anonymized text) |
Endpoints
| Method | Path | Description |
|---|---|---|
| GET | / |
Web dashboard |
| GET | /health |
Service + model status |
| GET | /models |
Default, suggested and currently cached models, per pipeline step |
| POST | /analyze |
Full pipeline on a single message |
| POST | /predict |
Sentiment only (backward compatible) |
| POST | /ingest |
Start a background batch-ingest job for a CSV/JSON file, returns a job_id |
| GET | /ingest/status/{job_id} |
Progress (and, once done, the result) of a background ingest job |
| GET | /records |
Recent stored (anonymized) records |
| GET | /records/export |
Stored records exported as XML, JSON or CSV (?format=) |
| GET | /stats |
Aggregate statistics (includes the list of supported export formats) |
Example
curl -X POST https://<space-url>/analyze \
-H "Content-Type: application/json" \
-d '{"text": "Hi, John Smith here, my order never arrived. Email john@example.com"}'
curl -X POST https://<space-url>/ingest -F "file=@sample_chats.csv"
# -> {"job_id": "...", "total": 8}
curl https://<space-url>/ingest/status/<job_id>
# -> {"status": "running", "processed": 4, "total": 8, "elapsed_seconds": 3.2}
# poll again once "status" is "done" (or "error") for the final result
curl "https://<space-url>/records/export?format=xml" -o records.xml
curl "https://<space-url>/records/export?format=json" -o records.json
curl "https://<space-url>/records/export?format=csv" -o records.csv
Exporting stored records
Since V4, GET /records/export accepts a format query parameter (xml,
json or csv, default xml) and an optional limit. All three formats
share the same underlying data (db._fetch_export_rows); adding a new
format only requires one small function in db.py plus an entry in
db.EXPORT_FORMATS - main.py and the dashboard pick it up automatically.
CSV flattens the nested entity list into a single "TYPE:text; ..." cell
per record and is written with a UTF-8 BOM so it opens correctly in Excel.
On the dashboard (/static/records.html), an Export ▾ dropdown button
lists the formats returned by GET /stats (export_formats); picking one
downloads the file via Content-Disposition: attachment.
Choosing models from the Hugging Face Hub
Since V3, each of the three ML steps can use a different Hugging Face Hub model, selected per request:
| Field | Task | Default |
|---|---|---|
sentiment_model |
sentiment-analysis |
vojmahdal/roberta-sentiment-3labels |
ner_model |
token-classification |
dslim/bert-base-NER |
topic_model |
zero-shot-classification |
facebook/bart-large-mnli |
All three are accepted by /analyze (JSON body) and /ingest (form
fields); /predict only accepts sentiment_model since it is
sentiment-only. If a field is omitted, that step's default model is used.
Requested models are downloaded and cached in memory on first use
(processors/model_registry.py), namespaced by task with a small FIFO
cache (6 pipelines) to bound memory usage. GET /models lists, per step,
the default model, a few suggested models, and which ones are currently
cached.
The web dashboard exposes this as a dropdown (<select>) for each step,
listing the suggested models plus a "Custom model" option. The free-text
field for typing any other Hugging Face repo id is hidden by default and
only appears once "Custom model" is selected in the dropdown.
Security note: loaded pipelines never use trust_remote_code=True, so an
arbitrary/untrusted model id supplied by a caller cannot execute custom
Python code inside the server process - it is limited to standard
transformers inference for the given task. An invalid or incompatible
model id results in a clean HTTP 400 response instead of crashing the
server.
Batched inference for large /ingest batches
Since V6, pipeline.process_batch (used only by /ingest) no longer loops
over process_message once per message. Instead it calls each processor's
batched function - ner.extract_entities_batch, topics.classify_topic_batch,
sentiment.analyze_sentiment_batch - once for the whole list of texts, and
reassembles the per-message results afterwards. Each batched function passes
batch_size=16 to the underlying transformers pipeline so the model
itself processes several texts per forward pass instead of one at a time.
This matters most for topic classification: zero-shot classification scores
every text against every candidate label as a separate NLI pass through
facebook/bart-large-mnli (~400M parameters), so for the default 8 labels
that's 8 passes per message. Batching those calls together is the difference
between a 200-message /ingest request taking minutes rather than tens of
minutes on CPU-only hosting (e.g. the free tier of Hugging Face Spaces).
process_message (used by /analyze and /predict, always a single
message) is unchanged.
Background ingest jobs (progress bar + timer)
Since V7, POST /ingest no longer blocks until the whole batch is
processed. It parses/validates the uploaded file synchronously (fast, no
model calls), registers a job via jobs.py and starts the actual NLP
processing in a background thread, returning {"job_id", "total"}
immediately (HTTP 202).
pipeline.process_batch processes messages in fixed-size chunks (40 by
default, see pipeline._CHUNK_SIZE) instead of one call covering the whole
batch, and reports cumulative progress via an on_progress callback after
each chunk - this is what gives GET /ingest/status/{job_id} something to
report before the whole job is done. The chunk size is a deliberate
trade-off: large enough to keep most of V6's batching speedup, small enough
to give a handful of progress updates instead of only 0% and 100%.
Job state (status, processed, total, timestamps, and the final result
or error) lives in an in-memory dict in jobs.py - intentionally simple,
consistent with the rest of this prototype. A server restart loses in-flight
job status; the dashboard treats an unknown job_id (HTTP 404) as an error
rather than hanging forever.
The dashboard polls /ingest/status/{job_id} once per second and updates a
<progress> bar (processed / total) plus a locally-ticking elapsed-time
timer (updated every 100ms from the browser's own clock, so it stays smooth
between polls). When the job reaches done or error, polling stops and
the final summary/error is shown, same as the old synchronous response.
Data protection
The original message text is never stored in readable form. Only a SHA-256 hash (for deduplication) and the anonymized text are persisted. Because the transformation is reversible in principle and re-identification could occur with additional information, the approach is pseudonymization under the GDPR (Art. 4(5)); stored data therefore remains personal data and is handled with data minimization in mind.
Local run
pip install -r requirements.txt
python -m spacy download en_core_web_lg
uvicorn main:app --reload --port 8000
requirements-dev.txt holds extra dependencies (pandas, seqeval) needed
only by the offline helper scripts prepare_dataset.py and evaluate.py -
not required to run the API itself.