Spaces:
Sleeping
Sleeping
File size: 9,709 Bytes
29bf0f0 f390b04 29bf0f0 f390b04 8dd6e85 f390b04 2e3f5aa caf4ed9 f390b04 2e3f5aa f390b04 caf4ed9 f390b04 caf4ed9 f390b04 caf4ed9 f390b04 2e3f5aa caf4ed9 2e3f5aa caf4ed9 2e3f5aa caf4ed9 2e3f5aa caf4ed9 f390b04 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 | ---
title: Conversation Data Extraction System
sdk: docker
emoji: 🗂️
colorFrom: indigo
colorTo: purple
app_port: 8000
---
# Conversation Data Extraction System
A system that extracts structured information from chat conversations:
**named entities (NER), topics and sentiment**, then **pseudonymizes** personal
data (GDPR) and stores the structured result.
## Project versions
This repository is developed incrementally as part of a diploma thesis. Each
version is tagged so the progression is visible in the git history.
| Version | Tag | Description |
|---------|-----|--------------|
| V1 | `v1.0` | Single fine-tuned RoBERTa model, `/predict` endpoint, simple SQLite logging with regex anonymization. |
| V2 | `v2.0` | Full extraction pipeline: batch `/ingest`, NER, zero-shot topic classification, Presidio-based anonymization, dashboard with stored records and statistics. |
| V3 | `v3.0` | Dynamic model selection from the Hugging Face Hub (sentiment, NER and topic classification) at request time, plus XML export of stored records. |
| V4 | `v4.0` | Export of stored records in a choice of formats (XML, JSON, CSV) via a single endpoint, picked from a dropdown button on the dashboard. |
| V5 | `v5.0` | Per-step model choice (sentiment, NER, topics) switched from radio buttons to a `<select>` dropdown; a free-text field for a custom model appears only when "Custom model" is selected. |
| V6 | `v6.0` | Batched inference for `/ingest`: NER, topic classification and sentiment analysis each run once over the whole batch instead of once per message, significantly reducing total processing time for large batches (especially on CPU-only hosting). |
| V7 | `v7.0` | `/ingest` runs as a background job instead of one long blocking request; the dashboard polls job status and shows a progress bar and elapsed-time timer. |
See [`docs/class_diagram.md`](docs/class_diagram.md) for the current architecture.
## Pipeline
```
ingest → NER → topic classification → sentiment → anonymization → storage
```
| Step | Method / model |
|------|----------------|
| NER | `dslim/bert-base-NER` (inference) |
| Topics | `facebook/bart-large-mnli` (zero-shot) |
| Sentiment | fine-tuned RoBERTa (`vojmahdal/roberta-sentiment-3labels`) |
| Anonymization | Microsoft Presidio (spaCy NER + regex), regex fallback |
| Storage | SQLite (hash of original + anonymized text) |
## Endpoints
| Method | Path | Description |
|--------|------|-------------|
| GET | `/` | Web dashboard |
| GET | `/health` | Service + model status |
| GET | `/models` | Default, suggested and currently cached models, per pipeline step |
| POST | `/analyze` | Full pipeline on a single message |
| POST | `/predict` | Sentiment only (backward compatible) |
| POST | `/ingest` | Start a background batch-ingest job for a CSV/JSON file, returns a `job_id` |
| GET | `/ingest/status/{job_id}` | Progress (and, once done, the result) of a background ingest job |
| GET | `/records` | Recent stored (anonymized) records |
| GET | `/records/export` | Stored records exported as XML, JSON or CSV (`?format=`) |
| GET | `/stats` | Aggregate statistics (includes the list of supported export formats) |
### Example
```bash
curl -X POST https://<space-url>/analyze \
-H "Content-Type: application/json" \
-d '{"text": "Hi, John Smith here, my order never arrived. Email john@example.com"}'
```
```bash
curl -X POST https://<space-url>/ingest -F "file=@sample_chats.csv"
# -> {"job_id": "...", "total": 8}
curl https://<space-url>/ingest/status/<job_id>
# -> {"status": "running", "processed": 4, "total": 8, "elapsed_seconds": 3.2}
# poll again once "status" is "done" (or "error") for the final result
```
```bash
curl "https://<space-url>/records/export?format=xml" -o records.xml
curl "https://<space-url>/records/export?format=json" -o records.json
curl "https://<space-url>/records/export?format=csv" -o records.csv
```
## Exporting stored records
Since V4, `GET /records/export` accepts a `format` query parameter (`xml`,
`json` or `csv`, default `xml`) and an optional `limit`. All three formats
share the same underlying data (`db._fetch_export_rows`); adding a new
format only requires one small function in `db.py` plus an entry in
`db.EXPORT_FORMATS` - `main.py` and the dashboard pick it up automatically.
CSV flattens the nested entity list into a single `"TYPE:text; ..."` cell
per record and is written with a UTF-8 BOM so it opens correctly in Excel.
On the dashboard (`/static/records.html`), an **Export ▾** dropdown button
lists the formats returned by `GET /stats` (`export_formats`); picking one
downloads the file via `Content-Disposition: attachment`.
## Choosing models from the Hugging Face Hub
Since V3, each of the three ML steps can use a different Hugging Face Hub
model, selected per request:
| Field | Task | Default |
|-------|------|---------|
| `sentiment_model` | `sentiment-analysis` | `vojmahdal/roberta-sentiment-3labels` |
| `ner_model` | `token-classification` | `dslim/bert-base-NER` |
| `topic_model` | `zero-shot-classification` | `facebook/bart-large-mnli` |
All three are accepted by `/analyze` (JSON body) and `/ingest` (form
fields); `/predict` only accepts `sentiment_model` since it is
sentiment-only. If a field is omitted, that step's default model is used.
Requested models are downloaded and cached in memory on first use
(`processors/model_registry.py`), namespaced by task with a small FIFO
cache (6 pipelines) to bound memory usage. `GET /models` lists, per step,
the default model, a few suggested models, and which ones are currently
cached.
The web dashboard exposes this as a **dropdown (`<select>`)** for each step,
listing the suggested models plus a "Custom model" option. The free-text
field for typing any other Hugging Face repo id is hidden by default and
only appears once "Custom model" is selected in the dropdown.
**Security note:** loaded pipelines never use `trust_remote_code=True`, so an
arbitrary/untrusted model id supplied by a caller cannot execute custom
Python code inside the server process - it is limited to standard
`transformers` inference for the given task. An invalid or incompatible
model id results in a clean HTTP 400 response instead of crashing the
server.
## Batched inference for large `/ingest` batches
Since V6, `pipeline.process_batch` (used only by `/ingest`) no longer loops
over `process_message` once per message. Instead it calls each processor's
batched function - `ner.extract_entities_batch`, `topics.classify_topic_batch`,
`sentiment.analyze_sentiment_batch` - once for the whole list of texts, and
reassembles the per-message results afterwards. Each batched function passes
`batch_size=16` to the underlying `transformers` pipeline so the model
itself processes several texts per forward pass instead of one at a time.
This matters most for topic classification: zero-shot classification scores
every text against every candidate label as a separate NLI pass through
`facebook/bart-large-mnli` (~400M parameters), so for the default 8 labels
that's 8 passes per message. Batching those calls together is the difference
between a 200-message `/ingest` request taking minutes rather than tens of
minutes on CPU-only hosting (e.g. the free tier of Hugging Face Spaces).
`process_message` (used by `/analyze` and `/predict`, always a single
message) is unchanged.
## Background ingest jobs (progress bar + timer)
Since V7, `POST /ingest` no longer blocks until the whole batch is
processed. It parses/validates the uploaded file synchronously (fast, no
model calls), registers a job via `jobs.py` and starts the actual NLP
processing in a background thread, returning `{"job_id", "total"}`
immediately (HTTP 202).
`pipeline.process_batch` processes messages in fixed-size chunks (40 by
default, see `pipeline._CHUNK_SIZE`) instead of one call covering the whole
batch, and reports cumulative progress via an `on_progress` callback after
each chunk - this is what gives `GET /ingest/status/{job_id}` something to
report before the whole job is done. The chunk size is a deliberate
trade-off: large enough to keep most of V6's batching speedup, small enough
to give a handful of progress updates instead of only 0% and 100%.
Job state (`status`, `processed`, `total`, timestamps, and the final result
or error) lives in an in-memory dict in `jobs.py` - intentionally simple,
consistent with the rest of this prototype. A server restart loses in-flight
job status; the dashboard treats an unknown `job_id` (HTTP 404) as an error
rather than hanging forever.
The dashboard polls `/ingest/status/{job_id}` once per second and updates a
`<progress>` bar (`processed` / `total`) plus a locally-ticking elapsed-time
timer (updated every 100ms from the browser's own clock, so it stays smooth
between polls). When the job reaches `done` or `error`, polling stops and
the final summary/error is shown, same as the old synchronous response.
## Data protection
The original message text is **never stored in readable form**. Only a SHA-256
hash (for deduplication) and the anonymized text are persisted. Because the
transformation is reversible in principle and re-identification could occur with
additional information, the approach is **pseudonymization** under the GDPR
(Art. 4(5)); stored data therefore remains personal data and is handled with
data minimization in mind.
## Local run
```bash
pip install -r requirements.txt
python -m spacy download en_core_web_lg
uvicorn main:app --reload --port 8000
```
`requirements-dev.txt` holds extra dependencies (`pandas`, `seqeval`) needed
only by the offline helper scripts `prepare_dataset.py` and `evaluate.py` -
not required to run the API itself.
|