File size: 9,709 Bytes
29bf0f0
f390b04
29bf0f0
f390b04
8dd6e85
 
 
f390b04
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2e3f5aa
caf4ed9
 
 
 
f390b04
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2e3f5aa
f390b04
 
caf4ed9
 
f390b04
caf4ed9
 
f390b04
 
 
 
 
 
 
 
 
 
 
caf4ed9
 
 
 
f390b04
 
2e3f5aa
caf4ed9
 
 
2e3f5aa
 
caf4ed9
 
 
 
 
 
 
 
 
 
 
 
 
 
2e3f5aa
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
caf4ed9
 
 
 
2e3f5aa
 
 
 
 
 
 
 
caf4ed9
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f390b04
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
---
title: Conversation Data Extraction System
sdk: docker
emoji: 🗂️
colorFrom: indigo
colorTo: purple
app_port: 8000
---

# Conversation Data Extraction System

A system that extracts structured information from chat conversations:
**named entities (NER), topics and sentiment**, then **pseudonymizes** personal
data (GDPR) and stores the structured result.

## Project versions

This repository is developed incrementally as part of a diploma thesis. Each
version is tagged so the progression is visible in the git history.

| Version | Tag | Description |
|---------|-----|--------------|
| V1 | `v1.0` | Single fine-tuned RoBERTa model, `/predict` endpoint, simple SQLite logging with regex anonymization. |
| V2 | `v2.0` | Full extraction pipeline: batch `/ingest`, NER, zero-shot topic classification, Presidio-based anonymization, dashboard with stored records and statistics. |
| V3 | `v3.0` | Dynamic model selection from the Hugging Face Hub (sentiment, NER and topic classification) at request time, plus XML export of stored records. |
| V4 | `v4.0` | Export of stored records in a choice of formats (XML, JSON, CSV) via a single endpoint, picked from a dropdown button on the dashboard. |
| V5 | `v5.0` | Per-step model choice (sentiment, NER, topics) switched from radio buttons to a `<select>` dropdown; a free-text field for a custom model appears only when "Custom model" is selected. |
| V6 | `v6.0` | Batched inference for `/ingest`: NER, topic classification and sentiment analysis each run once over the whole batch instead of once per message, significantly reducing total processing time for large batches (especially on CPU-only hosting). |
| V7 | `v7.0` | `/ingest` runs as a background job instead of one long blocking request; the dashboard polls job status and shows a progress bar and elapsed-time timer. |

See [`docs/class_diagram.md`](docs/class_diagram.md) for the current architecture.

## Pipeline

```
ingest → NER → topic classification → sentiment → anonymization → storage
```

| Step | Method / model |
|------|----------------|
| NER | `dslim/bert-base-NER` (inference) |
| Topics | `facebook/bart-large-mnli` (zero-shot) |
| Sentiment | fine-tuned RoBERTa (`vojmahdal/roberta-sentiment-3labels`) |
| Anonymization | Microsoft Presidio (spaCy NER + regex), regex fallback |
| Storage | SQLite (hash of original + anonymized text) |

## Endpoints

| Method | Path | Description |
|--------|------|-------------|
| GET | `/` | Web dashboard |
| GET | `/health` | Service + model status |
| GET | `/models` | Default, suggested and currently cached models, per pipeline step |
| POST | `/analyze` | Full pipeline on a single message |
| POST | `/predict` | Sentiment only (backward compatible) |
| POST | `/ingest` | Start a background batch-ingest job for a CSV/JSON file, returns a `job_id` |
| GET | `/ingest/status/{job_id}` | Progress (and, once done, the result) of a background ingest job |
| GET | `/records` | Recent stored (anonymized) records |
| GET | `/records/export` | Stored records exported as XML, JSON or CSV (`?format=`) |
| GET | `/stats` | Aggregate statistics (includes the list of supported export formats) |

### Example

```bash
curl -X POST https://<space-url>/analyze \
  -H "Content-Type: application/json" \
  -d '{"text": "Hi, John Smith here, my order never arrived. Email john@example.com"}'
```

```bash
curl -X POST https://<space-url>/ingest -F "file=@sample_chats.csv"
# -> {"job_id": "...", "total": 8}
curl https://<space-url>/ingest/status/<job_id>
# -> {"status": "running", "processed": 4, "total": 8, "elapsed_seconds": 3.2}
# poll again once "status" is "done" (or "error") for the final result
```

```bash
curl "https://<space-url>/records/export?format=xml" -o records.xml
curl "https://<space-url>/records/export?format=json" -o records.json
curl "https://<space-url>/records/export?format=csv" -o records.csv
```

## Exporting stored records

Since V4, `GET /records/export` accepts a `format` query parameter (`xml`,
`json` or `csv`, default `xml`) and an optional `limit`. All three formats
share the same underlying data (`db._fetch_export_rows`); adding a new
format only requires one small function in `db.py` plus an entry in
`db.EXPORT_FORMATS` - `main.py` and the dashboard pick it up automatically.
CSV flattens the nested entity list into a single `"TYPE:text; ..."` cell
per record and is written with a UTF-8 BOM so it opens correctly in Excel.

On the dashboard (`/static/records.html`), an **Export ▾** dropdown button
lists the formats returned by `GET /stats` (`export_formats`); picking one
downloads the file via `Content-Disposition: attachment`.

## Choosing models from the Hugging Face Hub

Since V3, each of the three ML steps can use a different Hugging Face Hub
model, selected per request:

| Field | Task | Default |
|-------|------|---------|
| `sentiment_model` | `sentiment-analysis` | `vojmahdal/roberta-sentiment-3labels` |
| `ner_model` | `token-classification` | `dslim/bert-base-NER` |
| `topic_model` | `zero-shot-classification` | `facebook/bart-large-mnli` |

All three are accepted by `/analyze` (JSON body) and `/ingest` (form
fields); `/predict` only accepts `sentiment_model` since it is
sentiment-only. If a field is omitted, that step's default model is used.
Requested models are downloaded and cached in memory on first use
(`processors/model_registry.py`), namespaced by task with a small FIFO
cache (6 pipelines) to bound memory usage. `GET /models` lists, per step,
the default model, a few suggested models, and which ones are currently
cached.

The web dashboard exposes this as a **dropdown (`<select>`)** for each step,
listing the suggested models plus a "Custom model" option. The free-text
field for typing any other Hugging Face repo id is hidden by default and
only appears once "Custom model" is selected in the dropdown.

**Security note:** loaded pipelines never use `trust_remote_code=True`, so an
arbitrary/untrusted model id supplied by a caller cannot execute custom
Python code inside the server process - it is limited to standard
`transformers` inference for the given task. An invalid or incompatible
model id results in a clean HTTP 400 response instead of crashing the
server.

## Batched inference for large `/ingest` batches

Since V6, `pipeline.process_batch` (used only by `/ingest`) no longer loops
over `process_message` once per message. Instead it calls each processor's
batched function - `ner.extract_entities_batch`, `topics.classify_topic_batch`,
`sentiment.analyze_sentiment_batch` - once for the whole list of texts, and
reassembles the per-message results afterwards. Each batched function passes
`batch_size=16` to the underlying `transformers` pipeline so the model
itself processes several texts per forward pass instead of one at a time.

This matters most for topic classification: zero-shot classification scores
every text against every candidate label as a separate NLI pass through
`facebook/bart-large-mnli` (~400M parameters), so for the default 8 labels
that's 8 passes per message. Batching those calls together is the difference
between a 200-message `/ingest` request taking minutes rather than tens of
minutes on CPU-only hosting (e.g. the free tier of Hugging Face Spaces).
`process_message` (used by `/analyze` and `/predict`, always a single
message) is unchanged.

## Background ingest jobs (progress bar + timer)

Since V7, `POST /ingest` no longer blocks until the whole batch is
processed. It parses/validates the uploaded file synchronously (fast, no
model calls), registers a job via `jobs.py` and starts the actual NLP
processing in a background thread, returning `{"job_id", "total"}`
immediately (HTTP 202).

`pipeline.process_batch` processes messages in fixed-size chunks (40 by
default, see `pipeline._CHUNK_SIZE`) instead of one call covering the whole
batch, and reports cumulative progress via an `on_progress` callback after
each chunk - this is what gives `GET /ingest/status/{job_id}` something to
report before the whole job is done. The chunk size is a deliberate
trade-off: large enough to keep most of V6's batching speedup, small enough
to give a handful of progress updates instead of only 0% and 100%.

Job state (`status`, `processed`, `total`, timestamps, and the final result
or error) lives in an in-memory dict in `jobs.py` - intentionally simple,
consistent with the rest of this prototype. A server restart loses in-flight
job status; the dashboard treats an unknown `job_id` (HTTP 404) as an error
rather than hanging forever.

The dashboard polls `/ingest/status/{job_id}` once per second and updates a
`<progress>` bar (`processed` / `total`) plus a locally-ticking elapsed-time
timer (updated every 100ms from the browser's own clock, so it stays smooth
between polls). When the job reaches `done` or `error`, polling stops and
the final summary/error is shown, same as the old synchronous response.

## Data protection

The original message text is **never stored in readable form**. Only a SHA-256
hash (for deduplication) and the anonymized text are persisted. Because the
transformation is reversible in principle and re-identification could occur with
additional information, the approach is **pseudonymization** under the GDPR
(Art. 4(5)); stored data therefore remains personal data and is handled with
data minimization in mind.

## Local run

```bash
pip install -r requirements.txt
python -m spacy download en_core_web_lg
uvicorn main:app --reload --port 8000
```

`requirements-dev.txt` holds extra dependencies (`pandas`, `seqeval`) needed
only by the offline helper scripts `prepare_dataset.py` and `evaluate.py` -
not required to run the API itself.