File size: 13,609 Bytes
732b14f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
be9fd4a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
732b14f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
# Data Privacy — What We Have In Place Right Now

**For:** Behrang
**From:** Huzaifa
**Date:** 14 May 2026
**Companion document:** `DATA_SECURITY.md` (the full technical version)
**This document:** The short, direct answers to your questions, in your words.

**Accuracy note (May 2026):** Claims about **OpenAI** (training, retention, ZDR) depend on **OpenAI’s published terms and policies at the time** — always confirm the live pages before customer commitments. In-app behaviour is verified against this repository’s code.

---

## Your questions, answered yes/no

| Your question | Yes / No / Partial | One-line answer |
|---|---|---|
| **Do we have anything in place at all?** | **YES** | Multiple layers. Listed below in §1. |
| **Is each customer's data private to them?** | **YES** | Hard-enforced **per-tenant** isolation in the API, database, file paths, and vector retrieval. §2. |
| **Are they in their "enclosed space"?** | **PARTIAL** | Customer **content** (files, DB, FAISS index) stays on infrastructure they control. **OpenAI’s API** receives prompts, retrieved text, and (when enabled) images for model inference. Third-party **libraries** may also use HTTPS (e.g. first-time download of the local embedding model). §3. |
| **Can OpenAI use our customers' data to train AI?** | **NO** *(per OpenAI’s public API / enterprise privacy statements)* | OpenAI documents that API inputs/outputs are not used for training by default — confirm current wording on their site. §4. |
| **Does OpenAI keep our customers' data forever?** | **NO** *(per OpenAI’s public retention statements)* | Public docs describe time-limited retention for API logs unless legally required otherwise; **Zero Data Retention** is a separate contractual option. §5. |
| **Can we tell customers their data is private?** | **YES** *(with honest scope)* | Use the scripts in §7; distinguish **tenant vs tenant** (enforced in code) from **user vs user inside one tenant** (see §6). |
| **Can a junior surveyor at Firm A see reports at Firm B?** | **NO — not via normal API paths** | Different `tenant_id`s; requests are scoped to the authenticated tenant and vector search drops non-matching chunks. §2. |
| **Can a junior surveyor at Firm A see another colleague's reports at Firm A?** | **Currently yes** | Per-user permissions inside a firm are **not yet** built. This is a known gap. §6. |

---

## §1a — Personalised report style (RAG, not fine-tuning)

Users are **not** training OpenAI’s base models. They upload past reports; we extract text, **redact PII/confidential fields**, embed chunks into a **tenant-private FAISS index**, and at generate time retrieve **only that tenant’s** chunks to steer tone and wording (plus optional KB fallback before any uploads).

| Step | What happens in this app |
|---|---|
| 1. Upload previous reports | `POST /upload` (PDF/DOCX) → disk under `upload_dir/{tenant_id}/` |
| 2. Extract text | `app/ingest/pipeline.py` (Docx/PDF loaders) |
| 3. Anonymise / redact | `app/services/document_sanitiser.py` before indexing (regex by default; optional LLM chunks in dev) |
| 4. Private vector store | FAISS (or Qdrant) with `tenant_id` on every chunk |
| 5. New report generation | `resolve_retrieval_doc_allowlist` → full tenant library (`app/services/personalised_rag.py`); `retrieve_for_report` ranks attached docs first |
| 6. Style + examples in prompts | `_get_or_build_style_profile` + RAG passages in `app/generator/adapter.py` (no KB once uploads exist) |

Controlled by **`PERSONALISED_STYLE_RAG_ENABLED=true`** (default). Set **`strict_uploaded_only=true`** on `/generate` to limit retrieval to documents attached to one report only.

**Scope:** isolation is **per tenant (firm)**, not per individual surveyor inside a firm — see §6.

**On-disk originals:** sanitisation applies to **indexed** text; uploaded files on disk are **not** rewritten. Firms that need disk-level redaction should add their own pre-upload process or a future feature.

---

## §1 — What we actually have in place today

This reflects **this repository’s application code** (not every possible deployment wrapper such as cloud logging or reverse proxies).

### Data that stays on the firm’s infrastructure (under normal deployment)

1. **Uploaded reports** (PDF/DOCX) → written to configured local disk (`upload_dir`, default under the service user’s home directory). Not uploaded to a separate SaaS vector DB by this app.
2. **Inspector notes** (bullets) → stored with the report in the **SQL** database (SQLite by default; Postgres supported via `DATABASE_URL`).
3. **Vector index****FAISS** on disk (`faiss_index_path`, default under `~/.report_genius/`).
4. **Embeddings (default)****`prefer_local_embeddings=True`** uses **Hugging Face sentence-transformers** (`local_embedding_model`, default `all-MiniLM-L6-v2`) on the **same machine** as the app. **Exception:** if that backend cannot be loaded and an **OpenAI API key** is set, the factory **falls back to OpenAI embeddings**; operators can also set **`prefer_local_embeddings=False`** to prefer OpenAI. See `app/embeddings/factory.py` and `app/config.py`.
5. **Database** → not replicated to a third party by the application itself; backup/replication is the **deployer’s** responsibility.

### What leaves the customer environment (product-relevant)

1. **OpenAI HTTPS API** — used for section generation, proofreading/enhancement, notes expansion (when a key is set), style analysis, agentic inspector loop, risk assessment helper, optional vision/photo analysis, optional LLM grounding, canonical rollout LLM merge, etc. Payloads include prompts and **retrieved or user-supplied text** (and **images** when vision/photo features are used).
2. **Dependency network use (not “our” `requests` code)** — e.g. **downloading** the Hugging Face embedding model weights on first use, **OpenAI SDK** traffic, and any other libraries your environment pulls at runtime. Our `app/` tree does **not** use `requests` / `httpx` / `urllib.request` directly for custom integrations; that does **not** prove “no other HTTPS on the machine.”

**Do not claim:** “Zero outbound network except OpenAI.” **Do claim:** “We do not ship separate analytics or vector-Database SaaS integrations in this codebase; the **designed** model-data path is local storage + OpenAI inference, with an optional OpenAI embedding path if configured.”

### Contractual / vendor protections (OpenAI — verify live)

6. **OpenAI Business / API terms and DPA** — relationship and processor terms are defined by **OpenAI’s** published agreements (links in `DATA_SECURITY.md` §11).
7. **Training opt-out for API data** — described in OpenAI’s enterprise/privacy documentation; **re-read the live page** before sales use.
8. **Retention** — public docs describe limited retention for abuse/safety; exact periods and ZDR eligibility are **contract- and product-specific**.
9. **TLS** — HTTPS to `api.openai.com` (and to other hosts used by installed libraries).

---

## §2 — "Is each firm's data private and isolated?"

**YES for tenant-to-tenant isolation.** The API attaches a **`tenant_id`** to work; ORM queries and file/photo paths are scoped; FAISS metadata must match the request tenant.

Concrete example — vector retrieval (`app/vectorstore/faiss_wrapper.py`; line numbers can shift between commits):

```python
if meta.get("tenant_id") != tenant_id:
    continue
```

Every FAISS hit with a mismatched `tenant_id` is skipped before it becomes a search result.

**Do not use a made-up count** (“19 places / 10 files”). The right statement: **`tenant_id` appears throughout `app/` (API, ingestion, cache, generation, vectorstore, etc.)** — we can show representative files on request.

**Customer-facing (tenant vs tenant):**

> "Two surveying firms are isolated by **tenant**: API and retrieval paths are built so one tenant cannot read another tenant’s documents or index entries. That is enforced in code, not only in policy."

---

## §3 — "Are they in their enclosed space?"

**Honest answer:** **Customer-controlled storage** (uploads, DB, FAISS files, default local embeddings) stays on **their** servers or VPC **when they host the stack**. **Content-bearing requests** for LLM features are sent to **OpenAI** (and optionally use **OpenAI embeddings**) under the configuration described in §1. **Images** used for vision go to the same class of API when those features are enabled.

**Suggested wording:**

> "Your reports, notes, database, and search index remain on infrastructure you control when you deploy our application there. For AI drafting and related features, the application calls **OpenAI’s API** over HTTPS with the data needed for that request (text, and images if you use photo/vision features). OpenAI’s **training and retention rules** are defined in their published terms — we cite those links in our technical pack and keep them updated."

---

## §4 — "Can OpenAI use our customers' data to train AI?"

**Per OpenAI’s public documentation** (wording can change — verify before webinars):

> "By default, OpenAI does not use data from … **their API platform** — including inputs or outputs — for training or improving their models."
>
> — [openai.com/enterprise-privacy](https://openai.com/enterprise-privacy/) (confirm live text)

**Do not paraphrase as “we verified in code that OpenAI never trains.”** That is a **vendor** commitment.

---

## §5 — "Does OpenAI keep our customers' data forever?"

**Per public docs**, API logs are not kept indefinitely; exact retention and exceptions are in OpenAI’s materials. **Zero Data Retention (ZDR)** is a **separate programme** (application/approval, endpoint restrictions) — see [platform.openai.com/docs/guides/your-data](https://platform.openai.com/docs/guides/your-data).

**Recommendation:** pursue ZDR if legal/commercial teams want the strongest **contractual** minimisation of retention — then quote **only** what OpenAI confirms in writing for your org.

---

## §6 — What we DON'T yet have (being honest)

| Gap | Severity | Notes |
|---|---|---|
| **Zero Data Retention** | Ops / legal | Apply and validate against actual endpoints used. |
| **Azure OpenAI** | Product | Not implemented in this repo today; roadmap item if customers require UK residency on Microsoft’s stack. |
| **Self-hosted LLM** | Product | Not implemented; possible future `local_llm_base_url`–style integration. |
| **Per-user ACL inside one tenant** | Privacy | Same `tenant_id` can still see all firm reports until row-level or chunk-level user scoping exists. |
| **Public security page** | Marketing | Still worth publishing; base it on `DATA_SECURITY.md` after edits. |
| **Photo / vision payloads** | Disclosure | `app/services/photo_vision.py` and `app/generator/vision_analyzer.py` send images to the configured OpenAI-compatible vision API when enabled. |

---

## §7 — What you can tell customers (fact-checked scripts)

Use **§1–§3** above as the source of truth for “what stays local” vs “what goes to OpenAI.” For **training/retention**, read OpenAI’s **current** page aloud or quote verbatim.

### The 30-second pitch (code-accurate)

> "When you host our application on your infrastructure, your uploaded reports, your database, and your semantic index stay with you. We use a **local embedding model by default** so vectorisation normally stays on the same server; you can change that in configuration. For AI writing and analysis, the app calls **OpenAI** over HTTPS with the minimum context needed for each request—including photos if you turn on vision. **Between firms**, the product isolates data by **tenant** in the API and in vector search. **Inside one firm**, finer-grained per-user locks are still on our roadmap."

### When they ask "where is my data?"

1. **On your deployment** — files, SQLite/Postgres, FAISS files, default local embeddings.  
2. **With OpenAI for model calls** — prompts, snippets, and optional images, under their terms.  
3. **Elsewhere** — this codebase does not add its own analytics export; your **hosting provider’s** logs and backups are outside this app’s control—disclose them in your deployment runbook.

### When they ask "what about Firm B?"

> "Another customer’s tenant id cannot read your tenant’s rows or vector metadata; FAISS results are filtered by tenant. That does not replace your own **access control and IAM** on who can log in as your tenant."

---

## §8 — Friday follow-ups

Align engineering tasks with **`DATA_SECURITY.md`** (implementation checklist and risk register), after removing references to **non-existent** modules in that file.

---

## §9 — Two documents

| Document | Audience |
|---|---|
| **`BOSS_BRIEF_DATA_POLICY.md`** | Short internal brief / sales prep |
| **`DATA_SECURITY.md`** | Engineers — technical detail and sources |

---

## §10 — After improvements (truthful sequencing)

- **Today:** Local-first storage and indexing; OpenAI for LLM (and optionally embeddings); tenant isolation in code.  
- **After ZDR (if approved):** Stronger **contractual** retention story per OpenAI’s ZDR programme.  
- **After Azure / self-hosted work (if built):** Additional deployment options—describe only once shipped.

---

*Code pointer for tenant filtering in FAISS: `app/vectorstore/faiss_wrapper.py` (search for `tenant_id`). OpenAI policy links: see `DATA_SECURITY.md` §11 — refresh before customer-facing use.*