File size: 45,548 Bytes
732b14f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
# Data Security & Privacy — RICS Report Generator

**Audience:** Internal engineering (DevOps + backend), founder for client conversations and webinars.
**Status:** Working document. Decisions outstanding for the Friday meeting are flagged with **[DECIDE]**.
**Last reviewed:** 14 May 2026

### How to read this document

| Term | Meaning here |
|---|---|
| **This repository / this app** | The FastAPI Python service under `app/` in this git repo—not the customer’s VPN, cloud control plane, or desktop browsers unless we say so. |
| **Customer / firm** | The organisation that deploys the software and owns the disk and database. |
| **Tenant** | Logical partition in our data model (`tenant_id`). Isolation between tenants is enforced in API and FAISS code paths; **per-user ACL inside one tenant is not implemented** (see §9). |
| **Tier A** | Direct OpenAI API with standard commercial logging/retention as published by OpenAI at the time. **Matches shipped code today** when `OPENAI_API_KEY` is set. |
| **Tier B** | Tier A plus **Zero Data Retention** on the OpenAI **organisation** (OpenAI programme + endpoint rules). **Not a separate code path yet** beyond what OpenAI enforces; optional `zdr_enforced`-style guardrails are listed in §10 and are **not** present in code as of last review. |
| **Tier C** | Customer inference against **Azure OpenAI** in the customer’s Microsoft tenant (e.g. UK South). **Roadmap: not implemented** in this repo (no Azure client or settings fields). |
| **Tier D** | Inference against a **self-hosted** OpenAI-compatible HTTP API (e.g. Ollama, vLLM). **Roadmap: not implemented** in this repo. |
| **Vendor claims** | Anything about training, retention, regions, or subprocessors is **only** as accurate as the linked page on the day you read it—this file is not legal advice. |

---

## 1. Two-paragraph summary you can use in any conversation

Surveyors trust us with property addresses, owner names, valuations, and condition findings — every report contains personal data and commercially sensitive information. Our job is to give them a clear, defensible answer to one question: **where does this data go?**

**Customer content (uploads, database, FAISS index) stays on infrastructure the customer controls when they self-host or run in their own cloud account.** **OpenAI’s HTTPS API** is used for LLM features (and **optionally for embeddings** under the exact rules in **§2.2**). Under OpenAI’s published terms, **API data is not used to train models by default**, and **retention is time-limited** (see **§3** — always confirm the live policy pages). **Other HTTPS traffic** can still occur from **dependencies** (for example, downloading HuggingFace model weights on first use); our `app/` package does not add separate `requests`/`httpx` integrations for analytics or third-party vector DBs (**§2.1** search proof). For stronger guarantees we describe **roadmap tiers**: Zero Data Retention (OpenAI programme), Azure OpenAI (not implemented in this repo yet), and self-hosted models (not implemented yet)—each needs explicit engineering and operations work beyond flipping a single flag.

---

## 2. Where data goes in this repository (code-backed)

This section answers: **which bytes leave the host because our application asked them to**, and **where we store data locally**. It is based on reading `app/` and `app/embeddings/factory.py` at **last review (14 May 2026)**. It **excludes** infrastructure you add yourself (cloud logging, APM, reverse proxies, SIEM, Windows Update, Docker registry pulls, etc.).

### 2.1 Data this app keeps on the host (no OpenAI client for these artefacts)

These features read and write customer data **only** on the deployment host using local disk + local SQL + in-process Python. They do **not** import `openai.OpenAI`, `langchain_openai.ChatOpenAI`, or `OpenAIEmbeddings` for the listed job.

| Component | Where in code | Default on-disk / DB location (from `app/config.py` unless env overrides) |
|---|---|---|
| Uploaded reports (PDF/DOCX) | `app/api/upload.py`, `app/ingest/*` | `upload_dir``~/.report_genius/uploads` |
| Inspector notes | `GenerateRequest.bullets` and related ORM models | Same database as other report fields |
| Vector index | `app/vectorstore/faiss_wrapper.py` | `faiss_index_path``~/.report_genius/faiss_index` |
| SQL database | All persistence through SQLAlchemy / settings | `database_url``sqlite+aiosqlite:///./dev.db`; production usually overrides with Postgres via `DATABASE_URL` |
| Ingestion, chunking, FAISS search, retrieval orchestration | `app/ingest/*`, `app/retrieval/*`, vectorstore | Same host; **no** `requests` / `httpx` / `urllib.request` usage in `app/**/*.py` for these code paths |

**What we proved by search:** `rg "requests\\.(get|post)|httpx\\.|urllib\\.request" app/`**zero matches**. So this repo does **not** ship first-party calls to random analytics hosts, ad networks, or hosted vector DBs.

**What that search does *not* prove:** the Python process cannot make HTTPS calls. Dependencies (`openai`, `langchain_openai`, `langchain_huggingface`, `sentence-transformers`, etc.) use their own TLS stacks.

### 2.2 Embeddings — exact decision table (`app/embeddings/factory.py`)

`get_embedding_client()` builds **one** global `Embeddings` instance using `Settings.prefer_local_embeddings` and `Settings.openai_api_key`:

| `prefer_local_embeddings` | Order tried | OpenAI receives text chunks for embedding when |
|---|---|---|
| `True` (default) | ① `HuggingFaceEmbeddings(model_name=local_embedding_model)` — default model **`all-MiniLM-L6-v2`**. ② If ① returns `None` (missing deps / import failure): `OpenAIEmbeddings(model=embedding_model)` **if** `openai_api_key` is non-empty. ③ If still nothing: `FakeEmbeddings` (test / broken install). | Only when ① failed **and** ② succeeded. |
| `False` (legacy) | ① `OpenAIEmbeddings` if key present. ② else HuggingFace. ③ else Fake. | Whenever ① succeeds with a key. |

When OpenAI embeddings run, the model name is `Settings.embedding_model`, default **`text-embedding-3-small`**.

**First-time HuggingFace use:** downloading `all-MiniLM-L6-v2` weights is **HTTPS to Hugging Face Hub (or a mirror)**. That traffic is **public model weights**, not a copy of the customer’s report archive sent as “training data” to OpenAI.

### 2.3 OpenAI chat, vision, and tools — complete production module list

Every **non-test** file under `app/` that constructs `openai.OpenAI` or `langchain_openai.ChatOpenAI` for real inference:

| File | When it runs | What OpenAI sees (plain English) |
|---|---|---|
| `app/generator/adapter.py` | `openai_api_key` non-empty and caller uses generate / proofread / enhance / weave | Prompts: section skeleton, bullets, RAG standard passages, instructions from `app/generator/prompts.py` |
| `app/generator/notes_expander.py` | Key set; LLM expansion path chosen | Raw inspector bullets |
| `app/generator/style_analyzer.py` | Key set; style profile requested | Paragraph samples from uploaded firm documents |
| `app/generator/vision_analyzer.py` | Key set; vision path not skipped | **Pixels** (encoded image) + vision instructions |
| `app/generator/postprocess.py` | Key set; LLM grounding branch taken | Draft section text + grounding snippets |
| `app/services/photo_vision.py` | Key set; batch photo analysis runs | **Pixels** + instructions |
| `app/services/canonical_rollout.py` | Key set; LLM merge helper runs | Two competing text bodies for merge |
| `app/agentic/inspector_loop.py` | Key set; agentic pipeline enabled | Full inspector loop: notes, seeds, tool JSON, partial completions |
| `app/agentic/agents.py` | Key set; JSON risk helper used | Bullets / short context block passed into that helper |

`app/services/generation.py` **orchestrates** the above; it does not construct its own parallel OpenAI client for those features.

**Model names in config (`app/config.py`):** `chat_model` default **`gpt-4o-mini`** (main adapter). `inspector_body_model` default **`gpt-4o`** (inspector loop body drafting).

**Azure:** There is **no** `AzureOpenAI` client and **no** Azure endpoint settings in this repository today. Tier C is documentation only until we ship that code.

### 2.4 Network egress diagram (application-designed paths only)

```
┌─────────────────────────────────────────────────────────────┐
│  Surveyor firm's environment                                 │
│                                                              │
│  ┌─ Web UI                                                   │
│  │                                                           │
│  ▼                                                           │
│  FastAPI ──► SQLite/Postgres                                 │
│         │                                                    │
│         ├──► File storage (PDFs, DOCX)                       │
│         │                                                    │
│         ├──► FAISS index (local disk)                        │
│         │                                                    │
│         ├──► HuggingFace MiniLM-L6-v2 (local, on CPU/GPU)    │
│         │      (first run may download weights — HTTPS)       │
│         │                                                    │
│         └──► OpenAI Python client ───────┐ TLS 1.2+          │
│                                          │ (primary LLM API)  │
└──────────────────────────────────────────│──────────────────┘


                                  api.openai.com (OpenAI-hosted)
```

**Primary egress for customer *content* designed by this app:** TLS to **OpenAI** for chat/vision/tools when `OPENAI_API_KEY` is set and those code paths run; plus **OpenAI** for embeddings **only** in the §2.2 cases. **Other TLS the app indirectly causes:** Hugging Face Hub (or a mirror) when sentence-transformers downloads `local_embedding_model` weights.

**Explicitly out of this diagram:** cloud provider metadata APIs, log shippers, antivirus updates, DNS, NTP, container registries—**your** platform team owns that surface.

---

## 3. OpenAI's published data policy (vendor — re-read before contracts)

The bullets and quotes below are **summaries or verbatim excerpts** from OpenAI’s public pages linked in **§11**, current as of **May 2026** when this file was written. **They are not legal advice** and they **will** drift—before a customer signature, open each URL and read the live text.

**Relationship to our code:** Training/retention rules are enforced by **OpenAI’s platform and your contract**, not by anything we can “prove” inside `app/`.

### 3.1 Training

> "By default, OpenAI does not use data from ChatGPT Enterprise, ChatGPT Business, ChatGPT Edu, ChatGPT for Healthcare, ChatGPT for Teachers, or their API platform — including inputs or outputs — for training or improving their models."
> — [openai.com/enterprise-privacy](https://openai.com/enterprise-privacy/)

OpenAI states this default applied from **1 March 2023** onward (API platform included in the quoted list). **Practical meaning for us:** any customer using our integration against the **OpenAI API** should verify their **exact** OpenAI product (Business vs consumer ChatGPT vs API) on the live page—this document treats **API / Business** as the relevant class for this codebase.

### 3.2 Retention

> "After 30 days, API inputs and outputs are removed from OpenAI logs, unless legally required to retain them."
> — [openai.com/enterprise-privacy](https://openai.com/enterprise-privacy/)

OpenAI describes this window as tied to **abuse monitoring** and trust-and-safety workflows. **Operational detail:** assume prompts/responses **may be processed or sampled** by automated systems and, where their policy allows, humans during the stated retention—read their current abuse-monitoring explanation rather than assuming “nobody ever sees it.”

### 3.3 Zero Data Retention (ZDR)

> "If you are a business customer that uses OpenAI's Zero Data Retention (ZDR) API, OpenAI never retains the prompts you send or the answers returned."
> — [platform.openai.com/docs/guides/your-data](https://platform.openai.com/docs/guides/your-data)

**Operational caveats (read OpenAI’s ZDR page before promising anything):**
1. **Enrolment:** ZDR is **not** automatic for every API key—it is a **programme** you apply for on the OpenAI org.
2. **Endpoint compatibility:** Some platform features (e.g. extended prompt caching, background mode, Assistants API with threads, certain file/Batch flows) are called out by OpenAI as **incompatible** or restricted with ZDR—our engineering checklist is in **§10**.
3. **Legal overrides:** Court orders, litigation holds, or government demands can still force retention **outside** normal commercial deletion timelines. ZDR narrows **contractual** logging; it does not create immunity from law.

**Why we still want ZDR:** when eligible and correctly configured, it removes the **routine 30-day log retention** described for standard API processing—again, verify the **current** ZDR description on OpenAI’s site.

### 3.4 Data Processing Addendum

The DPA text is published by OpenAI and is **legally binding only after** your organisation completes whatever acceptance / order flow OpenAI requires for **Business / API** use. This document does **not** record your signature status—legal owns that record.

The published DPA covers (high level; read the PDF):
- GDPR Article 28 processor obligations (OpenAI = Processor, our firm clients = Controller)
- Standard Contractual Clauses (SCCs) for EU transfers
- UK IDTA Addendum for UK transfers
- Sub-processor general authorisation with change-notification
- Breach notification, data-subject-request assistance
- HIPAA BAA available on Enterprise/Scale tiers (not relevant to RICS but indicates depth)

Source: [openai.com/policies/data-processing-addendum](https://openai.com/policies/data-processing-addendum/), PDF v.010126 (effective 1 Jan 2026).

### 3.5 Sub-processors (where data physically lives)

OpenAI uses Microsoft Azure as its primary infrastructure, with Google Cloud, Oracle Cloud, and CoreWeave for additional compute. Other sub-processors handle support, payments, and trust-and-safety review. The authoritative list is [openai.com/policies/sub-processor-list](https://openai.com/policies/sub-processor-list/) — we should reference it as a URL rather than copying it (the list changes).

**Material point for UK surveyors (direct OpenAI API — Tier A):** OpenAI’s **sub-processor list** shows US and other jurisdictions. A UK controller using the **public** OpenAI API therefore typically relies on **SCCs + UK IDTA** (or equivalent) for **international transfers** to OpenAI US operations—exact mechanism is in OpenAI’s DPA.

**What Tier C changes (once we ship it):** inference runs in the **customer’s Microsoft Azure subscription** in a region the customer selects (e.g. **UK South** for London). Data still goes to **Microsoft** as processor; it does **not** remove the need for a DPA—it **changes** the processor and the **geographic / network boundary** compared with calling `api.openai.com` from Tier A. Always validate region, private link, and key management against the **current** Azure OpenAI documentation before promising a map pin.

### 3.6 Encryption in transit and at rest (OpenAI direct API — Tier A)

- **Transit:** TLS 1.2+ between our Python process and OpenAI’s API hosts (per OpenAI’s security documentation).
- **At rest (on OpenAI’s side):** OpenAI documents **AES-256** for stored artefacts **during whatever retention window applies** to your account tier (standard vs ZDR). **Customer-managed keys for data at rest inside OpenAI’s direct API** are **not** a documented feature today—if a customer demands CMK for inference payloads, that is a **Tier C (Azure)** conversation.

---

## 4. The four-tier privacy ladder

| Tier | Shipped in this repository today? | What changes vs previous tier |
|---|---|---|
| **A** | **Yes** — direct OpenAI HTTPS calls when `OPENAI_API_KEY` is set | Baseline commercial API. |
| **B** | **Partially** — still Tier A code paths; ZDR is an **OpenAI org configuration** + legal paperwork, not a different Python SDK | Same HTTPS destination; **contractual** logging/retention rules change per OpenAI ZDR programme. Optional app guardrails in §10 are **not implemented** yet. |
| **C** | **No** — requires new settings + client factory + refactors listed in §10 | HTTPS goes to **customer’s Azure OpenAI endpoint** instead of `api.openai.com`. |
| **D** | **No** — requires local OpenAI-compatible HTTP endpoint + same factory refactors | HTTPS goes to **customer-controlled** inference server (or stays on loopback). |

The narrative below (pros/cons/Q&A) is for **sales and architecture planning**. **Engineering truth** is the table above.

### Tier A — Standard OpenAI API (where we are today)

**Posture (plain English):** Our Python service sends prompts to **OpenAI’s API** under whatever **OpenAI Business / API terms** you have signed. OpenAI’s public pages state that **API inputs/outputs are not used to train models by default** and describe **time-bounded retention** for logs—see **§3** for quotes and links.

**Pros:** No Azure or GPU capex for LLM inference; swap `chat_model` / `inspector_body_model` in `app/config.py` when OpenAI releases newer chat models; billing is usage-based.

**Cons:** Prompts and completions leave the firm’s network boundary and are processed in OpenAI’s multi-tenant cloud. OpenAI’s documentation and sub-processor list describe **US and other** processing locations for parts of the service—**not** “UK-only” on Tier A.

**What a surveyor's compliance team will ask (short answers — verify on live OpenAI pages):**
- Is it used to train models? → **OpenAI says “not by default” for the API platform** (see §3.1 quote).
- How long is it stored? → **OpenAI publishes ~30-day log retention for standard API processing** with carve-outs for legal holds (see §3.2 quote).
- Is there a DPA? → **OpenAI publishes a DPA for business use** — your legal team must confirm it is **executed for your org** (§3.4).
- Where physically? → **See OpenAI’s current sub-processor / infrastructure disclosures** — not a single static country list in this file.
- Can humans see prompts? → **During any retention window OpenAI defines for abuse monitoring, assume automated and (where their policy permits) human review is possible** — read their abuse-monitoring pages; do not promise “no human ever.”

**Cost:** Order-of-magnitude **$0.001–0.005 per section generation** at public `gpt-4o-mini` list pricing—recalculate from OpenAI’s **current** pricing page before quoting a customer.

### Tier B — Standard OpenAI API + Zero Data Retention + signed DPA

**Posture:** Identical **Python code paths** as Tier A. The difference is **contract + OpenAI org settings**: you must (a) have a **signed** Business DPA on file, (b) be **accepted into OpenAI’s ZDR programme** for that org, and (c) restrict product features to **ZDR-allowed** endpoints only.

**Pros:** If OpenAI’s ZDR terms apply to your org as configured, **routine long-lived log retention of prompts/completions is removed** per OpenAI’s ZDR description. Tokens are still billed the same way as Tier A.

**Cons:** You must maintain eligibility (endpoint restrictions). Legal/compliance must read OpenAI’s **current** ZDR rules—this file is not the contract.

**Implementation work (split by owner):**

| Owner | Task |
|---|---|
| **Legal / admin** | Apply for ZDR in the OpenAI org dashboard; store executed DPA + ZDR confirmation. |
| **Engineering (recommended, not done yet)** | Add `openai_zdr_enforced` (name TBD) boolean + startup validation that refuses known ZDR-incompatible SDK features (checklist §10). |
| **Customer success** | Publish which OpenAI org ID / region / model names the deployment uses so customers can mirror it in their vendor reviews. |

**Technical note:** ZDR enforcement happens **inside OpenAI’s platform** once enabled. Our repository **does not yet contain** the guardrail flag described in §10—today we rely on OpenAI’s org settings plus careful endpoint choice.

**Recommendation:** **Pursue ZDR for every production OpenAI org** that serves paying surveyors. Cost at inference time is **$0 extra**; benefit is **narrower retained footprint on OpenAI’s side** per their ZDR materials.

### Tier C — Azure OpenAI Service (roadmap — not in this repository)

**What Tier C is:** After we implement it, the Python service will call **`{customer}.openai.azure.com`** (or regional equivalent) with the customer’s **own** API key / managed identity, instead of calling **`api.openai.com`**. The **customer’s** Microsoft Customer Agreement / Online Services Terms + Microsoft’s Azure OpenAI privacy pages govern inference—not OpenAI’s **direct** consumer API terms.

**Relationship to Microsoft 365 Copilot:** Copilot is a **separate Microsoft 365 product** with its own privacy statement. Microsoft also offers **Azure OpenAI** as a platform service. **Marketing must not claim “we are Copilot.”** Accurate line: *“Same **platform family** Microsoft uses for enterprise AI in Azure; Copilot-specific data flows are documented on Microsoft’s Copilot privacy pages.”*

**Microsoft-published commitments (Azure OpenAI — read the live page; excerpt below can drift):**

> "Your prompts (inputs) and completions (outputs), your embeddings, and your training data:
> - are NOT available to other customers.
> - are NOT available to OpenAI or other Azure Direct Model providers.
> - are NOT used by Azure Direct Model providers to improve their models or services.
> - are NOT used to train any generative AI foundation models without your permission or instruction."
> — [Azure OpenAI — Data, privacy, and security](https://learn.microsoft.com/en-us/azure/ai-foundry/responsible-ai/openai/data-privacy)

**Infrastructure fact (still read Microsoft docs):** Model weights for Azure OpenAI deployments are served from **Microsoft-operated** Azure regions. Microsoft documents that this path does **not** send your prompts to the **public ChatGPT service** or to **OpenAI’s consumer endpoints** as part of that product architecture—verify wording on the page above.

**Regions (examples Microsoft lists for deployments — confirm in Azure portal before sales):** UK South (London), West Europe, North Europe, Sweden Central, France Central, Germany West Central, Switzerland North, etc. **Strict “UK-only processing”** means: create the Azure OpenAI resource in **UK South**, pick a **regional** (non-global) deployment SKU per Microsoft’s current guidance, and review networking diagrams with the customer’s cloud architect—**we do not guarantee routing** in this markdown file.

**Compliance Q&A (each answer must be checked against the linked Microsoft page the week you pitch it):**

| Question | Where to read the authoritative answer | Short pointer (not a substitute for the doc) |
|---|---|---|
| Training use of customer content? | Same “Data, privacy, and security” page (quoted block above) | Microsoft states **no training** on customer prompts/completions without permission, for the described Azure OpenAI / Azure Direct Models configuration. |
| Physical region? | Azure region picker + Microsoft region docs | Customer chooses region (e.g. **UK South**). |
| OpenAI employees viewing prompts? | Same privacy page + OpenAI’s role statement there | Microsoft states prompts/completions are **not** made available to OpenAI for the Azure OpenAI service model described. |
| Abuse monitoring / human review? | [Abuse monitoring for Azure OpenAI](https://learn.microsoft.com/en-us/azure/ai-foundry/openai/concepts/abuse-monitoring) | Microsoft may perform automated abuse detection; **Modified Abuse Monitoring** is the documented programme to reduce certain reviews—eligibility is Microsoft-gated. |
| Encryption at rest + CMK? | [Encrypt Azure OpenAI data at rest](https://learn.microsoft.com/en-us/azure/ai-foundry/openai/encrypt-data-at-rest) | Microsoft documents **AES-256** by default and **customer-managed keys** via Key Vault for supported scenarios—confirm feature flags for the exact SKU. |
| Private network? | [Network isolation / private endpoint docs](https://learn.microsoft.com/en-us/azure/ai-foundry/openai/how-to/on-your-data-configuration) | Microsoft supports **private endpoints** and disabling public access **when configured**—not automatic. |
| Contract? | Customer’s existing Microsoft agreements | Often the **Microsoft Products and Services Data Protection Addendum (DPA)** applies—legal must confirm which agreement is in force. |

**Pros (why customers ask for Tier C):** Data processing moves to **Microsoft Azure** under customer-controlled regions and optional private networking/CMK—reduces reliance on **OpenAI direct US API** subprocessors from §3.5.

**Cons:** ~**20–30% higher per-token list price** than direct OpenAI (historical ballpark—re-quote from Azure pricing). Customer must own Azure subscription + quota approvals. We must maintain a **second HTTP client** (Azure OpenAI) and regression-test every call site in §2.3.

**Implementation work (engineering backlog — not started):**
1. Add `AzureOpenAI` client module + Pydantic settings (`azure_openai_endpoint`, `azure_openai_api_version`, `azure_openai_deployment_name`, optional `azure_openai_api_key` or managed-identity flags—**exact field names to be chosen at implementation time**).
2. Central factory returns either `OpenAI` or `AzureOpenAI` compatible chat/embeddings clients.
3. Refactor every file listed in **§2.3** to consume the factory.
4. Write customer runbook: resource group, UK South OpenAI resource, private endpoint, Key Vault CMK, RBAC, logging exclusions.

**Cost model (illustrative only — recompute before contracts):** Same rough order of magnitude as earlier draft (~$25/mo inference + ~$15/mo private endpoint/Key Vault for a small tenant) — **treat numbers as Fermi estimates**, not quotes.

**This is the architecture to recommend** when a prospect’s written requirement is **“model calls must land in our Microsoft tenant / UK Azure region.”** It is **not** something they can buy from us **today** out of this repository without custom integration work.

### Tier D — Self-hosted open-weight model inside the firm's perimeter

This is what many **on-premises LLM** pitches advocate: **customer report content** is not sent to OpenAI or another cloud LLM API when inference runs entirely on infrastructure the customer controls.

**Models to consider** (as of May 2026):
| Model | Params | Notes |
|---|---|---|
| Google Gemma 3 27B / 4 26B (when GA) | 27B / 26B | Strong general reasoning, instruction-tuned, permissive Gemma licence |
| Meta Llama 3.3 70B | 70B | Excellent quality, Llama Community licence (free for <700M MAU) |
| Mistral Large 2 / Codestral | varies | EU-headquartered vendor, Mistral Research Licence |
| Qwen 2.5 72B / Qwen 3 | 72B / varies | Strong Chinese-vendor model, Apache 2.0 for most variants |
| Phi-3 / Phi-4 (small) | 3.8B / 14B | Microsoft, MIT licence, fast |

For RICS report generation specifically, **the model needs to handle ~20-30k tokens of context** (uploaded report excerpts + notes + schema). Gemma 3 27B, Llama 3.3 70B, and Qwen 2.5 72B all handle this comfortably with 128k context windows.

**Target posture:** No cloud **LLM vendor** receives customer prompts or completions. Model weights are local; inference is a local process. **Operational reality:** OS updates, package mirrors, NTP, monitoring agents, or downloading weights can still create network traffic unless you engineer an air-gapped image—that is **deployment** work, not something this application disables by itself.

**What a surveyor's compliance team will ask:**
- Where does inference run? → **On VMs or bare-metal the customer pays for**, inside their network boundary (or a colo they contract with).
- Is it used to train anything? → **Not in the SaaS sense** — open-weight files are static; **no** gradient updates to the base model happen during normal chat inference. (If you fine-tune on customer data, that is a **separate** data decision.)
- Who can see prompts? → **Anyone with OS / hypervisor / backup access to the host** can, in principle, read RAM or disk—same as any sensitive workload. **No** separate OpenAI/Microsoft LLM API receives the text **when** you truly keep calls on-loopback to your own server.
- Encryption? → **Full-disk encryption, TLS inside the cluster, and key management** are the customer’s controls—same as running Postgres or HR systems on-prem.

**Pros:**
- **Sales-accurate headline:** Customer report text is **not** sent to OpenAI’s public API **for inference** when every LLM call stays on an endpoint the customer controls.
- **Cost shape:** Predictable GPU hire or amortised hardware vs per-token cloud bills—magnitude depends on concurrency; **do not** treat viral “10× cheaper” posts as benchmarks without your own measurement run.
- **Vendor surface:** No OpenAI/Azure LLM **subscription** for inference; you still owe diligence on the model **licence**, security patches, and supply chain of the weight files.

**Cons:**
- Significant infrastructure investment: a single NVIDIA H100 or two A100s per tenant for a 27B model at INT4 quantisation. ~£25k–£50k capex per host, or ~£3–6/hour on cloud GPU rental.
- We become the MLOps team — model patching, scaling, monitoring, drift detection, evaluation runs.
- Quality risk: frontier cloud models still win on **some** long reasoning and edge-case refusal tests. For **RICS section drafting with heavy RAG**, the practical gap must be measured with **your** prompts—run a blinded QA review before promising parity.
- Cold-start onboarding: each firm needs hardware procurement or a managed deployment.

**Hybrid Tier D variant — managed-but-isolated (commercial concept, not shipped):** A hosting provider runs **dedicated** GPU + DB + vector index **per customer** so prompts never cross tenant boundaries. Vendors such as RunPod, Together, or Modal sell variations of this **isolated tenancy** idea. **This repository does not include** Kubernetes manifests, per-tenant GPU schedulers, or billing for that model—it is listed here only so sales knows the category exists.

**Implementation work:**
1. Add a **`LocalLLM` client path** (thin wrapper around an OpenAI-compatible HTTP API such as **Ollama** or **vLLM**).
2. Add config: `local_llm_endpoint`, `local_llm_model_name`, `local_llm_api_key` (if required).
3. Use the same **factory refactor** described for Tier C so all LLM calls can target local base URLs.
4. Re-tune prompts for smaller models; run evaluation against the current GPT-4o-mini baseline before promising parity.
5. Document the per-tenant deployment topology (hardware spec, network policy, model patching).

**Cost (order-of-magnitude planning numbers — re-price before any quote):**
- **Capex path:** A single high-memory GPU server suitable for a 27B-class INT4 model is commonly quoted in the **low tens of thousands of GBP** capital range before power and datacentre fit-out.
- **Rent path:** Cloud GPU **spot / on-demand** pricing fluctuates daily; budget **single-digit GBP per GPU-hour** for H100-class cards as a Fermi check, then multiply by expected surveyor concurrency and session length.
- **Throughput rule of thumb:** One H100-class GPU can comfortably serve **roughly single-digit concurrent interactive surveyors** for ~20–30k-token contexts **if** batching is tuned—run your own load test; do not treat this as an SLA.

**When Tier D wins on total cost:** Usually when **many** tenants or **very high** monthly token volume amortises fixed GPU cost. Below that crossover, Tier A/B stays cheaper because you are not paying for idle silicon.

---

## 5. Recommended path (decision tree for sales)

Use the ASCII tree below as a **conversation map**, not a promise of SKU availability:

```
                "Where does my data go?"

       ┌──────────────────┼──────────────────┐
       ▼                  ▼                  ▼
"We need it gone   "We use Microsoft        "Standard SaaS,
 from the          365 / Azure already.      cost-effective,
 internet."        Compliance team will      compliance team
                   prefer that boundary."    is comfortable
                                              with OpenAI."
       │                  │                  │
       ▼                  ▼                  ▼
   TIER D            TIER C                TIER A + ZDR
   Self-hosted       Azure OpenAI          (Tier B)
   open-weight       (UK South)
   (Gemma/Llama)     + CMK + private
                     endpoint
```

**Default commercial posture to aim for:** **Tier B (Standard OpenAI + ZDR)** once OpenAI approves ZDR for your org and you have audited endpoint compatibility. Until then, describe Tier B as the **target** baseline, not something already switched on in every deployment.

**Premium offering (roadmap): Tier C.** Pitch this as the **Microsoft/Azure boundary** pattern for firms already on M365/Azure—**after** the Azure OpenAI integration exists in product.

**Enterprise / on-premises offering (roadmap): Tier D.** For firms who insist on no public-cloud LLM. Higher setup cost; quote per deployment.

---

## 6. Microsoft 365 Copilot vs Azure OpenAI vs this product

**Three different things:**

| Name | What it is | Relationship to this repo |
|---|---|---|
| **Microsoft 365 Copilot** | A **Microsoft 365** end-user product (Word/Outlook/etc. integrations) with its own privacy statement. | **We are not Copilot.** We do not ship Office add-ins or Microsoft Graph connectors for Copilot in this repository. |
| **Azure OpenAI Service** | A **Microsoft Azure** platform API for GPT-family models inside a customer subscription. | **Roadmap integration (Tier C).** When shipped, our FastAPI service would call the customer’s Azure endpoint instead of `api.openai.com`. |
| **This RICS report generator** | A **Python FastAPI** application in this git repo. | **Shipped today:** Tier A paths to OpenAI when a key is set (see **§2.3**). **Not shipped:** Azure adapter, local LLM adapter, ZDR guardrail flag. |

**Approved explanation for non-technical buyers:** “Microsoft publishes **separate** privacy documentation for **Copilot** and for **Azure OpenAI**. If your procurement rule is *‘model calls must stay inside our Azure tenant,’* that maps to **Azure OpenAI (Tier C)**—an integration we have **specified** in §4 but **not coded** yet. If your rule is *‘must match Copilot,’* read the **Copilot** privacy page literally; it is a different product.”

**Approved explanation once Tier C exists:** “We call **your** Azure OpenAI deployment in **your** chosen region (for example UK South), optionally with private endpoints and customer-managed keys exactly as described in Microsoft’s Azure OpenAI documentation linked in §4.”

---

## 7. Friday meeting — concrete decisions to make

Bring printed copies of this document. Walk through the four tiers. Ask the engineers:

1. **[DECIDE]** Do we apply for OpenAI Zero Data Retention this week? (Tier B). Estimated effort: 30 minutes of paperwork + a one-day audit to confirm we use no ZDR-incompatible endpoints. **Recommendation: yes, no downside.**

2. **[DECIDE]** Do we build **Azure OpenAI support** (client factory + config) and start offering Tier C as a paid premium tier? Estimated effort: 1 engineer × ~1–2 weeks for the factory + call-site refactor + docs + Azure subscription cost during testing. **Recommendation: yes, but commit only after we have 1 prospect who has asked for it.**

3. **[DECIDE]** Do we prototype Tier D with Ollama + a Gemma or Llama model? Estimated effort: 1 engineer × 2 weeks for the router + 2 weeks of evaluation runs to confirm quality is acceptable. **Recommendation: prototype now, productise later — start with a sandbox H100 instance from RunPod (~£200/month) so the work is reversible.**

4. **[DECIDE]** Do we add a configuration flag (`llm_provider: openai | azure | local`) at the global / per-tenant level? Estimated effort: 2 days, mostly tests. **Recommendation: yes, regardless — even if we only ship Tier B in the first version, this flag is the foundation for everything else.**

5. **[DECIDE]** Do we publish a public-facing "Data Security" page that quotes from this document? Estimated effort: 1 day. **Recommendation: yes, by end of next sprint. It pre-empts the question we get in every webinar.**

---

## 8. Webinar talking points

Five sentences that match **§2 (code)** + **§3 (OpenAI vendor text)** as of last review. **Replace** sentences 3–4 after Azure/local LLM code ships.

1. *"Your uploaded reports, database, and FAISS search index live on the infrastructure that runs our app—typically your server or your cloud account. **By default** we compute embeddings locally with a Hugging Face model; we can optionally call **OpenAI embeddings** depending on configuration. For drafting, proofreading, and optional photo analysis, we call **OpenAI’s API** over HTTPS with the prompts and context each feature needs."*

2. *"OpenAI’s **public** API and enterprise privacy pages describe that API inputs and outputs are **not used to train** their models by default, and describe **limited log retention** for safety and abuse review. For the strongest contractual minimisation of retention, customers should ask us about **Zero Data Retention** once we have completed OpenAI’s programme enrolment—always read the **live** vendor page before signing."*

3. *"For firms that need Microsoft’s UK boundary, **Azure OpenAI** is the industry pattern Microsoft documents for regulated workloads—it is on our **roadmap** as an integration option, not a switch that exists in this repository yet."*

4. *"For firms that require **no** public-cloud LLM, we can target a **self-hosted** open-weight model behind an OpenAI-compatible endpoint—that is also **roadmap** engineering, not the default install today."*

5. *"Between surveying firms, we enforce **tenant isolation** in the API and in vector search so one tenant’s chunks are not returned to another. **Inside** one firm, finer-grained per-user document ACLs are still a roadmap item—your IAM and VPN policies still matter."*

---

## 9. Risk register

Residual risks after you have handled **LLM vendor choice**. Each row states **what can go wrong**, **severity**, and **what we already do in code** (if anything):

| Risk | Severity | Mitigation |
|---|---|---|
| **Indirect prompt injection** via uploaded reports — instructions hidden in a PDF/DOCX | Medium | Defence in depth: grounded generation, optional **LLM grounding** in `app/generator/postprocess.py`, tenant-scoped retrieval. No separate “Pebblo” integration in this repo—evaluate third-party input filters if needed. |
| **RAG access control** — benign question returns chunks the caller should not see | High | **Between tenants:** FAISS filters on `tenant_id` (`app/vectorstore/faiss_wrapper.py`). **Within one tenant:** per-user ACL is **not** implemented—any user who can act as that tenant can access that tenant’s indexed documents. Treat as roadmap. |
| **Sub-processor change** at OpenAI / Azure | Low | Subscribe to vendor change notifications; update this doc. |
| **Model updates** breaking output quality | Medium | `chat_model` in settings pins a model name; re-run regression tests before changing. |
| **OpenAI court orders / litigation holds** affecting retention | Medium (Tier A) | ZDR and legal review if a customer needs stricter handling than standard API terms. |
| **Token logging in our own infrastructure** — bodies in load-balancer or app logs | Medium | Audit deployment logging; confirm prompts/responses are not written at INFO in production. |
| **Photo data** to OpenAI vision | Medium | `app/services/photo_vision.py` and `app/generator/vision_analyzer.py` when enabled. |
| **Backup / replica copies** of SQLite/Postgres | Low–Medium | Deployer responsibility; document encryption at rest for backups. |

---

## 10. Implementation checklist per tier

### Tier A → Tier B (Zero Data Retention)
- [ ] Apply for ZDR on the OpenAI org dashboard
- [ ] Audit code for ZDR-incompatible endpoints (extended caching, background mode, Assistants API, Batch API). Confirm against **actual** call sites in this repository.
- [ ] Add config flag `openai_zdr_enforced: bool` and a startup check that refuses to fall back
- [ ] Add a `Data-Handling-Tier: zdr` HTTP response header (so compliance teams can verify)
- [ ] Update website / public-facing data security page

### Tier C (Azure OpenAI)
- [ ] Build **Azure OpenAI** client support (factory + settings).
- [ ] Add config: `llm_provider`, `azure_openai_endpoint`, `azure_openai_api_version`, `azure_openai_deployment_name`, `azure_use_managed_identity`
- [ ] Refactor OpenAI call sites (adapter, inspector loop, notes expander, style analyser, postprocess, vision, photo vision, canonical rollout, agents risk helper, generation orchestration) to use the factory
- [ ] Document tenant Azure resource provisioning (Azure OpenAI resource + private endpoint + Key Vault + RBAC)
- [ ] Document the "use your own Azure subscription" model vs "we provision one per tenant"
- [ ] Apply for **Modified Abuse Monitoring** when a customer requires Microsoft’s documented reduction of certain abuse-monitoring reviews — confirm scope in [Azure OpenAI abuse monitoring](https://learn.microsoft.com/en-us/azure/ai-foundry/openai/concepts/abuse-monitoring) (not a blanket “no humans” guarantee).
- [ ] Add tests: connection works, falls back gracefully if Azure auth fails
- [ ] Update website

### Tier D (Self-hosted open-weight)
- [ ] Build **local LLM** path (OpenAI-compatible HTTP API — Ollama is the simplest backend).
- [ ] Add config: `local_llm_endpoint`, `local_llm_model_name`, `local_llm_concurrent_requests`
- [ ] Stand up a sandbox: RunPod H100 or A100 instance with Ollama serving Gemma 3 27B or Llama 3.3 70B (INT4 quantised)
- [ ] Run evaluation against the local model. Compare to GPT-4o-mini baseline. Tune prompts if quality gap is unacceptable.
- [ ] Document the per-tenant deployment topology (hardware spec, network policy, model patching)
- [ ] Add a `vLLM` deployment as an alternative for higher-concurrency tenants
- [ ] Document the "managed-but-isolated" middle option (we operate the model on RunPod or Modal per tenant)
- [ ] Vision model: stand up LLaVA or Pixtral on the same host

---

## 11. References (vendor URLs — open before relying on any claim above)

These URLs are the **authoritative** sources for OpenAI/Microsoft training, retention, regions, abuse monitoring, and subprocessors. If a sentence in §3–§4 disagrees with the live page, **the live page wins**.

- [Enterprise privacy at OpenAI](https://openai.com/enterprise-privacy/)
- [Business data privacy, security, and compliance | OpenAI](https://openai.com/business-data/)
- [Data controls in the OpenAI platform](https://platform.openai.com/docs/guides/your-data)
- [OpenAI Data Processing Addendum](https://openai.com/policies/data-processing-addendum/)
- [OpenAI Sub-processor List](https://openai.com/policies/sub-processor-list/)
- [OpenAI Response to NYT data demands](https://openai.com/index/response-to-nyt-data-demands/)
- [Azure OpenAI / Azure Direct Models: Data, privacy, and security](https://learn.microsoft.com/en-us/azure/ai-foundry/responsible-ai/openai/data-privacy)
- [Azure OpenAI abuse monitoring](https://learn.microsoft.com/en-us/azure/ai-foundry/openai/concepts/abuse-monitoring)
- [Azure OpenAI encryption of data at rest (CMK / BYOK)](https://learn.microsoft.com/en-us/azure/ai-foundry/openai/encrypt-data-at-rest)
- [Azure OpenAI network & access configuration](https://learn.microsoft.com/en-us/azure/ai-foundry/openai/how-to/on-your-data-configuration)
- [Foundry Models sold directly by Azure — including UK South](https://learn.microsoft.com/en-us/azure/foundry/foundry-models/concepts/models-sold-directly-by-azure)
- [Microsoft 365 Copilot — Data, Privacy, and Security](https://learn.microsoft.com/en-us/copilot/microsoft-365/microsoft-365-copilot-privacy)

---

*Circulate internally and reuse text for customer-facing pages only after legal review. When OpenAI or Microsoft changes policy, re-read **§11** and update **§3–§4**; when our code changes egress, re-audit **§2** and update **§10**.*