# Data Privacy — What We Have In Place Right Now **For:** Behrang **From:** Huzaifa **Date:** 14 May 2026 **Companion document:** `DATA_SECURITY.md` (the full technical version) **This document:** The short, direct answers to your questions, in your words. **Accuracy note (May 2026):** Claims about **OpenAI** (training, retention, ZDR) depend on **OpenAI’s published terms and policies at the time** — always confirm the live pages before customer commitments. In-app behaviour is verified against this repository’s code. --- ## Your questions, answered yes/no | Your question | Yes / No / Partial | One-line answer | |---|---|---| | **Do we have anything in place at all?** | **YES** | Multiple layers. Listed below in §1. | | **Is each customer's data private to them?** | **YES** | Hard-enforced **per-tenant** isolation in the API, database, file paths, and vector retrieval. §2. | | **Are they in their "enclosed space"?** | **PARTIAL** | Customer **content** (files, DB, FAISS index) stays on infrastructure they control. **OpenAI’s API** receives prompts, retrieved text, and (when enabled) images for model inference. Third-party **libraries** may also use HTTPS (e.g. first-time download of the local embedding model). §3. | | **Can OpenAI use our customers' data to train AI?** | **NO** *(per OpenAI’s public API / enterprise privacy statements)* | OpenAI documents that API inputs/outputs are not used for training by default — confirm current wording on their site. §4. | | **Does OpenAI keep our customers' data forever?** | **NO** *(per OpenAI’s public retention statements)* | Public docs describe time-limited retention for API logs unless legally required otherwise; **Zero Data Retention** is a separate contractual option. §5. | | **Can we tell customers their data is private?** | **YES** *(with honest scope)* | Use the scripts in §7; distinguish **tenant vs tenant** (enforced in code) from **user vs user inside one tenant** (see §6). | | **Can a junior surveyor at Firm A see reports at Firm B?** | **NO — not via normal API paths** | Different `tenant_id`s; requests are scoped to the authenticated tenant and vector search drops non-matching chunks. §2. | | **Can a junior surveyor at Firm A see another colleague's reports at Firm A?** | **Currently yes** | Per-user permissions inside a firm are **not yet** built. This is a known gap. §6. | --- ## §1 — What we actually have in place today This reflects **this repository’s application code** (not every possible deployment wrapper such as cloud logging or reverse proxies). ### Data that stays on the firm’s infrastructure (under normal deployment) 1. **Uploaded reports** (PDF/DOCX) → written to configured local disk (`upload_dir`, default under the service user’s home directory). Not uploaded to a separate SaaS vector DB by this app. 2. **Inspector notes** (bullets) → stored with the report in the **SQL** database (SQLite by default; Postgres supported via `DATABASE_URL`). 3. **Vector index** → **FAISS** on disk (`faiss_index_path`, default under `~/.report_genius/`). 4. **Embeddings (default)** → **`prefer_local_embeddings=True`** uses **Hugging Face sentence-transformers** (`local_embedding_model`, default `all-MiniLM-L6-v2`) on the **same machine** as the app. **Exception:** if that backend cannot be loaded and an **OpenAI API key** is set, the factory **falls back to OpenAI embeddings**; operators can also set **`prefer_local_embeddings=False`** to prefer OpenAI. See `app/embeddings/factory.py` and `app/config.py`. 5. **Database** → not replicated to a third party by the application itself; backup/replication is the **deployer’s** responsibility. ### What leaves the customer environment (product-relevant) 1. **OpenAI HTTPS API** — used for section generation, proofreading/enhancement, notes expansion (when a key is set), style analysis, agentic inspector loop, risk assessment helper, optional vision/photo analysis, optional LLM grounding, canonical rollout LLM merge, etc. Payloads include prompts and **retrieved or user-supplied text** (and **images** when vision/photo features are used). 2. **Dependency network use (not “our” `requests` code)** — e.g. **downloading** the Hugging Face embedding model weights on first use, **OpenAI SDK** traffic, and any other libraries your environment pulls at runtime. Our `app/` tree does **not** use `requests` / `httpx` / `urllib.request` directly for custom integrations; that does **not** prove “no other HTTPS on the machine.” **Do not claim:** “Zero outbound network except OpenAI.” **Do claim:** “We do not ship separate analytics or vector-Database SaaS integrations in this codebase; the **designed** model-data path is local storage + OpenAI inference, with an optional OpenAI embedding path if configured.” ### Contractual / vendor protections (OpenAI — verify live) 6. **OpenAI Business / API terms and DPA** — relationship and processor terms are defined by **OpenAI’s** published agreements (links in `DATA_SECURITY.md` §11). 7. **Training opt-out for API data** — described in OpenAI’s enterprise/privacy documentation; **re-read the live page** before sales use. 8. **Retention** — public docs describe limited retention for abuse/safety; exact periods and ZDR eligibility are **contract- and product-specific**. 9. **TLS** — HTTPS to `api.openai.com` (and to other hosts used by installed libraries). --- ## §2 — "Is each firm's data private and isolated?" **YES for tenant-to-tenant isolation.** The API attaches a **`tenant_id`** to work; ORM queries and file/photo paths are scoped; FAISS metadata must match the request tenant. Concrete example — vector retrieval (`app/vectorstore/faiss_wrapper.py`; line numbers can shift between commits): ```python if meta.get("tenant_id") != tenant_id: continue ``` Every FAISS hit with a mismatched `tenant_id` is skipped before it becomes a search result. **Do not use a made-up count** (“19 places / 10 files”). The right statement: **`tenant_id` appears throughout `app/` (API, ingestion, cache, generation, vectorstore, etc.)** — we can show representative files on request. **Customer-facing (tenant vs tenant):** > "Two surveying firms are isolated by **tenant**: API and retrieval paths are built so one tenant cannot read another tenant’s documents or index entries. That is enforced in code, not only in policy." --- ## §3 — "Are they in their enclosed space?" **Honest answer:** **Customer-controlled storage** (uploads, DB, FAISS files, default local embeddings) stays on **their** servers or VPC **when they host the stack**. **Content-bearing requests** for LLM features are sent to **OpenAI** (and optionally use **OpenAI embeddings**) under the configuration described in §1. **Images** used for vision go to the same class of API when those features are enabled. **Suggested wording:** > "Your reports, notes, database, and search index remain on infrastructure you control when you deploy our application there. For AI drafting and related features, the application calls **OpenAI’s API** over HTTPS with the data needed for that request (text, and images if you use photo/vision features). OpenAI’s **training and retention rules** are defined in their published terms — we cite those links in our technical pack and keep them updated." --- ## §4 — "Can OpenAI use our customers' data to train AI?" **Per OpenAI’s public documentation** (wording can change — verify before webinars): > "By default, OpenAI does not use data from … **their API platform** — including inputs or outputs — for training or improving their models." > > — [openai.com/enterprise-privacy](https://openai.com/enterprise-privacy/) (confirm live text) **Do not paraphrase as “we verified in code that OpenAI never trains.”** That is a **vendor** commitment. --- ## §5 — "Does OpenAI keep our customers' data forever?" **Per public docs**, API logs are not kept indefinitely; exact retention and exceptions are in OpenAI’s materials. **Zero Data Retention (ZDR)** is a **separate programme** (application/approval, endpoint restrictions) — see [platform.openai.com/docs/guides/your-data](https://platform.openai.com/docs/guides/your-data). **Recommendation:** pursue ZDR if legal/commercial teams want the strongest **contractual** minimisation of retention — then quote **only** what OpenAI confirms in writing for your org. --- ## §6 — What we DON'T yet have (being honest) | Gap | Severity | Notes | |---|---|---| | **Zero Data Retention** | Ops / legal | Apply and validate against actual endpoints used. | | **Azure OpenAI** | Product | Not implemented in this repo today; roadmap item if customers require UK residency on Microsoft’s stack. | | **Self-hosted LLM** | Product | Not implemented; possible future `local_llm_base_url`–style integration. | | **Per-user ACL inside one tenant** | Privacy | Same `tenant_id` can still see all firm reports until row-level or chunk-level user scoping exists. | | **Public security page** | Marketing | Still worth publishing; base it on `DATA_SECURITY.md` after edits. | | **Photo / vision payloads** | Disclosure | `app/services/photo_vision.py` and `app/generator/vision_analyzer.py` send images to the configured OpenAI-compatible vision API when enabled. | --- ## §7 — What you can tell customers (fact-checked scripts) Use **§1–§3** above as the source of truth for “what stays local” vs “what goes to OpenAI.” For **training/retention**, read OpenAI’s **current** page aloud or quote verbatim. ### The 30-second pitch (code-accurate) > "When you host our application on your infrastructure, your uploaded reports, your database, and your semantic index stay with you. We use a **local embedding model by default** so vectorisation normally stays on the same server; you can change that in configuration. For AI writing and analysis, the app calls **OpenAI** over HTTPS with the minimum context needed for each request—including photos if you turn on vision. **Between firms**, the product isolates data by **tenant** in the API and in vector search. **Inside one firm**, finer-grained per-user locks are still on our roadmap." ### When they ask "where is my data?" 1. **On your deployment** — files, SQLite/Postgres, FAISS files, default local embeddings. 2. **With OpenAI for model calls** — prompts, snippets, and optional images, under their terms. 3. **Elsewhere** — this codebase does not add its own analytics export; your **hosting provider’s** logs and backups are outside this app’s control—disclose them in your deployment runbook. ### When they ask "what about Firm B?" > "Another customer’s tenant id cannot read your tenant’s rows or vector metadata; FAISS results are filtered by tenant. That does not replace your own **access control and IAM** on who can log in as your tenant." --- ## §8 — Friday follow-ups Align engineering tasks with **`DATA_SECURITY.md`** (implementation checklist and risk register), after removing references to **non-existent** modules in that file. --- ## §9 — Two documents | Document | Audience | |---|---| | **`BOSS_BRIEF_DATA_POLICY.md`** | Short internal brief / sales prep | | **`DATA_SECURITY.md`** | Engineers — technical detail and sources | --- ## §10 — After improvements (truthful sequencing) - **Today:** Local-first storage and indexing; OpenAI for LLM (and optionally embeddings); tenant isolation in code. - **After ZDR (if approved):** Stronger **contractual** retention story per OpenAI’s ZDR programme. - **After Azure / self-hosted work (if built):** Additional deployment options—describe only once shipped. --- *Code pointer for tenant filtering in FAISS: `app/vectorstore/faiss_wrapper.py` (search for `tenant_id`). OpenAI policy links: see `DATA_SECURITY.md` §11 — refresh before customer-facing use.*