Buckets:
| Name | Size | Uploaded | Xet hash |
|---|---|---|---|
| code | 5 items | ||
| marginals | 2 items | ||
| rubrics | 1 items | ||
| synthetic | 2 items | ||
| .gitattributes | 2.5 kB xet | 738f1125 | |
| README.md | 4.9 kB xet | 64521229 | |
| model.npz | 14.8 MB xet | 9e050b15 |
MHUsage
Differentially private synthetic conversation vectors approximating the patterns in 8,237 real user conversations with ChatGPT and Claude relevant to mental health. This dataset accompanies Transluce's Mental Health Behavior Report; see the report's Privacy Statement for the full disclosure-risk assessment.
The underlying "real conversation vectors" were generated by OpenAI and Anthropic using LLM judges to grade each conversation on rubrics provided by Transluce, giving each conversation a vector of 190 summary binary features plus two numeric length columns. These vectors contain no free-text content, no verbatim quotations, and no usernames, session IDs, geographic information, or timestamps. The raw vectors are not released; this dataset is a DP-protected synthetic counterpart.
Privacy
Individual conversation vectors are protected with (ε, δ)-differential privacy at ε = 10, δ = 10⁻⁵ (accounted internally in zero-concentrated DP, ρ ≈ 1.78). The implementation relies on OpenDP for privacy-critical operations. The mechanism is based on AIM, adapted for this use case:
- A fixed workload of marginals chosen from column names alone is measured with the Gaussian mechanism.
- A latent class model (essentially MBI's mixture-of-products estimator) is initialized and trained on these noisy marginals, followed by AIM's iterative method: each round selects a pair of columns with the exponential mechanism (scored by how badly the model currently captures the pair) and measures it with the Gaussian mechanism.
- The model generates synthetic records via randomized rounding.
- Records are post-processed so they do not contain incompatible feature
values: for an exclusion (e.g. no record has both
pre_adult_life_stageandlate_life_stage) one of the incompatible features is cleared with probability proportional to prevalence; for an implication (e.g. abucket/column may be 1 only if its applicability criterion is 1) the consequent is set whenever the antecedent holds. Fewer than 0.1% of cells are affected.
Files
synthetic/synthetic-8239.parquet— the canonical output: 8,239 synthetic conversation vectors (Source,n_messages,n_words, and the binary feature columns).synthetic/synthetic-50k.parquet— 50,000 additional synthetic conversation vectors with the same columns, generated frommodel.npzwithcode/generate.py --rows 50000 --seed 0.marginals/marginals.csv(+.parquetmirror) — the noisy values of every measured marginal: 420 one-way cells, 31,564 fixed-workload cells, and 12,156 adaptively selected cells, with each measurement's noise scale (sigma) and measurement order (marginal). These are the measurements before model fitting — substantially noisier than the same marginals computed over the synthetic data — released so other models can be trained from the same DP information.rubrics/rubrics.parquet— the definition of every one of the 190 feature columns: the 11 applicability criteria and the 179 additional user-property features. Each row carries the judging text for that feature — the full LLM-judge rubric where the feature was judged by a standalone rubric, or the category description (plus parent dimension description) where it was judged as one category of a taxonomy dimension. Thebucket/columns use the correspondingly named applicability-criterion rubric.model.npz— the fitted mixture-of-products model, released so it can be used directly to generate additional synthetic records (including with i.i.d. sampling instead of randomized rounding; seecode/generate.py --sampler iid).code/— the mechanism (dp_synth.py), the shared schema (schema.py), the generator (generate.py), and the incompatible-flag repair (smooth.py,smoothing_rules.csv).
Together these contain the full information purchased with the privacy budget: anyone can regenerate a synthetic dataset from the same distribution, or fit a different model to the same noisy marginals.
Contact
If you have any concerns about the way your data might have been handled, please contact info@transluce.org.
- Total size
- 24 MB
- Files
- 13
- Last updated
- Aug 31
- Pre-warmed CDN
- US EU US EU