Spaces:
Running
Running
| # Roadmap | |
| ## Backlog (pick one at a time) | |
| Working list as of 2026-07-31 - features, bugs, and chores in rough priority | |
| order. Detailed specs for the numbered features live in the sections below. | |
| **Features** | |
| - [x] PsyEmbedding models in the picker (PI short-term ask): all four models | |
| in the hub collection are in the registry and live on dev - verified | |
| against huggingface.co/collections/Culture-and-Morality-Lab/psyembedding | |
| 2026-07-31. | |
| - [x] LLM-generated construct items (PI idea 2026-07-31): BUILT 2026-08-05 | |
| (see ITEM_GENERATION.md). Design note + PI-approved prompt and | |
| cautionary wording; POST /api/constructs/generate-items (signed-in, | |
| 20/day cap, Claude Haiku, preview-only into the existing review-and-save | |
| flow); provenance on construct + run metadata; badges in picker/results. | |
| Remaining before Friday launch: ANTHROPIC_API_KEY on the Space (billing), | |
| live smoke test, then the SWLS/MFQ validation run (correlate per-text | |
| CCR scores from generated vs. validated items). | |
| - [x] Anchor vectors / bipolar constructs (feature 2 below): SHIPPED 2026-08-22 | |
| (spec 0006). AV = C - C_opposite centroids with a cosine default and a | |
| dot-product toggle, both item sets entered or uploaded, bipolar results | |
| and export at output_schema_version 1.2. Live on dev; NOT yet on the lab | |
| production Space. | |
| - [ ] Automatic chunking for over-limit rows (feature 1 below). PI re-sent | |
| the spec 2026-07-31; matches feature 1 (optional, warn about | |
| truncation without it, chunk-embedding vs chunk-similarity averaging | |
| still an open question for him). | |
| - [ ] Dark mode for the app - /guide and /product already have it; the React | |
| app doesn't. | |
| - [ ] Decide the default model with the PI: MiniLM (CCR reference) vs. a | |
| PsyEmbedding model ("prominent or default" was the PI's ask). | |
| **Bugs / fixes** | |
| - [x] PsyEmbedding HF repos: DONE 2026-08-22. The missing 1_Pooling config | |
| was pushed to all four repos with lab credentials; models.yaml re-pins | |
| each to the fixed revision (one commit above the old pin, identical | |
| weights) and the pooling_fallback workaround is removed. | |
| - [x] PsyEmbedding benchmarks: DONE 2026-08-26. All SEVEN models were | |
| unbenchmarked, not just the four. Each now records short/medium/long | |
| seconds per 1k texts, measured on an Apple M4 Pro at 2 threads and | |
| labelled as such. Still worth a re-run on the Space before the numbers | |
| are shown to users: a shared vCPU is roughly 2-4x slower. | |
| **Updates / chores** | |
| - [x] Open-source warning text from the PI: received and placed verbatim on | |
| the landing page (2026-07-31), along with his hero copy and "Who runs | |
| this" text. | |
| - [x] LICENSE: DONE. MIT, single-licensed (DECISIONS 2026-08-26); the two | |
| untracked drafts (a duplicate MIT and an Apache-2.0 alternative) are | |
| removed and README now states the license and separates it from the | |
| questionnaire items, which belong to their original authors. | |
| - [x] Construct library verification pass: DONE 2026-08-27 (spec 0007). Noor's | |
| review of all 525 items applied, and her follow-up answers resolved every | |
| open question except the two PI decisions; 88 of 94 constructs verified. | |
| - [ ] Construct library: 6 constructs still unverified, all waiting on the PI: | |
| the IPIP "I" prefix (50 items, 5 constructs) and restoring the K10 stem | |
| (10 items). Everything else from the review is resolved. | |
| - [ ] Durable source links: the Grit-S entries point at personal Dropbox URLs, | |
| dirty_dozen_* traded a working ResearchGate link for a paywalled PsycNET | |
| one (the reviewer's link carried a session token that could not be | |
| committed), cage_questionnaire uses a token-style USPSTF file URL, and | |
| mfq_care/mfq_fairness point at a third-party mirror of the cited paper. | |
| All four want a stable replacement. | |
| - [ ] Verify Dr. Chen's maintainer pre-assignment exists on /admin and that | |
| she can sign in. | |
| - [ ] Self-service password reset (currently admin-only; tied to the planned | |
| Supabase auth swap). | |
| - [ ] Invite links: currently ON HOLD behind CCR_INVITES_ENABLED - decide | |
| keep/kill permanently. | |
| - [ ] Celery + Redis job queue - only when multi-instance deployment or retry | |
| semantics are needed (see jobs.py docstring). | |
| PI-requested features (Mohammad, 2026-07-18), grounded in Teitelbaum & Simchon | |
| (2025), *Neural Text Embeddings in Psychological Research*, Psychological | |
| Methods, https://doi.org/10.1037/met0000768. | |
| ## 1. Automatic text chunking for over-limit rows | |
| **Problem.** Every model has a token window (MiniLM 256, the others 512 tokens | |
| = roughly 350-400 English words). Rows beyond it are silently truncated today; | |
| we only warn (TEXTS_MAYBE_TRUNCATED). In a typical upload most rows fit and a | |
| handful do not (e.g. 190 of 200 under the limit, 10 over). | |
| **Feature.** An optional per-run "Split long texts into chunks" toggle | |
| (default OFF - never changes existing behavior silently). | |
| - Detection: count tokens with the SELECTED model's own tokenizer (exact, not | |
| a word-count estimate). The Step 3 card shows the toggle only when the | |
| corpus has over-limit rows: "N rows exceed this model's 512-token window." | |
| - Off (default): current behavior, plus the existing truncation warning, with | |
| hint text: "text beyond the model's window is ignored." | |
| - On: each over-limit row is split into sequential chunks of at most | |
| max_seq_length tokens (example: 1,200 tokens -> 512 + 512 + 176). Rows | |
| within the limit are untouched. | |
| - Row-level result from chunk results - two candidate aggregations (PI listed | |
| both; decide before implementation): | |
| a) average the chunk EMBEDDINGS (optionally length-weighted), then score | |
| the averaged embedding once - keeps one scoring path; | |
| b) score each chunk, then average the chunk SIMILARITIES per row. | |
| - Warnings: chunked runs report TEXTS_CHUNKED (count + affected rows) instead | |
| of TEXTS_MAYBE_TRUNCATED for those rows. | |
| - Reproducibility: chunking config (on/off, chunk size, aggregation) goes into | |
| run metadata AND the generated reproduction script, which must implement the | |
| identical split so exported scores reproduce offline. | |
| **Open questions for the PI** | |
| - Aggregation default: mean of embeddings vs mean of similarities? | |
| - Length-weight the chunk average (176-token tail counts less) or plain mean? | |
| - Any chunk overlap (e.g. 50 tokens) to avoid cutting sentences, or none? | |
| ## 2. Anchor vectors (bipolar constructs) | |
| **Problem.** Plain CCR scores similarity to a single construct C. Constructs | |
| with a natural opposite (happiness vs sadness, internal vs external locus of | |
| control) are better measured along the direction BETWEEN the poles - this also | |
| cancels shared confounds like "questionnaire-ness" (both poles are worded as | |
| questionnaire items, so their difference subtracts that style component; | |
| see the paper's Appendix B). | |
| **Feature.** Optional second item set on a run: | |
| - C = centroid of the target construct's item embeddings | |
| - C_opp = centroid of the opposing construct's item embeddings | |
| - AV = C - C_opp | |
| - loading = cos(T, AV) for each text embedding T | |
| Higher = toward the target pole, negative = toward the opposing pole. | |
| - UX: in the construct picker, an "Add contrasting construct (anchor vector)" | |
| option opens the SAME selection flows for the opposite pole (library / | |
| typed / file upload). Both item sets show side by side before running. | |
| - Data model: Job gains an optional opposite_construct_id. Metadata records | |
| BOTH construct snapshots + item hashes and a scoring block | |
| {"method": "anchored_vector", "similarity": "cosine"}. | |
| - Results page: the score is now bipolar - histogram centered on 0, negative | |
| scores meaningful (toward the opposite pole), top/bottom texts labeled | |
| "most <target>" / "most <opposite>". Per-item loadings shown per pole. | |
| - Reproduction script: embeds both item sets verbatim and reproduces AV math. | |
| - Reverse-scored items: unchanged in v1 (the paper's footnote 27 suggests | |
| negating reverse items; our (R) flags already carry the information). | |
| **Open questions for the PI** | |
| - Similarity metric: cosine(T, AV) per the PI's formula; the paper found | |
| dot(T, AV) sometimes better (Appendix B). Config flag or fixed cosine? | |
| - Should anchored runs also report the plain per-pole similarities | |
| (cos(T, C), cos(T, C_opp)) in the export for transparency? | |
| ## 3. Multi-construct runs - IMPLEMENTED | |
| **Problem.** Each construct had to be run separately, which is slow when the | |
| goal is to see how two or more constructs are interrelated in the same texts. | |
| **Implemented (2026-07-23).** A run now accepts up to 10 constructs | |
| (`POST /api/jobs` takes `construct_ids`; `construct_id` still works). The | |
| corpus is embedded ONCE and every construct is scored against the same | |
| document embeddings, so N constructs cost barely more than one - and the | |
| per-text scores are row-aligned by construction. | |
| - Results add a "Construct interrelations" card: Pearson r between per-text | |
| CCR scores, plus a collapsible per-construct section (histogram, item | |
| loadings, top/bottom texts). | |
| - Export: per-construct prefixed columns ({slug}_sim_item_N, {slug}_ccr_score) | |
| under output_schema_version 1.1; single-construct exports unchanged at 1.0. | |
| - Metadata records every construct snapshot + item hash and the correlation | |
| matrix; the reproduction script embeds the corpus once, scores all | |
| constructs, and prints the same correlations. | |
| - A multi-construct run counts once toward the anonymous/saved-run limits; | |
| the constructs-per-run cap (10) bounds export width, not compute. | |
| ## Sequencing | |
| Anchor vectors first (pure scoring change, no ingestion changes), then | |
| chunking (touches ingestion, warnings, cache keys - chunked and unchunked | |
| embeddings must not share a cache entry). Multi-construct runs (3) shipped | |
| first: no scoring or ingestion changes, and the embedding-reuse plumbing it | |
| added (one corpus pass, N scorings) is what chunking's cache work builds on. | |