foresight / server /README.md
umangchaudhry's picture
Deploy 13220d7 from GitHub
acb3189 verified
|
Raw
History Blame Contribute Delete
30.6 kB

Foresight app server

Serves the frontend, the knowledge base, per-student data, and the companion. FastAPI, small on purpose.

Job Where
Serve app/ and knowledge-base/ (at /kb/…) app.py
Sign-up / sign-in and sessions auth.py
Read/write per-student JSON in a private HF Dataset repo storage.py
Index the knowledge base for retrieval (in-memory BM25, stdlib only) kb/
Ask Foresight β€” the LangGraph agent, its tools, and chat threads agent/

The last two are documented in depth in ../docs/agent-architecture.md. The short version: kb/ normalizes all 9 sources into one Doc shape and builds a BM25 index on a background thread at startup; agent/ is a ReAct loop over eight typed campus tools plus OpenAI's hosted web search, streaming its answer over SSE.

Run it locally

pip install -r requirements.txt
uvicorn server.app:app --reload --port 7860

Open http://localhost:7860, create an account. With no HF_TOKEN set, accounts and saves go to a git-ignored .data/ directory instead of the dataset repo, so the whole app runs with no credentials. With no OPENAI_API_KEY, Ask Foresight reports itself as unconfigured and every other screen works normally β€” that degradation is deliberate and covered by tests.

Run the tests:

pip install -r requirements-dev.txt
python -m pytest tests/ -q

They never call a model or the network: retrieval is checked against the real committed knowledge base, and the agent runs against a scripted fake model. That's also why they can gate deployment β€” .github/workflows/test.yml runs the suite on every pull request, and deploy-space.yml won't publish to the Space unless it passes.

Build the container the same way it deploys:

docker build -t foresight .
docker run --rm -p 7860:7860 foresight

Environment

Variable Required What it does
FORESIGHT_SESSION_SECRET in deployment Signs the session cookie. Unset β‡’ random per process β‡’ every restart signs everyone out
HF_TOKEN in deployment Write token for the dataset repo. Unset β‡’ local-directory fallback, which a container loses on restart
FORESIGHT_DATASET_REPO no Defaults to umangchaudhry/foresight
FORESIGHT_LOCAL_DATA_DIR no Fallback directory when there's no token (default .data)
FORESIGHT_HTTPS_ONLY no Set truthy in deployment to mark the session cookie Secure
OPENAI_API_KEY for Ask Foresight and syllabus parsing Unset β‡’ both are disabled and say so; the rest of the app is unaffected
FORESIGHT_CHAT_MODEL no Defaults to gpt-5.6-sol. Swap tiers without a code change
FORESIGHT_KB_DIR no Where the index reads from (default knowledge-base/)
FORESIGHT_ALLOW_TIME_TRAVEL no Truthy lets GET /api/survey?today= shift the date, so the May/December check-ins can be exercised. Never set in deployment
FORESIGHT_MAX_UPLOAD_BYTES no Syllabus upload cap (default 10 MB)
FORESIGHT_MAX_PDF_PAGES no Pages read per PDF (default 40)
FORESIGHT_MAX_DOC_CHARS no Characters of document text kept (default 120,000)
FORESIGHT_MIN_DOC_CHARS no Below this a PDF counts as having no text layer and goes to the model as a file (default 200)
FORESIGHT_MIN_LOCATION_CONFIDENCE no Gazetteer floor for pinning a building_slug (default 0.5)
FORESIGHT_SYLLABUS_MODEL no Defaults to gpt-5.6-sol. Not the nano tier β€” see syllabus/parse.py
FORESIGHT_SYLLABUS_MAX_TOKENS no Output cap for one parse (default 8000)

GET /healthz reports which storage backend is live, whether chat is enabled, and the knowledge-base index status β€” the fastest way to catch a misconfigured Space.

Accounts

Anyone can sign up with a username and password; no invite or code needed. A signed cookie keeps them signed in for 30 days.

  • Usernames are 3–32 characters of letters, numbers, dots, dashes or underscores, case-insensitive. The character set is tight because the username becomes a filename in the dataset repo.
  • Passwords are at least 8 characters, hashed with hashlib.scrypt and a random per-user salt. Hashed, not encrypted β€” encryption is reversible, so a copy of the dataset would hand over working credentials, and people reuse passwords across accounts. The scrypt cost is stored on each account, so SCRYPT can be raised later without invalidating existing passwords.
  • Wrong password and unknown username return the same message, so the login form can't be used to find out who has an account.
  • student_id is a random value assigned at signup, not the username, so a username never appears in a storage path.

auth.py is the seam to replace when Vanderbilt SSO arrives; everything else only depends on current_student(request) returning an id.

Storage layout

auth/users/{username}.json                      # credentials + student_id
students/{student_id}/profile.json
students/{student_id}/syllabi.json              # added by the syllabus work
students/{student_id}/chats/index.json          # thread list
students/{student_id}/chats/{thread_id}.json    # one conversation

One file per user, never a shared mutable file. A dataset repo is git, not a database β€” there's no row locking, so two writers on one file lose a race. Per-user files make that impossible between users, and write_json retries the same-file case.

Every write is a network commit, so the client debounces (app/store.js) rather than saving on each keystroke, chat writes once per completed turn rather than per token, and the survey writes once per wizard page rather than per Likert click. Repo history is permanent, which is why the app stores parsed fields rather than uploaded documents β€” and why deleting a conversation writes a tombstone rather than removing the file.

profile.json is a free-form blob: no schema, no key allowlist, and PUT is a whole-document replace with the merge happening on the client. Adding a field needs no migration β€” every reader treats a missing key as empty. The keys in use:

Key Written by
majors, minors lists β€” double and triple majors are ordinary. major (singular string) is the legacy shape, still written for compatibility and still read as a one-item list by survey.programs()
still_exploring, strengths, interests, goals, study_time, activities intake wizard + My Story
class_year, class_year_as_of the year the student says they're in, plus the academic year they said it in β€” stamped server-side, and the pair is what advances the answer as terms pass
grad_year intake wizard + My Story, optional; stored if offered and nothing reads it
background list of first_view / international / transfer
orgs campus groups the student says they've joined β€” [{id, name}] from the AnchorLink catalog. The Today feed stars events whose host matches, and the id is what re-resolves a group that has since been renamed. Never required by the intake gate: a first-year at move-in has joined nothing
orgs_prompt_dismissed the Today card asking for the above, waved off
interest_topics, interest_topics_of server-stamped. The free-text interests mapped onto the shared topic vocabulary, plus the exact string it was derived from. The _of half is what makes this cost one model call per real interest change rather than one per save, and what lets a student's hand-edited chips survive every unrelated save β€” see kb/topics.py
survey_runs one entry per survey wave, responses keyed by item id β€” see survey/
onboarding_opted_out, checkin_opted_out the two per-prompt opt-outs
onboarding_seen_at first interaction with the wizard; floors the check-in schedule
worries legacy β€” a positional array from the pre-instrument prototype, still read by the agent prompt, never written

The survey (server/survey/)

The intake and end-of-semester questionnaire β€” Dr. London's First View instrument. Two files, no I/O beyond reading the bank, so the whole "which questions does this student see today" question is answerable in a unit test.

Module What it does
survey/items.json The 38 items with stable slug ids, constructs, the reverse-coded flag, and which 12 suit a first-year. The single source of truth: Python reads it, and the browser gets it over GET /api/survey.
survey/schedule.py Standing from an expected graduation year, which version of the instrument applies, when a check-in wave opens (May 1 / December 5), and intake_complete β€” the computed gate that decides whether the overlay shows.
Route What it does
GET /api/survey The whole bank plus the scale, which version the student's intake uses, whether the intake is complete, and which check-in wave (if any) is open. ?today=YYYY-MM-DD shifts the date for testing, and only when FORESIGHT_ALLOW_TIME_TRAVEL is truthy.

Ids rather than positions is the load-bearing decision: the prototype stored responses in an array indexed into a hardcoded four-question list, so any change to the question set silently reinterpreted every stored profile β€” and varying the set by year makes that certain. Full background, including what the source document does not specify: ../docs/first-view-survey.md.

Syllabi (server/syllabus/)

A student's syllabi are the only source of their real schedule β€” YES is off the table β€” so the calendar, the campus map, the Today feed, Prep and My VU all read this one file.

Module What it does
syllabus/extract.py Routes an upload: extracted text, or the file itself for the model to read.
syllabus/parse.py The one function that calls a model β€” one syllabus, one call, strict schema out. The only network access in the package and the only place the model name appears.
syllabus/terms.py Derives a term's anchors from knowledge-base/academics/calendar.json β€” first/last day of classes, the exam period, breaks β€” so "Week 6" can become a real ISO date (week_day), and an implausible one can be caught (classify_date). prompt_context() is the block handed to the model.
syllabus/schema.py EXTRACTION_SCHEMA, the strict JSON schema the model must return, and build_record(), which turns model output into the stored record.

Endpoints

Route What it does
GET /api/syllabi The saved list, the default term, and whether parsing is configured at all
POST /api/syllabi/parse One uploaded file β†’ one parsed record. Saves nothing. Reports duplicate_of when the student already has that course this term
PUT /api/syllabi Save one reviewed/corrected record, replacing any earlier copy of the same course and term
DELETE /api/syllabi/{id} Remove one β€” and with it every downstream appearance, since the record is the only copy
GET /api/schedule The confirmed ones expanded into dated calendar occurrences β€” see My classes

Parsing and saving are separate requests. A parse is a proposal until the student confirms it on the review screen, which is what makes reviewed a real gate rather than a field nobody sets. The trade-off is that an unreviewed parse is lost on a refresh; that matches the status flow the issue describes (queued β†’ parsing β†’ needs review β†’ saved).

One request per file. Several syllabi upload concurrently from the browser with genuinely independent status and failures, and no job queue is needed in a container that sleeps. A malformed file can't block the others.

Corrections are re-normalized, never trusted. schema.from_client() runs an edited record through exactly the same normalization as a fresh parse β€” times, dates, days, course code, building resolution, flags β€” because "the student typed it" is not a reason to store due: "next friday" or a building_slug that doesn't exist. reviewed is the one field the client may set.

Failures map to status codes the UI can act on: 400 for a file the student should swap (with the reason shown as-is), 502 when the model call fails, 503 when the server has no OPENAI_API_KEY β€” checked before reading the file, so nothing is uploaded that can't be processed.

Routing an upload

Students upload PDFs, Word documents and photos of printouts, and nothing reads all three well, so extract() routes by what the file actually is:

Input Read by Where
PDF with a text layer pypdf in the container
.docx / .txt / .md python-docx in the container
PDF with no text layer the model, as application/pdf vision
.png / .jpg / .webp / .gif the model, as an image vision

The text path is preferred and covers the normal case. Both libraries are pure Python, so the image needs no system packages; only the text needed for parsing leaves the container, not the whole document; and a bad parse can be traced to exactly what the model was given.

Photos and text-less PDFs take the vision path because there is nothing to extract. The alternative was an OCR system dependency (tesseract, ~100 MB), and OCR on a skewed, badly-lit photo of a table produces confident-looking garbage β€” which becomes confident-looking wrong exam dates, the failure this feature exists to avoid. A vision model reads that photo far better. Nothing is rasterized here: a scanned PDF is sent as a PDF.

The cost is made explicit rather than hidden β€” the record carries text_source: "vision" and a from_photo flag, so the review screen can tell the student this one came from a photo and is worth checking against the original.

Refusals still exist for what neither side can read: .doc, Pages, HEIC (an iPhone default β€” Safari usually converts to JPEG on upload; if that turns out not to hold, pillow-heif is a pip-only converter), an empty or oversized file, a corrupt or password-protected PDF, and a text-less PDF longer than the page cap (every page would become an image). Each raises extract.Unsupported with a message written to be shown to the student as-is and telling them what to do instead β€” a per-file failure with a retry, never a silent drop.

Table text is pulled out of .docx explicitly: syllabus schedules are nearly always tables, and python-docx keeps table text out of paragraphs, so skipping them would lose most of the due dates.

Three properties of the calendar data shape drive terms.py, all covered by tests/test_syllabus_terms.py: multi-day ranges exist only in event titles (end_iso is null on every all-day event); a range's date_iso can disagree with its title (Thanksgiving is dated the Saturday before), in which case the title wins; and a term can be published without class dates (spring-2027 today) β€” complete() reports that, and nothing is guessed, because a wrong start date shifts every relative date in the syllabus silently.

students/{student_id}/syllabi.json holds a list of records:

{
  "id": "bsci-1510-fall-2026",
  "course_code": "BSCI 1510",
  "course_title": "Introduction to Biological Sciences",
  "term": "fall-2026",
  "instructor": { "name": "…", "email": "…",
                  "office": "Stevenson Center 5726", "office_building_slug": "sc-chemistry",
                  "office_room": "5726", "office_location_confidence": 0.95,
                  "office_hours": "Tue 2–4pm, or by appointment",
                  "office_hours_slots": [ { "days": ["Tue"], "start": "14:00", "end": "16:00",
                                            "location_raw": "Stevenson Center 5726",
                                            "building_slug": "sc-chemistry", "room": "5726" } ] },
  "meetings": [ { "days": ["Mon","Wed","Fri"], "start": "10:00", "end": "10:50",
                  "location_raw": "Stevenson Center 4327", "building_slug": "sc-lecture",
                  "room": "4327", "location_confidence": 0.95 } ],
  "exams":       [ { "title": "Midterm 1", "date": "2026-09-24", "start": "19:00",
                     "end": "21:00", "location_raw": "Wilson Hall 103",
                     "building_slug": "wilson-hall", "room": "103", "weight": "20%" } ],
  "assignments": [ { "title": "Problem Set 3", "due": "2026-09-12", "weight": "5%" } ],
  "grading":     [ { "component": "Class participation", "weight": "10%" } ],
  "source_file": "bsci1510-syllabus.pdf",
  "parsed_at":   "2026-08-20T12:00:00+00:00",
  "text_source": "extracted",
  "confidence":  { "meetings": 0.9, "exams": 0.6, "assignments": 0.4 },
  "flags":       [ { "path": "assignments[0].due", "code": "unresolved_date", "value": "Week 14" } ],
  "reviewed":    false
}
  • reviewed is a hard gate. Only the student sets it, from the review screen; schema.ship_ready() is the only thing a downstream surface should read. A silently mis-parsed exam date is worse than no exam date.
  • The model never fills the server-owned fields β€” id, source_file, parsed_at, text_source, reviewed, building_slug, room, flags. Asked for a building slug, a model invents one that doesn't exist.
  • location_raw is kept verbatim and resolved through kb/buildings.py (aliases, the seven numbered Stevenson sub-buildings, room extraction). Below MIN_LOCATION_CONFIDENCE the slug stays null β€” a wrong pin on the map is worse than a missing one. Rooms are never validated: there is no list of every classroom and the feature doesn't need one, so room is carried as text and only an unresolved building is flagged. That resolver lives under server/ precisely because the image doesn't ship collectors/ β€” see its docstring before moving it back.
  • The instructor's office is resolved like any other location β€” going to office hours is one of the things this product exists to make ordinary, so the office gets a building_slug and can be pinned on the map. An office that isn't a place ("by appointment") resolves to nothing and is not flagged; that would be noise.
  • Office hours are stored twice, on purpose. office_hours is the verbatim text the student reads, including "or by appointment". office_hours_slots is the same information as recurring slots a calendar can place, in the same shape as meetings β€” so anything that can draw a class can draw an office hour without a second code path. A slot that names no location inherits the office; one with neither a day nor a start time is dropped rather than stored as a row of nulls.
  • Weights live in two places, and both are needed. Exams and assignments carry their own weight, and grading holds the whole breakdown verbatim. Components like "Class participation 10%" are neither an exam nor a dated assignment and would otherwise be dropped, and the parts can't be checked against 100% without the table. Weights stay strings β€” "one letter grade" and "10% each" are real. When every weight is a plain percentage and the total is more than a point off 100, a grading_total flag says so; it's advisory, since extra credit exists.
  • confidence is per section, because assignment tables parse far less reliably than a meeting time, and the review screen needs to know where to draw attention. flags is derived and advisory, for the same purpose.
  • Dedupe on course_code + term (schema.dedupe_key) so re-uploading a course offers to replace rather than silently adding a second copy.

The Today dashboard (server/today.py)

GET /api/today returns everything the landing screen renders: classes, week, feed, and a needs block the browser uses for its empty states. ?today= shifts the date under FORESIGHT_ALLOW_TIME_TRAVEL, same as /api/survey.

payload() is a pure function of (profile, syllabi, index, today), so the ranking is testable without storage, a session, or a clock β€” see tests/test_today.py.

  • Server-side, unlike the calendar. app/calendar.js reads /kb/ in the browser, which is right for pure date filtering. Today can't: the feed ranks free-text interests with BM25 and gates the schedule on syllabus.schema.ship_ready. Reimplementing either in JS means two rankers to keep in step and 2 MB of JSON ranked in a tab. Today's first real consumer of ship_ready β€” before this, the reviewed gate had no caller outside its own tests.
  • classes is suppressed outside the class period. During fall break the honest answer is "no classes", and an empty card reads as a bug β€” so terms.classify_date drives a note instead ("Fall 2026 classes begin Wednesday, August 26.").
  • week is coursework only. Campus deadlines live in the feed. The mock listed add/drop in both places, which is what made the old strip feel like filler.
  • The feed mixes three pools β€” events hosted by a group the student joined, academic key dates, and events matching their majors and interests β€” ranked by a handful of named constants at the top of the module. Every weight maps to a why string, so "why am I being shown this" always has an answer.
  • Key dates are filtered and capped. The academic calendar publishes cohort deadlines with no audience tag (Fall 2026-AC is on 78 of 85 records), so the title is the only signal: an A&S student is not shown "Half-term Business Module I", and registrar-internal milestones like "Discrepancy reporting begins" are never shown to anyone. Uncapped, a September window returned fifteen deadlines and buried every event. The real fix is per-audience tagging in the collector β€” the same gap the event domain tags have.
  • Interest matching runs on tags first, words second. The tag pool intersects the student's interest_topics with an event's topics β€” an agreement on a shared vocabulary, which is why it outranks any lexical match. Underneath it, the BM25 pool requires a multi-word phrase to match all of its content words and falls back to the phrase's rarest word: that makes "public health" find the public-health congress instead of all 51 events mentioning health, and keeps "brain science" returning nothing rather than a polymer-chemistry conference. Frequency can't separate every case β€” "science" is in 12.2% of events and "music" in 15.0% β€” so the lexical pool stays deliberately sparse-and-right, and the tags carry the recall the words can't.
  • Repeats collapse to one card. normalize.py collapses a campus event's occurrences, but AnchorLink publishes a recurring drop-in as separate events with separate ids β€” 15 of them for "Academic Wellbeing Drop-In". Ranking by topic surfaces all at once, so _collapse_repeats keeps the best-scoring, counts the rest, and shows the soonest date. Keyed on title and host, so two groups running a "Regular Meeting" stay separate.

My classes (server/schedule.py)

GET /api/schedule serves the My classes layer on both surfaces that need it: items are the student's confirmed syllabi expanded into dated occurrences for the Grand Calendar's grid β€” class meetings, office hours, exams and due dates, for every term they have a syllabus for β€” and places are the same syllabi grouped by building for the campus map's pins. payload() is a pure function of (syllabi, calendar) with no clock and no storage, so both are tested at fixed dates against fixed anchors (tests/test_schedule.py).

  • The one layer that isn't a /kb/ read. Every other calendar layer is date filtering in the browser, which is right. This can't be: placing a weekly meeting needs syllabus/terms.py's anchors, derived from event title text ("Undergraduate examinations and reading days, Dec. 11-19") by rules a JavaScript copy would drift from within one term. The frontend receives date + HH:MM and joins them itself, offset-free, so calendar.js keeps reading times out of strings rather than parsing them β€” see its timezone rule.
  • Recurrence is never guessed. Meetings are placed only between the first and last day of classes (not the exam period β€” a final has its own date), skipping breaks. A term with no published class dates yields no meetings at all and an unplaced row saying why, because guessing a start date puts every meeting in the wrong week.
  • Dated things are absolute. An exam or assignment carries its own ISO date and is placed as written, even outside the anchors: the student confirmed it, and the review screen already questioned the implausible ones.
  • Nothing is dropped in silence. A meeting with no days, an exam with no date, a whole term with no anchors, a syllabus saved but never confirmed, a room that matched no building β€” each is reported (unplaced, unmapped, unreviewed) for the screen to explain. "My midterm isn't on here" with no reason given is the failure that costs a student's trust in the schedule.
  • places come off the records, not off items. A room is a fact the syllabus states; it doesn't depend on the university having published the term's dates, so a course with no placeable meetings still has a pin. Classes and office hours only β€” an exam hall is one date the calendar already carries, and pinning it would put a room a student visits once beside the three they visit weekly. building_slug stays null whenever the gazetteer wasn't confident, and a null is reported rather than guessed: a wrong pin is worse than a missing one, the same rule the directory collector follows.
  • Separate from today.py on purpose. Both read syllabi through ship_ready, but Today answers "what is on today, what is due this week" with a dashboard's suppression and urgency rules, while this places every occurrence on a grid. The overlap is the gate and the HH:MM convention, both shared.

The interest vocabulary (server/kb/topics.py)

~33 slugs that events, organizations and students are all tagged into, so the Today feed matches on a shared vocabulary rather than on wording.

Derived, not invented. AnchorLink's taxonomy has 75 values; strip the administrative buckets (Registered Student Organization, POM Level 1 Compliance, two dozen Advised by …) and what remains already tags 686 of 1,014 organizations for free. CATEGORIES maps those onto slugs and IGNORED_CATEGORIES names the rest, so a new category from the collector fails a test rather than being silently dropped.

Two axes

topics is what a thing is about. audience is who it's for β€” undergraduate, graduate, faculty-staff, alumni, public β€” and it exists because topic tagging worked and exposed the next problem: 170 events came back tagged research, and 112 of them were IRB office hours and grants-management training. All genuinely about research, none of it for a first-year.

is_for_students() is the gate, applied in today.py's add() so no pool can forget it. It drops an event only when the audience is known and names neither undergraduate nor public β€” and never for an event hosted by a group the student said they joined, because membership is a fact they stated and the audience tag is our inference. It filters what the app pushes, not what a student pulls: Ask Foresight still answers about anything, since a student asking directly should find it.

AUDIENCE_BY_CALENDAR settles two of the eight publishing calendars deterministically β€” Office of the Provost (165 events, 159 of them three recurring administrative series) and Alumni. The rest are read by the model, deliberately: Owen and the Law School look graduate-only but publish prospective-student sessions an undergraduate should see.

Beware that extra["audience"] means three different things depending on kind β€” an event's controlled vocabulary, an office's list of who it serves, and a housing process's bare cohort string. Three collectors picked the same obvious word. Anything reading it must scope by kind; there's a note in normalize.py.

Three ways a thing gets tagged, in descending order of trust:

  1. for_categories() β€” deterministic, from AnchorLink's own taxonomy. Free.
  2. the collectors' classify.py scripts, and classify_interests() β€” a nano model, for what carries no taxonomy (campus events have none; 328 orgs resolve to nothing).
  3. from_text() β€” a hand-written alias table, the no-key fallback. It alone resolves "brain science" β†’ health-medicine and "music production" β†’ music, which is why a keyless deployment still ranks on tags.

It lives under server/ and is imported by the collectors β€” the same direction as the building gazetteer, and collectors/directory/collect_directory.py explains why the reverse would fail in deployment. A vocabulary that drifted between build time and runtime would break every match silently, so the copy-with-a-comment pattern the four-domain definitions use is not safe for this one.

GET /api/topics serves the slugs and labels to the My Story chips for the same reason.

Deploying

.github/workflows/deploy-space.yml uploads app/, server/, knowledge-base/, the Dockerfile, and requirements.txt to the Space on every push to main that touches them, generating the Spaces YAML frontmatter so the project README doesn't have to carry it. Meeting notes, literature, docs, and collectors stay in GitHub.

Two separate secrets, easy to conflate:

  • the GitHub Actions HF_TOKEN deploys the code to the Space;
  • a Space secret HF_TOKEN (set in the Space's own settings) is what lets the running app read and write accounts and student data.

Setting only the first gives a Space that builds and serves but silently falls back to the ephemeral container directory β€” accounts would disappear on every restart.

Set these three in the Space's own settings: HF_TOKEN, FORESIGHT_SESSION_SECRET, and OPENAI_API_KEY. OPENAI_API_KEY has the same trap as HF_TOKEN β€” a repository secret does nothing for the running app. Then check /healthz reports "backend": "huggingface-dataset", chat.enabled: true, and knowledge_base.ready: true.