Spaces:
Running
Running
File size: 33,918 Bytes
2fc729c ddfd32f d53482f 2fc729c d53482f 2fc729c d53482f ddfd32f 2fc729c ddfd32f 2fc729c d53482f ddfd32f d53482f 2fc729c d99ff72 2fc729c d99ff72 2fc729c d99ff72 2fc729c ce7686b 2fc729c ce7686b 2fc729c ce7686b d99ff72 ce7686b d99ff72 ce7686b d53482f d99ff72 d53482f ce7686b d53482f d99ff72 d53482f d99ff72 ce7686b 2fc729c ce7686b d53482f ce7686b d53482f ce7686b d53482f ce7686b d53482f ce7686b d53482f ce7686b 2fc729c d53482f 2fc729c d53482f 2fc729c d53482f 2fc729c 3b6e0e5 dddbb8b 3b6e0e5 9ac8a8e ce7686b d53482f 2fc729c 3b6e0e5 dddbb8b ce7686b d53482f ce7686b d99ff72 ce7686b d53482f ce7686b d53482f ce7686b d53482f ce7686b d53482f 2c859ea d53482f 2c859ea d53482f 2c859ea d53482f 2c859ea d53482f ddfd32f d53482f ddfd32f d53482f ddfd32f d53482f 2c859ea d53482f 2c859ea d53482f 2c859ea d53482f 2c859ea d53482f 2c859ea d53482f 2c859ea d99ff72 ddfd32f 3b6e0e5 2fc729c | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 853 854 855 856 857 858 859 860 861 862 863 864 865 866 867 868 869 870 871 872 873 874 875 876 877 878 879 880 881 882 883 884 885 886 887 | # Design Proposal (DP) โ forecaster-agent
**Version:** 0.11 (Self-Evolving Transition Recommendations)
**Last updated:** 2026-06-24
Architectural design implementing [PRD.md](./PRD.md) under Harness Engineering standards.
---
## 1. Layout & import strategy
The project uses a **flat root layout** (not `src/` yet). All modules live at repository root. Imports follow one rule:
```python
try:
from schemas import Prediction # script / pytest entry
except ImportError:
from .schemas import Prediction # optional package import
```
**Never** add `/home/sean` to `sys.path` (symlink collision with duplicate SQLAlchemy metadata). Use `paths.PROJECT_ROOT` instead.
```
forecaster-agent/
โโโ config.yaml
โโโ paths.py # PROJECT_ROOT constant
โโโ run.py # CLI entry
โโโ loop.py # orchestrator
โโโ forecast.py # LLM seam + prediction/judge
โโโ evolution.py # case library, PCA/GMM, OOD
โโโ crowd.py # crowd gate
โโโ job_radar.py # hybrid RAG + search retrieval + transition paths
โโโ registry.py # SQLite registry
โโโ schemas.py # SQLModel entities + engine
โโโ services/
โ โโโ read_model.py # read-only API seam (MCP/REST target)
โ โโโ job_query_agent/ # Phase 9: retrieval QA loop
โ โ โโโ discover.py
โ โ โโโ evaluate.py
โ โ โโโ propose.py
โ โ โโโ simulate.py
โ โ โโโ apply.py
โ โ โโโ loop.py
โ โ โโโ traces.py
โ โ โโโ audit.py
โ โโโ transition_evaluator/ # Phase 10: LLM-as-judge self-evolution
โ โโโ __init__.py
โ โโโ cache.py # persistent data/transition_eval_cache.json
โ โโโ evaluate.py # evaluate_pair + run_evaluation_pass + _promote_to_kb
โโโ dashboard.py # Streamlit UI
โโโ tests/ # offline harness (HR-1), 190 tests
โโโ data/
โ โโโ forecaster.db
โ โโโ jobs_kb.json
โ โโโ query_seed.json # Phase 9 seed queries (115+)
โ โโโ transition_eval_cache.json # Phase 10 cached LLM scores
โโโ pending/
โ โโโ job_calibration/ # Phase 9 human-review proposals
โโโ docs/
โโโ PRD.md
โโโ DP.md
```
---
## 2. Orchestrator data flow (v0.5)
```
ingest.gather_signals()
โ
โผ
evolution.build_prior(scenario from config)
โ EvolutionPrior.to_prompt_context()
โผ
forecast.generate_predictions(signals, track, evolution_prior=...)
โ
โผ
registry.add_many() โโ dedup by fingerprint
โ
โผ
publish.publish_or_queue(require_review)
```
### Loop pseudocode
```python
def run_cycle(cfg):
reg = Registry()
resolve_due(reg, cfg["model"])
signals = ingest.gather_signals(...)
prior_ctx = evolution.build_prior(
current_scenario=cfg["evolution"].get("scenario") or CURRENT_AI_SCENARIO,
n_bootstrap=cfg["evolution"]["n_bootstrap"],
).to_prompt_context()
preds = forecast.generate_predictions(
signals, reg.track_record_summary(),
evolution_prior=prior_ctx, ...)
new = reg.add_many(preds)
publish.publish_or_queue(..., require_review=cfg["require_review"])
```
**Not yet in loop:** Discord/Telegram bots โ REST API only for Phase 2.
---
## 3. Harness interfaces
### 3.1 Embedder (`crowd.py` + `job_radar.py`)
```python
class Embedder(Protocol):
def embed(self, texts: list[str]) -> list[list[float]]: ...
```
**Crowd gate (`crowd.py`):**
- **Production:** Voyage / OpenAI / sentence-transformers
- **Harness:** `HashingEmbedder` (deterministic MD5 bag-of-tokens)
**Job Radar search (`job_radar.resolve_embedder`):**
| `job_radar.search.embedder` | Class | Use |
|-----------------------------|-------|-----|
| `sentence_transformers` (default in `config.yaml`) | `SentenceEmbedder` | HF Space / local dashboard โ `paraphrase-multilingual-MiniLM-L12-v2` |
| `hashing` (default in `config.ci.yaml`) | `RadarHashingEmbedder` | CI / pytest (HR-1 offline, no model download) |
Semantic path embeds **`_job_embed_title(job)`** (title + `title_zh` + `search_aliases` only).
Lexical path still uses the full **`_job_embed_text(job)`** document (description, skills, industry).
### 3.2 SoundnessJudge (`crowd.py`)
- **Production:** `LLMSoundnessJudge`
- **Harness:** `HeuristicSoundnessJudge`
### 3.3 LLM (`forecast.py`)
Single seam:
```python
def call_llm(system: str, user: str, *, max_tokens: int, model: str | None) -> str: ...
```
Priority: `GROQ_API_KEY` โ `ANTHROPIC_API_KEY` โ `RuntimeError`.
Future: `MockLLMClient` behind same signature for fully offline `run.py once --mock`.
---
## 4. Persistence
### 4.1 SQLite (`schemas.py`)
```python
DATABASE_URL = "sqlite:///data/forecaster.db"
```
| Table | Purpose |
|-------|---------|
| `prediction` | Core forecasts |
| `contribution` | Crowd submissions (Phase 2) |
| `predictionbet` | Virtual prediction market |
| `jobfeedback` | Career survey + 6-month follow-up (`JobFeedback`; Phase 8 adds `experience_level`) |
`registry_path` in old configs is **deprecated**; use `database_path` for documentation only.
### 4.2 Static publish outputs
| Backend | Output |
|---------|--------|
| `FilePublisher` | `site/index.html`, `site/feed.json` |
| `WebhookPublisher` | Slack/Discord JSON |
| `GitPublisher` | commit + push Pages repo |
Review gate writes to `pending/*.md` + sidecar `.json`.
---
## 5. Services layer (read model)
`services/read_model.py` is the **only** module future MCP/REST servers should call for reads:
| Function | Backing |
|----------|---------|
| `get_scoreboard()` | `Registry.scoreboard()` |
| `get_ood_assessment(scenario)` | `evolution.build_prior()` |
| `search_jobs(query, industry, scenario)` | `job_radar.*` |
| `list_open_predictions()` | `Registry.open_predictions()` |
**Rule:** No Streamlit imports in `services/`. No LLM calls in read model.
---
## 6. Phase 3 โ MCP tools (implemented)
| MCP Tool | Handler | Maps to |
|----------|---------|---------|
| `get_calibration_scoreboard` | `handle_get_calibration_scoreboard` | `services.get_scoreboard()` |
| `get_ood_assessment` | `handle_get_ood_assessment` | `evolution.build_prior()` |
| `search_jobs` | `handle_search_jobs` | `job_radar.*` |
| `list_open_predictions` | `handle_list_open_predictions` | `Registry.open_predictions()` |
Entry point: `mcp_server.py` (stdio). Configuration: `docs/MCP.md`.
Handlers live in `services/mcp_handlers.py` โ testable without MCP runtime (HR-1).
### Phase 3b โ REST (implemented)
| REST endpoint | Handler |
|---------------|---------|
| `GET /v1/scoreboard` | `handle_get_calibration_scoreboard` |
| `GET/POST /v1/ood` | `handle_get_ood_assessment` |
| `GET/POST /v1/jobs/search` | `handle_search_jobs` |
| `GET /v1/predictions/open` | `handle_list_open_predictions` |
Entry point: `api_server.py`. Configuration: `docs/API.md`, `config.yaml` โ `api`.
Auth: `FORECASTER_API_KEY` env (optional). Rate limit: `api.rate_limit_per_minute`.
Write endpoints (`POST /contributions`, `POST /forecast`) require Phase 2 + BUSL review.
---
## 7. Job Radar โ search vs recommendation (HR-12)
Job Radar exposes **two independent scoring pipelines**. They must not be merged in UI or API responses.
### 7.1 Search retrieval (query โ occupation anchor)
**Purpose:** find which KB occupation best matches a free-text query (user's current role, industry keyword, etc.).
Location: `job_radar._score_job_match()`, `_apply_search_scores()`, `find_best_match()`
```
combined = w_emb ยท cosine(q, job_title_vec) + w_lex ยท lexical_overlap(q, job_full_text)
ร multi_word_penalty (if โฅ2 query tokens and too few hits)
+ industry_only_boost (optional, industry-only queries)
```
Where `job_title_vec` comes from `SentenceEmbedder(_job_embed_title)` in production, or
`RadarHashingEmbedder(_job_embed_text)` when `embedder: hashing`.
Current defaults live in `config.yaml` โ `job_radar.search.*` (HR-3):
| Key | Default | Meaning |
|-----|---------|---------|
| `embedder` | `sentence_transformers` | `hashing` in `config.ci.yaml` for offline CI |
| `embed_weight` / `lex_weight` | 0.45 / 0.55 | Blend semantic cosine + lexical overlap |
| `tier_no_match` | 0.42 | Below โ no confident KB match; LLM KB expansion path |
| `tier_weak` | 0.55 | Weak tier ceiling; also `query-agent` `min_sim_after` default |
| `tier_strong` | 0.65 | Strong search hit banner |
| `title_aliases` | `{}` + code `_QUERY_TITLE_ALIASES` | `normalize_search_query()` merges file + built-in map |
Multi-word penalty: queries with โฅ2 tokens require hits on at least `max(2, โn/2โ)` tokens or score ร `multi_word_penalty` (0.45).
**Lexical tokens:** `_SEARCH_TOKEN` matches ASCII alphanumerics and CJK runs (`[\u4e00-\u9fff]{2,}`).
`_query_tokens()` keeps tokens **โฅ 2 characters** (preserves `ml`, `hr`, `qa`, `gp`; drops single-char ASCII only).
**Bilingual semantic:** `paraphrase-multilingual-MiniLM-L12-v2` embeds query and title documents in a shared space โ Chinese queries like `ๆคๅฃซ` / `ไบบๅทฅๆบ่ฝๅทฅ็จๅธ` match without per-query manual aliases.
**Hot-role guardrail:** `CORE_HOT_ROLE_QUERIES` โ CI + query-agent audit assert each pair hits `expected_id` with `sim โฅ tier_weak`.
**Not a transition recommendation.** A high `combined_similarity` only means text overlap with an occupation profile.
### 7.2 Scenario impact (structural + semantic hybrid)
Unchanged โ ranks occupations under an AI diffusion scenario:
$$\text{hybrid}_j = \alpha \cdot \text{impact}_j + \beta \cdot \text{sim}(q, j)$$
$$\text{impact}_j = \text{base\_demand\_trend}_j + \sum_i s_i \cdot w_{j,i}$$
Config: `job_radar.alpha`, `job_radar.beta` (HR-3).
Used for timeline / impact views, not for "should I switch to this role?".
### 7.3 Transition recommendation (`compute_transition_paths`)
**Purpose:** given an **anchor occupation** (selected from search or dropdown), rank target roles by career-switch feasibility.
Default weights (`_TRANSITION_WEIGHTS`, override via `config.yaml` โ `job_radar.transition.*`):
| Component | Weight | Signal |
|-----------|--------|--------|
| `skill` | 0.40 | 1 โ normalised Euclidean distance on 8-dim skill vectors |
| `overlap` | 0.15 | Jaccard on `required_skills` |
| `risk` | 0.25 | `displacement_risk_current โ displacement_risk_target` (โฅ 0) |
| `demand` | 0.20 | Normalised scenario `impact_score` of target |
Filters: exclude self; exclude `category == at_risk` targets.
Phase 8 personalization (implemented): session profile `{experience_level, max_retrain_months}` in `ui/sidebar.py` adjusts weights via `personalization_weights()` and filters candidates whose `retrain_months` exceed the cap; juniors up-weight low skill-gap targets.
**UI flow (implemented):**
```
User query โโโบ find_best_match() โโโบ anchor role
โ
โผ
compute_transition_paths(anchor, all_jobs, scenario)
โ
โผ
transition cards (top_k)
```
Search result lists (browse mode) re-rank by `transition_score` **after** anchor is fixed โ not by `combined_similarity` alone.
**Anchored-search at-risk suppression (HR-12 UX, v0.10):**
When `search_query` resolves to a strong anchor (`anchor_job` set), `ui/tabs/radar.py` clears
`at_risk_list` โ weak text scores must not show unrelated at-risk occupations (regression:
`software engineer` โ Logistics Dispatcher at simโ0.17). User sees
`radar_matrix_at_risk_anchor_skip` caption and transition paths from the anchor instead.
Test: `tests/test_radar_render.py::test_anchored_search_hides_unrelated_at_risk`
### 7.4 Field feedback crowd (`JobFeedback`) โ HR-13
Separate from prediction crowd (`contribution` table + `services/crowd_service.py`).
| Field | Phase 8 |
|-------|---------|
| `job_title` | โ
KB dropdown โ canonical English title |
| `industry`, `status`, `confidence`, `transition_target` | โ
|
| `experience_level` | โ
`junior` \| `mid` \| `senior` |
Aggregation: `get_empirical_metrics()` โ `compute_field_calibration()` in `services/job_market.py`.
Today: prefer `(title, experience_level)` cell when nโฅ`min_stratified_responses`; else fall back to title-only pool.
Survey UI: `ui/tabs/radar.py` form โ map localized labels back to canonical titles (existing pattern).
**Prediction crowd does not collect title/tenure** unless a future expert-weighting feature is scoped separately.
## 8. Test harness
```
tests/
โโโ test_crowd.py # gate decisions with behavioural assertions
โโโ test_evolution.py # case library, OOD, PCA determinism
โโโ test_registry.py # dedup, Brier, due(), scoreboard
โโโ test_job_radar.py # search tiers, transition paths, CORE hot roles
โโโ test_job_query_agent.py # Phase 9 audit / simulate / apply / loop
โโโ test_radar_render.py # Streamlit smoke + anchored at-risk regression
โโโ test_mcp_handlers.py
โโโ test_api.py # REST /v1 via TestClient
```
Run: `python -m pytest tests/ -v` โ **no network, no API key**.
CI: `.github/workflows/ci.yml` โ Python 3.11 + 3.12 matrix; post-pytest `query-agent audit`.
Daily: `.github/workflows/daily-pages.yml` โ forecast cycle + `query-agent run` + KB/config commit.
---
## 9. Known technical debt
| Item | Priority | Status |
|------|----------|--------|
| `dashboard.py` monolith (~1150 LOC) | P1 | โ
104-LOC orchestrator + `ui/tabs/*` |
| Discord/Telegram crowd bots | P2 | โ
`bots/` |
| No `MockLLMClient` for offline `run.py once` | P2 | open |
| `src/` package layout + Poetry lock | P3 | open |
| Alembic migrations | P3 | open |
| SQLite engine global singleton, poor test isolation | P3 | โ
`make_engine()` + `Registry(engine=โฆ)` + `isolated_registry` fixture |
| Search thresholds hard-coded in `job_radar.py` | P0 | โ
Phase 8 โ `config.yaml` (HR-3) |
| Search UI conflated with transition ranking | P1 | โ
Phase 8 (HR-12) |
| `JobFeedback` missing experience level | P1 | โ
Phase 8 (HR-13) |
| No session user profile for retrain cap | P2 | โ
Phase 8 |
| Hot-role search gaps caught only by user reports | P0 | โ
Phase 9 query-agent |
| No automated alias calibration loop | P1 | โ
`query-agent run` + daily cron |
| Anchored search shows unrelated at-risk roles | P1 | โ
HR-12 UX fix + render test |
---
## 10. v0.6 Design additions (Integrity & Learning Loop)
### 10.1 Citation sanitiser (HR-9)
Location: `forecast._sanitize_sources(sources: list[str]) -> list[str]`
Rules applied at `_parse_json` time before a `Prediction` is persisted:
| Check | Action |
|-------|--------|
| Not a string or empty | drop |
| Doesn't start with `http://` or `https://` | drop |
| Domain is `example.com`, `example.org`, or `placeholder` | drop |
| arXiv URL whose YYMM is in the future | drop (hallucinated) |
| All other URLs | keep as-is (no live HEAD check, stays offline-safe) |
Tests: `tests/test_forecast_quality.py::test_sanitize_sources`
### 10.2 Horizon normalisation (HR-consistent formatting)
Location: `forecast._normalize_horizon(h: str) -> str`
| Input | Output | Rule |
|-------|--------|------|
| `"2027"` | `"2027-Q4"` | bare year โ year-end quarter |
| `"2027-H1"` | `"2027-Q2"` | half-year โ closing quarter |
| `"2027-H2"` | `"2027-Q4"` | half-year โ closing quarter |
| `"2027-Q3"` | `"2027-Q3"` | already canonical, pass through |
Applied in `_parse_json` alongside `_sanitize_sources`.
Tests: `tests/test_forecast_quality.py::test_normalize_horizon`
### 10.3 Resolved-state durability (HR-10)
`run.py export` serialises the full `Prediction` via `model_dump_json()`, which
includes `status`, `outcome`, `brier`, `resolved_at`. On reload,
`Prediction.model_validate_json(line)` restores the resolved state. `_coerce`
in `dashboard_seed.py` back-fills missing `brier` / `resolved_at` only when the
status already indicates resolution, so it is idempotent.
Invariant test: `tests/test_dashboard_seed.py::test_ensure_demo_registry_loads_seed_plus_live`
already asserts both seed and live statements are present after a round-trip.
### 10.4 Groq rate-limit degradation (P2-A)
`forecast.call_llm` wraps the Groq call in a retry block:
```
try:
return _call_once(system, user, max_tokens, model)
except RateLimitError:
time.sleep(60)
return _call_once(system, user, max_tokens, model)
except Exception as exc:
raise RuntimeError(f"LLM call failed: {exc}") from exc
```
- One retry with 60-second back-off avoids wasted runs on transient 429.
- Still raises after one retry so CI never hangs forever.
- `RateLimitError` import: `from groq import RateLimitError` (guarded with
`ImportError` fallback to keep the harness offline-safe).
### 10.5 User feedback visibility (P2-B)
`ui/tabs/radar.py` shows a `st.info` badge with count of calibrated jobs when
`_field_recs` (from `compute_field_calibration`) is non-empty:
```
โน๏ธ 3 occupations calibrated from real user feedback (N=47 responses)
```
This closes the perception gap: users know their survey answers matter.
---
## 12. v0.8 Design additions (Live Track Record Credibility)
### 12.1 Origin classification (HR-11)
Location: `services/track_record.py`
| Function | Purpose |
|----------|---------|
| `seed_prediction_ids(seed_path)` | Load fingerprint set from `predictions_seed.json` |
| `prediction_origin(p, seed_ids)` | Returns `"seed"` if `p.id` โ seed_ids else `"live"` |
| `partition_by_origin(preds, seed_ids)` | Split into `(seed_preds, live_preds)` |
| `scoreboard_subset(preds)` | Same shape as `Registry.scoreboard()` for a filtered list |
| `upcoming_resolutions(preds, *, limit=10)` | Open/due preds sorted by `resolution_date` asc |
**Rule:** seed IDs are computed once per dashboard render from the committed seed file;
live predictions are everything else in the registry (including cron-generated rows).
### 12.2 Live-only export
`run.py export` filters with `partition_by_origin` before writing JSONL.
Seed data never enters `predictions_live.jsonl` โ it remains in `predictions_seed.json`.
### 12.3 Export verification (CI guard)
`run.py verify-export`:
1. Load live predictions from DB (`partition_by_origin`)
2. Load `predictions_live.jsonl`
3. For each live id, assert matching `status`, `outcome`, `brier` in JSONL
4. Exit code 1 on any mismatch (prevents silent resolve/export regression)
Called in `.github/workflows/daily-pages.yml` immediately after `export`.
### 12.4 Track Record UI layout
```
โโ Curated benchmark (seed) โโโโโโโโโโโโโโโโโโ
โ resolved N | mean Brier X | table โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โโ Live LLM predictions โโโโโโโโโโโโโโโโโโโโโโ
โ open N | resolved N | mean Brier Y โ
โ โถ Upcoming resolutions (live only) โ
โ resolved table (live only) โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
```
CSV download includes `origin` column on every row.
---
## 13. v0.9 Design additions (Personalized Job Radar โ implemented)
### 13.1 Harness invariants
| ID | Enforcement location |
|----|---------------------|
| **HR-12** | `ui/tabs/radar.py` โ search banner uses `combined_similarity`; transition expander uses `compute_transition_paths` only; tests assert no cross-ranking |
| **HR-13** | `schemas.JobFeedback.experience_level`; survey + `get_empirical_metrics()` stratification |
| **HR-3** | New `config.yaml` keys under `job_radar.search` and `job_radar.personalization` |
### 13.2 Config schema (implemented)
```yaml
job_radar:
alpha: 0.6
beta: 0.4
search:
embedder: sentence_transformers # hashing in config.ci.yaml (HR-1)
embed_weight: 0.45
lex_weight: 0.55
multi_word_penalty: 0.45
industry_only_boost: 0.08
tier_no_match: 0.42
tier_weak: 0.55
tier_strong: 0.65
title_aliases: {} # merged with _QUERY_TITLE_ALIASES in code
transition:
skill: 0.40
overlap: 0.15
risk: 0.25
demand: 0.20
personalization:
junior_retrain_cap_months: 6
mid_retrain_cap_months: 12
senior_retrain_cap_months: 24
min_stratified_responses: 5
```
### 13.3 Experience level โ transition weights (sketch)
Pure function in `job_radar.py` (HR-2 testable):
```python
def personalization_weights(
base: dict,
*,
experience_level: str,
max_retrain_months: int | None = None,
) -> dict:
...
```
- **Junior:** increase `overlap` weight, filter targets with `retrain_months > cap`
- **Senior:** increase `demand` + `risk` weights slightly; allow higher retrain cap
- Defaults when profile unset: current `_TRANSITION_WEIGHTS`
### 13.4 Test plan (offline, HR-1)
| Test | Asserts |
|------|---------|
| `test_score_job_match_multi_word_penalty` | โ
exists |
| `test_find_best_match_finance_not_ai_trading` | finance query โ unrelated AI role (regression) |
| `test_transition_paths_respect_retrain_cap` | Phase 8 โ filtered when profile set |
| `test_empirical_metrics_stratified_by_experience` | Phase 8 โ `(title, level)` grouping |
| `test_search_tiers_from_config` | Phase 8 โ thresholds read from config stub |
### 13.5 i18n keys (Phase 8)
Add to `ui/i18n.py`: `radar_fb_experience`, `radar_match_tier_*`, `radar_profile_*`, `radar_matrix_at_risk_anchor_skip`.
---
## 14. v0.10 Design additions (Job Query Calibration Agent โ implemented)
### 14.1 Module layout
```
services/job_query_agent/
โโโ discover.py # core_hot, query_seed.json, JobFeedback titles
โโโ evaluate.py # tier + P0 regression classification (HR-2 pure)
โโโ propose.py # CalibrationProposal โ pending/job_calibration/
โโโ search_log.py # record + aggregate Radar/HF search JSONL (P2)
โโโ simulate.py # dry-run alias/title_alias/kb_profile_new before apply
โโโ apply.py # patch KB/config; kb_profile_new via LLM cache; run_apply_pending
โโโ loop.py # multi-round discover โ auto-apply โ re-check
โโโ traces.py # JSONL append-only (gitignored path)
โโโ audit.py # single-pass audit; raises on P0 failure
```
### 14.2 Evaluation verdicts
`evaluate_query(discovered, jobs)` โ `QueryVerdict`:
| `status` | Condition | Agent action |
|----------|-----------|--------------|
| `ok` | Match tier acceptable | trace only |
| `p0_regression` | CORE query: `best_id โ expected_id` | **CI fail** |
| `weak_core` | CORE query: correct id but `sim < tier_weak` | propose `alias_patch`; auto-apply if sim improves |
| `kb_gap` | Non-core: `sim < tier_no_match` | propose `kb_profile_new` โ **always queue** |
| `weak_match` | Non-core: weak tier | propose `alias_patch` to `best_id` |
### 14.3 Proposal types
| type | payload | persist target | auto-apply |
|------|---------|----------------|------------|
| `alias_patch` | `add_aliases: [query]` | `jobs_kb.json` โ `search_aliases` | โ
gated |
| `title_alias` | `canonical: <KB title>` | `config.yaml` โ `job_radar.search.title_aliases` | โ
gated |
| `kb_profile_new` | `query`, `nearest_id` | `generate_job_profile_via_llm` โ `jobs_kb.json` | โ
gated (`min_search_log_occurrences`) or `query-agent apply` |
### 14.4 Closed-loop flow (`loop.run_calibration_cycle`)
```
for round in 1..max_rounds:
jobs = load_knowledge_base()
for item in discover_queries(cfg):
verdict = evaluate_query(item, jobs)
append_trace(...)
if verdict.ok: continue
proposal = propose_from_verdict(verdict)
sim_after = simulate(proposal)
if can_auto_apply(sim_before, sim_after, expected_id):
apply_proposal() # writes KB or config
else:
queue_proposal(pending/) # skipped when --dry-run
break if no auto_applies this round
return final audit summary (no raise; caller may exit 1 on P0)
```
`can_auto_apply` rules (`apply.py`):
- `auto_apply.enabled` and `proposal.type` โ `auto_apply.types`
- `sim_after โฅ min_sim_after` (default = `tier_weak`)
- CORE: `best_id_after == expected_id`; `target_id == expected_id`
- `alias_patch`: `sim_after > sim_before`
### 14.5 Audit flow (CI โ read-only)
```
discover_queries(cfg)
โ for each query: find_best_match + search_match_tier
โ evaluate_query โ append_trace
โ optional queue_proposal (once subcommand only)
โ raise AssertionError if p0_regression or weak_core
```
### 14.6 CLI
```bash
python run.py query-agent audit # CI: exit 1 on P0 / weak-core
python run.py query-agent once # audit + queue proposals (no auto-apply)
python run.py query-agent run # closed loop: simulate + auto-apply safe fixes
python run.py query-agent run --dry-run
python run.py query-agent apply # apply all pending/job_calibration/*.json
python run.py query-agent ingest-logs path.jsonl # merge HF export into search log
python run.py query-agent transition-eval [N] # Phase 10: LLM-evaluate up to N pairs
```
| Workflow | Step |
|----------|------|
| `.github/workflows/ci.yml` | `pytest` โ `query-agent audit` |
| `.github/workflows/daily-pages.yml` | `query-agent run` with `GROQ_API_KEY` โ commit `jobs_kb.json` / `config.yaml` / search log |
### 14.7 Config (`config.yaml`)
```yaml
job_query_agent:
enabled: true
traces_path: data/query_agent_traces.jsonl # .gitignore
discover:
include_core: true
include_seed: true
seed_path: data/query_seed.json
include_feedback_titles: true
feedback_min_responses: 1
max_queries_per_run: 200
evaluate:
fail_on_weak_core: true
review:
require_review: true
pending_dir: pending/job_calibration
auto_apply:
enabled: true
types: [alias_patch, title_alias]
min_sim_after: 0.55
max_rounds: 3
```
`run.py load_config()` sets `cfg["_config_path"]` so `apply.py` can write title aliases back to the loaded config file.
### 14.8 Test plan (offline, HR-1)
| Test | Asserts |
|------|---------|
| `test_run_audit_passes_on_real_kb` | CORE guard green on committed KB |
| `test_run_audit_fails_on_regression` | P0 poisons CI |
| `test_simulate_alias_patch_improves_match` | simulate monotonic improvement |
| `test_can_auto_apply_requires_improvement` | gate blocks no-op patches |
| `test_run_calibration_cycle_dry_run` | loop completes without writes |
| `test_anchored_search_hides_unrelated_at_risk` | HR-12 UX regression |
| `test_discover_from_search_log_weighted` | P2 frequency-weighted discovery |
| `test_run_apply_pending_kb_profile` | P3 pending โ KB append |
### 14.9 Search log ingest (P2)
**Record:** `ui/tabs/radar.py` calls `record_search_log()` on every non-empty search (tier, `best_id`, `sim`).
**Storage:** `data/radar_search_log.jsonl` (append-only JSONL; committed by daily cron when present).
**Discover:** `discover_from_search_log()` โ queries with `occurrences โฅ search_log_min_occurrences`, sorted by weight (frequency).
**HF export:** copy Space log file โ `python run.py query-agent ingest-logs export.jsonl` merges into the repo log for the next `query-agent run`.
---
## 15. v0.10b Semantic job search (implemented)
### 15.1 Dual embedder config (HR-1)
| File | `job_radar.search.embedder` | Runtime |
|------|----------------------------|---------|
| `config.yaml` | `sentence_transformers` | HF Space, local dashboard |
| `config.ci.yaml` | `hashing` | GitHub Actions pytest + `query-agent audit` |
`job_radar.resolve_embedder(search_cfg)` returns `SentenceEmbedder` or `RadarHashingEmbedder`.
### 15.2 Title-only semantic vectors
```python
def _job_embed_title(job) -> str:
# title + title_zh + search_aliases ONLY
def _job_embed_text(job) -> str:
# full document for lexical overlap (description, skills, industry, โฆ)
```
Rationale: embedding the full description diluted role-name similarity (e.g. unrelated jobs sharing โAIโ, โdataโ, โmanagementโ vocabulary).
### 15.3 Representative calibration lifts (prod embedder)
| Query | Before (hashing / diluted) | After (MiniLM title-only) |
|-------|---------------------------|---------------------------|
| `software developer` | ~0.60 | **0.845** |
| `lawyer` | ~0.00 | **0.944** |
| `ๆคๅฃซ` | ~0.00 | **0.869** |
| `ไบบๅทฅๆบ่ฝๅทฅ็จๅธ` | โ | **0.956** |
### 15.4 Lexical abbreviation fix
`_query_tokens()` filters to `len(token) >= 2`, preserving `ml`, `hr`, `qa`, `gp` for overlap scoring.
### 15.5 Deployment
- `requirements.txt`: `sentence-transformers>=3.0`
- `Dockerfile`: pre-download `paraphrase-multilingual-MiniLM-L12-v2`
- Release notes: [RELEASE_v0.10.md](./RELEASE_v0.10.md)
---
## 16. Phase 10 โ Self-Evolving Transition Recommendations (v0.11, implemented)
### 16.1 Problem with prior transition scoring
Three compounding defects caused poor career-path recommendations:
| Defect | Root cause | Symptom |
|--------|------------|---------|
| Dead overlap signal | `_skill_jaccard` always returned 0 (BLS skill vectors โ actual skill text) | 15% weight completely ignored |
| Demand/risk dominance | `risk_reduction + demand_norm` pushed all at-risk roles toward highest-demand targets | Accountant โ AI Engineer, regardless of industry |
| Missing LLM validation | Curated `transition_targets` existed for ~50 roles; new roles had empty targets | No guidance for new KB entries |
### 16.2 Scoring formula (v0.11)
```python
skill_sem = _skill_semantic_sim(current_job, tgt) # [0, 1]
domain_prox = _domain_proximity(current_job, tgt) # 0.0 / 0.5 / 1.0
overlap_signal = 0.65 * skill_sem + 0.35 * domain_prox
score = (
w["skill"] * proximity # 40%: economic skill sensitivity
+ w["overlap"]* overlap_signal # 15%: real skill + domain adjacency
+ w["risk"] * risk_reduction # 25%: displacement risk improvement
+ w["demand"] * demand_norm # 20%: forward demand forecast
)
if llm_conf is not None:
score *= (0.5 + llm_conf * 0.5) # LLM confidence gate
```
### 16.3 `_skill_semantic_sim(a, b)`
```python
def _skill_semantic_sim(a: dict, b: dict) -> float:
# text = ", ".join(required_skills) for each job
# embed with SentenceEmbedder (lazy-loaded singleton)
# cosine similarity; cached by sorted (id_a, id_b)
# returns 0.0 when no embedder (CI path, HR-1)
```
Cache key: `(min(id_a, id_b), max(id_a, id_b))` โ float in `_SKILL_SEM_CACHE`.
### 16.4 `_domain_proximity(a, b)`
Adjacency graph (`_ADJACENT_INDUSTRIES`) encodes industry clusters:
```python
_ADJACENT_INDUSTRIES = {
"Finance": {"Legal", "Government"},
"Legal": {"Finance", "Government"},
"Tech": {"Media"},
"Manufacturing": {"Construction", "Logistics"},
...
}
# same industry โ 1.0; adjacent โ 0.5; unrelated โ 0.0
```
FinanceโFinance transitions are naturally domain-proximate (1.0); FinanceโTech jumps pay a 0.5 penalty even if skill similarity is moderate.
### 16.5 `services/transition_evaluator/` architecture
```
cache.py
_load(path) / _save(path)
key(a, b) = "{anchor_id}โ{candidate_id}"
get(aid, cid) / put(aid, cid, feasibility, reasoning)
missing_pairs(jobs, top_k=5) # only jobs with empty transition_targets
evaluate.py
_SYSTEM โ 5-tier feasibility prompt (see ยง16.6)
_prompt(anchor, candidate) โ injects required_skills + skill gap
evaluate_pair(anchor, candidate, force=False) โ dict | None
run_evaluation_pass(jobs, max_pairs=40, min_feasibility=0.55) โ summary
_promote_to_kb(anchor_id, cand_id, result, kb_path)
โ appends {target_id, retrain_months, skill_bridge, confidence, _source: "llm_eval"}
โ retrain_months = max(2, round(24 ร (1 โ feasibility)))
```
### 16.6 LLM judge system prompt (5-tier scale)
```
0.85โ1.0 Natural progression โ same domain, 0โ6 months upskilling
0.60โ0.84 Adjacent move โ related domain, 6โ18 months
0.35โ0.59 Deliberate pivot โ adjacent industry, 18โ36 months
0.10โ0.34 Major reskill โ different domain, substantial retraining
0.00โ0.09 Extreme leap โ almost no skill overlap
```
Returns `{"feasibility": float, "reasoning": "max 20 words", "recommended": bool}`.
### 16.7 Integration into calibration loop
```python
# services/job_query_agent/loop.py โ after query calibration rounds
if not dry_run and auto_apply.enabled:
transition_summary = run_evaluation_pass(
fresh_jobs, kb_path=kb_path,
max_pairs=cfg.get("transition_eval_max_pairs", 40),
)
return {..., "transition_eval": transition_summary}
```
The daily CI workflow thus runs:
```
pytest โ query-agent audit โ (daily cron) query-agent run
โ round 1..N: query calibration
โ post-calibration: transition-eval (fill empty targets)
โ commit jobs_kb.json if any promotions
```
### 16.8 Config
```yaml
job_query_agent:
auto_apply:
enabled: true
transition_eval_max_pairs: 40 # pairs per daily run
```
No new harness invariant required โ evaluation uses existing Groq LLM seam (HR-6) and falls back silently when key absent (non-fatal, `skipped` count in summary).
---
## 11. Security & compliance defaults
- `require_review: true` โ never change default in repo
- Disclaimer in every published report (`publish.DISCLAIMER`)
- OOD warning must appear in forecast prompt when `is_ood`
- BUSL-1.1 โ commercial API deployment needs license
|