relo-calculator / docs /ARCHITECTURE.md
sabilmakbar's picture
serve pairs from per-city index (numbeo rankings): 361 cities, up to 10/country, all 92 capitals; myanmar->yangon
11b094a
|
Raw
History Blame Contribute Delete
9.37 kB
# Architecture & how it works
## Tools
A landing page (`/`) offers two independent tools:
- **Relocation calculator** (`/relo`, computes via `POST /compare`) β€” cost of living, FX, and projected savings between two cities.
- **Take-home estimator** (`/tax`) β€” rough gross β†’ net (after-tax) by country.
## Relocation calculator flow
```
User opens / (landing) β†’ picks "Relocation calculator" (/relo)
β”‚
β–Ό
POST /compare
β”‚
β”œβ”€β–Ί currencies.py
β”‚ Infers the ISO 4217 currency for each country
β”‚ (Malaysia β†’ MYR, Singapore β†’ SGD). No manual input.
β”‚
β”œβ”€β–Ί data_sources.py ┐
β”‚ Scrapes Numbeo β”‚ run concurrently
β”‚ cost/rent diffsβ”‚ (asyncio.gather)
β”‚ β”‚
└─► fx.py β”˜
β”‚ If cross-currency, fetches ~200d of daily rates and
β”‚ computes a blended next-month rate forecast (+ band).
β”‚
└─► model.py
Splits income by the budget sliders, scales costs by the
Numbeo percentages (FX-adjusted cross-country), and returns
monthly savings, the savings delta vs. home, and the
break-even salary.
```
## What you input
| Field | Required | Notes |
|---|---|---|
| From Country / City | Yes | Country is a **dropdown**; selecting one auto-fills the city with its capital (editable). City is otherwise free text |
| To Country / City | Yes | Where you're considering moving |
| Current net salary | Yes | Monthly take-home, in home currency (label shows the inferred code) |
| New net salary | One of these | Offer amount, **in destination currency** |
| Expected increase % of savings | One of these | Target savings growth vs. now; the app back-solves the salary needed |
| Savings rate (slider + number) | No | % of income saved (default 20%) |
| Rent share (slider + number) | No | Rent's % of the **remaining** spend (default 25%); the complement is other living costs |
Currencies are **inferred automatically** from the country names. Cross-country comparisons trigger an FX lookup; same-country (or same-currency) comparisons skip it. Submitted form values persist across submits, and all monetary figures (inputs and results) render with thousand separators. The two budget sliders show a **live breakdown** as you drag.
The country→currency and country→capital data live in [`../app/data/`](../app/data/) as JSON (`currency_by_country.json`, `capital_city.json`). Cities stay free text — there's no reliable offline list of Numbeo cities to populate a dropdown from.
## What you get
- How much more/less expensive daily life and rent are in the destination (Numbeo)
- **FX prediction** for next month with a sensitivity band, shown both directions
- Estimated monthly savings at home vs. destination β€” in both currencies, plus a % delta
- The break-even salary needed in the destination to maintain your current savings
- If you gave a savings target, the destination salary required to hit it
## The savings model
The budget split is **controlled by two sliders**:
- **Savings rate** β€” % of net income saved (default **20%**)
- **Rent share** β€” rent's % of the remaining spend (default **25%**)
With the defaults this works out to **20% rent Β· 60% other living Β· 20% savings** of income (`savings + rent + other = 100%` always). The weights flow into the model per request; `W_RENT` / `W_NON_RENT` in `model.py` are only the fallback defaults when the sliders aren't sent.
The model assumes you **replicate your home lifestyle** in the new city β€” each spending bucket is scaled by its own Numbeo index (rent by the Rent index, other costs by the Cost-of-Living-excl-rent index), independently. For cross-country moves the home spending baseline is converted to the destination currency before the Numbeo percentages are applied (the percentages are already FX-adjusted), which avoids mixing currency scales.
## FX prediction
The next-month exchange-rate forecast is a **weighted blend of three EMAs** over the last ~200 trading days:
```
forecast = 0.3Β·EMA-30 + 0.3Β·EMA-90 + 0.4Β·EMA-180
```
(renormalised over whatever horizons the history supports). EMA-30 adds recent responsiveness; EMA-180 anchors the long-term trend. The spread across the EMAs forms a sensitivity band, re-run through the model to give a savings range. A green β–² / red β–Ό arrow shows whether the forecast sits above or below the EMA-180 anchor. Very small rates are shown scaled (e.g. `1,000,000 IDR = 49 EUR`).
**Data sources** (tried in order, automatic fallback):
1. **Yahoo Finance** β€” daily history, broadest currency coverage; primary source. The client seeds Yahoo's consent cookies (via `fc.yahoo.com`) and retries across both API hosts to reduce 429s β€” though IP-based throttling on shared/free hosts can't be fully avoided, which is why the fallbacks exist.
2. **Frankfurter / ECB** β€” free, no API key, ~31 major currencies (incl. MYR, SGD); used if Yahoo is unavailable or rate-limited (HTTP 429).
3. **currency-api (spot)** — free, keyless, ~150 currencies (incl. Gulf currencies like SAR that ECB omits). No daily history → **spot rate only** (no EMA forecast). Spot lookups are cached in-process for an hour and are reversible (a cached A→B also answers B→A as 1/rate).
The blend renormalises when history is short. If **all** sources fail, the cost-of-living comparison still renders β€” only the savings estimate (which needs currency conversion) is skipped, with a clear notice.
> Not TradingView / XE / Wise: none offer a free, keyless public API for rates β€” TradingView means scraping (against ToS), and XE/Wise require paid or authenticated business accounts.
> EMA is a smoothing/trend tool, not a precise forecaster β€” FX is famously hard to predict and a random walk beats most models. Treat the numbers as a plausible range, not a guarantee.
## Take-home pay estimator (`/tax`)
A separate page estimates **gross β†’ net** take-home from a flat per-country effective rate ([`../app/data/tax_rates.json`](../app/data/tax_rates.json)). It's deliberately decoupled from the relocation calculator (which stays net-only and precise). The rates are **approximate, indicative figures with cited provenance** (a `_meta` block in the JSON): OECD members align with the OECD *net personal average tax rate* (Taxing Wages), zero-income-tax jurisdictions per PwC, others rounded from public tax summaries. The page shows the metric, sources, and links to PwC / Numbeo to verify.
## Limitations
- **Numbeo data quality** β€” indices are crowd-sourced and may lag reality, especially for smaller cities. Missing or mis-named cities return a clear error.
- **Budget weights are a proxy** β€” the split is configurable but still a simplification of real spending.
- **FX is a forecast, not a guarantee** β€” see above.
- **Currency coverage** β€” an unknown country falls back to no FX conversion (with a warning) rather than failing.
- **Tax estimates are indicative** β€” flat effective rates, not payroll-grade.
- **Scraping dependency** β€” if Numbeo/Yahoo change structure, the relevant fetch breaks; errors are surfaced, not swallowed. The Numbeo fetch uses `curl_cffi` to impersonate a real Chrome TLS/JA3 fingerprint (plus its header set), primes cookies with a warm-up request, and retries 429/503 with backoff. This beats Cloudflare *fingerprint*-based bot-detection but **not** a datacenter IP-reputation block β€” free hosts (Render, HF Spaces) sit on ranges Numbeo/Cloudflare 503s regardless of fingerprint. Production therefore serves from an **offline cache** (see below); only un-cached pairs attempt a live scrape.
## Numbeo cache
Because deployed free hosts get 503'd, cross-country comparisons in production are served from a **per-city cost index**, [`../app/data/numbeo_index.json`](../app/data/numbeo_index.json) β€” built **locally from a clean residential IP** by [`../scripts/build_cache.py`](../scripts/build_cache.py) and committed to the repo (so it survives the container's ephemeral filesystem and every spin-down). `get_percentage_diff` computes the pair from the two cities' indices when both are known, and only falls back to a live scrape otherwise β€” which then degrades to a "temporarily unavailable" notice if the host is blocked.
**How it's built β€” the index trick.** Numbeo's pairwise percentages are *reciprocal and transitive* (verified), i.e. they all derive from a single cost index per city. So instead of caching NΓ—(Nβˆ’1) pairs (hundreds of thousands across the whole dropdown), we store **one index per city** and compute any pair on the fly: `valuePct(Aβ†’B) = (index_B / index_A βˆ’ 1) Γ— 100` (ratios are basis-independent). The builder pulls Numbeo's **"Cost of Living Index by City"** ranking in a *single* request β€” ~550 cities with their Cost-of-Living and Rent indices (NYC = 100 basis) β€” keeps up to **10 cities per country** (always including the capital, aliased when Numbeo names it differently, e.g. "New York" β†’ "New York, NY"), and backfills the handful of capitals the ranking omits (Vientiane, Yangon, Bandar Seri Begawan) from static snapshots. That's one request for full coverage β€” ~360 cities across all 92 countries. COL data moves slowly, so a snapshot stays valid for months; re-run the script to refresh.