Spaces:
Sleeping
Sleeping
File size: 28,030 Bytes
379bf78 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 | """ECL assembly + allowance movement decomposition -- the deterministic engine.
THE FORMULA (tests/fixtures/compute_ecl.py section 3 -- the golden convention)
------------------------------------------------------------------------------
Per exposure, over projected periods t = 1..T:
ECL = sum_t S(t-1) * lambda_t * LGD_t * EAD_t * (1 + EIR)^-t
S(t-1) competing-risk survival to the START of period t,
S(t) = prod_{k<=t} (1 - lambda_default_k - lambda_prepay_k)
(engine/hazard.py pd_term_structure algebra, reproduced exactly)
lambda_t conditional default hazard in period t
LGD_t expected loss-given-default for a default in period t
EAD_t CONTRACTUAL amortisation balance entering period t
EIR per-period effective interest rate; losses crystallise at the
END of period t, hence the (1+EIR)^-t discount factor.
The period is ABSTRACT in the kernel (ecl_schedule): the section-3 fixture
uses ANNUAL periods (hazards 1.5/2.0/2.2/2.0/1.8%, LGD 35%, EIR 6%,
straight-line EUR 200k repayments -> 12m ECL 4,952.83 / lifetime 16,571.39),
while the panel path (ecl_for_snapshot) uses QUARTERS. "12-month ECL" is the
sum of the first `periods_12m` periods (1 annual period for the fixture,
4 quarters on the panel).
CRITICAL -- NO PREPAYMENT DOUBLE COUNTING (restated; a reviewer will check)
------------------------------------------------------------------------------
S(t) above ALREADY includes prepayment survival (competing risk). EAD_t must
therefore be the CONTRACTUAL amortisation balance from engine/ead.py --
deliberately NOT prepay-scaled. Scaling EAD by prepayment expectations on top
of the prepay-inclusive S(t) would count prepayment twice and understate
lifetime ECL. ecl_for_snapshot consumes engine.ead.ead_matrix unmodified.
EIR / DISCOUNTING CONVENTION (documented)
------------------------------------------------------------------------------
IFRS 9 discounts at the ORIGINATION effective interest rate. The panel's live
per-quarter rate is the current note rate interest_rate_time (nominal annual
%, quarterly compounding) and the book is US fixed-rate mortgages, so the
current note rate is the disciplined EIR proxy: eir_q = rate / 400 -- the
same RATE_DIVISOR convention that drives the engine/ead.py annuity. The
orig_rate_missing flag concerns the ORIGINATION-rate snapshot only (see
engine/ead.py) -- the current note rate is populated on every panel row.
Defensive fallback, mirroring ead.py's straight-line branch: a NaN or
non-positive rate discounts at eir_q = 0 (undiscounted); no panel row
triggers it.
STAGE MAPPING (IFRS 9) AND THE STAGE-3 RULE
------------------------------------------------------------------------------
Both 12-month and lifetime ECL are ALWAYS computed per loan; the REPORTED
allowance is 12m for Stage 1 and lifetime for Stage 2/3 (reported_allowance).
Stages come from engine/staging.py (relative SICR test; Stage 3 = the
snapshot quarter's default rows).
Stage 3 (credit-impaired, default has HAPPENED): ECL = LGD(x_now) * EAD_now
with EAD_now = balance_time and LGD from engine/lgd.py at the default-quarter
covariates. No PD (the default is an event, not a probability) and no
further discounting: the loss crystallises now, and the workout-recovery
discounting is embedded in the vendor's realised lgd_time that the LGD model
is fitted on (engine/lgd.py documented simplification). ecl_12m =
ecl_lifetime = that value for Stage-3 rows.
PROJECTION CONVENTIONS (inherit engine/staging.py's rung-1 choices)
------------------------------------------------------------------------------
* Horizon: R = mat_time - t remaining contractual quarters, floored at 1
(at/past-maturity loans due within one quarter), optionally capped by
EclConfig.max_horizon.
* Hazards: the exact age-offset factorisation of engine/staging.py -- all
covariates FROZEN at their snapshot values, only the loan-age spline
advances along the projection (rung-1 tail assumption; no macro path,
no LTV amortisation; scenario conditioning is a later rung).
* LGD_t: same clock. A default in projected quarter k happens at age
age0 + k; the LGD model's loan_age covariate advances with that clock
(exact within the fitted logit-linear stages -- eta(k) = eta0 +
beta_age * k -- cross-checked against predict_components in
crosscheck_lgd_grid), all other LGD covariates frozen at the snapshot.
LGD can exceed 1 by the excess-loss loading (beyond-EAD workout costs are
real data, never silently clipped -- engine/lgd.py).
* Population: panel rows at time t EXCLUDING payoff_event == 1 rows
(payoff quarter = derecognition, not a stageable exposure -- same
convention as the staging exhibits). Default rows stay (Stage 3).
MOVEMENT DECOMPOSITION (movement_decomposition)
------------------------------------------------------------------------------
Waterfall between two allowance snapshots, SEQUENTIAL attribution on a
documented ordering (attribution is ORDER-DEPENDENT: each step is measured
holding the previous steps' state; a different order allocates the
interaction terms differently -- stated here, not hidden):
opening sum of reported allowance at t0
1. stage_migration surviving (common) loans: reported allowance at t0
MARKS under the t1 STAGE minus under the t0 stage --
the pure effect of 12m <-> lifetime re-classification
at frozen parameters (computable because both ECLs are
always carried per loan)
2. remeasurement surviving loans, already at the t1 stage: t1 marks
minus t0 marks -- macro / parameter / amortisation /
discount-unwind re-measurement, i.e. "same loans,
re-marked"
3. derecognitions loans present at t0 only (prepaid, matured, or t0
defaults resolved out of the panel): minus their
opening reported allowance
4. new_loans loans present at t1 only (post-t0 originations /
first observations): plus their closing reported
allowance
closing sum of reported allowance at t1
The five components sum EXACTLY to closing - opening (algebraic identity,
asserted; unit-tested in tests/test_ecl.py).
DOCUMENTED SIMPLIFICATIONS
------------------------------------------------------------------------------
* EIR proxy = current note rate / 400 per quarter (fixed-rate book; IFRS 9's
origination EIR coincides up to repricing, which is out of scope).
* Frozen covariates over the projection on both the hazard and LGD legs;
only the loan-age clock advances (hazard spline + LGD loan_age term).
* Stage-3 ECL = LGD * current balance, undiscounted (loss crystallised;
workout discounting embedded in the vendor lgd_time).
* 12-month ECL = first 4 projected quarters (EclConfig.quarters_12m).
* Movement attribution is sequential and order-dependent (ordering above).
* Payoff-quarter rows are derecognitions, excluded from the ECL book.
Deterministic throughout: no LLM, no network, no randomness (the only RNG,
in crosscheck_lgd_grid, is fixed-seed sampling for a numerical cross-check).
"""
from __future__ import annotations
from dataclasses import dataclass
import numpy as np
import pandas as pd
from patsy import build_design_matrices
from scipy.special import expit
from engine.ead import RATE_DIVISOR, SNAPSHOT_COLS, ead_matrix
from engine.hazard import HazardModel
from engine.lgd import LGD_REQUIRED_COLS, LgdModels, predict_components
from engine.lgd import _prepare as _lgd_prepare
from engine.staging import (
_CHUNK_LOANS,
_EXIT_CAP,
StagingConfig,
_age_curve,
_design_matrix,
_spline_cols,
assign_stages,
)
#: quarters that make up the 12-month ECL window on the quarterly panel
QUARTERS_12M = 4
#: columns movement_decomposition requires on each allowance frame
ALLOWANCE_COLS = ["id", "stage", "ecl_12m", "ecl_lifetime"]
#: components of the movement waterfall, in attribution order
WATERFALL_COMPONENTS = ["opening", "stage_migration", "remeasurement",
"derecognitions", "new_loans", "closing"]
@dataclass(frozen=True)
class EclConfig:
"""ECL assembly configuration.
staging : StagingConfig driving the Stage-1/2/3 assignment
(engine/staging.py; defaults = doubling + 0.5pp p.a. add-on).
quarters_12m : projected quarters in the "12-month" ECL window (4).
max_horizon : optional cap on the lifetime projection horizon R
(quarters); None = full contractual remaining life.
"""
staging: StagingConfig = StagingConfig()
quarters_12m: int = QUARTERS_12M
max_horizon: int | None = None
# ---------------------------------------------------------------------------
# the shared survival / marginal-PD algebra (one implementation, two callers)
# ---------------------------------------------------------------------------
def _survival_marginal(lam_d: np.ndarray,
lam_p: np.ndarray) -> tuple[np.ndarray, np.ndarray]:
"""(n, T) hazards -> (survival entering t, marginal PD in t), both (n, T).
survival_start[:, t-1] = S(t-1) = prod_{k<t} (1 - lam_d_k - lam_p_k);
marginal[:, t-1] = S(t-1) * lam_d_t. Identical algebra (including the
per-period exit-probability cap) to engine.hazard.pd_term_structure.
"""
if lam_d.shape != lam_p.shape or lam_d.ndim != 2:
raise ValueError("hazard grids must be equal-shape 2-D arrays")
exit_p = np.clip(lam_d + lam_p, 0.0, _EXIT_CAP)
surv = np.cumprod(1.0 - exit_p, axis=1)
surv_start = np.hstack([np.ones((surv.shape[0], 1)), surv[:, :-1]])
return surv_start, surv_start * lam_d
# ---------------------------------------------------------------------------
# the ECL kernel (single exposure; fixture-facing)
# ---------------------------------------------------------------------------
def _as_period_vector(x, T: int, name: str) -> np.ndarray:
v = np.asarray(x, dtype=float)
if v.ndim == 0:
v = np.full(T, float(v))
v = v.ravel()
if v.size != T:
raise ValueError(f"{name} must be scalar or length {T}, got {v.size}")
return v
def ecl_schedule(hazard_default, lgd, ead, eir,
hazard_prepay=None) -> pd.DataFrame:
"""Per-period ECL table for ONE exposure -- the compute_ecl section-3 kernel.
Parameters
----------
hazard_default : per-period conditional default hazards lambda_t,
t = 1..T (the length defines T). Periods are abstract: annual for
the fixture, quarterly on the panel.
lgd : LGD_t, scalar or length-T (may exceed 1 -- beyond-EAD workout
costs; engine/lgd.py).
ead : EAD_t, scalar or length-T -- the CONTRACTUAL balance entering
period t (never prepay-scaled; module docstring, CRITICAL section).
eir : per-period effective interest rate; discount = (1+eir)^-t
(end-of-period loss crystallisation).
hazard_prepay : optional per-period competing prepayment hazards;
default all-zero. Survival is S(t) = prod (1 - lam_d - lam_p).
Returns
-------
DataFrame, one row per period t = 1..T:
period, hazard_default, hazard_prepay, survival_start (= S(t-1)),
marginal_pd (= S(t-1)*lambda_t), lgd, ead, discount, ecl.
Sum the `ecl` column (ecl_totals) for 12m / lifetime ECL.
"""
lam_d = np.asarray(hazard_default, dtype=float).ravel()
T = lam_d.size
if T == 0:
raise ValueError("need at least one projection period")
lam_p = (np.zeros(T) if hazard_prepay is None
else _as_period_vector(hazard_prepay, T, "hazard_prepay"))
lgd_v = _as_period_vector(lgd, T, "lgd")
ead_v = _as_period_vector(ead, T, "ead")
eir = float(eir)
if not (np.all(np.isfinite(lam_d)) and np.all(np.isfinite(lam_p))
and np.all(np.isfinite(lgd_v)) and np.all(np.isfinite(ead_v))
and np.isfinite(eir)):
raise ValueError("non-finite inputs")
if np.any(lam_d < 0) or np.any(lam_d > 1) or np.any(lam_p < 0) \
or np.any(lam_p > 1):
raise ValueError("hazards must lie in [0, 1]")
if np.any(lgd_v < 0) or np.any(ead_v < 0):
raise ValueError("lgd and ead must be non-negative")
if eir <= -1.0:
raise ValueError("eir must exceed -100%")
surv_start, marginal = _survival_marginal(lam_d[None, :], lam_p[None, :])
periods = np.arange(1, T + 1)
discount = (1.0 + eir) ** -periods.astype(float)
ecl = marginal[0] * lgd_v * ead_v * discount
return pd.DataFrame({
"period": periods,
"hazard_default": lam_d,
"hazard_prepay": lam_p,
"survival_start": surv_start[0],
"marginal_pd": marginal[0],
"lgd": lgd_v,
"ead": ead_v,
"discount": discount,
"ecl": ecl,
})
def ecl_totals(schedule: pd.DataFrame,
periods_12m: int) -> tuple[float, float]:
"""(12-month ECL, lifetime ECL) from an ecl_schedule table.
periods_12m: leading periods that span 12 months -- 1 for the annual
fixture convention, QUARTERS_12M = 4 on the quarterly panel.
"""
if periods_12m < 1:
raise ValueError("periods_12m must be >= 1")
e = schedule["ecl"].to_numpy()
return float(e[:periods_12m].sum()), float(e.sum())
# ---------------------------------------------------------------------------
# vectorised per-loan grids (snapshot path)
# ---------------------------------------------------------------------------
def _marginal_pd_grid(models: dict[str, HazardModel], frame: pd.DataFrame,
R: np.ndarray) -> np.ndarray:
"""(n, max R) unconditional default probabilities S(t-1)*lambda_d(t).
Same exact age-offset factorisation as engine.staging._cum_default_pd
(whose row sums this grid reproduces to float noise -- gated in
analysis/run_ecl.py), but keeping the per-period marginals the ECL sum
needs instead of collapsing to the cumulative PD. Cells beyond a loan's
R are exactly 0 (off-window hazards are set to 0).
"""
ages0 = frame["loan_age"].to_numpy()
if not np.allclose(ages0, np.round(ages0)):
raise ValueError("loan_age must be integral quarters")
ages0 = np.round(ages0).astype(int)
R = np.asarray(R, dtype=int)
if (R < 1).any():
raise ValueError("horizons must be >= 1 (floor R upstream)")
n, Tmax = len(R), int(R.max())
max_age = int(ages0.max() + Tmax)
offsets, curves = {}, {}
for name in ("default", "prepay"):
m = models[name]
beta = m.result.params.to_numpy()
idx = _spline_cols(m)
X = _design_matrix(m, frame)
offsets[name] = X @ beta - X[:, idx] @ beta[idx]
curves[name] = _age_curve(m, max_age)
marginal = np.zeros((n, Tmax))
for lo in range(0, n, _CHUNK_LOANS):
sl = slice(lo, min(lo + _CHUNK_LOANS, n))
a0, r = ages0[sl], R[sl]
k = np.arange(int(r.max()))
ages = a0[:, None] + 1 + k[None, :]
on = k[None, :] < r[:, None]
lam = {}
for name in ("default", "prepay"):
eta = offsets[name][sl][:, None] + curves[name][ages]
with np.errstate(over="ignore"):
lam[name] = np.where(on, 1.0 - np.exp(-np.exp(eta)), 0.0)
_, marg = _survival_marginal(lam["default"], lam["prepay"])
marginal[sl, :marg.shape[1]] = marg
return marginal
def _lgd_eta0(result, frame: pd.DataFrame) -> np.ndarray:
"""Linear predictor of one LGD stage at the snapshot covariates."""
prep = _lgd_prepare(frame)
bad = int(prep[LGD_REQUIRED_COLS].isna().any(axis=1).sum())
if bad:
raise ValueError(f"{bad} rows carry NaN in required LGD covariates")
di = result.model.data.design_info
X = np.asarray(build_design_matrices([di], prep)[0], dtype=float)
if X.shape[0] != len(prep):
raise RuntimeError("LGD design matrix dropped rows; misaligned")
return X @ result.params.to_numpy()
def _lgd_grid(models: LgdModels, frame: pd.DataFrame,
Tmax: int) -> np.ndarray:
"""(n, Tmax) expected LGD for a default in projected quarter k = 1..Tmax.
Both LGD stages are logit-linear, so advancing ONLY the loan_age clock
(default in quarter k -> age age0 + k; every other covariate frozen at
the snapshot -- module docstring) is the exact shift
eta(k) = eta0 + beta_age * k on each stage:
lgd(k) = (1 - expit(eta_cure(k))) * (expit(eta_sev(k)) + loading)
identical to engine.lgd.predict_components on a frame with
loan_age = age0 + k (cross-checked by crosscheck_lgd_grid).
"""
k = np.arange(1, Tmax + 1, dtype=float)[None, :]
eta_c = _lgd_eta0(models.cure_result, frame)[:, None] \
+ float(models.cure_result.params["loan_age"]) * k
eta_s = _lgd_eta0(models.severity_result, frame)[:, None] \
+ float(models.severity_result.params["loan_age"]) * k
return (1.0 - expit(eta_c)) * (expit(eta_s) + models.excess_loading)
def crosscheck_lgd_grid(models: LgdModels, snapshot_df: pd.DataFrame,
n_sample: int = 20, k_max: int = 40,
seed: int = 0) -> float:
"""Max |_lgd_grid - predict_components| over a fixed-seed loan sample.
Rebuilds, for each sampled loan, the k = 1..k_max default-quarter frames
with loan_age = age0 + k and scores them through the public
engine.lgd.predict_components (the slow reference path). Expected ~1e-15
(identical linear predictors). Deterministic (fixed seed).
"""
rng = np.random.default_rng(seed)
pick = rng.choice(len(snapshot_df), size=min(n_sample, len(snapshot_df)),
replace=False)
sample = snapshot_df.iloc[pick].reset_index(drop=True)
fast = _lgd_grid(models, sample, k_max)
worst = 0.0
for i in range(len(sample)):
rep = pd.DataFrame([sample.iloc[i][LGD_REQUIRED_COLS]] * k_max)
rep = rep.reset_index(drop=True).astype(float)
rep["loan_age"] = rep["loan_age"] + np.arange(1, k_max + 1)
ref = predict_components(models, rep)["lgd"].to_numpy()
worst = max(worst, float(np.abs(fast[i] - ref).max()))
return worst
# ---------------------------------------------------------------------------
# per-loan ECL at a reporting snapshot
# ---------------------------------------------------------------------------
def ecl_for_snapshot(panel_df: pd.DataFrame, t: int, models: dict,
config: EclConfig | None = None) -> pd.DataFrame:
"""Stage + 12-month ECL + lifetime ECL per loan at reporting quarter t.
Parameters
----------
panel_df : loan-quarter rows containing quarter t AND history
(time <= t rows feed the staging origination-macro map and the
probation lookback; rows after t are never used).
t : reporting quarter (end-of-quarter convention).
models : {'default': HazardModel, 'prepay': HazardModel,
'lgd': LgdModels} -- the fitted rung-1 engines.
config : EclConfig (module docstring for every convention).
Returns
-------
One row per NON-PAYOFF loan at t (payoff rows are derecognitions):
id, time, stage, trigger_reason, loan_age, R, balance, eir_q,
lgd_current, lifetime_pd_now, lifetime_pd_engine (the ECL grid's own
cumulative PD -- cross-check column), pd_ratio, default_event,
ecl_12m, ecl_lifetime, ecl_reported, coverage.
ecl_12m and ecl_lifetime are BOTH always populated; ecl_reported is the
IFRS 9 allowance (12m for Stage 1, lifetime for Stage 2/3). Stage-3 rows
carry ecl_12m = ecl_lifetime = LGD(x_now) * balance (module docstring).
EAD paths are CONTRACTUAL (engine/ead.py) -- never prepay-scaled: the
survival weights already carry prepayment (CRITICAL section).
"""
cfg = config if config is not None else EclConfig()
for key in ("default", "prepay", "lgd"):
if key not in models:
raise KeyError(f"models dict must contain '{key}'")
hz = {"default": models["default"], "prepay": models["prepay"]}
t = int(t)
hist = panel_df.loc[panel_df["time"] <= t]
if not (hist["time"] == t).any():
raise ValueError(f"no rows at time {t}")
staged = assign_stages(hist, hz, cfg.staging, snapshot_time=t)
snap = (hist.loc[hist["time"] == t]
.sort_values("id").reset_index(drop=True))
if snap["id"].duplicated().any():
raise ValueError(f"duplicate loan ids at time {t}")
snap = snap.loc[snap["payoff_event"] == 0].reset_index(drop=True)
if snap.empty:
raise ValueError(f"no stageable (non-payoff) rows at time {t}")
snap = snap.merge(
staged[["id", "stage", "trigger_reason", "lifetime_pd_now",
"lifetime_pd_orig", "pd_ratio"]],
on="id", how="left", validate="one_to_one")
if snap["stage"].isna().any():
raise RuntimeError("staging table did not cover every snapshot loan")
R = np.maximum(snap["mat_time"].to_numpy(dtype=int) - t, 1)
if cfg.max_horizon is not None:
R = np.minimum(R, int(cfg.max_horizon))
Tmax = int(R.max())
marginal = _marginal_pd_grid(hz, snap, R) # S(t-1)*lambda_t
ead_g = ead_matrix(snap[SNAPSHOT_COLS], Tmax).to_numpy() # contractual
lgd_g = _lgd_grid(models["lgd"], snap, Tmax) # age clock only
comp0 = predict_components(models["lgd"], snap) # default-now LGD
rate = snap["interest_rate_time"].to_numpy(dtype=float)
usable = np.isfinite(rate) & (rate > 0.0)
eir_q = np.where(usable, rate, 0.0) / RATE_DIVISOR # documented proxy
periods = np.arange(1, Tmax + 1, dtype=float)
discount = (1.0 + eir_q)[:, None] ** -periods[None, :]
cells = marginal * lgd_g * ead_g * discount # THE formula
q12 = min(int(cfg.quarters_12m), Tmax)
ecl_12m = cells[:, :q12].sum(axis=1)
ecl_life = cells.sum(axis=1)
# Stage 3: default happened this quarter -- LGD-based loss on current EAD
balance = snap["balance_time"].to_numpy(dtype=float)
lgd_now = comp0["lgd"].to_numpy()
is_def = snap["default_event"].to_numpy() == 1
stage = snap["stage"].to_numpy().astype(int)
if not np.array_equal(is_def, stage == 3):
raise RuntimeError("stage-3 rows must be exactly the default rows")
s3 = lgd_now * balance
ecl_12m = np.where(is_def, s3, ecl_12m)
ecl_life = np.where(is_def, s3, ecl_life)
reported = np.where(stage == 1, ecl_12m, ecl_life)
if not (np.all(np.isfinite(ecl_life)) and np.all(ecl_12m >= 0.0)
and np.all(ecl_12m <= ecl_life + 1e-9 * np.maximum(ecl_life, 1))):
raise AssertionError("ECL sanity violated: need 0 <= 12m <= lifetime")
return pd.DataFrame({
"id": snap["id"].to_numpy(),
"time": t,
"stage": stage,
"trigger_reason": snap["trigger_reason"].to_numpy(),
"loan_age": snap["loan_age"].to_numpy(),
"R": R,
"balance": balance,
"eir_q": eir_q,
"lgd_current": lgd_now,
"lifetime_pd_now": snap["lifetime_pd_now"].to_numpy(),
"lifetime_pd_engine": marginal.sum(axis=1),
"pd_ratio": snap["pd_ratio"].to_numpy(),
"default_event": snap["default_event"].to_numpy(),
"ecl_12m": ecl_12m,
"ecl_lifetime": ecl_life,
"ecl_reported": reported,
"coverage": np.divide(reported, balance,
out=np.full_like(reported, np.nan),
where=balance > 0.0),
})
# ---------------------------------------------------------------------------
# allowance movement decomposition
# ---------------------------------------------------------------------------
def reported_allowance(allowance_df: pd.DataFrame) -> np.ndarray:
"""IFRS 9 reported allowance per row: 12m if Stage 1, else lifetime."""
missing = [c for c in ALLOWANCE_COLS if c not in allowance_df.columns]
if missing:
raise KeyError(f"allowance frame is missing columns: {missing}")
return np.where(allowance_df["stage"].to_numpy() == 1,
allowance_df["ecl_12m"].to_numpy(dtype=float),
allowance_df["ecl_lifetime"].to_numpy(dtype=float))
def _pick(df: pd.DataFrame, stage: np.ndarray, suffix: str) -> np.ndarray:
"""Reported allowance from `suffix` marks under an arbitrary stage."""
return np.where(stage == 1,
df[f"ecl_12m{suffix}"].to_numpy(dtype=float),
df[f"ecl_lifetime{suffix}"].to_numpy(dtype=float))
def movement_decomposition(allowance_t0_df: pd.DataFrame,
allowance_t1_df: pd.DataFrame,
label_t0: str = "t0",
label_t1: str = "t1") -> pd.DataFrame:
"""Allowance waterfall between two ecl_for_snapshot outputs.
Sequential attribution, ORDER-DEPENDENT by construction (module
docstring states the convention): on surviving loans, (1) stage
migration is measured FIRST, at frozen t0 marks (t0's own 12m/lifetime
ECLs re-picked under the t1 stage); (2) re-measurement second, at the
already-migrated stage (t1 marks minus t0 marks under the t1 stage);
portfolio change enters as (3) derecognitions at their t0 reported
allowance and (4) new loans at their t1 reported allowance.
Input frames need ALLOWANCE_COLS = [id, stage, ecl_12m, ecl_lifetime]
with unique ids (any ecl_reported column is ignored and re-derived).
Returns a DataFrame with one row per WATERFALL_COMPONENTS entry:
component | amount | n_loans | kind ('level' or 'delta')
where amount sums EXACTLY (asserted) as
closing = opening + stage_migration + remeasurement
+ derecognitions + new_loans.
"""
frames = {}
for name, df in (("t0", allowance_t0_df), ("t1", allowance_t1_df)):
missing = [c for c in ALLOWANCE_COLS if c not in df.columns]
if missing:
raise KeyError(f"{name} frame is missing columns: {missing}")
if df["id"].duplicated().any():
raise ValueError(f"{name} frame has duplicate loan ids")
frames[name] = df[ALLOWANCE_COLS]
a0, a1 = frames["t0"], frames["t1"]
opening = float(reported_allowance(a0).sum())
closing = float(reported_allowance(a1).sum())
m = a0.merge(a1, on="id", how="outer", suffixes=("_0", "_1"),
indicator=True)
gone = m["_merge"] == "left_only"
born = m["_merge"] == "right_only"
both = m["_merge"] == "both"
derecognitions = -float(_pick(m[gone], m.loc[gone, "stage_0"].to_numpy(),
"_0").sum())
new_loans = float(_pick(m[born], m.loc[born, "stage_1"].to_numpy(),
"_1").sum())
c = m[both]
st0 = c["stage_0"].to_numpy().astype(int)
st1 = c["stage_1"].to_numpy().astype(int)
at_t0_old_stage = _pick(c, st0, "_0")
at_t0_new_stage = _pick(c, st1, "_0")
at_t1_new_stage = _pick(c, st1, "_1")
stage_migration = float((at_t0_new_stage - at_t0_old_stage).sum())
remeasurement = float((at_t1_new_stage - at_t0_new_stage).sum())
total = (opening + stage_migration + remeasurement + derecognitions
+ new_loans)
if abs(total - closing) > 1e-6 * max(abs(closing), 1.0):
raise AssertionError(
f"waterfall identity violated: components give {total:.6f}, "
f"closing is {closing:.6f}")
out = pd.DataFrame({
"component": WATERFALL_COMPONENTS,
"amount": [opening, stage_migration, remeasurement, derecognitions,
new_loans, closing],
"n_loans": [len(a0), int((st0 != st1).sum()), int(both.sum()),
int(gone.sum()), int(born.sum()), len(a1)],
"kind": ["level", "delta", "delta", "delta", "delta", "level"],
})
out.attrs["label_t0"] = label_t0
out.attrs["label_t1"] = label_t1
return out
|