double-exposure / docs /WP9_DESIGN.md
Eddie Faillace
double-exposure app deploy snapshot 2026-07-21 (WP-24 calibration pass)
7dff04f
|
Raw
History Blame Contribute Delete
17.7 kB

A newer version of the Gradio SDK is available: 6.24.0

Upgrade

WP-9 Design Doc — Projection-Guided Diffusion Sampling (L4)

Author: Claude Fable 5 (frontier design pass), 2026-07-07 Status: DESIGN ONLY — implementation is gated (see §1). This document is WP-9's first deliverable ("Design doc first"). It makes the hard algorithmic decisions up front so the eventual implementation WP is execution, not research. Prereqs to read: MASTERPLAN Part I (I.5, I.5a, I.6), Part III L4. This doc does not restate the physics; it builds on it.


1. The gate is not yet open (honest precondition)

WP-9's dependency line reads: "Only start when L0–L3 metrics plateau on the benchmark." They have not plateaued, and this doc does not claim otherwise:

  • WP-6 (DIP) just became the strongest offline source at 256px (0.6712 vs 0.7838 LPIPS) — a source that is still improving as we fix the ruler, not a plateau.
  • WP-7 (refinement) is net-neutral (4/6 improved, flat mean) and its ≥70% accept item is still open pending the 50-case 256px run.
  • The 50-case benchmark at 256px — the actual instrument that would show a plateau — has not been run.
  • 50-case 256px verdict (2026-07-07): top sources have NOT plateaued (best-heu=0.7720, DIP=0.6031 pulls clearly ahead, refined-heu=0.8143; spread >>0.05); WP-9 remains gated, cheaper-layer work continues.

Therefore: write the design now (cheap, frontier-appropriate), but do NOT implement until the 50-case 256px benchmark is run and shows L1.2/L1.3/L3 within noise of each other across two successive improvements. Diffusion guidance is the most expensive layer in the plan (GPU, per-step VAE round-trips); spending that budget before the cheaper layers have demonstrably topped out inverts the plan's whole cost discipline. This doc's acceptance criteria (§9) are written so the implementation WP can be picked up unchanged whenever the gate opens.


2. Problem, stated as a linear inverse problem

From I.5: densitometry gives us the total exposure field Ĥ = H₁ + H₂ directly (per pixel, on VALID mask pixels). We want two plausible scene images A, B. The measurement operator is:

y = M · x       where   x = (H₁; H₂) stacked,   M = [I  I],   y = Ĥ

M is linear in exposure space — this is the entire reason DDNM-style null-space projection applies. For a single pixel the operator is the 1×2 matrix [1 1]:

  • Range space (what the measurement fixes): the sum h₁ + h₂.
  • Null space (what the prior must fill): the difference h₁ − h₂, a 1-D line per pixel.

The pseudo-inverse M⁺ = [I; I]/2 distributes a measured sum equally. DDNM's core identity, per pixel:

(h₁, h₂)_consistent  =  M⁺·ĥ  +  (I − M⁺M)·(h₁, h₂)_prior
                     =  (ĥ/2, ĥ/2)  +  ( (h₁−h₂)/2, (h₂−h₁)/2 )_prior

i.e. keep the prior's difference, replace its sum with the measured sum. Equivalently (the form we implement): compute residual r = ĥ − (h₁ + h₂) and split it, h₁ += r/2, h₂ += r/2. These are identical; the residual form is what handles non-negativity and the inequality (shoulder) case cleanly (§6).

The catch that makes this non-trivial (and why it's a T3, not a bolt-on): the diffusion prior does not live in exposure space H. It generates display-referred sRGB images. The map from a display image to its exposure contribution is φ(img) = luminance_from_linear(srgb_to_linear(img)) — pointwise, monotone, nonlinear. So the operator is linear in H but the sampled variable is img, and H = φ(img). DDNM strictly requires a linear M∘(sampling variable). We recover applicability because φ is pointwise monotone and invertible: the projection is done in H-space and mapped back through φ⁻¹. This is exactly L4's "map into exposure space (pointwise monotone maps), project, map back." §5 handles the one place this bites (the VAE).


3. Core algorithm: two coupled chains + per-step exposure-space projection

Two independent Stable-Diffusion sampling chains (same UNet weights, different latents, different text prompts). They are coupled only through a per-step projection — no cross-attention, no joint model, no UNet backprop.

prompts_a, prompts_b ← analyze_scan(observed_positive).prompt_a / prompt_b   (reuse WP-5 VLM)
z_a, z_b ← random latents (or img2img-encoded from a heuristic split as a warm start)

for t in reverse diffusion schedule:
    # 1. Standard DDIM/DDPM predicted-x0 from each chain (NO grad through UNet)
    x0_a_lat = predict_x0(unet, z_a, t, prompts_a)
    x0_b_lat = predict_x0(unet, z_b, t, prompts_b)

    # 2. Decode to pixel RGB, map to exposure (luminance) space
    A = vae_decode(x0_a_lat);  B = vae_decode(x0_b_lat)          # sRGB [0,1]
    h1 = φ(A);  h2 = φ(B)                                         # exposure (luminance)

    # 3. Fit nuisance gains g1,g2 (product=1) by LS on VALID (every K steps; §7)
    g1, g2 = fit_gains(h1, h2, Ĥ, mask)         # cheap, no backprop

    # 4. DDNM range-space projection with noisy-consistency scaling (§5)
    r = Ĥ − (g1·h1 + g2·h2)                      # residual on VALID
    step = λ(t) · r                              # λ(t)∈[0,1] ramps 0→1 as t→0 (§5)
    (h1, h2) ← apply_projection(h1, h2, step, mask, curve)   # equal split + non-neg + shoulder ineq (§6)

    # 5. Map back to images, re-encode, and re-noise to level t-1
    A' = render(h1 · scene chroma from A);  B' = render(h2 · chroma from B)   (§4)
    x0_a_lat ← vae_encode(A');  x0_b_lat ← vae_encode(B')
    z_a ← renoise(x0_a_lat, t→t-1);  z_b ← renoise(x0_b_lat, t→t-1)

return A, B  (final decode); register both as one candidate pair

The projection is the physics; the UNet is the prior; they meet once per step in exposure space. This is DDNM (Wang et al. 2022) adapted from "one image, global linear A" to "two images, per-pixel additive A with a pointwise nonlinear sampling map."


4. What the constraint applies to: luminance only, chroma from the prior

Decision: the sum constraint is enforced on the luminance/exposure channel only. Chroma (the null space of a luminance-only measurement) is left entirely to the diffusion prior. Rationale:

  • The physics constraint Ĥ = H₁ + H₂ is a statement about exposure, which for a single B&W emulsion is luminance, and for the green-channel-driven color path (WP-8) is the green exposure. It says nothing about hue. Forcing chroma would be inventing physics.
  • This matches the plan's "B&W first" cross-cutting rule and keeps WP-9 compatible with the color pipeline without a per-channel diffusion model (a future WP if ever).
  • Concretely: render(h · chroma) scales each chain's current decoded chroma to the projected luminance (the same luminance-ratio chroma-carry trick already in _replicate_separation / demix's _render_h_to_positive — reuse it, do not reinvent).

So the projected images keep the UNet's colors and textures; only their brightness is pulled onto the physical sum. This is the single most important scoping decision — it is what makes the projection a cheap per-pixel operation instead of an intractable joint constraint over RGB.


5. The VAE round-trip is the real cost (and the PSLD caveat)

DDNM's range-space correction must be applied to x̂₀ in the space where the operator is defined = pixel/exposure space. SD samples in a 4-channel latent. So every projected step pays vae_decode → project → vae_encode. Two independent problems:

  1. Cost. Decode+encode per step × 2 chains × ~30–50 steps is the dominant runtime. On the user's MPS hardware this is minutes/image (acceptable for a T3 offline candidate, not for interactive use — hence off-by-default, §8).
  2. The VAE is not a perfect autoencoder (this is the "PSLD caveat in latent space" the plan names). Re-encoding a projected image injects reconstruction error, and doing it on an early, still-noisy x̂₀ over-commits to a bad estimate.

Mitigations (decided):

  • Noisy-consistency scaling λ(t) (from DDNM+): scale the projection step by a schedule that is ~0 at high noise and ramps to 1 as t→0. Early steps let the prior explore; late steps enforce the constraint hard. Concretely λ(t) = (1 − ᾱ_t)^(−½)-normalised to [0,1], or the simpler linear λ = 1 − t/T. Start with linear; the schedule is a tuning knob, not a design fork.
  • Sparse early projection: decode/project/encode only every K steps (K≈5) for the first ~70% of the schedule, then every step in the final ~30%. Cuts VAE round-trips ~3× with negligible quality loss (the early corrections are down-weighted by λ anyway).
  • The final returned images are a clean decode of the last projected latents (no re-encode after the last projection), so VAE error is not doubled at the output.

6. Confidence mask & non-negativity (the honest projection)

The equal-split projection is modified per pixel by the WP-2 confidence mask (TOE=0, VALID=1, SHOULDER=2) — because the measurement Ĥ means different things in each region:

  • VALID: full equality projection h₁+h₂ = ĥ, split residual equally.
  • Non-negativity: if the equal split would push either hᵢ < 0, clamp that layer to 0 and give the entire residual to the other (a plausible split cannot have negative exposure). This is the "with non-negativity handling" clause in L4.
  • SHOULDER (saturation): ĥ is only a lower bound on the true sum (density saturated at D_max). Projection becomes an inequality: only correct upward — if g₁h₁+g₂h₂ < ĥ_lowerbound, push up; if already ≥, leave the prior's values alone. Never pull a shoulder pixel down to a saturated lower bound.
  • TOE (fog/noise): ĥ is poorly constrained; apply no projection (mask the residual to 0). Let the prior fill these pixels. This mirrors how the WP-3 physics loss already down-weights TOE via _valid_weight.

All four cases are a single masked, clamped residual update — the implementation is ~15 lines, but getting the shoulder inequality and non-negativity right is the correctness core and must be unit-tested independently (§9).


7. Nuisance gains g₁, g₂ (I.5a) inside the loop

The true constraint is g₁·φ(A) + g₂·φ(B) = Ĥ with free per-layer gains (product fixed to 1 to remove redundancy with the curve speed point) absorbing unknown relative exposure / latent-image fading / scanner calibration. Reuse WP-7's exact least-squares gain fit (the learnable-g machinery already in hybrid_loss/latent_optimizer), not a new derivation: fit log₁₀g by LS on VALID pixels against the current (h₁, h₂) estimates, clamp to GAIN_GRID_RANGE. Decision: refit every K steps (same K as the sparse projection), not every step — the gains are slowly varying and the fit is noisy on early x̂₀. Fold g₁, g₂ into the residual (r = Ĥ − (g₁h₁ + g₂h₂)) and into the map-back.


8. Architecture, API surface, integration

New module guided_sampling.py (top level, sibling to latent_optimizer.py / demix.py). Mirrors the injected-callable pattern that makes demix/DIP testable offline:

  • GuidedSamplingConfig — steps, guidance schedule λ, projection stride K, chroma-carry flag, fallback-mode flag, seed.
  • predict_x0: Callable[(latent, t, prompt) -> latent]injected diffusion denoiser. Real impl wraps an SD-1.5 pipeline from diffusers (MPS-feasible; SDXL optional); the offline stub is a trivial denoiser (e.g. identity-toward-a-fixed-image or a blur) so the projection loop is fully unit-testable with zero weights and no network, exactly like stub_cleanup for demix.
  • projection_step(h1, h2, h_total, mask, curve, gains) -> (h1', h2') — pure numpy, the §6 correctness core, independently testable.
  • guided_separate(positive_rgb, h_total, confidence_mask, film_curve, predict_x0, analysis, config) -> Optional[SeparationResult] — the §3 loop; returns one candidate pair with method="guided_diffusion", candidate_id="guided_i{steps}", and diagnostics carrying final masked sum-constraint error (reuse the structured-diagnostics pattern from WP-6).
  • Reuse, do not reinvent: densitometry.srgb_to_linear/luminance_from_linear/ linear_to_srgb for φ and render; PiecewiseFilmCurve.inverse if a density-space variant is needed; app.demix.analyze_scan for prompts; latent_optimizer's VAE encode/decode helpers (now public) and gain-fit; SeparationResult as the return type.
  • Integration: register as a candidate source in generate_candidates behind an include_guided: bool = False flag (off by default — GPU/slow), via the same _append_*-style helper WP-5.1/WP-6 established. One UI checkbox "Guided diffusion (very slow, GPU)". Density path required (skip if density is None), same rule as demix/DIP.

Everything offline-testable except the real-SD run, which is manual/GPU-gated and documented in the Result note — the standing pattern for token/GPU-gated work in this repo.


9. Acceptance criteria for the implementation WP (write these into the WP-9 impl spec)

Offline (CI, stub denoiser — the real gate):

  1. projection_step unit tests: (a) idempotent on an already-consistent pair; (b) strictly reduces |g₁h₁+g₂h₂ − ĥ| on VALID; (c) non-negativity — never returns a negative layer; (d) shoulder pixels are only pushed up, never down; (e) TOE pixels untouched. Each with teeth (construct an input that violates it, prove the naive version fails).
  2. Full loop with the stub denoiser drives two chains to satisfy the masked sum constraint to < 2% (same bar as demix) on the two committed 64×64 fixtures AND the 256px generated cases. This proves the projection math is correct independent of any diffusion prior.
  3. Registration: generate_candidates(include_guided=True) (tiny step count, stub denoiser, monkeypatched) yields a method=="guided_diffusion" candidate that rank_candidates scores. Goes through the real generate_candidates surface (per the WP-8 lesson — no direct-call theater).

Manual (GPU-gated, documented, NOT CI): 4. Real SD-1.5 run on the 256px cases: guided candidates ranked against DIP (current best) on per-layer LPIPS. Target: competitive with or beating DIP on ≥50% of cases. Report the real table verbatim; if it loses, that is an honest finding (the fixture-scale lesson applies — judge at 256px, and note real-scan behavior is the true test).

Invariant: the projection must be a genuine projection — verify (2) by asserting the residual is monotone-non-increasing across the last 30% of steps (where λ≈1).


10. Fallback path (behind a flag, not primary)

If the VAE re-encode (§5) proves too lossy to hit the <2% offline bar, fall back to gradient guidance (ΠGDM/DPS-style): instead of hard range-space replacement, add ∇_z ‖g₁φ(A)+g₂φ(B) − Ĥ‖²_VALID to each chain's score at each step. This backprops through the VAE decode only (not the UNet — far cheaper than full DPS), which is the "PSLD caveats in latent space" the plan flags. Decision: implement projection as primary, gradient guidance behind config.mode="gradient", and let the offline bar-2 test decide which ships. Do not implement gradient-through-UNet DPS (Chung et al. 2022) — the plan explicitly deprioritizes it as expensive, and the linear operator makes projection the right tool.


11. Open questions / risks (for the implementer to surface, not silently resolve)

  • VAE reconstruction floor: if vae_encode(vae_decode(x)) ≠ x error exceeds the 2% sum bar on its own, no amount of projection helps — measure this first, before building the loop, and STOP/report if the floor is above bar.
  • Warm start: random latents vs img2img-encoding a heuristic split as the starting point. Recommendation: img2img warm start from the best demo/DIP candidate — turns WP-9 into a refiner of the existing pipeline rather than a cold generator, which is both cheaper and more likely to beat the baseline. Decide empirically.
  • Schedule λ(t): linear is the starting point; if late-step hard projection causes visible seams, soften. This is tuning, in-scope for the impl WP, not a redesign.
  • Model licensing / weights: SD-1.5 weights + diffusers is a real dependency and a first network/GPU dependency for the repo — must stay lazy-imported and fully absent from the offline path, like replicate/anthropic today.

12. One-paragraph summary for the WP-9 impl kickoff

Two independent Stable-Diffusion chains, prompted by the WP-5 VLM scene descriptions, coupled only by a per-step DDNM-style null-space projection performed in exposure space on the luminance channel: decode each chain's predicted-x₀ to RGB, map to exposure via the densitometry sRGB→luminance transform, fit the I.5a nuisance gains by least squares on VALID pixels, split the residual Ĥ − (g₁h₁+g₂h₂) equally with non-negativity clamping and a shoulder-inequality/TOE-skip per the confidence mask, carry chroma from the prior, re-encode, and renoise — with a noise-level-scaled guidance strength and sparse early projection to control the VAE round-trip cost. No UNet backprop. Offline-tested to <2% sum-constraint error with a stub denoiser; real SD run is GPU-gated and judged at 256px against the DIP baseline. Gated: do not implement until the 50-case 256px benchmark shows L0–L3 have plateaued.