Show the distribution of probabilities per column, not their mean

#11
by lvwerra HF Staff - opened
Hugging Face Biology Research org

These probabilities are bimodal: a base is confidently coding or confidently not. Averaging a bin therefore lands on a value almost no base holds. A column that is 20% exon reads as 0.2, which looks like background, while the binary view fires on the same column because some base there really is near 1. The two views contradicted each other and both were right โ€” neither could say "mostly background, with a real peak in it".

What changed

The binned probability view is now a 2D histogram: one column per position bin, colour showing how many bases fall in each probability band. It bins max(P_positive, P_negative) โ€” the same value the threshold and the binary labels are computed from โ€” so the two views finally agree. A thin line keeps the mean for continuity.

Below the binning threshold nothing changes: every base still gets its own point, per strand.

Evidence

On a real segment at 578 bases per column:

position mean line actual bases
50,286 0.215 max 0.965, 23.5% above 0.5
72,828 0.177 max 0.840, 23.7% above 0.5
180,336 0.105 max 0.836, 10.9% above 0.5

And the reason the binary view looked saturated โ€” any over a bin, measured on the same segment where 3.05% of bases are coding:

bases per bin bins drawn as 1
1 3.0%
1,000 9.9%
10,000 40.0%
60,000 100.0%

Cost

Counting runs in chunks. A whole-chromosome window is 100M+ bases and an index array over all of them at once costs more than the probabilities themselves. Chunked, 100M bases take 1.0s with no measurable overhead above the array already in memory, and every base is counted exactly once โ€” which is a test.

Thirty-six tests pass.

๐Ÿค– Generated with Claude Code

lvwerra changed pull request status to merged

Sign up or log in to comment