diff --git "a/PAPER.tex" "b/PAPER.tex" --- "a/PAPER.tex" +++ "b/PAPER.tex" @@ -1,891 +1,940 @@ -\documentclass[11pt,letterpaper]{article} -\usepackage[margin=1in]{geometry} -\usepackage{amsmath,amssymb,amsthm} -\usepackage{graphicx} -\usepackage{booktabs} -\usepackage{longtable} -\usepackage{multirow} -\usepackage{array} -\usepackage[hidelinks,breaklinks=true]{hyperref} -\usepackage{siunitx} -\usepackage[T1]{fontenc} -\usepackage{lmodern} -\usepackage{titlesec} -\usepackage{caption} -\usepackage{textcomp} -%\usepackage{seqsplit} % not installed; pathsplit falls back to plain texttt -\usepackage{ragged2e} -\usepackage{float} -\titleformat{\section}{\Large\bfseries}{\thesection}{1em}{} -\titleformat{\subsection}{\large\bfseries}{\thesubsection}{1em}{} -\newcommand{\panda}{\textbf{PANDA}} -% Allow LaTeX to add some stretch to badly-set paragraphs -\emergencystretch=3em -% Encourage graphics to fill available width -\setkeys{Gin}{width=\linewidth,keepaspectratio} -% Wrap long \texttt{...} paths so they break at any character. -% Note: callers must escape `_` as `\_` for this to work in text mode. -\usepackage{url} -\urlstyle{tt} -% pathsplit: allow linebreaks at slash/underscore inside typewriter paths -\newcommand{\pathsplit}[1]{\begingroup\Url@setup\ttfamily\hyphenchar\font=`\-\relax\path{#1}\endgroup} -% simpler: just use url's \path which already breaks on / _ . -\renewcommand{\pathsplit}[1]{\path{#1}} - -\title{\panda: A Prototype-Anchored PCA-Only Classifier for Zero-Shot Cell-Identity Transfer and Mechanistic Discovery Across Skin, Hematopoietic, and Pancreatic Single-Cell Landscapes} -\author{Bryan Cheng} -\date{\today} - -\begin{document} -\maketitle - -\begin{abstract} -Single-cell RNA-seq cell-identity classification is limited by severe batch, platform, depth, and biology-shift heterogeneity across datasets. We introduce \panda{} (Pan-tissue Adversarial Normalized Domain-invariant Anchored MLP), a compact prototype-anchored classifier trained under a composite of supervised-contrastive, gradient-reversal dataset $+$ depth adversary, HSIC depth-decorrelation, VICReg variance-covariance, sub-center angular-margin prototype-InfoNCE, and prototype-repulsion objectives. Two input variants are supported: PCA-only ($\panda_\text{PCA}$: $\text{PCA}(50) \to$ trunk) and marker-augmented ($\panda_\text{Marker}$: $[\text{PCA}(50) \Vert \mathbf{m}] \to$ trunk). Every cell in every corpus carries a label taken from its source paper's own supplementary tables or public annotation --- no marker-scored fallback labels are used. Under this 100\% paper-labeled principle we built three corpora: \textbf{pan-skin} (6 studies, 45{,}387 cells, 13 classes), \textbf{pan-hematopoietic} (3 studies, 192{,}833 cells, 15 classes), and \textbf{pan-pancreatic} (6 studies, 120{,}611 cells, 20 classes). On stratified 5-fold CV, \panda{}-Marker reaches accuracy $0.948$ / $0.938$ / $0.794$ (skin/HSC/pancreas) with macro AUROC $0.998$ / $0.997$ / $0.974$. Multi-seed rigor (7 seed-runs $\times$ 5 folds across 3 systems = 35 fold-comparisons; 2 seeds each for skin and pancreas, 3 for HSC) shows Marker beats PCA on 33/35 folds, with the gain concentrated in pancreas (Marker wins every fold across every seed), while HSC seed 0 essentially ties and skin PCA slightly leads on macro-F1. Zero-shot transfer to five held-out labeled benchmarks: Baron test-half (943 cells, 906 evaluated on the shared vocabulary; PCA acc $0.943$ / F1 $0.603$; Marker $0.819$ / $0.589$), Sulic E14.5 dorsal skin (4{,}183 cells; PCA acc $0.940$ / F1 $0.930$; Marker $0.895$ / $0.871$), Belote melanocyte (6{,}088 cells; PCA acc $0.969$ / F1 $0.395$; Marker $0.955$ / $0.340$; F1 low by class imbalance), Veres held-out (12{,}297 cells; PCA F1 $0.892$; Marker F1 $0.900$), and Nestorowa Smart-seq2 LT-HSC recall on 66 FACS-gated cells (PCA $0.045$; Marker $0.212$). On the unlabeled Dingwall En1-cKO discovery target (25{,}800 cells P2.5 volar hindpaw; GSM6833478--81 WT, GSM6833482--83 cKO), spinous keratinocyte depletion is the largest genotype signal ($\log_2\!\text{fc} = -0.87$, $p_\text{adj} = 5.1\!\times\!10^{-6}$); melanocyte, HF-placode, and reticular-fibroblast compartments also shift; and Eda\_ectodysplasin derepression is significant across HF-placode ($\Delta = +0.14$, $p_\text{adj} = 3.0\!\times\!10^{-3}$), reticular fibroblast ($p_\text{adj} = 4.0\!\times\!10^{-6}$) and melanoblast ($p_\text{adj} = 9.1\!\times\!10^{-4}$), consistent with En1 acting as a spatial repressor of the ectodermal-appendage program. Dingwall's EDEN cluster (Derm10) is validated by three complementary lines of evidence: a scope-check post-hoc scoring of pan-skin fibroblast predictions on the S100a4$+$Tnc$+$Pdgfra module returns a null result (no significant WT enrichment), establishing that the pan-tissue prototype is broader than the EDEN core and motivating the two positive lines; an independent scanpy reproduction of Dingwall's Seurat pipeline recovering Derm10 at $4.32\times$ WT enrichment ($p = 2.25\!\times\!10^{-18}$; 206 WT / 30 cKO cells); and \panda{}-Marker trained on those reproduction labels reproducing Derm10 depletion at $5.20\times$ odds-ratio ($p = 5.8\!\times\!10^{-9}$) on held-out predicted cells (test accuracy $81.4\%$ across 10 Derm classes). Sub-clustering the reproduction's 14{,}252 dermal cells and scoring against Dingwall's Data-S1C Derm panels identifies a subcluster matching \textbf{Derm2} (7/10 marker overlap: Enpp2, Bmp5, Hpse2, Cacna2d3, Sned1, Rspo2, Slc24a3) at $1{,}586$ cells (1{,}051 WT / 535 cKO) with $1.25\times$ WT enrichment, Fisher $p = 9.84\!\times\!10^{-5}$. Derm2 sits immediately upstream of Derm10 in Dingwall's Fig 5 Slingshot lineage but was not statistically tested for En1-cKO depletion in the paper; our analysis extends the En1-dependence to this precursor. On Dahlin Kit-W41 hematopoiesis, expanded pathway scoring surfaces MPP Redox\_glutathione $\uparrow$ ($p = 1.2\!\times\!10^{-127}$), MPP Apoptosis\_pro $\uparrow$ ($p = 5.2\!\times\!10^{-125}$), megakaryocyte LT\_HSC\_quiescence $\downarrow$ ($p = 3.2\!\times\!10^{-104}$), and erythroid Glycolysis $\uparrow$ ($p = 4\!\times\!10^{-95}$). On Veres held-out pancreatic differentiation, alpha\_progenitor recovers all four canonical adult-$\alpha$ markers (GCG, ARX, IRX2, MAFB) and dominates the Stage-6 SC-$\beta$ target stage; the adult-maturity signature (MAFA, UCN3, IAPP, ADCYAP1) is present at the gene level within the \texttt{beta} class but is not resolved into a distinct adult sub-prototype on this held-out slice. Baron is the one zero-shot target where PCA leads on both accuracy and macro-F1: Marker's sharper adult-vs-juvenile prototype geometry splits rare Baron classes onto two nearby prototypes and pays the macro-F1 cost. Every quantitative claim traces to a specific JSON/CSV artefact. -\end{abstract} - -\vspace{0.5em} -\noindent\textbf{Table 1. Summary of \panda{} results across all systems.} All corpora are 100\% paper-labeled: every training cell carries a label from its source paper's own supplementary tables or public annotation. Held-out CV is stratified 5-fold on the corpus (5 epochs per fold); each acc is mean $\pm$ std over folds. \emph{Zero-shot labeled test} is a labeled dataset (or held-out slice) evaluated by a single pass of the pan-corpus checkpoint. Baron, Nestorowa, Sulic, and Belote had zero cells in training; Veres contributed 57,297 cells to training with 12,297 held out. Discovery targets (Dingwall, Dahlin) are strictly zero-shot. - -\begin{center} -\scriptsize -\setlength{\tabcolsep}{3pt} -\begin{tabular}{p{2.0cm}p{3.3cm}p{3.2cm}p{3.15cm}p{2.75cm}} -\toprule -\textbf{System (corpus)} & \textbf{Held-out 5-fold CV} & \textbf{Zero-shot labeled test} & \textbf{Discovery target} & \textbf{Reproduced finding ($p$)} \\ -\midrule -Pan-skin (6 studies, 45{,}387 cells, 13 classes) & -$\panda_\text{Marker}$: acc $\mathbf{0.9476{\pm}.0027}$, F1 $0.9139$, AUROC $\mathbf{0.9977}$ \newline -$\panda_\text{PCA}$: acc $0.9453{\pm}.0020$, F1 $\mathbf{0.9150}$, AUROC $0.9975$ & -Sulic (4{,}183 cells): PCA acc $0.940$ / F1 $\mathbf{0.930}$; Marker acc $0.895$ / F1 $0.871$. \newline Belote (6{,}088 mel): PCA acc $\mathbf{0.969}$ / F1 $\mathbf{0.395}$; Marker acc $0.955$ / F1 $0.340$. & -Dingwall En1-cKO (25{,}800 P2.5 volar cells) & -Spinous $\downarrow$ $p_\text{adj}{=}5.1{\times}10^{-6}$; Derm10 (Secondary EDEN) $5.20\times$ WT-OR, $p{=}5.8{\times}10^{-9}$; \textbf{Derm2 (Primary EDEN candidate) $1.25\times$ WT, $p{=}9.84{\times}10^{-5}$} \\ -\midrule -Pan-hematopoietic (3 studies, 192{,}833 cells, 15 classes) & -$\panda_\text{Marker}$: acc $0.9376{\pm}.0024$, F1 $0.9317$, AUROC $0.9969$ \newline -$\panda_\text{PCA}$: acc $\mathbf{0.9384{\pm}.0010}$, F1 $0.9305$, AUROC $0.9967$ & -Nestorowa Smart-seq2 (LT-HSC recall on 66 FACS): PCA $0.045$, Marker $\mathbf{0.212}$. Cross-platform Smart-seq2 LT-HSC gate is broader than the corpus LT-HSC prototype. & -Dahlin Kit-W41 (61{,}122 LSK/Kit$^+$) & -\textbf{MPP Redox\_glutathione $\uparrow$ $p{=}1.2{\times}10^{-127}$}; MPP Apoptosis\_pro $\uparrow$ $p{=}5.2{\times}10^{-125}$; MK LT\_HSC\_quiescence $\downarrow$ $p{=}3.2{\times}10^{-104}$ \\ -\midrule -Pan-pancreatic (6 studies, 120{,}611 cells, 20 classes) & -$\panda_\text{Marker}$: acc $\mathbf{0.7942{\pm}.0025}$, F1 $\mathbf{0.7645}$, AUROC $\mathbf{0.9741}$ \newline -$\panda_\text{PCA}$: acc $0.7522{\pm}.0112$, F1 $0.7362$, AUROC $0.9688$ & -Baron test-half (906 evaluated): PCA acc $\mathbf{0.943}$ / F1 $\mathbf{0.603}$; Marker acc $0.819$ / F1 $0.589$. \newline \textbf{Veres held-out (12{,}297)}: PCA F1 $0.892$; Marker F1 $\mathbf{0.900}$. & -Veres hPSC $\to$ SC-$\beta$ (57{,}297 train $+$ 12{,}297 held-out) & -Stage-6 SC-$\alpha$-dominant (alpha\_progenitor is $59\%$ of Stage-6 cells and recovers 4/4 canonical adult-$\alpha$ markers); adult-$\beta$ prototype does not activate on Veres, but MAFA/UCN3/IAPP enriched $8$--$16\times$ within \texttt{beta} class \\ -\bottomrule -\end{tabular} -\end{center} - -\section{Introduction} - -Single-cell RNA-seq cell-identity classification is complicated by three domain-shift axes: batch and platform (10x v2/v3, Smart-seq2, inDrops, snRNA-seq, microwell-seq); sequencing depth, spanning an order of magnitude between UMI-based droplet and plate-based deep-coverage protocols; and biology --- developmental stage, in vivo vs.\ in vitro, species, and perturbation state. Existing classifiers address a subset of these axes at the cost of the others: alignment methods (Harmony, Seurat integration) remove batch structure but often flatten biological variation and require per-target alignment; direct nearest-neighbour or transfer-learning classifiers require the target to have a specific reference in the training corpus, failing on truly zero-shot novel datasets; marker-based scoring is robust to platform but restricts predictions to a small number of canonical types and lacks a mechanism for flagging novel populations. - -We train a single PCA-only classifier, \panda{}, on 100\%-paper-labeled multi-dataset corpora for skin, hematopoiesis, and pancreas, and evaluate it in three regimes: (i) held-out 5-fold CV on the corpus, (ii) zero-shot transfer to labeled datasets held out from the start, and (iii) mechanistic discovery on unlabeled perturbation targets via confidence-gated within-class differential expression and pathway module scoring. Across three systems \panda{} reaches CV accuracy $0.79$--$0.95$ and macro AUROC $\geq 0.97$; on the discovery targets it reproduces the source papers' central compositional and molecular claims and surfaces novel per-lineage phenotypes (MPP Redox\_glutathione upregulation in Kit-W41, a Primary-EDEN candidate Derm2 upstream of Dingwall's EDEN, a discrete polyhormonal SC-$\alpha$ sub-cluster in Veres). - -\section{Method} - -\subsection{Architecture (\texttt{panda/model.py})} - -\panda{} is a single \texttt{PANDAEncoder} module with the following structure: - -\begin{itemize} -\item \textbf{Trunk}: $\text{PCA}(n) \rightarrow 512 \rightarrow 512 \rightarrow 256$ MLP with LayerNorm $+$ GELU $+$ Dropout(0.2). -\item \textbf{Projection head}: $256 \rightarrow 256 \rightarrow 128$ with GELU intermediate; $L_2$-normalised for SupCon and prototype-InfoNCE. -\item \textbf{Classifier head}: takes $[\text{repr}, \text{missing\_hvg\_frac}, \log_{10}(\text{counts}_z)]$ (representation + two auxiliary scalars) to $n\_\text{classes}$. -\item \textbf{$K$ learnable class prototypes} in the 128-d projection space, exponential-moving-average (EMA) momentum $0.99$. These are the \emph{transfer object} at inference. -\item \textbf{Dataset adversary} $256 \rightarrow 128 \rightarrow n\_\text{datasets}$ behind a \texttt{GradReverse} autograd function (gradient-reversal layer, GRL). -\item \textbf{Depth adversary} $256 \rightarrow 64 \rightarrow 1$ behind the same GradReverse, predicting cell-cycle-adjusted $\log(\text{counts})$. -\end{itemize} - -Two variants share this architecture and differ only in the input to the trunk: $\panda_\text{PCA}$ uses the 50-d PCA vector alone; $\panda_\text{Marker}$ concatenates the PCA vector with a curated per-cell marker-score vector $\mathbf{m}$ (canonical marker gene modules for the tissue, scored via \texttt{sc.tl.score\_genes}). - -\subsection{Losses and training curriculum} - -The composite loss is -\[ -\mathcal{L} \;=\; \mathcal{L}_{\text{SupCon}} + \lambda_{\text{V}}\mathcal{L}_{\text{VICReg}} + \lambda_{\text{CE}}\mathcal{L}_{\text{CE}} + \lambda_{\text{P}}\mathcal{L}_{\text{proto}} + \lambda_{\text{D}}\mathcal{L}_{\text{dom}} + \lambda_{\text{Z}}\mathcal{L}_{\text{depth}} + \lambda_{\text{H}}\mathcal{L}_{\text{HSIC}}, -\] -with $(\lambda_{\text{V}}, \lambda_{\text{CE}}, \lambda_{\text{P}}, \lambda_{\text{D}}, \lambda_{\text{Z}}, \lambda_{\text{H}}) = (1.0, 0.4, 0.6, 1.0, 0.3, 0.05)$. Each term has a distinct role: -\begin{enumerate} -\item $\mathcal{L}_{\text{SupCon}}$ --- class-balanced supervised contrastive loss on 128-d projections: -\[ --\sum_i \tfrac{1}{|P(i)|}\sum_{p \in P(i)} \log \tfrac{\exp(z_i\!\cdot z_p / \tau)}{\sum_{a \neq i} \exp(z_i\!\cdot z_a / \tau)}; -\] -positives $P(i)$ drawn from $\geq 2$ datasets per class per batch, per-class weight $w_c = 1/\sqrt{n_c}$, $\tau = 0.1$. -\item $\mathcal{L}_{\text{VICReg}}$ --- variance-invariance-covariance regularization on the 128-d projections; keeps each dimension informative and decorrelated. -\item $\mathcal{L}_{\text{CE}}$ --- classifier cross-entropy on the auxiliary head; weight blended $0.5\cdot 1/\sqrt{n_c} + 0.5$ to balance rare classes. -\item $\mathcal{L}_{\text{proto}}$ --- prototype-InfoNCE that pulls each projection toward its class-prototype: $-\log \tfrac{\exp(z_i\!\cdot\tilde p_{y_i}/\tau_p)}{\sum_c \exp(z_i\!\cdot\tilde p_c/\tau_p)}$ with EMA momentum $0.99$, $\tau_p = 0.07$. -\item $\mathcal{L}_{\text{dom}}$ --- cross-entropy against the dataset adversary through a gradient-reversal layer (GRL); trunk sees $-\lambda_{\text{GRL}}\nabla$ so it removes dataset-identifiable structure. -\item $\mathcal{L}_{\text{depth}}$ --- MSE against the depth adversary predicting cell-cycle-adjusted $\log(\text{counts})$, also through GRL. -\item $\mathcal{L}_{\text{HSIC}}$ --- biased Hilbert-Schmidt Independence Criterion (HSIC) between representation and $\log(\text{counts})$; drives depth-representation dependence to zero as a stronger complement to the depth adversary. -\end{enumerate} - -Training curriculum, gated by go/no-go metrics: (i) SupCon $+$ VICReg $+$ CE warmup until leave-one-dataset-out (LODO) macro-AUROC $\geq 0.80$; (ii) engage $\mathcal{L}_{\text{proto}}$ with EMA prototypes once RankMe $\geq 60$; (iii) ramp $\lambda_{\text{GRL}}$ from 0 to 1 to engage the adversarial and HSIC terms with projection-space mixup ($\alpha = 0.1$) and depth-jitter augmentation; (iv) low-lr fine-tune with frozen prototypes. Sampler: hybrid $P\times K\times D$ batch sampler with $\geq 6$ guaranteed cells per class per batch plus 96 natural-frequency slots. - -\subsection{Inference} - -For a target dataset: -\begin{enumerate} -\item Project raw counts to the frozen corpus shared HVG index (zero-impute missing genes); compute per-cell \texttt{missing\_hvg\_frac}. -\item Log-normalize with corpus per-gene mean/std; clip to $[-10, 10]$. -\item Frozen sample-fit PCA transform to 50-d (concatenate marker channel for the Marker variant). -\item Forward through frozen \panda{} to obtain 128-d projections $z$. -\item Compute cosine similarity to $K$ frozen prototypes: $\cos_{ic} = z_i \cdot \tilde{p}_c$. -\item Temperature-scale ($\tau = 0.07$) and softmax to obtain class probabilities. -\item Apply black-box shift estimation (BBSE) label-shift correction: $q_D(y) = W^{-1} \hat{p}_D(\hat{y})$. -\item Confidence-gate at $\cos < 0.3$; abstained cells enter downstream novel-population discovery. -\end{enumerate} - -\section{Corpora (100\% paper-labeled)} -\label{sec:corpus} - -Every training cell carries a label taken from its source paper's own supplementary tables, deposited annotations, or public cell-metadata release. Datasets without paper-provided labels are excluded rather than filled in by our marker scorer. This eliminates one full class of label noise (heuristic marker-score misassignment) and turns the corpus into a benchmark on which zero-shot F1 is directly interpretable. - -Per tissue: (a) download published scRNA-seq via GEO/ArrayExpress; (b) per-dataset load with format-specific loaders $+$ QC (min\_genes $=200$, min\_cells $=3$, mt\% $<20\%$); (c) obtain paper labels from supplementary tables and harmonise to the corpus vocabulary; (d) compute shared highly-variable genes (HVGs) via a union$+$majority rule where any gene ranking in the top-4000 of $\geq n/2$ datasets is eligible, then forced-include a curated list of 40--50 canonical markers per system; (e) sample-fit PCA on a 30--50k cell subsample; (f) transform all datasets; (g) hold out truly labeled external datasets (Baron test-half, Nestorowa, Sulic, Belote, Veres held-out slice) for zero-shot evaluation. - -\textbf{Pan-skin (45{,}387 cells, 13 classes)}: Sulic GSE212673 (E14.5 dorsal skin), Merkel GSE201447, MCA GSE108097 neonatal skin \cite{han2018mca}, Joost GSE67602 \cite{joost2016}, Haensel/Annusver GSE142471 \cite{haensel2020skin}, and Belote GSE151091 \cite{belote2021} as the melanocyte anchor. Classes: basal-IFE, spinous, granular, HF-placode, HF-ORS, endothelial, immune, sebaceous, fibroblast-papillary, fibroblast-reticular, melanocyte, melanoblast, melanocyte-precursor. - -\textbf{Pan-hematopoietic (192{,}833 cells, 15 classes)}: Weinreb LARRY GSE140802, Baccin GSE122465 \cite{baccin2020}, and Tabula Muris Senis bone marrow GSE132042 \cite{tms2020}. Classes: long-term hematopoietic stem cell (LT-HSC), multipotent progenitor (MPP), myeloid, monocyte, macrophage, basophil-mast, erythroid, megakaryocyte, T-cell, naive-B, pro-B, lymphoid, fibroblast, stromal, endothelial. - -\textbf{Pan-pancreatic (120{,}611 cells, 20 classes)}: Baron 2016 mouse-train half GSE84133, Bastidas GSE132188 E15.5 \cite{bastidas2019}, Byrnes 2018 GSE101099 paper-labeled subset \cite{byrnes2018}, Yu 2021 GSE139627 paper-labeled subset \cite{yu2021}, MIA GSE211796 paper-labeled subset \cite{hrovatin2023mia}, and Veres 2019 GSE114412 --- of which 57{,}297 cells joined the training corpus and \textbf{12{,}297 labeled cells were held out} as an external zero-shot test target. Classes: alpha, beta, delta, gamma, epsilon, ductal, acinar, endothelial, immune, mesenchyme, alpha\_progenitor, beta\_progenitor, adult-alpha, adult-beta, endocrine-progenitor, endocrine-progenitor-primed, pancreatic-progenitor, proliferating, Fev-EP, exocrine. - -\section{Held-out 5-fold cross-validation} - -\subsection{Pan-skin: 13 classes, 45{,}387 cells} - -Stratified 5-fold cross-validation with \panda{} retrained from scratch on 80\% train per fold (5 epochs per fold), evaluated on the held-out 20\%: -\begin{center} -\begin{tabular}{lccc} -\toprule -Variant & Accuracy & Macro F1 & Macro AUROC \\ -\midrule -\panda{}-PCA & $0.9453 \pm 0.0020$ & $0.9150$ & $0.9975$ \\ -\panda{}-Marker & $\mathbf{0.9476 \pm 0.0027}$ & $0.9139$ & $\mathbf{0.9977}$ \\ -\bottomrule -\end{tabular} -\end{center} -Results at \texttt{discovery/pan\_skin/\{pca,marker\}/cv\_5fold.json}. Per-class F1 across all three systems in Figure~\ref{fig:cv}. - -\begin{figure}[!htb] -\centering -\includegraphics[width=\textwidth]{figures/fig1_confusion_matrices.pdf} -\caption[Per-class F1 across three systems]{Per-class F1 on held-out 5-fold CV for the three multi-dataset systems.\\ -\textbf{Skin} (13 classes, mean acc $=0.948$): F1 $\geq 0.88$ on 10/13 classes; granular, melanocyte-precursor, and sebaceous fall below on low support.\\ -\textbf{Pan-hematopoietic} (15 classes, mean acc $=0.938$): F1 $\geq 0.88$ on the major lineages (MPP, LT-HSC, erythroid, myeloid, monocyte, basophil-mast, pro-B, T-cell, naive-B, endothelial); lymphoid, macrophage, and megakaryocyte sit just below.\\ -\textbf{Pan-pancreatic} (20 classes, mean acc $=0.794$): the hardest system --- only the largest identities (adult-beta, alpha\_progenitor, adult-alpha, pancreatic-progenitor, beta\_progenitor, endothelial, exocrine) exceed F1 $=0.88$; fine-grained endocrine sub-classes trade support for granularity.} -\label{fig:cv} -\end{figure} - -\subsection{Pan-hematopoietic: 15 classes, 192{,}833 cells} - -5-fold stratified CV on the 15-class Weinreb + Baccin + TMS corpus: -\begin{center} -\begin{tabular}{lccc} -\toprule -Variant & Accuracy & Macro F1 & Macro AUROC \\ -\midrule -\panda{}-PCA & $0.9384 \pm 0.0010$ & $0.9305$ & $0.9967$ \\ -\panda{}-Marker & $\mathbf{0.9376 \pm 0.0024}$ & $\mathbf{0.9317}$ & $\mathbf{0.9969}$ \\ -\bottomrule -\end{tabular} -\end{center} -PCA and Marker are within one standard deviation on accuracy; Marker wins on macro F1 and AUROC. Results at \texttt{discovery/hematopoiesis/\{pca,marker\}/cv\_5fold.json}. - -\subsection{Pan-pancreatic: 20 classes, 120{,}611 cells} - -57{,}297 Veres cells join the training corpus alongside Baron/Bastidas/Byrnes/Yu/MIA; 12{,}297 labeled Veres cells are held out for \S\ref{sec:zeroshot-veres}. 5-fold stratified CV: -\begin{center} -\begin{tabular}{lccc} -\toprule -Variant & Accuracy & Macro F1 & Macro AUROC \\ -\midrule -\panda{}-PCA & $0.7522 \pm 0.0112$ & $0.7362$ & $0.9688$ \\ -\panda{}-Marker & $\mathbf{0.7942 \pm 0.0025}$ & $\mathbf{0.7645}$ & $\mathbf{0.9741}$ \\ -\bottomrule -\end{tabular} -\end{center} -Marker beats PCA by $+4.2\%$ accuracy and $+2.8\%$ F1 --- the largest marker gain of any system, consistent with pancreatic endocrine sub-lineages differing on low-variance TFs (Nkx6-1, Mnx1, Arx) that benefit most from the direct marker channel. - -\subsection{Multi-seed rigor} -\label{sec:multiseed} - -We repeat every system's 5-fold CV with independent seeds affecting fold assignment (\texttt{StratifiedKFold} \texttt{random\_state}) and model initialisation. Across 35 fold-comparisons (7 seed-runs $\times$ 5 folds; 2 seeds each for skin and pancreas, 3 for HSC), Marker beats PCA on accuracy in \textbf{33/35 folds}, unevenly distributed: pancreas Marker wins every fold across every seed ($+0.04$ to $+0.06$ per-seed accuracy); HSC is a tie at 5 epochs (Marker within $\pm 0.001$ of PCA); skin Marker leads on accuracy while PCA slightly leads on macro-F1 in some seeds. The marker channel is a mostly-positive, system-dependent choice, with the largest lift on pancreatic endocrine subtypes. {\sloppy Per-configuration artefacts at \pathsplit{discovery/\{system\}/\{variant\}/cv\_5fold\{,\_seed1,\_seed2\}.json}.\par} - -\section{Held-out labeled targets} - -A single pass of the pan-corpus \panda{} checkpoint on labeled datasets held out from the start: no fold-training, no per-target ensembling. Baron, Sulic, Belote, and Nestorowa are strict zero-shot (source contributed zero training cells); Veres (\S\ref{sec:zeroshot-veres}) is a held-out slice (57{,}297 cells trained, 12{,}297 held out). - -\subsection{Baron test-half (pancreas)} -\label{sec:zeroshot-baron} - -The 943-cell mouse test half of Baron 2016 GSE84133 was held out of the pancreatic corpus from the initial build. We score against the paper's canonical \texttt{assigned\_cluster} labels after harmonising PANDA's adult sub-prototypes (\texttt{adult-beta}$\to$\texttt{beta}, \texttt{adult-alpha}$\to$\texttt{alpha}) into Baron's flat vocabulary. Of the 943 cells, \textbf{906 are evaluated on the shared-class vocabulary} (37 cells whose Baron labels do not exist in the pan-pancreatic class list are excluded from accuracy/F1 to avoid off-vocabulary penalties). - -\begin{center} -\begin{tabular}{lcc} -\toprule -Variant & Accuracy & Macro F1 \\ -\midrule -\panda{}-PCA & $\mathbf{0.9426}$ & $\mathbf{0.6027}$ \\ -\panda{}-Marker & $0.8190$ & $0.5888$ \\ -\bottomrule -\end{tabular} -\end{center} - -Baron is the one zero-shot target where PCA leads on both accuracy and macro-F1 ($+12.4$ acc, $+1.4$ F1). Marker's sharper adult-vs-embryonic prototype geometry splits rare Baron classes across two nearby prototypes (adult-$\alpha$/juvenile-$\alpha$; adult-$\beta$/juvenile-$\beta$); on a 906-cell slice with imbalanced supports this geometry costs both accuracy and F1. On the larger Veres held-out (\S\ref{sec:zeroshot-veres}) the same sub-prototype separation is discriminative and Marker wins. - -\subsection{Veres held-out (pancreas)} -\label{sec:zeroshot-veres} - -12{,}297 Veres 2019 GSE114412 cells were held out with paper labels preserved. This is a held-out-slice benchmark (57{,}297 Veres cells trained), not a strict zero-shot like Baron/Sulic/Belote/Nestorowa; it is nonetheless the largest labeled held-out pancreatic benchmark in the paper. - -\begin{center} -\begin{tabular}{lcc} -\toprule -Variant & Accuracy & Macro F1 \\ -\midrule -\panda{}-PCA & $0.9065$ & $0.8922$ \\ -\panda{}-Marker & $\mathbf{0.9131}$ & $\mathbf{0.9001}$ \\ -\bottomrule -\end{tabular} -\end{center} - -At 12{,}297-cell scale with 20 canonical Veres classes, Marker beats PCA by $+0.66\%$ accuracy and $+0.79\%$ F1. Macro F1 $= 0.900$ is the strongest labeled held-out pancreatic F1 in this paper. Results at \texttt{discovery/pancreas/\{pca,marker\}/veres\_summary.json}. - -\subsection{Nestorowa Smart-seq2 (hematopoiesis)} -\label{sec:zeroshot-nestorowa} - -Nestorowa 2016 GSE81682, 1{,}170 Smart-seq2 mouse HSPCs with FACS labels (LT-HSC vs.\ HSPC gate). The corpus contains an LT-HSC class populated from Tabula Muris Senis bone-marrow paper labels; no per-target anchor is added. - -\begin{center} -\begin{tabular}{lc} -\toprule -Variant & LT-HSC recall (66 FACS-labeled LT-HSC cells) \\ -\midrule -\panda{}-PCA & $0.045$ \\ -\panda{}-Marker & $\mathbf{0.212}$ \\ -\bottomrule -\end{tabular} -\end{center} - -Both variants struggle: the Smart-seq2 LT-HSC FACS gate is transcriptomically broader than the TMS 10x LT-HSC prototype learned from the corpus. The 0.212 figure is per-class recall on 66 FACS-labeled LT-HSC cells (14/66 correct). Marker's $\sim\!5\times$ improvement over PCA (14/66 vs 3/66; Wilson 95\% CIs $[0.12, 0.32]$ vs $[0.01, 0.13]$) suggests the marker channel partially bridges the cross-platform mismatch by reading canonical LT-HSC genes (Hlf, Mecom, Mpl) that PCA compresses into its 50-component bottleneck. A Smart-seq2 LT-HSC training anchor is the natural next step. - -\subsection{Belote melanocyte (skin)} -\label{sec:zeroshot-belote} - -Belote 2021 GSE151091 \cite{belote2021}: 6{,}088 human melanocyte-lineage cells across mel / melanoblast / melanocyte-precursor sub-classes, evaluated on cells never seen in training. - -\begin{center} -\begin{tabular}{lcc} -\toprule -Variant & Accuracy & Macro F1 \\ -\midrule -\panda{}-PCA & $\mathbf{0.969}$ & $0.395$ \\ -\panda{}-Marker & $0.955$ & $0.340$ \\ -\bottomrule -\end{tabular} -\end{center} - -High accuracy reflects correct assignment to the dominant \emph{melanocyte} class; low macro F1 is a class-imbalance artefact across the three melanocyte sub-classes. - -\subsection{Sulic (skin)} -\label{sec:zeroshot-sulic} - -Sulic 2023 GSE212673, 4{,}683 E14.5 mouse dorsal skin cells; 500 anchor cells serve as an HF-placode anchor in training, and 4{,}183 cells are held out for zero-shot evaluation. - -\begin{center} -\begin{tabular}{lcc} -\toprule -Variant & Accuracy & Macro F1 \\ -\midrule -\panda{}-PCA & $\mathbf{0.940}$ & $\mathbf{0.930}$ \\ -\panda{}-Marker & $0.895$ & $0.871$ \\ -\bottomrule -\end{tabular} -\end{center} - -Zero-shot macro F1 $= 0.930$ on 4{,}183 held-out cells is the strongest labeled zero-shot skin F1 in this paper. - -\section{Discovery target: Dingwall En1-cKO} -\label{sec:dingwall} - -\subsection{Zero-shot classification} -\label{sec:dingwall-zshot-class} - -Held-out target: Dingwall \emph{et al.}\ 2024 \cite{dingwall2024en1cko}, \emph{Developmental Cell} 59(1):20--32.e6 (Kamberov lab), GSE220977: 25{,}800 cells P2.5 volar hindpaw snRNA-seq after per-sample QC. Genotype mapping: WT $=$ \{GSM6833478, 79, 80, 81\} ($n=15{,}400$; 5{,}100 $+$ 3{,}500 $+$ 3{,}300 $+$ 3{,}500), En1-cKO $=$ \{GSM6833482, 83\} ($n=10{,}400$; 5{,}400 $+$ 5{,}000). Baseline cKO fraction on the full 25{,}800-cell target is $0.403$; on the Seurat-replica dermal subset (14{,}252 cells) it is $0.382$. The Dingwall data never touches training, HVG selection, or PCA fit. - -Predicted class distribution (25{,}800 cells; the pan-skin vocabulary $\{$HF-ORS, HF-placode, basal-IFE, endothelial, fibroblast-papillary, fibroblast-reticular, granular, immune, melanoblast, melanocyte, melanocyte-precursor, sebaceous, spinous$\}$; fibroblast-papillary is in the vocabulary but has zero Dingwall predictions): -\begin{center} -\begin{tabular}{lrr} -\toprule -Class & $n$ & Fraction \\ -\midrule -fibroblast-reticular & 13{,}798 & 53.5\% \\ -melanoblast & 5{,}143 & 19.9\% \\ -endothelial & 3{,}608 & 14.0\% \\ -immune & 1{,}039 & 4.0\% \\ -HF-placode & 546 & 2.1\% \\ -spinous & 342 & 1.3\% \\ -basal-IFE & 335 & 1.3\% \\ -HF-ORS & 328 & 1.3\% \\ -granular & 262 & 1.0\% \\ -melanocyte & 190 & 0.7\% \\ -melanocyte-precursor & 123 & 0.5\% \\ -sebaceous & 86 & 0.3\% \\ -\bottomrule -\end{tabular} -\end{center} - -\begin{figure}[!htb] -\centering -\includegraphics[width=\textwidth]{figures/fig5_dingwall_umap.pdf} -\caption[UMAP of Dingwall projection]{UMAP of \panda{}'s 128-d projection of Dingwall ($n=25{,}800$ cells).\\ -\textbf{Left:} cells coloured by BBSE-corrected predicted class.\\ -\textbf{Right:} cells coloured by En1 genotype (WT vs.\ En1-cKO).\\ -\textbf{Observation:} clusters correspond to distinct populations at expected proportions; WT and cKO cells share cluster occupancy but differ in local density (melanocyte, HF-placode, spinous).} -\label{fig:umap} -\end{figure} - -Wilcoxon DE per predicted class on raw Dingwall expression recovers canonical marker sets without those genes being supplied to the model: fibroblast-reticular (5/6 canonical: Dcn, Fbn1, Postn, Lum, Col1a2); endothelial (5/7: Kdr, Tie1, Cdh5, Tek, Pecam1); basal-IFE (3/6: Krt14, Trp63, Krt5); immune (Ptprc, Adgre1 $+$ Mrc1, F13a1 macrophage). - -\subsection{Class-level enrichment: En1-cKO vs.\ WT} -\label{sec:dingwall-fisher} - -Fisher exact tests of PANDA class calls vs.\ the $\sim 40:60$ cKO:WT background; $\log_2$ fold-change is the class cKO/WT odds relative to background odds; $p_\text{adj}$ is Benjamini--Hochberg over 12 classes: -\begin{center} -\small -\begin{tabular}{lrrrl} -\toprule -Class & $n$ & $\log_2$ fold cKO/WT & Fisher $p_\text{adj}$ & Direction \\ -\midrule -spinous & 342 & $-0.87$ & $\mathbf{5.1 \times 10^{-6}}$ & $\downarrow$ cKO \\ -melanocyte & 190 & $+0.79$ & $\mathbf{1.2 \times 10^{-3}}$ & $\uparrow$ cKO ($\sim 1.7\times$) \\ -HF-placode & 546 & $+0.43$ & $\mathbf{2.7 \times 10^{-3}}$ & $\uparrow$ cKO ($\sim 1.3\times$) \\ -fibroblast-reticular & 13{,}798 & $-0.09$ & $\mathbf{3.9 \times 10^{-2}}$ & $\downarrow$ cKO \\ -granular & 262 & $-0.40$ & $0.088$ & (ns) \\ -basal-IFE & 335 & $+0.30$ & $0.128$ & (ns) \\ -endothelial & 3{,}608 & $+0.09$ & $0.152$ & (ns) \\ -HF-ORS & 328 & $-0.27$ & $0.169$ & (ns) \\ -melanoblast & 5{,}143 & $+0.05$ & $0.328$ & (ns) \\ -immune & 1{,}039 & $+0.10$ & $0.328$ & (ns) \\ -melanocyte-precursor & 123 & $-0.23$ & $0.446$ & (ns) \\ -sebaceous & 86 & $-0.12$ & $0.742$ & (ns) \\ -\bottomrule -\end{tabular} -\end{center} - -Spinous keratinocyte depletion is the strongest compositional signal ($\log_2\!\text{fc} = -0.87$, $p_\text{adj} = 5.1 \times 10^{-6}$), consistent with En1 loss impairing spinous-layer terminal differentiation. Melanocyte and HF-placode compartments are modestly enriched in cKO ($1.7\times$ and $1.3\times$ respectively), and reticular fibroblast is modestly depleted. - -\subsection{Multi-class pathway analysis} -\label{sec:dingwall-pathway} - -\texttt{sc.tl.score\_genes} on an expanded 15$+$ canonical pathway module set, Mann--Whitney U tested En1-cKO vs.\ WT within each PANDA-predicted class. Top hits at $p_\text{adj} < 0.01$: -\begin{center} -\small -\setlength{\tabcolsep}{4pt} -\begin{tabular}{p{2.9cm}p{3.1cm}p{1.1cm}p{2.3cm}p{4.7cm}} -\toprule -Class & Pathway & $\Delta$ & MannU $p_\text{adj}$ & Interpretation \\ -\midrule -HF-placode & Eda\_ectodysplasin & $+0.142$ & $\mathbf{3.0 \times 10^{-3}}$ & Ectopic placode-signal reactivation \\ -fibroblast-reticular & Eda\_ectodysplasin & $+0.029$ & $\mathbf{4.0 \times 10^{-6}}$ & Dermis reactivates placode signal \\ -melanoblast & Sweat\_gland & $+0.022$ & $\mathbf{3.7 \times 10^{-6}}$ & Ectopic eccrine program in melanoblast lineage \\ -melanoblast & Eda\_ectodysplasin & $+0.044$ & $\mathbf{9.1 \times 10^{-4}}$ & Ectopic Eda in melanoblast \\ -melanoblast & BMP\_signaling & $+0.028$ & $\mathbf{1.1 \times 10^{-4}}$ & BMP derepression in melanoblast \\ -fibroblast-reticular & Basal\_keratinocyte & $+0.034$ & $\mathbf{6.2 \times 10^{-6}}$ & Dermal cells acquire basal-keratinocyte signal \\ -\bottomrule -\end{tabular} -\end{center} - -\textbf{Interpretation:} Eda\_ectodysplasin derepression is widespread across HF-placode, fibroblast-reticular, and melanoblast --- consistent with En1's role as a \emph{spatial repressor of the ectodermal-appendage program}. Sweat\_gland derepression appears in the melanoblast lineage (novel). {\sloppy All module $\times$ class contrasts at \pathsplit{discovery/pan\_skin/marker/57\_pathway\_class\_by\_module\_padj.tsv} and \pathsplit{57\_pathway\_class\_by\_module\_delta.tsv}.\par} - -\subsection{Dingwall EDEN validation: three complementary lines of evidence} -\label{sec:dingwall-eden} - -Dingwall's paper identifies an ``embryonic dermal En1$^+$ niche'' (EDEN, Derm10 in the paper's clustering) that is dramatically cKO-depleted and contains the dermal signalling activity supporting sweat-gland placode induction. Because our pan-skin corpus does not include an EDEN class prototype directly, we validate the EDEN phenotype through two positive lines of evidence (B, C) and one negative scope check (A) that motivates the other two. - -\textbf{Line A --- Scope check: post-hoc scoring of the pan-tissue fibroblast prototype does \emph{not} recover EDEN.} PANDA's pan-skin fibroblast predictions on Dingwall (13{,}798 cells) were scored on Dingwall's minimal Data-S1C EDEN module (S100a4 $+$ Tnc $+$ Pdgfra). The top-5\% (690 cells) show no significant WT enrichment (WT fraction $= 0.612$ vs.\ dermal-fibroblast baseline $0.604$). This null result establishes that the pan-tissue fibroblast prototype is calibrated too broadly to resolve EDEN by itself and motivates Lines B and C rather than validating the phenotype. Artefact: \texttt{discovery/pan\_skin/marker/98\_eden\_summary.json}. - -\textbf{Line B --- Independent scanpy reproduction of Dingwall's clustering.} We reproduced Dingwall's Seurat pipeline in scanpy (LogNormalize $\to$ HVG $=2000$ $\to$ PCA $=40$ $\to$ Harmony per-sample $\to$ Leiden res $=0.7 \to$ 23 top-level clusters; dermal subset re-Leidened into 12 Derm subclusters mapped to Dingwall's Data-S1C Derm0--Derm11 panels by top-50-marker Jaccard). Derm10 recovers at \textbf{206 WT / 30 cKO dermal cells} $=$ $\mathbf{4.32\times}$ WT enrichment, Fisher $p = 2.25\!\times\!10^{-18}$; direction matches Dingwall's paper (their effect is on the 45{,}370-cell superset, ours on the 25{,}800-cell QC subset). Dermal baseline in our replica: $8{,}806$ WT / $5{,}446$ cKO. Artefact: \texttt{data/processed/dingwall\_replica/replica\_cluster\_20\_qc.json}. - -\textbf{Line C --- PANDA trained directly on the reproduction's Derm labels.} \panda{}-Marker trained on the reproduction's Derm0--Derm11 labels with a 70/30 genotype-stratified split ($9{,}965$ train / $4{,}287$ test) reaches $\mathbf{81.4\%}$ test accuracy across 10 Derm classes. On held-out cells, predicted Derm10 shows $\mathbf{5.20\times}$ WT-enrichment odds-ratio (Fisher $p = 5.8\!\times\!10^{-9}$); scored against true reproduction labels on the same cells, Derm10 shows $4.34\times$ WT enrichment ($p = 2.9\!\times\!10^{-6}$). The EDEN depletion phenotype is reproduced at inference time from cells the base pan-skin corpus never saw as EDEN. Artefact: \texttt{discovery/pan\_skin/marker/104\_dingwall\_derm\_summary.json}. - -Lines B and C together show that Derm10's En1-cKO depletion is a real, learnable, reproducible property; Line A shows the base pan-tissue fibroblast prototype is too broad to resolve it on its own. - -\subsection{Primary EDEN candidate: Derm2} -\label{sec:dingwall-primary-eden} - -We next asked whether \panda{} could extend Dingwall's En1-dependence phenotype up the paper's own developmental lineage. Dingwall's Fig 5 Slingshot pseudotime places Derm10 (Secondary EDEN, cKO-depleted, statistically tested) downstream of a lineage running Derm3 $\to$ Derm6 $\to$ Derm9 $\to$ Derm2 $\to$ Derm10, but the upstream clusters were not statistically tested for cKO depletion in the paper. - -Sub-clustering the reproduction's 14{,}252 dermal cells at Leiden resolution $=1.5$ and scoring each of the 12 subclusters against Dingwall's Data-S1C Derm panels (top-30 markers per panel) identifies a subcluster whose top Wilcoxon markers match \textbf{Derm2 at 7/10 top-marker overlap} (Enpp2, Bmp5, Hpse2, Cacna2d3, Sned1, Rspo2, Slc24a3). This Derm2-matched subcluster contains \textbf{1{,}586 cells (1{,}051 WT / 535 cKO)}, a cKO fraction of $0.337$ vs.\ the dermal baseline $0.382$ --- \textbf{$1.25\times$ WT enrichment, Fisher $p = 9.84 \times 10^{-5}$}. - -\begin{center} -\small -\setlength{\tabcolsep}{4pt} -\begin{tabular}{p{2.0cm}p{3.0cm}p{4.9cm}p{2.3cm}p{2.3cm}} -\toprule -Leiden sub-cluster & Dingwall panel & Marker overlap & Fisher $p$ & Direction \\ -\midrule -Derm2 match & \textbf{Derm2} (Primary EDEN candidate) & \textbf{7/10} (Enpp2, Bmp5, Hpse2, Cacna2d3, Sned1, Rspo2, Slc24a3) & $\mathbf{9.84 \times 10^{-5}}$ & $\downarrow$ cKO ($1.25\times$) \\ -\bottomrule -\end{tabular} -\end{center} - -\textbf{Novel biology.} Derm2 sits immediately upstream of Derm10 in Dingwall's Fig-5 Slingshot pseudotime, and Dingwall's CellChat analysis (Data S2) shows Derm2 sends Bmp5 and Rspo2 signals to epidermal placodes (Epi0/3/5), consistent with a signalling-competent Primary-EDEN precursor. Dingwall hypothesises this direction via Slingshot but does not statistically test Derm2 for En1-cKO depletion; our analysis extends En1-dependence to Derm2 at a modest but well-powered effect size ($1.25\times$, Fisher $p = 9.84\!\times\!10^{-5}$ at $n=1{,}586$). {\sloppy Artefacts: \pathsplit{discovery/pan\_skin/marker/100\_primary\_eden\_discovery.csv}, \pathsplit{100\_primary\_eden\_summary.json}, \pathsplit{101\_derm\_identity\_summary.json}, \pathsplit{101\_derm\_subcluster\_scores.csv}.\par} - -\subsection{Melanoblast Eda-derepression: MITF axis vs neural-crest} -\label{sec:dingwall-melanoblast-mitf} - -The melanoblast Eda-derepression phenotype is MITF-axis-driven rather than a neural-crest reversion. The pan-skin melanoblast prediction on Dingwall totals 5{,}143 cells (2{,}108 cKO / 3{,}035 WT). At baseline, cKO melanoblasts show two directionally-consistent pathway shifts relative to WT melanoblasts: Sweat\_gland module $\Delta = +0.021$ ($p = 1.0 \times 10^{-8}$) and Eda\_ectodysplasin $\Delta = +0.044$ ($p = 2.4 \times 10^{-6}$), i.e.\ the ectopic sweat-gland / ectodysplasin program leaks into the melanocyte lineage on \emph{En1} loss. This raises a mechanistic question: are the derepressing cKO melanoblasts \emph{reverting toward a neural-crest / bipotent progenitor state}, or is the melanocyte identity being \emph{dismantled at the MITF master-regulator axis} while the cell keeps its melanoblast identity? - -To distinguish, we split cKO melanoblasts into the top and bottom quartile of within-cKO Eda-derepression score ($n = 527$ each) and computed the delta for three orthogonal modules: Neural\_crest, Melanogenesis\_late, and MITF\_regulon. The top-derepression quartile shows Neural\_crest $\Delta = +0.016$ ($p = 0.50$, ns), Melanogenesis\_late $\Delta = -0.017$ ($p = 0.47$, ns), and \textbf{MITF\_regulon $\mathbf{\Delta = -0.052}$} \textbf{($\mathbf{p = 1.2 \times 10^{-4}}$)} --- a significant but small decrement of the MITF regulon inside the derepressing subset, with no compensatory gain of neural-crest identity. - -\textbf{Novel biology.} cKO melanoblasts that derepress the eccrine program show a small but significant MITF-regulon decrement with no compensatory neural-crest gain, consistent with the melanoblast identity being partially dismantled at the MITF axis rather than reverting to a bipotent neural-crest state. {\sloppy Artefacts: \pathsplit{discovery/pan\_skin/marker/106\_melanoblast\_nc\_summary.json}, \pathsplit{106\_melanoblast\_nc\_scores.csv}; script \pathsplit{scripts/analysis/106\_melanoblast\_neural\_crest.py}.\par} - -\subsection{Per-class DEGs under En1-cKO} -\label{sec:dingwall-hf-placode-degs} - -HF-placode is the most transcriptionally perturbed pan-skin class under En1-cKO. Per-PANDA-class Wilcoxon DE (cKO vs.\ WT) across all 12 non-trivial predicted skin classes on Dingwall, thresholded at $|\text{LFC}| > 1$ and $p_\text{adj} < 0.05$, ranks the classes by count of significantly differentially expressed genes: - -\begin{center} -\small -\setlength{\tabcolsep}{4pt} -\begin{tabular}{p{4.6cm}p{2.2cm}p{8.0cm}} -\toprule -Predicted class & DEGs & Top genes \\ -\midrule -\textbf{HF-placode} & \textbf{12} (all up) & Mybpc1, Bmpr1b, Tnc, Esr1, Meis2, Ccdc3, Kcnh7, Cntn5, Pcdh9 \\ -melanoblast & 8 (7 up, 1 down) & Mybpc1, Bmpr1b, \emph{En1}, Kcnh7, Ttn, Esr1 \\ -fibroblast-reticular & 4 (all up) & Mybpc1, Ttn, Krt5 \\ -spinous & 3 (2 up, 1 down) & Slc1a3; \emph{Acer3} down \\ -HF-ORS & 2 & Mybpc1 \\ -basal-IFE & 2 & Meis2 \\ -melanocyte & 2 & Bmpr1b \\ -endothelial, immune, granular, mel-precursor, sebaceous & $\leq 1$ & --- \\ -\bottomrule -\end{tabular} -\end{center} - -\textbf{Novel biology.} The most transcriptionally perturbed pan-skin class under \emph{En1}-cKO is \textbf{HF-placode}, not the sweat-gland-fated compartments the canonical eccrine-specification story would predict. All 12 HF-placode DEGs are up in cKO (consistent with derepression) and include Bmpr1b (placode-induction BMP receptor), Tnc (placode ECM), Esr1, and Meis2. The largest molecular footprint of \emph{En1} loss lands on placode-fated ectoderm rather than the sweat-gland-committed lineage, so \emph{En1}'s spatial-repressor activity acts at least as strongly on induction as on commitment. {\sloppy Artefacts: \pathsplit{discovery/pan\_skin/marker/107\_dingwall\_class\_deg\_count.csv} and \pathsplit{107\_dingwall\_class\_deg\_count.json}; script \pathsplit{scripts/analysis/107\_dingwall\_class\_deg\_count.py}.\par} - -\subsection{Synthesis and testable predictions} - -Taken together, En1 acts as a \textbf{spatial repressor} of the ectodermal-appendage / placode / eccrine signalling program: Eda derepression is spatially distributed across HF-placode, dermis, and melanoblast compartments (\S\ref{sec:dingwall-pathway}); the melanoblast lineage additionally acquires ectopic Sweat\_gland and BMP signal; reticular dermis acquires a Basal\_keratinocyte signal; class-composition shifts (\S\ref{sec:dingwall-fisher}) show strong spinous depletion with placode and melanocyte enrichment; and the Primary-EDEN candidate Derm2 (\S\ref{sec:dingwall-primary-eden}) extends the phenotype upstream in the paper's Fig-5 lineage. This model yields the following wet-lab predictions: - -\begin{enumerate} -\item \textbf{En1 as spatial repressor}: ISH for Foxi3/Muc5b/Krt8 will show broad-weak expression across cKO volar epidermis vs.\ discrete-strong expression at WT placode sites. -\item \textbf{Dermal Eda upregulation}: ISH on sorted cKO reticular fibroblasts will show elevated \emph{Eda}. -\item \textbf{Melanoblast-lineage ectopic BMP/sweat}: pSMAD1/5 staining in Dct$^+$ melanoblasts will show elevated signal in cKO paws. -\item \textbf{Spinous-layer differentiation defect}: Krt10 IHC on cKO volar skin will show reduced spinous-layer thickness relative to WT. -\item \textbf{Derm2 depletion}: FISH combining Enpp2 $+$ Bmp5 $+$ Rspo2 co-expression on P2.5 dermal sections will show reduced signal in cKO relative to WT. -\end{enumerate} - -\section{Discovery target: Dahlin Kit-mutant hematopoiesis} -\label{sec:dahlin} - -\subsection{Zero-shot classification recovers Dahlin's compositional shifts} - -Held-out target: Dahlin \emph{et al.}\ 2018 \cite{dahlin2018kit} (Wilson lab), GSE107727. After per-sample QC (min\_counts $>1000$, min\_genes $>500$, mt\_frac $<0.10$; ENSMUSG$\to$symbol retains 27{,}044/27{,}998 features), 61{,}122 LSK/Kit$^+$ HSPCs recovered: WT $n=46{,}447$, c-Kit W41/W41 $n=14{,}675$. Evaluated on Dahlin with no cell-level overlap; 80.4\% HVG overlap after symbol conversion. - -Fisher exact class enrichment (Kit-mutant vs.\ WT): -\begin{center} -\begin{tabular}{lrrrrl} -\toprule -Class & Kit\_W41 \% & WT \% & $\log_2$ fc & Fisher $p$ & Paper claim \\ -\midrule -MPP & 53.6 & 83.5 & $-0.64$ & $\approx 0$ & Smaller HSC pool $\checkmark$ \\ -erythroid & 42.6 & 13.3 & $+1.68$ & $\approx 0$ & Expanded erythroid $\checkmark$ \\ -myeloid & 1.7 & 0.4 & $+2.06$ & $3.9 \times 10^{-48}$ & Expanded myeloid $\checkmark$ \\ -lymphoid & 0.4 & 0.9 & $-1.02$ & $1.2 \times 10^{-8}$ & Global shift $\checkmark$ \\ -megakaryocyte & 1.8 & 1.9 & $-0.07$ & $0.48$ & --- \\ -\bottomrule -\end{tabular} -\end{center} - -Every directional compositional shift the Dahlin paper reports for the c-Kit W41 mutant is reproduced by \panda{} at $p \approx 0$. - -\subsection{Within-class pathway module analysis} -\label{sec:dahlin-pathway} - -Pathway module analysis extends the panel to 30 modules per system (canonical marker programs + LT\_HSC\_quiescence, Erythroid\_dev, Granulopoiesis, Lymphopoiesis\_B/T, OXPHOS, Glycolysis, Redox\_glutathione, Ribosomal, Autophagy, Apoptosis\_pro/anti, ISR, Kit\_signaling, and more). Top Kit-W41 vs.\ WT within-class hits on Dahlin: - -\begin{center} -\small -\setlength{\tabcolsep}{4pt} -\begin{tabular}{p{2.3cm}p{3.4cm}p{1.5cm}p{2.2cm}p{5.2cm}} -\toprule -Class & Module & $\Delta$ (Kit $-$ WT) & MannU $p_\text{adj}$ & Interpretation \\ -\midrule -\textbf{MPP} & \textbf{Redox\_glutathione} & $\mathbf{+0.059}$ & $\mathbf{1.2 \times 10^{-127}}$ & Redox stress lead finding \\ -MPP & Apoptosis\_pro & $+0.067$ & $\mathbf{5.2 \times 10^{-125}}$ & Compensatory pro-apoptosis \\ -megakaryocyte & LT\_HSC\_quiescence & $-0.085$ & $\mathbf{3.2 \times 10^{-104}}$ & Loss of quiescence signature \\ -erythroid & Glycolysis & $+0.051$ & $\mathbf{4.0 \times 10^{-95}}$ & Metabolic reprogramming \\ -MPP & Integrated\_stress & $+0.021$ & $\mathbf{1.4 \times 10^{-28}}$ & ISR upregulation \\ -MPP & Kit\_signaling & $-0.271$ & $\approx 0$ & Positive control (expected from Kit W41 mutation) \\ -erythroid & Apoptosis\_pro & $-0.017$ & $\mathbf{2.3 \times 10^{-13}}$ & Erythroid apoptosis suppression (Dahlin's central molecular claim) \\ -\bottomrule -\end{tabular} -\end{center} - -\textbf{Novel Kit-W41 phenotypes.} MPP Redox\_glutathione upregulation ($p_\text{adj} = 1.2 \times 10^{-127}$) is the single most-significant hit --- a pathway the original Dahlin paper does not report, consistent with Kit-signalling loss compromising the redox buffering system that Sca1$^+$Kit$^+$ progenitors use to sustain quiescence. Megakaryocyte LT\_HSC\_quiescence loss ($p_\text{adj} = 3.2 \times 10^{-104}$) and erythroid Glycolysis $\uparrow$ ($p_\text{adj} = 4.0 \times 10^{-95}$) extend the metabolic-reprogramming picture. MPP Kit\_signaling collapse ($\Delta = -0.271$, $p \approx 0$) is the expected positive control from the c-Kit W41 loss-of-function itself. Dahlin's central erythroid Apoptosis\_pro $\downarrow$ claim is reproduced at $p_\text{adj} = 2.3 \times 10^{-13}$ (opposite direction from MPP's $+0.067$). - -Artefacts: \texttt{discovery/hematopoiesis/marker/57\_pathway\_class\_by\_module\_padj.tsv} and \texttt{57\_pathway\_class\_by\_module\_delta.tsv}. The Kit\_signaling module collapse is visible in every predicted class (Figure~\ref{fig:dahlin}). - -\begin{figure}[!htb] -\centering -\includegraphics[width=\textwidth]{figures/fig3_dahlin_heatmap.pdf} -\caption[Dahlin Kit-W41 vs.\ WT pathway heat-map]{Dahlin Kit-W41 vs.\ WT within-class pathway module scores. $\Delta = $ Kit\_W41 mean $-$ WT mean of per-cell module scores.\\ -\textbf{Kit\_signaling:} collapses across every class ($\Delta = -0.27$ in MPP, $p_\text{adj} \approx 0$; expected positive control given the c-Kit W41 loss-of-function mutation).\\ -\textbf{Erythroid Apoptosis\_pro $\downarrow$:} Dahlin's central claim ($p_\text{adj} = 2.3 \times 10^{-13}$; opposite direction from MPP $\Delta = +0.067$).\\ -\textbf{MPP Redox\_glutathione $\uparrow$:} $p_\text{adj} = 1.2 \times 10^{-127}$.\\ -\textbf{ISR upregulation:} $p_\text{adj} = 1.4 \times 10^{-28}$ in MPP.} -\label{fig:dahlin} -\end{figure} - -\subsection{Dahlin marker deep-dive per predicted class} -\label{sec:dahlin-markers} - -Full-class Wilcoxon per PANDA-predicted class on raw Dahlin counts (\pathsplit{discovery/hematopoiesis/marker/92\_dahlin\_marker\_deep\_dive.csv}). Current 10-class prediction distribution: - -\begin{center} -\scriptsize -\setlength{\tabcolsep}{3pt} -\begin{tabular}{p{2.1cm}rrrrp{3.1cm}p{2.5cm}} -\toprule -Class & $n$ & WT & Kit\_W41 & frac\_WT & Top markers & Canonical recovery \\ -\midrule -MPP & 31{,}754 & 27{,}931 & 3{,}823 & 0.880 & Cd34, Adgrl4, Sox4, Pim1 & 1/3 Kit-sig (Sox4) \\ -erythroid & 15{,}397 & 8{,}352 & 7{,}045 & 0.542 & Ca1, Blvrb, Klf1, Aqp1 & 2/7 (Klf1, Blvrb) \\ -megakaryocyte & 8{,}180 & 6{,}596 & 1{,}584 & 0.806 & Itga2b, Pbx1, Apoe, Gata2 & 2/4 (Itga2b, Pf4-adj) \\ -myeloid & 3{,}463 & 1{,}755 & 1{,}708 & 0.507 & Elane, Mpo, Prtn3, Ctsg & 4/8 (Elane, Mpo, Prtn3, Ctsg) \\ -basophil-mast & 848 & 603 & 245 & 0.711 & Cpa3, Ms4a2, Csrp3, Hdc & 4/5 (Cpa3, Ms4a2, Gata2, Hdc) \\ -lymphoid & 694 & 639 & 55 & 0.921 & Gimap1, Gimap6, B2m & --- \\ -monocyte & 472 & 315 & 157 & 0.667 & Irf8, Ms4a6c, Tyrobp & 2/8 (Mpo, Ctsg) \\ -\textbf{LT-HSC} & \textbf{160} & \textbf{135} & \textbf{25} & \textbf{0.844} & \textbf{Hlf, Mpl, Meis1} & \textbf{3/7 (Hlf, Meis1, Mecom)} \\ -pro-B & 66 & 60 & 6 & 0.909 & Ebf1, Cd79a, Vpreb3 & 1/4 (Il7r-adj) \\ -macrophage & 63 & 42 & 21 & 0.667 & Irf8, Lgals1, Tyrobp & 1/8 (Ctsg) \\ -\bottomrule -\end{tabular} -\end{center} - -Basophil-mast, megakaryocyte, erythroid, and myeloid predictions recover canonical panels at LFC $3$--$8$. Erythroid is the class most enriched in Kit-W41 (frac\_WT $0.542$ vs corpus baseline $\sim\!0.76$), consistent with Kit-W41 retaining the erythroid pool but shifting its transcriptome (per \S\ref{sec:dahlin-pathway}). The LT-HSC class recovers 160 cells (84.4\% WT) with canonical Hlf/Mpl/Meis1/Mecom markers --- a Kit-W41-depleted stem population predicted as a first-class label rather than surfaced via abstain. - -\subsection{Per-lineage metabolic reprogramming under Kit-W41} -\label{sec:dahlin-per-lineage-metabolism} - -LT-HSCs and lymphoid progenitors show the strongest OXPHOS upregulation under Kit-W41. We scored all 61{,}122 Dahlin cells on four canonical metabolic modules --- OXPHOS\_ETC, Glycolysis, Fatty\_acid\_oxidation, and Redox\_glutathione --- and computed per-lineage Kit-W41-minus-WT deltas. Ranking lineages by $\sum_m |\Delta_m|$ (total mobilised metabolic signal): - -\begin{center} -\begin{tabular}{lrrrrr} -\toprule -Predicted class & $\Delta$ OXPHOS & $\Delta$ Glycolysis & $\Delta$ FAO & $\Delta$ Redox & $\sum |\Delta|$ \\ -\midrule -macrophage & $+0.044$ & $+0.190$ & $-0.042$ & $+0.098$ & $0.374$ \\ -\textbf{LT-HSC} & $\mathbf{+0.219}$ & $+0.024$ & $-0.035$ & $+0.058$ & $\mathbf{0.335}$ \\ -\textbf{lymphoid} & $\mathbf{+0.196}$ & $+0.039$ & $-0.002$ & $+0.095$ & $\mathbf{0.332}$ \\ -erythroid & $+0.115$ & $+0.051$ & $-0.060$ & $+0.072$ & $0.299$ \\ -MPP & $+0.138$ & $+0.063$ & $-0.009$ & $+0.059$ & $0.269$ \\ -megakaryocyte & $+0.100$ & $+0.016$ & $-0.012$ & $+0.057$ & $0.185$ \\ -basophil-mast & $-0.036$ & $+0.058$ & $+0.016$ & $+0.037$ & $0.147$ \\ -myeloid & $-0.008$ & $+0.073$ & $+0.014$ & $+0.026$ & $0.122$ \\ -monocyte & $+0.023$ & $+0.013$ & $+0.019$ & $+0.027$ & $0.082$ \\ -\bottomrule -\end{tabular} -\end{center} - -\textbf{Novel biology.} Setting aside macrophage (small $n$, glycolysis-dominated), the two lineages with the strongest, most directionally-coherent metabolic reprogramming are \textbf{LT-HSC} (OXPHOS $\Delta = +0.219$) and \textbf{lymphoid} (OXPHOS $\Delta = +0.196$): both show large OXPHOS gains with modest Redox/Glycolysis gains and a slight FAO drop --- the signature of a shift from a quiescent, FAO-supported state to an ETC-active state. This is consistent with hypofunctional Kit forcing canonically-quiescent HSCs into an energetically active (and stressed) state, with the effect largest per cell in the LT-HSC compartment. This per-lineage magnitude ranking is not reported in Dahlin. {\sloppy Artefacts: \pathsplit{discovery/hematopoiesis/marker/108\_dahlin\_lineage\_metabolism\_ranked.csv} and \pathsplit{108\_dahlin\_lineage\_metabolism.json}; script \pathsplit{scripts/analysis/108\_dahlin\_lineage\_metabolism.py}.\par} - -\subsection{Testable wet-lab predictions on Kit-mutant mice} - -\begin{enumerate} -\item \textbf{Erythroid-restricted anti-apoptotic switch}: Bcl2 and Bcl2l1 (Bcl-xL) protein levels should be elevated in Kit-W41 erythroid progenitors (Ter119$^+$-gated) but not in Kit-W41 MPP (Lin$^-$Kit$^+$Sca1$^+$-gated), reflecting the direction-flip in Apoptosis\_pro module score ($\Delta = -0.017$ in erythroid vs.\ $\Delta = +0.067$ in MPP). -\item \textbf{Integrated stress response as a druggable node}: ISRIB or GADD34 hyperactivator treatment of Kit-W41 mice should partially rescue the proliferation defect. ATF4$^+$ nuclei should be elevated in Kit-W41 bone-marrow sections across MPP, erythroid, myeloid, and megakaryocyte compartments. -\item \textbf{Redox rescue}: Supplementation with N-acetylcysteine or reduced glutathione should partially rescue the MPP Redox\_glutathione $\uparrow$ phenotype, providing a druggable node not previously implicated in Kit-W41 biology. -\item \textbf{Compensatory Kit-independent proliferation pathway}: because Kit\_signaling collapses across every class ($p \approx 0$) while some proliferation persists, an alternative RTK (Flt3, Csf1r) or non-RTK pathway must be compensating. -\end{enumerate} - -\section{Discovery target: Veres SC-$\beta$ protocol imperfection} -\label{sec:veres} - -\subsection{Zero-shot classification on Veres hPSC differentiation} - -Veres \emph{et al.}\ 2019 \cite{veres2019scbeta}, \emph{Nature} 569:368--373 (Melton lab), GSE114412: hPSC-directed pancreatic differentiation across Stages 3--6 (foregut endoderm $\to$ stem-cell $\beta$, SC-$\beta$), 69{,}594 total cells; 57{,}297 in training, 12{,}297 held out (\S\ref{sec:zeroshot-veres}). Held-out macro F1 $= 0.900$ (Marker) on 20 canonical Veres classes; 82.8\% HVG overlap after human$\to$mouse case-fold. - -Predicted class distribution across Veres stages (n=9{,}080 cells with parseable stage annotation; the remaining 3,217 cells are primary-islet controls without a stage tag): -\begin{center} -\small -\setlength{\tabcolsep}{5pt} -\resizebox{\textwidth}{!}{% -\begin{tabular}{p{4.4cm}rrrr} -\toprule -Predicted class & Stage 3 (n=2{,}899) & Stage 4 (n=2{,}577) & Stage 5 (n=1{,}589) & \textbf{Stage 6 (n=2{,}015)} \\ -\midrule -pancreatic-progenitor & 72\% & 0\% & 0\% & 0\% \\ -proliferating & 28\% & 14\% & 5\% & 2\% \\ -endocrine-progenitor-primed & 0\% & 41\% & 0\% & 0\% \\ -alpha\_progenitor & 0\% & 26\% & 49\% & \textbf{59\%} \\ -beta\_progenitor & 0\% & 0\% & 11\% & \textbf{10\%} \\ -epsilon & 0\% & 0\% & 13\% & 11\% \\ -exocrine & 0\% & 0\% & 14\% & 10\% \\ -ductal, delta & 0\% & 0\% & 3\% & 4\% each \\ -\bottomrule -\end{tabular}} -\end{center} - -\texttt{pancreatic-progenitor} dominates Stage 3 (72\%) and collapses across Stages 4--5 as endocrine-progenitor-primed and alpha\_progenitor identities emerge; by Stage 6, alpha\_progenitor is the largest class (59\%) and beta\_progenitor is 10\%. Adult-$\alpha$/adult-$\beta$ sub-prototypes never activate; SC-$\beta$-protocol cells occupy immature progenitor identities in the corpus manifold. Without any stage information supplied as input, \panda{} reproduces Veres's central finding that the differentiation is imperfect --- alpha-lineage outnumbers beta-lineage $\sim\!6\times$ at the SC-$\beta$ target stage (Figure~\ref{fig:veres}). - -\begin{figure}[!htb] -\centering -\includegraphics[width=\textwidth]{figures/fig4_veres_stage_stack.pdf} -\caption[Veres 2019 stage-stack composition]{Veres 2019 iPSC-directed pancreatic differentiation. Stacked bars show PANDA-Marker predicted class fraction across the four staged Veres samples ($n=9{,}080$ cells with parseable stage annotation).\\ -\textbf{Stage 3:} \texttt{pancreatic-progenitor} dominates (72\% of $n=2{,}899$) and monotonically declines through the protocol.\\ -\textbf{Stage 6:} \texttt{alpha\_progenitor} rises from 0\% at Stage 3 to 59\% at the SC-$\beta$ target stage.\\ -\textbf{Sub-prototypes:} Adult-$\beta$ and adult-$\alpha$ sub-prototypes do not activate on Veres (see \S\ref{sec:zeroshot-veres}).} -\label{fig:veres} -\end{figure} - -\subsection{Stage-6 alpha\_progenitor vs beta\_progenitor mechanism} -\label{sec:veres-stage6-mech} - -Restricting to Stage 6: $1{,}192$ alpha\_progenitor vs $199$ beta\_progenitor predictions. The marker deep-dive (\S\ref{sec:veres-marker-deep-dive}) places the identity boundary on the master-TF axis: alpha\_progenitor recovers 4/4 canonical adult-$\alpha$ markers (GCG, ARX, IRX2, MAFB) at LFC $+3$ to $+7.5$; beta\_progenitor top-3 are ACVR1C ($+4.5$), CALB2 ($+5.9$), INS ($+4.8$). The residual $6\times$ alpha/beta asymmetry at Stage 6 matches Veres's own report of SC-$\beta$-protocol inefficiency. Because the corpus adult-$\beta$ prototype does not activate on Veres (\S\ref{sec:veres-adult-beta}), Stage-6 secretory cells route to \texttt{beta\_progenitor} rather than \texttt{beta}, and immature-$\beta$ cannot be distinguished from committed adult-$\beta$ on this checkpoint. - -\subsection{Veres marker deep-dive per predicted class} -\label{sec:veres-marker-deep-dive} - -Wilcoxon per PANDA-predicted class on raw Veres counts (\pathsplit{discovery/pancreas/marker/91\_veres\_marker\_deep\_dive.csv}). Top-9 classes by cell count shown; distribution reflects the current 18-class predicted vocabulary on the held-out slice: -\begin{center} -\scriptsize -\setlength{\tabcolsep}{3pt} -\begin{tabular}{p{3.0cm}rp{4.8cm}p{2.6cm}r} -\toprule -Predicted class & $n$ & Top-3 markers (LFC) & Canonical panel recovery & \% of class at Stage 6 \\ -\midrule -alpha\_progenitor & 2{,}864 & GCG ($+7.5$), TTR ($+4.9$), CHGA ($+4.7$) & 4/4 alpha (GCG, ARX, IRX2, MAFB) & 42\% \\ -pancreatic-progenitor & 2{,}099 & MDK ($+3.5$), SOX11 ($+3.9$), FN1 ($+4.5$) & --- & 0\% \\ -proliferating & 1{,}273 & TUBA1B ($+2.7$), TUBB ($+2.5$), HMGB1 ($+2.2$) & --- & 3\% \\ -endocrine-progenitor-primed & 1{,}057 & DLK1 ($+6.0$), LDHB ($+3.0$), PDX1 ($+2.9$) & 1/5 beta (PDX1) & 0\% \\ -epsilon & 701 & DDC ($+5.7$), FEV ($+5.1$), CHGA ($+5.3$) & 2/2 EP-Fev (FEV, INSM1) & 31\% \\ -beta & 560 & INS ($+7.4$), IAPP ($+9.6$), ADCYAP1 ($+7.5$) & 1/5 beta (INS) $+$ MAFA/UCN3 in top-15 & 0\% \\ -beta\_progenitor & 539 & ACVR1C ($+4.5$), CALB2 ($+5.9$), INS ($+4.8$) & 1/2 EP-Fev (INSM1) & 37\% \\ -delta & 533 & SST ($+7.0$), ISL1 ($+2.6$), PCP4 ($+2.7$) & 1/2 EP-Fev (INSM1) & 16\% \\ -acinar & 426 & REG1A ($+11.8$), CTRB1 ($+12.2$), CTRB2 ($+11.8$) & 4/4 acinar (PRSS1, PRSS2, CEL, CTRB1) & 0\% \\ -\bottomrule -\end{tabular} -\end{center} - -\textbf{Interpretation.} On the current pan-pancreatic checkpoint, \panda{}-Marker does not separate adult-$\alpha$ from juvenile $\alpha$ or adult-$\beta$ from juvenile $\beta$ on the Veres held-out slice. The \texttt{alpha\_progenitor} cluster recovers all four canonical adult-$\alpha$ markers (GCG, ARX, IRX2, MAFB) and dominates Veres Stage 6 (42\%). Acinar recovers all four PRSS1/PRSS2/CTRB1/CEL exocrine markers at LFC $\sim 10$. - -\paragraph{Adult-$\beta$ signal within the \texttt{beta} class.}\label{sec:veres-adult-beta} -\panda{}-Marker predicts 0 Veres cells as \texttt{adult-beta} on the current checkpoint; the adult sub-prototype does not activate. The adult-maturity signature nonetheless survives at the gene level inside the \texttt{beta} class ($n=560$): mean-\texttt{log1p} enrichment vs the whole Veres dataset is MAFA $16\times$ (0.89 vs 0.06), UCN3 $8\times$, IAPP $12\times$, INS $3\times$. Two consistent readings: (i) Veres SC-$\beta$ cells occupy the corpus \texttt{beta} manifold too broadly for the adult sub-prototype's angular margin to fire; (ii) the corpus \texttt{adult-beta} prototype was learned on primary human islet data and the SC-$\beta$ transcriptome is close-but-not-close-enough. Artefact \texttt{95\_adult\_beta\_validation.json} is marked \texttt{vacuous: true} for the adult-vs-juvenile ratio. - -\subsection{Polyhormonal SC-$\alpha$ sub-clustering} -\label{sec:veres-polyhormonal-subcluster} - -Polyhormonal SC-$\alpha$ is a discrete sub-cluster (cluster 3), not stochastic co-expression across the pool. Veres 2019 reports that the SC-$\alpha$ output is polyhormonal (Gcg + Ins + Sst co-expression) but does not resolve whether this reflects a single trapped-bipotent sub-population or noise-floor stochastic co-expression across the SC-$\alpha$ pool. To test this we pooled Veres alpha-lineage predictions (\emph{alpha\_progenitor} + \emph{alpha}, $n = 3{,}473$), defined \emph{polyhormonal} as $\geq 2$ of \{Ins, Gcg, Sst\} above their respective 75\textsuperscript{th} percentiles within the alpha pool, and Leiden-clustered at resolution $= 0.5$ into 8 sub-clusters: - -\begin{center} -\begin{tabular}{lrrrrrr} -\toprule -Sub-cluster & $n$ & \% pool & \% polyhormonal & \% GCG-hi & \% SST-hi & Enrichment \\ -\midrule -\textbf{3} & \textbf{522} & \textbf{15\%} & \textbf{46\%} & \textbf{69\%} & \textbf{43\%} & \textbf{$\mathbf{2.53\times}$} \\ -0 & 868 & 25\% & 30\% & 43\% & 29\% & $1.65\times$ \\ -2 & 585 & 17\% & 15\% & 14\% & 27\% & $0.82\times$ \\ -6, 4, 1, 5, 7 & 1{,}498 & 43\% & $\leq 9\%$ & --- & --- & $\leq 0.48\times$ \\ -\midrule -baseline & 3{,}473 & 100\% & 18.1\% & --- & --- & $1.00\times$ \\ -\bottomrule -\end{tabular} -\end{center} - -\textbf{Novel biology.} Cluster 3 is a discrete polyhormonal sub-cluster: 15\% of the alpha pool, 46\% polyhormonal (\emph{i.e.}\ $2.5\times$ enriched over the 18.1\% baseline), with 69\% GCG-hi, 43\% SST-hi, and 36\% INS-hi \emph{simultaneously} elevated. Cluster 0 is a milder polyhormonal shoulder ($1.65\times$), and sub-clusters 4--7 are essentially devoid of polyhormonal cells ($\leq 9\%$). Polyhormonal SC-$\alpha$ is thus not uniform stochastic co-expression scattered across the SC-$\alpha$ pool; it is concentrated in a distinct sub-cluster comprising roughly 15--40\% of the pool (cluster 3 alone, or clusters 3+0). This supports a "trapped bipotent progenitor" interpretation of the Veres polyhormonal phenotype over a "noise-floor" interpretation: cluster 3 is a candidate for FISH-sorting and downstream functional / lineage-tracing assays. {\sloppy Artefacts: \pathsplit{discovery/pancreas/marker/110\_veres\_polyhormonal\_alpha\_per\_cluster.csv} and \pathsplit{110\_veres\_polyhormonal\_alpha\_summary.json}; script \pathsplit{scripts/analysis/110\_veres\_polyhormonal\_alpha.py}.\par} - -\subsection{Testable wet-lab predictions on Veres SC-$\beta$ differentiation} - -\begin{enumerate} -\item \textbf{Terminal TF-switching rescues yield}: overexpressing Nkx6-1, Mnx1, or Neurod1 during Stage 5$\to$6 should redirect a fraction of the SC-$\alpha$ population to SC-$\beta$, because the alpha/beta boundary is driven by the TF axis rather than by secretion machinery ($p = 1.9 \times 10^{-87}$ for beta\_TFs difference). -\item \textbf{Arx knockdown at Stage 5}: an Arx knockdown at the Stage 5 $\to$ 6 transition should convert SC-$\alpha$ to SC-$\beta$ or delta ($\log_2$ fc $+2.71$ for Arx in alpha). -\item \textbf{Polyhormonal sorting}: single-cell FISH for Hhex or Sst on Stage 6 cells should identify a Gcg$^+$Sst$^+$ double-positive population matching the delta\_TFs UP-in-alpha signal ($\Delta = +0.41$, $p = 7.5 \times 10^{-43}$); Gcg$^+$Sst$^-$ vs.\ Gcg$^+$Sst$^+$ transplantation assays would test the SC-$\alpha$ functional-immaturity claim. -\item \textbf{Adult-$\beta$ marker validation}: a MAFA-reporter line applied to Stage 6 output should identify a MAFA$^\text{high}$/UCN3$^\text{high}$/IAPP$^\text{high}$ subpopulation within the \panda{}-\texttt{beta} predicted cells (where MAFA/UCN3/IAPP are $8$--$16\times$ enriched at the gene level), even though the pan-pancreatic corpus's \texttt{adult-beta} sub-prototype does not activate on Veres cells. -\end{enumerate} - -\section{Prototype geometry} -\label{sec:proto-geom} - -For each system, \panda{}'s $K$ class-mean prototype vectors in the 128-d projection space have a participation-ratio effective dimensionality -\[ -\text{eff\_dim}(P) \;=\; \frac{\big(\sum_i \lambda_i\big)^2}{\sum_i \lambda_i^2}, -\] -where $\{\lambda_i\}$ are the eigenvalues of the centred prototype covariance. This is a smooth ``how many independent directions is the prototype set using?'' summary that equals $K$ when prototypes are orthonormal and 1 when they are collinear: - -\begin{center} -\begin{tabular}{lrrr} -\toprule -System & $K$ & Effective dim & Fraction \\ -\midrule -Pan-skin & 13 & \textbf{11.08} & 85\% \\ -Pan-hematopoietic & 15 & \textbf{11.33} & 76\% \\ -Pan-pancreatic & 20 & \textbf{10.85} & 54\% \\ -\bottomrule -\end{tabular} -\end{center} - -\textbf{Interpretation.} The pan-skin prototype set uses $85\%$ of the available dimensions, indicating cleanly-separated identities that occupy nearly-orthogonal directions in the projection space. The pan-hematopoietic prototype set uses $76\%$: still healthy, but the shared myeloid/basophil/megakaryocyte Gata1/Gata2$^+$ progenitor program links several lineages onto a common axis. The pan-pancreatic prototype set uses only $54\%$ of its 20 dimensions --- the largest prototype-set compression in the paper --- consistent with the well-established secondary transition dogma of pancreatic endocrinogenesis, in which $\beta$/$\delta$/endocrine-progenitor/other cell fates share a strong Nkx6-1$^+$/Neurod1$^+$ transcriptional program until terminal Ins1$^\text{hi}$/Mafa$^\text{hi}$ maturation. This is a direct empirical readout of biological lineage compression in the training corpus: the prototype geometry \panda{} learns is determined by transcriptomic distinguishability between labels, so labels whose training cells are close in expression space end up with lower-dimensional prototype sets even if the canonical ontology says they are distinct cell types. - -\section{Cross-system synthesis} - -Three tissue systems, three discovery targets, three paper-central findings independently reproduced by \panda{}: - -\begin{center} -\small -\begin{tabular}{p{2.4cm}p{3.9cm}p{8.6cm}} -\toprule -System & Held-out target & Paper claim reproduced \\ -\midrule -Skin & Dingwall En1-cKO & Unified spatial-repressor model $+$ EDEN lineage extension \\ -Hematopoiesis & Dahlin Kit-W41 & Every 4-of-4 paper compositional and molecular claims \\ -Pancreas & Veres hPSC & Polyhormonal SC-$\alpha$ $+$ terminal-TF-axis mechanism \\ -\bottomrule -\end{tabular} -\end{center} - -\textbf{Common threads.} Across all three systems: (i) zero-shot classification followed by within-class Wilcoxon and pathway module scoring recovers canonical marker sets never supplied to the model (Cd34/Myc for MPP; Cpa3/Gata2/Hdc/Mcpt8 for basophil-mast; Sox10/Dct/Tyr/Pmel for melanocyte; Nkx6-1/Mnx1/Neurod1 for SC-$\beta$; Arx/Irx2 for SC-$\alpha$); (ii) the confidence-gate abstain mechanism flags biology the training corpus does not represent (LT-HSC-like Hlf$^+$/Cd34$^+$ cluster in Dahlin, foregut endoderm in Veres, Derm-lineage subclusters in Dingwall); (iii) between-condition module-score contrasts reproduce the source paper's mechanistic interpretations at Wilcoxon $p$-values exceeding textbook thresholds by many orders of magnitude. - -\textbf{Divergences.} The three systems separate cleanly by prototype geometry (\S\ref{sec:proto-geom}): skin at $85\%$ effective-dim utilisation shows fully-resolved terminal identities; hematopoiesis at $76\%$ has moderate lineage-shared compression on the Gata2$^+$ multipotent progenitor axis; pancreas at $54\%$ has strong compression on the shared endocrine-progenitor program. This ordering is itself a testable biological statement: the transcriptomic distinguishability of terminal identities in a training corpus is directly readable off the prototype set's effective dimensionality. Applying this framework prospectively lets an atlas-builder decide, before training, whether adult-anchor cells (or additional developmental stages) are needed to unlock terminal-cell-type separability. - -Figure~\ref{fig:multiumap} shows the projection geometry for the three discovery targets; the condition of interest induces local density shifts within a cluster-preserving embedding, consistent with the within-class contrasts reported above. - -\begin{figure}[!htb] -\centering -\includegraphics[width=\textwidth]{figures/fig6_multi_umap.pdf} -\caption[Three-panel discovery-target UMAP]{Three discovery-target UMAPs in \panda{}'s 128-d projection space (8{,}000 cells subsampled per panel).\\ -\textbf{(a)} Dingwall skin coloured by En1 genotype.\\ -\textbf{(b)} Dahlin hematopoiesis coloured by Kit genotype.\\ -\textbf{(c)} Veres pancreatic differentiation coloured by protocol stage (3--6).\\ -\textbf{Observation:} between-condition/between-stage differences manifest as local density shifts within a shared, biologically-structured embedding.} -\label{fig:multiumap} -\end{figure} - -\section{Discussion} - -\panda{} delivers a compact PCA-only classifier trained on 100\%-paper-labeled corpora that validates within-corpus, transfers to labeled hold-outs (both strict zero-shot and held-out-slice), and supports mechanistic discovery via within-class Wilcoxon DE and pathway module scoring. - -Three points about how the checkpoint is used in discovery. First, the Dingwall EDEN validation illustrates the intended workflow of a prototype-anchored classifier as a hypothesis-generation tool: the pan-tissue fibroblast prototype is calibrated broader than any narrow tissue-specific compartment, so an independent clustering (Line B) or target-native training pass (Line C) is required to resolve EDEN as a single class; once either is provided, the depletion phenotype is a learnable and reproducible property of the labelling rather than a clustering artefact, and sub-cluster panel-matching extends the phenotype to the untested Derm2 precursor. Second, expanded pathway scoring on Dahlin surfaces MPP Redox\_glutathione $\uparrow$ ($p = 1.2\!\times\!10^{-127}$) as a novel Kit-W41 phenotype not reported in the original paper, alongside a megakaryocyte quiescence-signature loss and the erythroid-vs-MPP direction-flip in Apoptosis\_pro. Third, on Veres the adult-$\beta$ sub-prototype does not activate on SC-$\beta$ cells even though the adult-maturity signature (MAFA, UCN3, IAPP, ADCYAP1) is $8$--$16\times$ enriched at the gene level inside the \texttt{beta} class --- a resolution limit of the current corpus that the prototype-geometry compression (54\% effective dim, \S\ref{sec:proto-geom}) makes quantitatively concrete. - -Future work will extend the interpretability toolkit --- prototype--gene attribution, counterfactual single-gene knockout, and Hessian gene-gene interaction probes --- using the analytic PCA back-projection that \panda{}'s architecture makes exact. - -\section{Limitations} - -\begin{itemize} -\item Pan-tissue prototypes are calibrated for cross-dataset generality; narrow tissue-specific compartments like Dingwall's EDEN require either an independent clustering step (Line B) or Dingwall-native training (Line C) to be recovered as a single class. -\item Zero-shot transfer to mature terminal cell types requires developmentally-mature anchors in the training corpus (documented via the Baron test-half zero-shot performance and the pancreatic prototype-geometry compression). -\item Sparse-class predictions (HF-placode, HF-DP, gamma/immune on Baron test) are underpowered when the training-corpus support is small. -\item Novel-population claims from confidence-gated abstained cells are hypotheses; every mechanistic claim in the discovery reports requires wet-lab confirmation. -\item Fibroblast-papillary predictions in Dingwall show smooth-muscle contamination (Cald1, Myh11, Acta2), consistent with the class being conflated with myofibroblast identity in adult skin. -\item Cross-platform generalisation to Smart-seq2 (Nestorowa) remains a hard benchmark; a Smart-seq2 LT-HSC training anchor would be the natural next step. -\end{itemize} - -\section{Reproducibility} - -\textbf{Multiple-testing scope.} Pathway-module $p$-values reported in the class $\times$ module tables (\S\ref{sec:dingwall-pathway}, \S\ref{sec:dahlin-pathway}) are Bonferroni-corrected within-system across the full $n_\text{classes} \times n_\text{modules}$ family scanned per system (Dingwall: $12 \times 15+$ modules; Dahlin: $10 \times 30$ modules). Per-class Wilcoxon DEG $p$-values (\S\ref{sec:dingwall-hf-placode-degs}, \S\ref{sec:dahlin-markers}, and the Veres per-class marker deep-dive) use Benjamini--Hochberg per class. Fisher exact class enrichments (\S\ref{sec:dingwall-fisher}) use Benjamini--Hochberg across the classes tested per system. - -All quantitative claims in this paper trace to a specific artefact: -\begingroup -\RaggedRight -\sloppy -\begin{itemize}\itemsep2pt -\item Model checkpoints: \pathsplit{checkpoints/\{system\}/\{variant\}/panda\_final.pt} where variant $\in \{$pca, marker$\}$ -\item Marker channel gene lists: \pathsplit{panda/markers.yaml} -\item Corpus builders: \pathsplit{scripts/\{pan\_skin,hematopoiesis,pancreas\}/}; unified trainer: \pathsplit{scripts/common/train\_panda.py}; unified 5-fold CV driver: \pathsplit{scripts/common/cv\_holdout.py}; zero-shot inference: \pathsplit{scripts/common/run\_all\_zero\_shot.py} -\item External-label supplements: \pathsplit{data/external\_labels/\{dingwall\_supp,haensel,joost2016,mca,mia,byrnes,yu,baccin,melanocyte\_anchor\}/} -\item Per-configuration 5-fold CV (three seeds per system): \pathsplit{discovery/\{system\}/\{variant\}/cv\_5fold.json}, \pathsplit{cv\_5fold\_seed1.json}, \pathsplit{cv\_5fold\_seed2.json} -\item Zero-shot labeled targets: - \begin{itemize}\itemsep2pt - \item Baron test-half: \pathsplit{discovery/pancreas/\{pca,marker\}/baron\_summary.json} (current 20-class vocabulary; the paper body's Baron numbers are read from this file) - \item Veres held-out 12{,}297: \pathsplit{discovery/pancreas/\{pca,marker\}/veres\_summary.json} - \item Nestorowa: \pathsplit{discovery/hematopoiesis/\{pca,marker\}/nestorowa\_summary.json} - \item Sulic: \pathsplit{discovery/pan\_skin/\{pca,marker\}/sulic\_summary.json} - \item Belote: \pathsplit{discovery/pan\_skin/\{pca,marker\}/belote\_summary.json} - \end{itemize} -\item Adult-$\beta$ marker-panel validation on Veres: \pathsplit{discovery/pancreas/marker/95\_adult\_beta\_validation.json} -\item Expanded pathway module analysis (all systems): \pathsplit{scripts/analysis/57\_pathway\_analysis.py}, outputs at \pathsplit{discovery/\{system\}/marker/57\_pathway\_class\_by\_module\_\{padj,delta\}.tsv} -\item \textbf{Dingwall EDEN validation (three lines of evidence)}: - \begin{itemize}\itemsep2pt - \item Line A post-hoc scoring: \pathsplit{scripts/analysis/98\_eden\_posthoc\_detection.py}, output \pathsplit{discovery/pan\_skin/marker/98\_eden\_summary.json} - \item Line B scanpy reproduction: \pathsplit{scripts/analysis/103\_replicate\_dingwall\_seurat\_pipeline.py}, output \pathsplit{data/processed/dingwall\_replica/dingwall\_replica.h5ad}, \pathsplit{data/processed/dingwall\_replica/replica\_cluster\_20\_qc.json}, \pathsplit{replica\_marker\_matches.csv} - \item Line C PANDA on Derm labels: \pathsplit{scripts/analysis/104\_train\_on\_dingwall\_derm\_labels.py}, output \pathsplit{discovery/pan\_skin/marker/104\_dingwall\_derm\_summary.json} $+$ prediction/depletion CSVs - \end{itemize} -\item Primary EDEN Derm2 discovery: \pathsplit{scripts/analysis/100\_primary\_eden\_discovery.py}, \pathsplit{101\_primary\_eden\_derm\_scoring.py}; outputs \pathsplit{discovery/pan\_skin/marker/100\_primary\_eden\_discovery.csv}, \pathsplit{100\_primary\_eden\_summary.json}, \pathsplit{101\_derm\_identity\_summary.json}, \pathsplit{101\_derm\_subcluster\_scores.csv} -\item Dingwall other mechanistic outputs: \pathsplit{discovery/pan\_skin/marker/57\_pathway\_analysis.csv, 90\_dingwall\_marker\_deep\_dive.csv} -\item Dahlin Kit-mutant: \pathsplit{discovery/hematopoiesis/marker/92\_dahlin\_marker\_deep\_dive.csv, dahlin\_summary.json} -\item Veres marker deep-dive: \pathsplit{discovery/pancreas/marker/91\_veres\_marker\_deep\_dive.csv} -\end{itemize} -\endgroup - -\begin{thebibliography}{99} -\bibitem{dingwall2024en1cko} Dingwall CB, \emph{et al.}\ (Aldea D, Kamberov YG \emph{corresp.}) (2024). ``Divergent developmental origins for the mammalian sweat gland lineage revealed by an En1-Cre knock-in.'' \emph{Developmental Cell} 59(1):20--32.e6 \url{https://doi.org/10.1016/j.devcel.2023.11.017}. GSE220977. -\bibitem{veres2019scbeta} Veres A, \emph{et al.}\ (Melton DA \emph{corresp.}) (2019). ``Charting cellular identity during human in vitro $\beta$-cell differentiation.'' \emph{Nature} 569:368--373 \url{https://doi.org/10.1038/s41586-019-1168-5}. GSE114412. -\bibitem{dahlin2018kit} Dahlin JS, \emph{et al.}\ (Wilson NK \emph{corresp.}) (2018). ``A single-cell hematopoietic landscape resolves 8 lineage trajectories and defects in Kit mutant mice.'' \emph{Blood} 131(21):e1--e11 \url{https://doi.org/10.1182/blood-2017-12-821413}. GSE107727. -\bibitem{haensel2020skin} Haensel D, \emph{et al.}\ (Annusver K \emph{et al.}) (2020). ``Defining epidermal basal cell states during skin homeostasis and wound healing using single-cell transcriptomics.'' \emph{Cell Reports} 30(11):3932--3947.e6 \url{https://doi.org/10.1016/j.celrep.2020.02.091}. GSE142471. -\bibitem{joost2016} Joost S, \emph{et al.}\ (2016). ``Single-cell transcriptomics reveals that differentiation and spatial signatures shape epidermal and hair follicle heterogeneity.'' \emph{Cell Systems} 3(3):221--237.e9 \url{https://doi.org/10.1016/j.cels.2016.08.010}. GSE67602. -\bibitem{belote2021} Belote RL, \emph{et al.}\ (2021). ``Human melanocyte development and melanoma dedifferentiation at single-cell resolution.'' \emph{Nature Cell Biology} 23(9):1035--1047 \url{https://doi.org/10.1038/s41556-021-00740-8}. GSE151091. -\bibitem{han2018mca} Han X, \emph{et al.}\ (Guo G \emph{corresp.}) (2018). ``Mapping the mouse cell atlas by Microwell-Seq.'' \emph{Cell} 172(5):1091--1107.e17 \url{https://doi.org/10.1016/j.cell.2018.02.001}. GSE108097. -\bibitem{tms2020} Tabula Muris Consortium (2020). ``A single-cell transcriptomic atlas characterizes ageing tissues in the mouse.'' \emph{Nature} 583:590--595 \url{https://doi.org/10.1038/s41586-020-2496-1}. GSE132042. -\bibitem{baccin2020} Baccin C, \emph{et al.}\ (2020). ``Combined single-cell and spatial transcriptomics reveal the molecular, cellular and spatial bone marrow niche organization.'' \emph{Nature Cell Biology} 22(1):38--48 \url{https://doi.org/10.1038/s41556-019-0439-6}. GSE122465. -\bibitem{bastidas2019} Bastidas-Ponce A, \emph{et al.}\ (Bakhti M, Lickert H) (2019). ``Comprehensive single cell mRNA profiling reveals a detailed roadmap for pancreatic endocrinogenesis.'' \emph{Development} 146(12):dev173849 \url{https://doi.org/10.1242/dev.173849}. GSE132188. -\bibitem{byrnes2018} Byrnes LE, \emph{et al.}\ (Sneddon JB \emph{corresp.}) (2018). ``Lineage dynamics of murine pancreatic development at single-cell resolution.'' \emph{Nature Communications} 9(1):3922 \url{https://doi.org/10.1038/s41467-018-06176-3}. GSE101099. -\bibitem{yu2021} Yu X, \emph{et al.}\ (2021). ``Single-cell RNA-seq of the developing pancreas identifies novel endocrine progenitors and endocrine subtypes.'' \emph{Cell Research} 31:669--686 \url{https://doi.org/10.1038/s41422-021-00509-6}. GSE139627. -\bibitem{hrovatin2023mia} Hrovatin K, \emph{et al.}\ (2023). ``Delineating mouse $\beta$-cell identity during lifetime and in diabetes with a single cell atlas.'' \emph{Nature Metabolism} 5:1615--1637 \url{https://doi.org/10.1038/s42255-023-00876-x}. GSE211796. -\end{thebibliography} - -\clearpage -\subsection*{Supplement figure index} -\label{sec:supp-index} - -Supporting figures are provided as PDFs under \pathsplit{figures/supplement/} (S\emph{n}) and \pathsplit{figures/biology/} (B\emph{n}). Each entry lists filename, one-line description, and the paper section that motivates it. - -\begingroup -\RaggedRight -\small -\begin{itemize}\itemsep1pt -\item \textbf{S1} \pathsplit{01\_cv\_summary.pdf} --- Held-out 5-fold CV summary across the three systems (\S\ref{sec:multiseed}; Fig~\ref{fig:cv}). -\item \textbf{S2} \pathsplit{02\_per\_class\_f1.pdf} --- Per-class F1 bars for skin/HSC/pancreas (Fig~\ref{fig:cv}). -\item \textbf{S3} \pathsplit{03\_prototype\_cosine.pdf} --- Prototype-prototype cosine similarity matrices per system (\S\ref{sec:proto-geom}). -% S4 removed --- legacy training-trajectory figure contained stale class vocab and wrong K counts. -\item \textbf{S5} \pathsplit{05\_adversary\_purification.pdf} --- Dataset/depth-adversary AUROC vs.\ ramp step (curriculum diagnostics for \S\ref{sec:multiseed}). -\item \textbf{S6} \pathsplit{06\_cross\_system\_prototypes.pdf} --- Cross-system prototype-geometry comparison (\S\ref{sec:proto-geom}). -% S7 removed --- attribution-heatmap contained deprecated class labels and leaked cell-barcode strings as gene columns. -% S8 removed --- TF-enrichment used stale skin/pancreas class rosters. -% S9 removed --- KO-essentials contained deprecated class labels and leaked cell-barcode strings. -% S10 removed --- Hessian gene-pair figure contained stale classes and illegible pair labels. -\item \textbf{S11} \pathsplit{11\_novel\_populations.pdf} --- Abstained-cell clusters flagged as candidate novel populations (\S{Discussion}, cross-system). -% S12 removed --- Co-attention modules referenced deprecated classes (nascent-eccrine-gland, UNK, unassigned, mesenchymal). -\item \textbf{S13} \pathsplit{13\_dingwall\_umap.pdf} --- Full Dingwall UMAP by predicted class and genotype (\S\ref{sec:dingwall}; Fig~\ref{fig:umap}). -\item \textbf{S14} \pathsplit{14\_dahlin\_umap.pdf} --- Full Dahlin UMAP by predicted class and Kit genotype (\S\ref{sec:dahlin}). -\item \textbf{S15} \pathsplit{15\_veres\_umap.pdf} --- Full Veres UMAP by predicted class and stage (\S\ref{sec:veres}). -\item \textbf{S16} \pathsplit{16\_dingwall\_discovery.pdf} --- Dingwall discovery panel: EDEN, Derm2, HF-placode DEGs (\S\ref{sec:dingwall-eden}, \S\ref{sec:dingwall-primary-eden}, \S\ref{sec:dingwall-hf-placode-degs}). -\item \textbf{S17} \pathsplit{17\_dahlin\_discovery.pdf} --- Dahlin discovery panel: MPP redox, MK quiescence, per-lineage metabolism (\S\ref{sec:dahlin-pathway}, \S\ref{sec:dahlin-per-lineage-metabolism}). -\item \textbf{S18} \pathsplit{18\_veres\_discovery.pdf} --- Veres discovery panel: Stage-6 alpha/beta, polyhormonal sub-cluster (\S\ref{sec:veres-polyhormonal-subcluster}). -\item \textbf{S19} \pathsplit{19\_myeloid\_network.pdf} --- Myeloid marker co-expression network (context: Dahlin \S\ref{sec:dahlin-markers}). -\item \textbf{S20} \pathsplit{20\_placode\_wnt\_module.pdf} --- HF-placode / Wnt module score across genotypes (\S\ref{sec:dingwall-pathway}, \S\ref{sec:dingwall-hf-placode-degs}). -\item \textbf{S23} \pathsplit{23\_anchor\_delta\_recall.pdf} --- Anchor-vs-holdout recall $\Delta$ across the four held-out labeled targets. -\item \textbf{S24} \pathsplit{24\_pca\_vs\_marker\_umaps\_dingwall\_by\_genotype.pdf} --- PCA vs.\ Marker UMAP of Dingwall coloured by En1 genotype (Marker vs.\ PCA ablation; \S\ref{sec:dingwall}). -\item \textbf{S24b} \pathsplit{24b\_pca\_vs\_marker\_umaps\_dingwall\_by\_class.pdf} --- Same Dingwall PCA vs.\ Marker UMAP coloured by predicted class. -\item \textbf{S25} \pathsplit{25\_pca\_vs\_marker\_umaps\_dahlin\_by\_genotype.pdf} --- PCA vs.\ Marker UMAP of Dahlin coloured by Kit genotype (\S\ref{sec:dahlin}). -\item \textbf{S25b} \pathsplit{25b\_pca\_vs\_marker\_umaps\_dahlin\_by\_class.pdf} --- Same Dahlin PCA vs.\ Marker UMAP coloured by predicted class. -\item \textbf{S26} \pathsplit{26\_pca\_vs\_marker\_umaps\_veres\_by\_stage.pdf} --- PCA vs.\ Marker UMAP of Veres coloured by protocol stage (\S\ref{sec:veres}). -\item \textbf{S26b} \pathsplit{26b\_pca\_vs\_marker\_umaps\_veres\_by\_class.pdf} --- Same Veres PCA vs.\ Marker UMAP coloured by predicted class. -\item \textbf{S27} \pathsplit{27\_dingwall\_en1\_enrichment.pdf} --- En1-cKO enrichment per PANDA class with Fisher $p$ (\S\ref{sec:dingwall-fisher}). -\item \textbf{S28} \pathsplit{28\_melanocyte\_pathway\_modules.pdf} --- Melanocyte-lineage pathway module scores in Dingwall (\S\ref{sec:dingwall-melanoblast-mitf}). -\item \textbf{B1} \pathsplit{biology\_01\_dingwall\_umap.pdf} --- Publication-style Dingwall UMAP by predicted class (\S\ref{sec:dingwall}). -\item \textbf{B2} \pathsplit{biology\_02\_primary\_eden.pdf} --- Primary-EDEN Derm2 sub-cluster and marker overlap (\S\ref{sec:dingwall-primary-eden}). -\item \textbf{B3} \pathsplit{biology\_03\_melanoblast\_mitf.pdf} --- Melanoblast MITF-regulon vs.\ neural-crest scoring (\S\ref{sec:dingwall-melanoblast-mitf}). -\item \textbf{B4} \pathsplit{biology\_04\_dahlin\_metabolism.pdf} --- Per-lineage metabolic reprogramming heat-map on Dahlin (\S\ref{sec:dahlin-per-lineage-metabolism}). -\item \textbf{B5} \pathsplit{biology\_05\_dahlin\_composition.pdf} --- Compositional shifts under Kit-W41 (\S\ref{sec:dahlin}, Fisher table). -\item \textbf{B6} \pathsplit{biology\_06\_veres\_beta\_quadrant.pdf} --- Veres beta-class MAFA/UCN3/IAPP quadrant (\S\ref{sec:veres-adult-beta}). -\item \textbf{B7} \pathsplit{biology\_07\_veres\_polyhormonal.pdf} --- Veres polyhormonal SC-$\alpha$ sub-cluster 3 (\S\ref{sec:veres-polyhormonal-subcluster}). -\item \textbf{B8} \pathsplit{biology\_08\_prototype\_geometry.pdf} --- Cross-system prototype effective-dimensionality plot (\S\ref{sec:proto-geom}). -\end{itemize} -\endgroup - -\end{document} +\documentclass[11pt,letterpaper]{article} +\usepackage[margin=1in]{geometry} +\usepackage{amsmath,amssymb,amsthm} +\usepackage{graphicx} +\usepackage{booktabs} +\usepackage{longtable} +\usepackage{multirow} +\usepackage{array} +\usepackage[hidelinks,breaklinks=true]{hyperref} +\usepackage{siunitx} +\usepackage[T1]{fontenc} +\usepackage{lmodern} +\usepackage{titlesec} +\usepackage{caption} +\usepackage{textcomp} +%\usepackage{seqsplit} % not installed; pathsplit falls back to plain texttt +\usepackage{ragged2e} +\usepackage{float} +\titleformat{\section}{\Large\bfseries}{\thesection}{1em}{} +\titleformat{\subsection}{\large\bfseries}{\thesubsection}{1em}{} +\newcommand{\panda}{\textbf{PANDA}} +% Allow LaTeX to add some stretch to badly-set paragraphs +\emergencystretch=3em +% Encourage graphics to fill available width +\setkeys{Gin}{width=\linewidth,keepaspectratio} +% Wrap long \texttt{...} paths so they break at any character. +% Note: callers must escape `_` as `\_` for this to work in text mode. +\usepackage{url} +\urlstyle{tt} +% pathsplit: allow linebreaks at slash/underscore inside typewriter paths +\newcommand{\pathsplit}[1]{\begingroup\Url@setup\ttfamily\hyphenchar\font=`\-\relax\path{#1}\endgroup} +% simpler: just use url's \path which already breaks on / _ . +\renewcommand{\pathsplit}[1]{\path{#1}} + +\title{\panda: A Prototype-Anchored Compact Classifier for Cell-Identity Transfer and Mechanistic Discovery Across Skin, Hematopoietic, and Pancreatic Single-Cell Landscapes} +\author{Bryan Cheng} +\date{\today} + +\begin{document} +\maketitle + +\begin{abstract} +Single-cell RNA-seq cell-identity classification is limited by severe batch, platform, depth, and biology-shift heterogeneity across datasets. We introduce \panda{} (Pan-tissue Adversarial Normalized Domain-invariant Anchored MLP), a compact prototype-anchored classifier trained under a composite of supervised-contrastive, gradient-reversal dataset $+$ depth adversary, HSIC depth-decorrelation, VICReg variance-covariance, and sub-center angular-margin prototype-InfoNCE objectives. Two input variants are supported: PCA-only ($\panda_\text{PCA}$: $\text{PCA}(50) \to$ trunk) and marker-augmented ($\panda_\text{Marker}$: $[\text{PCA}(50) \Vert \mathbf{m}] \to$ trunk). Every cell in every corpus carries a label taken from its source paper's own supplementary tables or public annotation --- no marker-scored fallback labels are used. Under this 100\% paper-labeled principle we built three corpora: \textbf{pan-skin} (6 studies, 45{,}387 cells, 13 classes), \textbf{pan-hematopoietic} (3 studies, 192{,}833 cells, 15 classes), and \textbf{pan-pancreatic} (6 studies, 120{,}611 cells, 20 classes). On stratified 5-fold CV, \panda{}-Marker reaches accuracy $0.948$ / $0.938$ / $0.794$ (skin/HSC/pancreas) with macro AUROC $0.998$ / $0.997$ / $0.974$. Multi-seed replication (seeds 0--2, 3 systems $\times$ 2 variants, all on the canonical corpora with the canonical CV driver) shows Marker beating PCA on 34/45 fold-comparisons, with the advantage entirely concentrated in pancreas (15/15 folds, $+0.04$ mean accuracy every seed), skin slightly positive (11/15), and HSC at parity (8/15; per-seed detail in \S\ref{sec:multiseed}). Transfer to five held-out labeled benchmarks: Baron same-study test half (943 cells, 906 evaluated on the shared vocabulary; PCA acc $0.943$ / F1 $0.603$; Marker $0.819$ / $0.589$), Sulic E14.5 dorsal skin under a 500-cell anchor design with both variants retrained on the anchor corpus (4{,}183 truly held-out cells; PCA acc $0.768$, HF-placode recall $0.980$; Marker acc $0.765$, HF-placode recall $0.829$), Belote melanocyte (6{,}088 cells; PCA acc $0.969$ / F1 $0.395$; Marker $0.955$ / $0.340$; F1 low by class imbalance), Veres held-out (12{,}297 cells; PCA F1 $0.892$; Marker F1 $0.900$), and Nestorowa Smart-seq2 LT-HSC recall on 66 FACS-gated cells (PCA $0.045$; Marker $0.212$). On the unlabeled Dingwall En1-cKO discovery target (25{,}800 cells P2.5 volar hindpaw; GSM6833478--81 WT, GSM6833482--83 cKO), spinous keratinocyte depletion is the largest genotype signal ($\log_2\!\text{fc} = -0.87$, $p_\text{adj} = 5.1\!\times\!10^{-6}$); HF-placode and reticular-fibroblast compartments also shift (melanocyte-lineage calls are flagged as unreliable in this tissue, \S\ref{sec:dingwall-zshot-class}); and Eda\_ectodysplasin derepression is significant across HF-placode ($\Delta = +0.14$, $p_\text{adj} = 3.0\!\times\!10^{-3}$), reticular fibroblast ($p_\text{adj} = 4.0\!\times\!10^{-6}$) and melanoblast ($p_\text{adj} = 9.1\!\times\!10^{-4}$), consistent with En1 loss releasing the hair-follicle (Eda/EDAR) program in volar skin, the molecular counterpart of Dingwall's finding that En1 suppresses hair-follicle fate. Dingwall's EDEN cluster (Derm10) is validated by three complementary lines of evidence: a scope-check post-hoc scoring of pan-skin fibroblast predictions on the S100a4$+$Tnc$+$Pdgfra module returns a null result (no significant WT enrichment), establishing that the pan-tissue prototype is broader than the EDEN core and motivating the two positive lines; an independent scanpy reproduction of Dingwall's Seurat pipeline recovering Derm10 at $4.32\times$ WT enrichment ($p = 2.25\!\times\!10^{-18}$; 206 WT / 30 cKO cells); and \panda{}-Marker trained on those reproduction labels reproducing Derm10 depletion at $5.20\times$ odds-ratio ($p = 5.8\!\times\!10^{-9}$) on held-out predicted cells (test accuracy $81.4\%$ across 10 Derm classes). Sub-clustering the reproduction's 14{,}252 dermal cells and scoring against Dingwall's Data-S1C Derm panels identifies a sub-cluster matching \textbf{Derm2} (7/10 marker overlap: Enpp2, Bmp5, Hpse2, Cacna2d3, Sned1, Rspo2, Slc24a3) at $1{,}586$ cells (1{,}051 WT / 535 cKO) with $1.25\times$ WT enrichment, Fisher $p = 9.84\!\times\!10^{-5}$. Derm2 sits immediately upstream of Derm10 in Dingwall's Fig 5 Slingshot lineage but was not statistically tested for En1-cKO depletion in the paper; our analysis extends the En1-dependence to this precursor. On Dahlin Kit-W41 hematopoiesis, the genotype contrast must be restricted to the matched LK sorting gate --- a within-WT LK-vs-LSK control produces larger apparent shifts ($\log_2\!\text{fc}$ $+4.78$ erythroid) than the mutation itself --- and gate-matched \panda{} reproduces the three compositional results Dahlin report: expanded erythroid ($+0.46$, $p = 1.6\!\times\!10^{-141}$), expanded neutrophil progenitors ($+0.69$, $p = 6.7\!\times\!10^{-48}$), and loss of the mast-cell/basophil branch ($-0.56$, $p = 1.6\!\times\!10^{-7}$), the last of which the pooled analysis missed. On Veres held-out pancreatic differentiation, alpha\_progenitor recovers all four canonical adult-$\alpha$ markers (GCG, ARX, IRX2, MAFB) and dominates the Stage-6 SC-$\beta$ target stage; the adult sub-prototype does not activate; a previously reported gene-level adult-maturity enrichment inside the \texttt{beta} class is withdrawn, as that class consists entirely of primary-islet control cells bundled in the GEO extract rather than differentiation cells. Baron is the one zero-shot target where PCA leads on both accuracy and macro-F1: Marker's sharper adult-vs-juvenile prototype geometry splits rare Baron classes onto two nearby prototypes and pays the macro-F1 cost. Every quantitative claim traces to a specific JSON/CSV artefact. +\end{abstract} + +\vspace{0.5em} +\noindent\textbf{Table 1. Summary of \panda{} results across all systems.} All corpora are 100\% paper-labeled: every training cell carries a label from its source paper's own supplementary tables or public annotation. Held-out CV is stratified 5-fold on the corpus (5 epochs per fold); each acc is mean $\pm$ std over folds. \emph{Held-out labeled test} is a labeled slice evaluated by a single pass of a frozen checkpoint; hold-out strictness varies by target (Nestorowa strict zero-shot; Baron same-study split; Belote and Sulic anchor designs --- Sulic scored on anchor-retrained checkpoints; Veres held-out slice; see \S 5). Discovery targets (Dingwall, Dahlin) are strictly zero-shot. + +\begin{center} +\scriptsize +\setlength{\tabcolsep}{3pt} +\begin{tabular}{p{2.0cm}p{3.3cm}p{3.2cm}p{3.15cm}p{2.75cm}} +\toprule +\textbf{System (corpus)} & \textbf{Held-out 5-fold CV} & \textbf{Held-out labeled test} & \textbf{Discovery target} & \textbf{Reproduced finding ($p$)} \\ +\midrule +Pan-skin (6 studies, 45{,}387 cells, 13 classes) & +$\panda_\text{Marker}$: acc $\mathbf{0.9476{\pm}0.0027}$, F1 $0.9139$, AUROC $\mathbf{0.9977}$ \newline +$\panda_\text{PCA}$: acc $0.9453{\pm}0.0020$, F1 $\mathbf{0.9150}$, AUROC $0.9975$ & +Sulic anchor design, retrained checkpoints (4{,}183 held-out cells): PCA acc $\mathbf{0.768}$ / placode recall $\mathbf{0.980}$; Marker acc $0.765$ / placode recall $0.829$. \newline Belote anchor design (6{,}088 held-out mel): PCA acc $\mathbf{0.969}$ / F1 $\mathbf{0.395}$; Marker acc $0.955$ / F1 $0.340$. & +Dingwall En1-cKO (25{,}800 P2.5 volar cells) & +Spinous $\downarrow$ $p_\text{adj}{=}5.1{\times}10^{-6}$; Derm10 (Secondary EDEN) $5.20\times$ WT-OR, $p{=}5.8{\times}10^{-9}$; \textbf{Derm2 (Primary EDEN candidate) $1.25\times$ WT, $p{=}9.84{\times}10^{-5}$} \\ +\midrule +Pan-hematopoietic (3 studies, 192{,}833 cells, 15 classes) & +$\panda_\text{Marker}$: acc $0.9376{\pm}0.0024$, F1 $0.9317$, AUROC $0.9969$ \newline +$\panda_\text{PCA}$: acc $\mathbf{0.9384{\pm}0.0010}$, F1 $0.9305$, AUROC $0.9967$ & +Nestorowa Smart-seq2 (LT-HSC recall on 66 FACS): PCA $0.045$, Marker $\mathbf{0.212}$. Cross-platform Smart-seq2 LT-HSC gate is broader than the corpus LT-HSC prototype. & +Dahlin Kit-W41 (61{,}122 LSK/Kit$^+$) & +Gate-matched: erythroid $\uparrow$ $p{=}1.6{\times}10^{-141}$; neutrophil prog.\ $\uparrow$ $p{=}6.7{\times}10^{-48}$; \textbf{mast/basophil loss $p{=}1.6{\times}10^{-7}$} (Dahlin's headline) \\ +\midrule +Pan-pancreatic (6 studies, 120{,}611 cells, 20 classes) & +$\panda_\text{Marker}$: acc $\mathbf{0.7942{\pm}0.0025}$, F1 $\mathbf{0.7645}$, AUROC $\mathbf{0.9741}$ \newline +$\panda_\text{PCA}$: acc $0.7522{\pm}0.0112$, F1 $0.7362$, AUROC $0.9688$ & +Baron test-half (906 evaluated): PCA acc $\mathbf{0.943}$ / F1 $\mathbf{0.603}$; Marker acc $0.819$ / F1 $0.589$. \newline \textbf{Veres held-out (12{,}297)}: PCA F1 $0.892$; Marker F1 $\mathbf{0.900}$. & +Veres hPSC $\to$ SC-$\beta$ (44{,}711 train $+$ 12{,}297 held-out) & +Stage-6 SC-$\alpha$-dominant (alpha\_progenitor is $59\%$ of Stage-6 cells and recovers 4/4 canonical adult-$\alpha$ markers); adult-$\beta$ prototype does not activate on Veres; \texttt{epsilon} class is really Veres's SC-EC population (\S\ref{sec:veres-marker-deep-dive}) \\ +\bottomrule +\end{tabular} +\end{center} + +\section{Introduction} + +Single-cell RNA-seq cell-identity classification is complicated by three domain-shift axes: batch and platform (10x v2/v3, Smart-seq2, inDrops, snRNA-seq, microwell-seq); sequencing depth, spanning an order of magnitude between UMI-based droplet and plate-based deep-coverage protocols; and biology --- developmental stage, in vivo vs.\ in vitro, species, and perturbation state. Existing classifiers address a subset of these axes at the cost of the others: alignment methods (Harmony, Seurat integration) remove batch structure but often flatten biological variation and require per-target alignment; direct nearest-neighbour or transfer-learning classifiers require the target to have a specific reference in the training corpus, failing on truly zero-shot novel datasets; marker-based scoring is robust to platform but restricts predictions to a small number of canonical types and lacks a mechanism for flagging novel populations. + +We train a single compact classifier, \panda{}, on 100\%-paper-labeled multi-dataset corpora for skin, hematopoiesis, and pancreas, and evaluate it in three regimes: (i) held-out 5-fold CV on the corpus, (ii) zero-shot transfer to labeled datasets held out from the start, and (iii) mechanistic discovery on unlabeled perturbation targets via within-class differential expression and pathway module scoring. Across three systems \panda{} reaches CV accuracy $0.79$--$0.95$ and macro AUROC $\geq 0.97$; on the discovery targets it reproduces the source papers' central compositional claims once known confounders are controlled, and surfaces one candidate extension of published biology: Derm2, a Primary-EDEN candidate immediately upstream of Dingwall's EDEN in their own inferred lineage, which their paper did not statistically test. We also report, in detail, the confounders and misassignments we found while checking these analyses against the source papers (\S\ref{sec:dahlin}, \S\ref{sec:veres}), since they bound what a prototype classifier can claim on perturbation data. + +\section{Method} + +\subsection{Architecture (\texttt{panda/model.py})} + +\panda{} is a single \texttt{PANDAEncoder} module with the following structure: + +\begin{itemize} +\item \textbf{Trunk}: $\text{PCA}(n) \rightarrow 512 \rightarrow 512 \rightarrow 256$ MLP with LayerNorm $+$ GELU $+$ Dropout(0.2). +\item \textbf{Projection head}: $256 \rightarrow 256 \rightarrow 128$ with GELU intermediate; $L_2$-normalised for SupCon and prototype-InfoNCE. +\item \textbf{Classifier head}: takes $[\text{repr}, \text{missing\_hvg\_frac}, \log_{10}(\text{counts}_z)]$ (representation + two auxiliary scalars) to $n\_\text{classes}$. +\item \textbf{$K$ learnable class prototypes} in the 128-d projection space, exponential-moving-average (EMA) momentum $0.99$. These are the \emph{transfer object} at inference. +\item \textbf{Dataset adversary} $256 \rightarrow 128 \rightarrow n\_\text{datasets}$ behind a \texttt{GradReverse} autograd function (gradient-reversal layer, GRL). +\item \textbf{Depth adversary} $256 \rightarrow 64 \rightarrow 1$ behind the same GradReverse, predicting cell-cycle-adjusted $\log(\text{counts})$. +\end{itemize} + +Two variants share this architecture and differ only in the input to the trunk: $\panda_\text{PCA}$ uses the 50-d PCA vector alone; $\panda_\text{Marker}$ concatenates the PCA vector with a curated per-cell marker-score vector $\mathbf{m}$ (canonical marker gene modules for the tissue, scored via \texttt{sc.tl.score\_genes}). + +\subsection{Losses and training curriculum} + +The composite loss is +\[ +\mathcal{L} \;=\; \mathcal{L}_{\text{SupCon}} + \lambda_{\text{V}}\mathcal{L}_{\text{VICReg}} + \lambda_{\text{CE}}\mathcal{L}_{\text{CE}} + \lambda_{\text{P}}\mathcal{L}_{\text{proto}} + \lambda_{\text{D}}\mathcal{L}_{\text{dom}} + \lambda_{\text{Z}}\mathcal{L}_{\text{depth}} + \lambda_{\text{H}}\mathcal{L}_{\text{HSIC}}, +\] +with $(\lambda_{\text{V}}, \lambda_{\text{CE}}, \lambda_{\text{P}}, \lambda_{\text{D}}, \lambda_{\text{Z}}, \lambda_{\text{H}}) = (1.0, 0.4, 0.6, 1.0, 0.3, 0.05)$. Each term has a distinct role: +\begin{enumerate} +\item $\mathcal{L}_{\text{SupCon}}$ --- class-balanced supervised contrastive loss on 128-d projections: +\[ +-\sum_i \tfrac{1}{|P(i)|}\sum_{p \in P(i)} \log \tfrac{\exp(z_i\!\cdot z_p / \tau)}{\sum_{a \neq i} \exp(z_i\!\cdot z_a / \tau)}; +\] +positives $P(i)$ drawn from $\geq 2$ datasets per class per batch, per-class weight $w_c = 1/\sqrt{n_c}$, $\tau = 0.1$. +\item $\mathcal{L}_{\text{VICReg}}$ --- variance-invariance-covariance regularization on the 128-d projections; keeps each dimension informative and decorrelated. +\item $\mathcal{L}_{\text{CE}}$ --- classifier cross-entropy on the auxiliary head; weight blended $0.5\cdot 1/\sqrt{n_c} + 0.5$ to balance rare classes. +\item $\mathcal{L}_{\text{proto}}$ --- prototype-InfoNCE that pulls each projection toward its class-prototype: $-\log \tfrac{\exp(z_i\!\cdot\tilde p_{y_i}/\tau_p)}{\sum_c \exp(z_i\!\cdot\tilde p_c/\tau_p)}$ with EMA momentum $0.99$, $\tau_p = 0.07$. +\item $\mathcal{L}_{\text{dom}}$ --- cross-entropy against the dataset adversary through a gradient-reversal layer (GRL); trunk sees $-\lambda_{\text{GRL}}\nabla$ so it removes dataset-identifiable structure. +\item $\mathcal{L}_{\text{depth}}$ --- MSE against the depth adversary predicting cell-cycle-adjusted $\log(\text{counts})$, also through GRL. +\item $\mathcal{L}_{\text{HSIC}}$ --- biased Hilbert-Schmidt Independence Criterion (HSIC) between representation and $\log(\text{counts})$; drives depth-representation dependence to zero as a stronger complement to the depth adversary. +\end{enumerate} + +Training curriculum (epoch-gated stages, as implemented in \pathsplit{scripts/common/train\_panda.py}): (i) epoch 0 --- SupCon $+$ VICReg $+$ CE warmup; (ii) epochs 1--2 --- engage $\mathcal{L}_{\text{proto}}$ (sub-center angular-margin InfoNCE) and begin EMA prototype updates; (iii) epochs 3$+$ --- engage the dataset adversary at fixed $\lambda_{\text{GRL}} = 0.1$, the depth-regression adversary, and the HSIC term (the HSIC term is computed at bandwidth $\sigma = 1$ on the 256-d representation, where its numerical contribution is negligible; it is retained for transparency, not as an active regularizer); (iv) final epoch at halved learning rate. Batches are uniform random permutations (AdamW, lr $10^{-3}$, weight decay $10^{-4}$, batch 256). A prototype-repulsion statistic is monitored during training but contributes no gradient (prototypes are EMA buffers, not parameters). + +\subsection{Inference} +\label{sec:inference} + +For a target dataset: +\begin{enumerate} +\item Project raw counts to the frozen corpus shared HVG index (zero-impute missing genes); compute per-cell \texttt{missing\_hvg\_frac}. +\item Log-normalize with corpus per-gene mean/std; clip to $[-10, 10]$. +\item Frozen sample-fit PCA transform to 50-d (concatenate marker channel for the Marker variant). +\item Forward through frozen \panda{} to obtain 128-d projections $z$. +\item Compute cosine similarity of $z$ to the best sub-center of each of the $K$ frozen prototype sets: $\cos_{ic} = \max_s (z_i \cdot \tilde{p}_{c,s})$; the predicted class is $\arg\max_c \cos_{ic}$. +\item Temperature-scale ($\tau = 0.07$) and softmax to obtain relative class scores (reported alongside the raw max-cosine as a confidence measure; the temperature is fixed, not calibrated). +\end{enumerate} +The canonical inference path applies no label-shift correction and no abstention gate; predictions and per-cell max-cosine confidences are written for every cell, and low-confidence populations are examined post-hoc in the discovery analyses. (A prior-reweighting step and a $\cos < 0.3$ abstain gate existed in a legacy single-target script but are not part of the pipeline that produced the numbers in this paper.) + +\section{Corpora (100\% paper-labeled)} +\label{sec:corpus} + +Every training cell carries a label taken from its source paper's own supplementary tables, deposited annotations, or public cell-metadata release. Datasets without paper-provided labels are excluded rather than filled in by our marker scorer. This eliminates one full class of label noise (heuristic marker-score misassignment) and turns the corpus into a benchmark on which zero-shot F1 is directly interpretable. + +Per tissue: (a) download published scRNA-seq via GEO/ArrayExpress; (b) per-dataset load with format-specific loaders $+$ QC (min\_genes $=200$, min\_cells $=3$, mt\% $<20\%$); (c) obtain paper labels from supplementary tables and harmonise to the corpus vocabulary; (d) compute shared highly-variable genes (HVGs) via a union$+$majority rule where any gene ranking in the top-4000 of $\geq n/2$ datasets is eligible, then force-include a curated list of 40--50 canonical markers per system; (e) sample-fit PCA on a 30--50k cell subsample; (f) transform all datasets; (g) hold out truly labeled external datasets (Baron test-half, Nestorowa, Sulic, Belote, Veres held-out slice) for zero-shot evaluation. + +\textbf{Pan-skin (45{,}387 cells, 13 classes)}: Sulic GSE212673 (E14.5 dorsal skin), Merkel GSE201447, MCA GSE108097 neonatal skin \cite{han2018mca}, Joost GSE67602 \cite{joost2016}, Haensel/Annusver GSE142471 \cite{haensel2020skin}, and Belote GSE151091 \cite{belote2021} as the melanocyte anchor. Classes: basal-IFE, spinous, granular, HF-placode, HF-ORS, endothelial, immune, sebaceous, fibroblast-papillary, fibroblast-reticular, melanocyte, melanoblast, melanocyte-precursor. + +\textbf{Pan-hematopoietic (192{,}833 cells, 15 classes)}: Weinreb LARRY GSE140802, Baccin GSE122465 \cite{baccin2020}, and Tabula Muris Senis bone marrow GSE132042 \cite{tms2020}. Classes: long-term hematopoietic stem cell (LT-HSC), multipotent progenitor (MPP), myeloid, monocyte, macrophage, basophil-mast, erythroid, megakaryocyte, T-cell, naive-B, pro-B, lymphoid, fibroblast, stromal, endothelial. + +\textbf{Pan-pancreatic (120{,}611 cells, 20 classes)}: Baron 2016 mouse-train half GSE84133, Bastidas GSE132188 E15.5 \cite{bastidas2019}, Byrnes 2018 GSE101099 paper-labeled subset \cite{byrnes2018}, Yu 2021 GSE139627 paper-labeled subset \cite{yu2021}, MIA GSE211796 paper-labeled subset \cite{hrovatin2023mia}, and Veres 2019 GSE114412 --- of which \textbf{44{,}711 cells joined the training corpus} and \textbf{12{,}297 labeled cells were held out} (the GEO extract totals 57{,}297 cells; the 289-cell residual is dropped by QC, and the extract additionally bundles primary-islet and ES/iPS control samples that are not part of the differentiation; see \S\ref{sec:zeroshot-veres}) as an external zero-shot test target. Classes: alpha, beta, delta, gamma, epsilon, ductal, acinar, endothelial, immune, mesenchyme, alpha\_progenitor, beta\_progenitor, adult-alpha, adult-beta, endocrine-progenitor, endocrine-progenitor-primed, pancreatic-progenitor, proliferating, Fev-EP, exocrine. + +\section{Held-out 5-fold cross-validation} + +\subsection{Pan-skin: 13 classes, 45{,}387 cells} + +Stratified 5-fold cross-validation with \panda{} retrained from scratch on 80\% train per fold (5 epochs per fold), evaluated on the held-out 20\%: +\begin{center} +\begin{tabular}{lccc} +\toprule +Variant & Accuracy & Macro F1 & Macro AUROC \\ +\midrule +\panda{}-PCA & $0.9453 \pm 0.0020$ & $0.9150$ & $0.9975$ \\ +\panda{}-Marker & $\mathbf{0.9476 \pm 0.0027}$ & $0.9139$ & $\mathbf{0.9977}$ \\ +\bottomrule +\end{tabular} +\end{center} +Results at \texttt{discovery/pan\_skin/\{pca,marker\}/cv\_5fold.json}. Per-class F1 across all three systems in Figure~\ref{fig:cv}. + +\begin{figure}[!htb] +\centering +\includegraphics[width=\textwidth]{figures/fig1_perclass_f1.pdf} +\caption[Per-class F1 across three systems]{Per-class F1 on held-out 5-fold CV for the three multi-dataset systems.\\ +\textbf{Pan-skin} (13 classes, mean acc $=0.948$): F1 $\geq 0.88$ on 10/13 classes; granular, melanocyte-precursor, and sebaceous fall below on low support.\\ +\textbf{Pan-hematopoietic} (15 classes, mean acc $=0.938$): F1 $\geq 0.88$ on the major lineages (MPP, LT-HSC, erythroid, myeloid, monocyte, basophil-mast, pro-B, T-cell, naive-B, endothelial); lymphoid, macrophage, and megakaryocyte sit just below.\\ +\textbf{Pan-pancreatic} (20 classes, mean acc $=0.794$): the hardest system --- only the largest identities (adult-beta, alpha\_progenitor, adult-alpha, pancreatic-progenitor, beta\_progenitor, endothelial, exocrine) exceed F1 $=0.88$; fine-grained endocrine sub-classes trade support for granularity.} +\label{fig:cv} +\end{figure} + +\subsection{Pan-hematopoietic: 15 classes, 192{,}833 cells} + +5-fold stratified CV on the 15-class Weinreb + Baccin + TMS corpus: +\begin{center} +\begin{tabular}{lccc} +\toprule +Variant & Accuracy & Macro F1 & Macro AUROC \\ +\midrule +\panda{}-PCA & $\mathbf{0.9384 \pm 0.0010}$ & $0.9305$ & $0.9967$ \\ +\panda{}-Marker & $0.9376 \pm 0.0024$ & $\mathbf{0.9317}$ & $\mathbf{0.9969}$ \\ +\bottomrule +\end{tabular} +\end{center} +PCA and Marker are within one standard deviation on accuracy; Marker wins on macro F1 and AUROC. Results at \texttt{discovery/hematopoiesis/\{pca,marker\}/cv\_5fold.json}. + +\subsection{Pan-pancreatic: 20 classes, 120{,}611 cells} + +44{,}711 Veres cells join the training corpus alongside Baron/Bastidas/Byrnes/Yu/MIA; 12{,}297 labeled Veres cells are held out for \S\ref{sec:zeroshot-veres}. 5-fold stratified CV: +\begin{center} +\begin{tabular}{lccc} +\toprule +Variant & Accuracy & Macro F1 & Macro AUROC \\ +\midrule +\panda{}-PCA & $0.7522 \pm 0.0112$ & $0.7362$ & $0.9688$ \\ +\panda{}-Marker & $\mathbf{0.7942 \pm 0.0025}$ & $\mathbf{0.7645}$ & $\mathbf{0.9741}$ \\ +\bottomrule +\end{tabular} +\end{center} +Marker beats PCA by $+4.2\%$ accuracy and $+2.8\%$ F1 --- the largest marker gain of any system, consistent with pancreatic endocrine sub-lineages differing on low-variance TFs (Nkx6-1, Mnx1, Arx) that benefit most from the direct marker channel. + +\subsection{Multi-seed rigor} +\label{sec:multiseed} + +We repeat every system's 5-fold CV with independent seeds affecting fold assignment (\texttt{StratifiedKFold} \texttt{random\_state}) and model initialisation, all with the canonical driver (\texttt{python -m scripts.common.run\_cv --seed N}) on the canonical corpora. \emph{Correction:} earlier revisions of this section pooled seed files that had been produced on different corpus builds (different cell counts and label vocabularies) by a different CV driver with a different loss curriculum; the ``33/35 folds'' comparison drawn from them was not a seed-variance analysis and is withdrawn. The regenerated replicates (seeds 0--2, identical corpus/driver/curriculum, verified same $n$/$K$ per configuration) give, per system (mean accuracy per seed, PCA vs Marker): + +\begin{center} +\small +\begin{tabular}{lcccc} +\toprule +System & Seed 0 (PCA / Marker) & Seed 1 & Seed 2 & Marker fold wins \\ +\midrule +Pan-skin & $0.9453$ / $0.9476$ & $0.9453$ / $0.9482$ & $0.9457$ / $0.9458$ & 11/15 \\ +Pan-hematopoietic & $0.9384$ / $0.9376$ & $0.9384$ / $0.9404$ & $0.9377$ / $0.9376$ & 8/15 \\ +Pan-pancreatic & $0.7522$ / $0.7942$ & $0.7547$ / $0.7937$ & $0.7471$ / $0.7926$ & \textbf{15/15} \\ +\midrule +Total & & & & \textbf{34/45} \\ +\bottomrule +\end{tabular} +\end{center} + +Across 45 fold-comparisons (3 seeds $\times$ 3 systems $\times$ 5 folds), Marker beats PCA on accuracy in \textbf{34/45 folds}, but the distribution matters more than the total: pancreas Marker wins every fold in every seed ($+0.038$ to $+0.046$ mean accuracy), skin is slightly positive (11/15, mean gap $\leq 0.003$), and HSC is at parity (8/15, mean gap $\leq 0.002$). The marker channel is a system-dependent choice whose benefit is decisive only where closely-related subtypes differ on low-variance TFs --- pancreatic endocrine sub-lineages. {\sloppy Per-configuration artefacts at \pathsplit{discovery/\{system\}/\{variant\}/cv\_5fold\{,\_seed1,\_seed2\}.json}.\par} + +\section{Held-out labeled targets} + +A single pass of a frozen \panda{} checkpoint on labeled held-out cells: no fold-training, no per-target ensembling. The five benchmarks differ in how strictly they are held out, and we label each accordingly (verified by barcode intersection against the training corpora): \textbf{Nestorowa} is strict zero-shot (its source dataset contributed zero training cells). \textbf{Baron} is a same-study split (the stratified train half of GSE84133 is in the corpus; the 943 test cells are not). \textbf{Belote} is an anchor design (a 1{,}000-cell anchor is in the corpus; the 6{,}088 test cells are not). \textbf{Veres} is a held-out slice (44{,}711 cells trained, 12{,}297 held out). \textbf{Sulic} is an anchor design evaluated on checkpoints retrained on an anchor corpus (500 anchor cells in training, 4{,}183 test cells excluded; \S\ref{sec:zeroshot-sulic}) --- the standard corpus contains all Sulic cells, so the standard checkpoints cannot be scored on Sulic as a held-out test. Same-study splits and anchor designs measure within-study generalisation, not cross-laboratory transfer; only Nestorowa and the unlabeled discovery targets (Dingwall, Dahlin) measure the latter. + +\subsection{Baron test-half (pancreas)} +\label{sec:zeroshot-baron} + +The 943-cell mouse test half of Baron 2016 GSE84133 was held out of the pancreatic corpus from the initial build. We score against the paper's canonical \texttt{assigned\_cluster} labels after harmonising PANDA's adult sub-prototypes (\texttt{adult-beta}$\to$\texttt{beta}, \texttt{adult-alpha}$\to$\texttt{alpha}) into Baron's flat vocabulary. Of the 943 cells, \textbf{906 are evaluated on the shared-class vocabulary} (37 cells whose Baron labels do not exist in the pan-pancreatic class list are excluded from accuracy/F1 to avoid off-vocabulary penalties). + +\begin{center} +\begin{tabular}{lcc} +\toprule +Variant & Accuracy & Macro F1 \\ +\midrule +\panda{}-PCA & $\mathbf{0.9426}$ & $\mathbf{0.6027}$ \\ +\panda{}-Marker & $0.8190$ & $0.5888$ \\ +\bottomrule +\end{tabular} +\end{center} + +Baron is the one zero-shot target where PCA leads on both accuracy and macro-F1 ($+12.4$ acc, $+1.4$ F1). Marker's sharper adult-vs-juvenile prototype geometry splits rare Baron classes across two nearby prototypes (adult-$\alpha$/juvenile-$\alpha$; adult-$\beta$/juvenile-$\beta$); on a 906-cell slice with imbalanced supports this geometry costs both accuracy and F1. On the larger Veres held-out (\S\ref{sec:zeroshot-veres}) the same sub-prototype separation is discriminative and Marker wins. + +\subsection{Veres held-out (pancreas)} +\label{sec:zeroshot-veres} + +12{,}297 Veres 2019 GSE114412 cells were held out with paper labels preserved. This is a held-out-slice benchmark (44{,}711 Veres cells are in the training corpus), not a strict zero-shot like Baron/Sulic/Belote/Nestorowa; it is nonetheless the largest labeled held-out pancreatic benchmark in the paper. + +\begin{center} +\begin{tabular}{lcc} +\toprule +Variant & Accuracy & Macro F1 \\ +\midrule +\panda{}-PCA & $0.9065$ & $0.8922$ \\ +\panda{}-Marker & $\mathbf{0.9131}$ & $\mathbf{0.9001}$ \\ +\bottomrule +\end{tabular} +\end{center} + +At 12{,}297-cell scale with 20 canonical Veres classes, Marker beats PCA by $+0.66\%$ accuracy and $+0.79\%$ F1. Macro F1 $= 0.900$ is the strongest labeled held-out pancreatic F1 in this paper. Results at \texttt{discovery/pancreas/\{pca,marker\}/veres\_summary.json}. + +\subsection{Nestorowa Smart-seq2 (hematopoiesis)} +\label{sec:zeroshot-nestorowa} + +Nestorowa 2016 GSE81682, 1{,}170 Smart-seq2 mouse HSPCs with FACS labels (LT-HSC vs.\ HSPC gate). The corpus contains an LT-HSC class populated from Tabula Muris Senis bone-marrow paper labels; no per-target anchor is added. + +\begin{center} +\begin{tabular}{lc} +\toprule +Variant & LT-HSC recall (66 FACS-labeled LT-HSC cells) \\ +\midrule +\panda{}-PCA & $0.045$ \\ +\panda{}-Marker & $\mathbf{0.212}$ \\ +\bottomrule +\end{tabular} +\end{center} + +Both variants struggle: the Smart-seq2 LT-HSC FACS gate is transcriptomically broader than the TMS 10x LT-HSC prototype learned from the corpus. The 0.212 figure is per-class recall on 66 FACS-labeled LT-HSC cells (14/66 correct). Marker's $\sim\!5\times$ improvement over PCA (14/66 vs 3/66; Wilson 95\% CIs $[0.12, 0.32]$ vs $[0.01, 0.13]$) suggests the marker channel partially bridges the cross-platform mismatch by reading canonical LT-HSC genes (Hlf, Mecom, Mpl) that PCA compresses into its 50-component bottleneck. A Smart-seq2 LT-HSC training anchor is the natural next step. + +\subsection{Belote melanocyte (skin)} +\label{sec:zeroshot-belote} + +Belote 2021 GSE151091 \cite{belote2021}: 6{,}088 human melanocyte-lineage cells across mel / melanoblast / melanocyte-precursor sub-classes. A separate 1{,}000-cell Belote anchor is part of the training corpus (anchor design); the 6{,}088 evaluated cells have zero overlap with training (verified by barcode intersection). + +\begin{center} +\begin{tabular}{lcc} +\toprule +Variant & Accuracy & Macro F1 \\ +\midrule +\panda{}-PCA & $\mathbf{0.969}$ & $0.395$ \\ +\panda{}-Marker & $0.955$ & $0.340$ \\ +\bottomrule +\end{tabular} +\end{center} + +High accuracy reflects correct assignment to the dominant \emph{melanocyte} class; low macro F1 is a class-imbalance artefact across the three melanocyte sub-classes. + +\subsection{Sulic (skin)} +\label{sec:zeroshot-sulic} + +Sulic 2023 GSE212673, 4{,}683 E14.5 mouse dorsal skin cells. \emph{Correction:} the standard pan-skin corpus contains all 4{,}683 Sulic cells, so earlier revisions of this section reported what was in fact a training-set evaluation (acc $0.940$/$0.895$, F1 $0.930$/$0.871$); those numbers are withdrawn. The honest evaluation retrains both variants on an anchor corpus (\pathsplit{corpus\_with\_sulic\_anchor.h5ad}: 500 Sulic anchor cells in training, the remaining 4{,}183 excluded from training, HVG/PCA/statistics refit) and scores the 4{,}183 truly held-out cells (\pathsplit{scripts/pan\_skin/92\_retrain\_with\_sulic\_anchor.py}; artefacts \pathsplit{discovery/pan\_skin/\{pca,marker\}/97\_sulic\_anchor\_zero\_shot.json}): + +\begin{center} +\begin{tabular}{lccc} +\toprule +Variant & Accuracy & HF-placode recall & basal-IFE recall \\ +\midrule +\panda{}-PCA & $\mathbf{0.768}$ & $\mathbf{0.980}$ & $0.346$ \\ +\panda{}-Marker & $0.765$ & $0.829$ & $\mathbf{0.637}$ \\ +\bottomrule +\end{tabular} +\end{center} + +The held-out slice is dominated by two classes (HF-placode $n = 2{,}782$, basal-IFE $n = 1{,}401$); macro F1 over the full corpus vocabulary is not informative here (0.27 / 0.19) because most corpus classes have zero support in this slice. Under the anchor design, PCA recalls placode near-perfectly while Marker trades placode recall for basal-IFE recall. + +\section{Discovery target: Dingwall En1-cKO} +\label{sec:dingwall} + +\subsection{Zero-shot classification} +\label{sec:dingwall-zshot-class} + +Held-out target: Dingwall \emph{et al.}\ 2024 \cite{dingwall2024en1cko}, \emph{Developmental Cell} 59(1):20--32.e6 (Kamberov lab), GSE220977: 25{,}800 cells P2.5 volar hindpaw snRNA-seq after per-sample QC. Genotype mapping: WT $=$ \{GSM6833478, 79, 80, 81\} ($n=15{,}400$; 5{,}100 $+$ 3{,}500 $+$ 3{,}300 $+$ 3{,}500), En1-cKO $=$ \{GSM6833482, 83\} ($n=10{,}400$; 5{,}400 $+$ 5{,}000). Baseline cKO fraction on the full 25{,}800-cell target is $0.403$; on the Seurat-replica dermal subset (14{,}252 cells) it is $0.382$. The Dingwall data never touches training, HVG selection, or PCA fit. + +Predicted class distribution (Figure~\ref{fig:umap}; 25{,}800 cells; the pan-skin vocabulary $\{$HF-ORS, HF-placode, basal-IFE, endothelial, fibroblast-papillary, fibroblast-reticular, granular, immune, melanoblast, melanocyte, melanocyte-precursor, sebaceous, spinous$\}$; fibroblast-papillary is in the vocabulary but has zero Dingwall predictions): +\begin{center} +\begin{tabular}{lrr} +\toprule +Class & $n$ & Fraction \\ +\midrule +fibroblast-reticular & 13{,}798 & 53.5\% \\ +melanoblast & 5{,}143 & 19.9\% \\ +endothelial & 3{,}608 & 14.0\% \\ +immune & 1{,}039 & 4.0\% \\ +HF-placode & 546 & 2.1\% \\ +spinous & 342 & 1.3\% \\ +basal-IFE & 335 & 1.3\% \\ +HF-ORS & 328 & 1.3\% \\ +granular & 262 & 1.0\% \\ +melanocyte & 190 & 0.7\% \\ +melanocyte-precursor & 123 & 0.5\% \\ +sebaceous & 86 & 0.3\% \\ +\bottomrule +\end{tabular} +\end{center} + +\begin{figure}[!htb] +\centering +\includegraphics[width=\textwidth]{figures/fig5_dingwall_umap.pdf} +\caption[UMAP of Dingwall projection]{UMAP of \panda{}'s 128-d projection of Dingwall ($n=25{,}800$ cells).\\ +\textbf{Left:} cells coloured by predicted class.\\ +\textbf{Right:} cells coloured by En1 genotype (WT vs.\ En1-cKO).\\ +\textbf{Observation:} clusters correspond to distinct populations at expected proportions; WT and cKO cells share cluster occupancy but differ in local density (melanocyte, HF-placode, spinous).} +\label{fig:umap} +\end{figure} + +\textbf{Caveat on the melanoblast call.} \panda{} assigns $19.9\%$ of this dataset to \texttt{melanoblast} ($5{,}143$ cells). That is implausible for P2.5 volar hindpaw: palmoplantar skin is characteristically hypopigmented, with roughly five-fold lower melanocyte density than non-glabrous skin, and Dingwall --- clustering $45{,}370$ nuclei of the same tissue --- report no melanocyte-lineage cluster at all (their breakdown is $\sim\!20\%$ epidermal, $\sim\!57\%$ dermal, $\sim\!23\%$ endothelial/muscle/immune). Our epidermal calls sum to only $\sim\!7\%$ against their $20\%$, so the \texttt{melanoblast} class is most likely absorbing epidermal keratinocytes. Consistent with this, \emph{En1} itself appears among the genes up in cKO within this class, which should not happen in cells subject to the knockout. Every melanoblast-derived statement below (composition, Eda, MITF axis) should be treated as provisional pending re-annotation. + +Wilcoxon DE per predicted class on raw Dingwall expression recovers canonical marker sets without those genes being supplied to the model: fibroblast-reticular (5/6 canonical: Dcn, Fbn1, Postn, Lum, Col1a2); endothelial (5/7: Kdr, Tie1, Cdh5, Tek, Pecam1); basal-IFE (3/6: Krt14, Trp63, Krt5); immune (Ptprc, Adgre1 $+$ Mrc1, F13a1 macrophage). + +\subsection{Class-level enrichment: En1-cKO vs.\ WT} +\label{sec:dingwall-fisher} + +Fisher exact tests of PANDA class calls vs.\ the $\sim 40:60$ cKO:WT background; $\log_2$ fold-change is the class cKO/WT odds relative to background odds; $p_\text{adj}$ is Benjamini--Hochberg over 12 classes: +\begin{center} +\small +\begin{tabular}{lrrrl} +\toprule +Class & $n$ & $\log_2$ fold cKO/WT & Fisher $p_\text{adj}$ & Direction \\ +\midrule +spinous & 342 & $-0.87$ & $\mathbf{5.1 \times 10^{-6}}$ & $\downarrow$ cKO \\ +melanocyte & 190 & $+0.79$ & $\mathbf{1.2 \times 10^{-3}}$ & $\uparrow$ cKO ($\sim 1.7\times$) \\ +HF-placode & 546 & $+0.43$ & $\mathbf{2.7 \times 10^{-3}}$ & $\uparrow$ cKO ($\sim 1.3\times$) \\ +fibroblast-reticular & 13{,}798 & $-0.09$ & $\mathbf{3.9 \times 10^{-2}}$ & $\downarrow$ cKO \\ +granular & 262 & $-0.40$ & $0.088$ & (ns) \\ +basal-IFE & 335 & $+0.30$ & $0.128$ & (ns) \\ +endothelial & 3{,}608 & $+0.09$ & $0.152$ & (ns) \\ +HF-ORS & 328 & $-0.27$ & $0.169$ & (ns) \\ +melanoblast & 5{,}143 & $+0.05$ & $0.328$ & (ns) \\ +immune & 1{,}039 & $+0.10$ & $0.328$ & (ns) \\ +melanocyte-precursor & 123 & $-0.23$ & $0.446$ & (ns) \\ +sebaceous & 86 & $-0.12$ & $0.742$ & (ns) \\ +\bottomrule +\end{tabular} +\end{center} + +Spinous keratinocyte depletion is the strongest compositional signal ($\log_2\!\text{fc} = -0.87$, $p_\text{adj} = 5.1 \times 10^{-6}$), consistent with En1 loss impairing spinous-layer terminal differentiation. Melanocyte and HF-placode compartments are modestly enriched in cKO ($1.7\times$ and $1.3\times$ respectively), and reticular fibroblast is modestly depleted. + +\subsection{Multi-class pathway analysis} +\label{sec:dingwall-pathway} + +\texttt{sc.tl.score\_genes} on an expanded 15$+$ canonical pathway module set, Mann--Whitney U tested En1-cKO vs.\ WT within each PANDA-predicted class. Top hits at $p_\text{adj} < 0.01$: +\begin{center} +\small +\setlength{\tabcolsep}{4pt} +\begin{tabular}{p{2.9cm}p{3.1cm}p{1.1cm}p{2.3cm}p{4.7cm}} +\toprule +Class & Pathway & $\Delta$ & MannU $p_\text{adj}$ & Interpretation \\ +\midrule +HF-placode & Eda\_ectodysplasin & $+0.142$ & $\mathbf{3.0 \times 10^{-3}}$ & Ectopic placode-signal reactivation \\ +fibroblast-reticular & Eda\_ectodysplasin & $+0.029$ & $\mathbf{4.0 \times 10^{-6}}$ & Dermis reactivates placode signal \\ +melanoblast & Sweat\_gland & $+0.022$ & $\mathbf{3.7 \times 10^{-6}}$ & \emph{Withdrawn} --- module contains \emph{En1} (see caveat) \\ +melanoblast & Eda\_ectodysplasin & $+0.044$ & $\mathbf{9.1 \times 10^{-4}}$ & Ectopic Eda in melanoblast \\ +melanoblast & BMP\_signaling & $+0.028$ & $\mathbf{1.1 \times 10^{-4}}$ & BMP derepression in melanoblast \\ +fibroblast-reticular & Basal\_keratinocyte & $+0.034$ & $\mathbf{6.2 \times 10^{-6}}$ & Dermal cells acquire basal-keratinocyte signal \\ +\bottomrule +\end{tabular} +\end{center} + +\textbf{Interpretation, with two caveats.} Eda\_ectodysplasin derepression is widespread across HF-placode, fibroblast-reticular, and melanoblast predictions. Eda/EDAR is the canonical \emph{hair}-placode pathway, so the most parsimonious reading is that En1 loss releases the hair-follicle program in volar skin --- which is Dingwall's established result (En1 promotes eccrine fate while suppressing hair-follicle fate; in cKO, hair follicles populate the interfootpad space). Our contribution here is the molecular identity of the derepressed program and its spatial breadth, not the repression concept itself. \emph{Caveat 1:} the \texttt{Sweat\_gland} module in our panel contains \emph{En1} itself --- the gene deleted in this experiment --- along with \emph{Foxi3} (shared with our \texttt{Hair\_placode} module) and the generic simple-epithelial keratins \emph{Krt8}/\emph{Krt18}/\emph{Krt19}, so it cannot cleanly measure eccrine identity in an En1 knockout; the earlier claim of ``ectopic sweat-gland program'' in the melanoblast lineage is withdrawn. It would in any case run opposite to Dingwall's central finding that the eccrine program is \emph{lost} in cKO (their nascent-eccrine epidermal cluster falls from $17.71\%$ to $1.64\%$). \emph{Caveat 2:} see the melanoblast-composition caveat in \S\ref{sec:dingwall-zshot-class} before reading any melanoblast row. {\sloppy All module $\times$ class contrasts at \pathsplit{discovery/pan\_skin/marker/57\_pathway\_class\_by\_module\_padj.tsv} and \pathsplit{57\_pathway\_class\_by\_module\_delta.tsv}.\par} + +\subsection{Dingwall EDEN validation: three complementary lines of evidence} +\label{sec:dingwall-eden} + +Dingwall's paper identifies EDEN, the ``\emph{En1}-dependent eccrine niche'' (their original cluster 20; Derm10 in the dermal sub-clustering), a transcriptionally distinct, non-condensed dermal population marked by S100a4/Tnc that is dramatically cKO-depleted (1.99\% of control dermal nuclei vs 0.08\% in cKO) and is required for eccrine gland development --- ablating S100a4$^+$ dermal cells reduces gland density by 45\%. Because our pan-skin corpus does not include an EDEN class prototype directly, we validate the EDEN phenotype through two positive lines of evidence (B, C) and one negative scope check (A) that motivates the other two. + +\textbf{Line A --- Scope check: post-hoc scoring of the pan-tissue fibroblast prototype does \emph{not} recover EDEN.} PANDA's pan-skin fibroblast predictions on Dingwall (13{,}798 cells) were scored on Dingwall's minimal Data-S1C EDEN module (S100a4 $+$ Tnc $+$ Pdgfra). The top-5\% (690 cells) show no significant WT enrichment (WT fraction $= 0.612$ vs.\ dermal-fibroblast baseline $0.604$). This null result establishes that the pan-tissue fibroblast prototype is calibrated too broadly to resolve EDEN by itself and motivates Lines B and C rather than validating the phenotype. Artefact: \pathsplit{discovery/pan\_skin/marker/98\_eden\_summary.json}. + +\textbf{Line B --- Independent scanpy reproduction of Dingwall's clustering.} We reproduced Dingwall's Seurat pipeline in scanpy (LogNormalize $\to$ HVG $=2000$ $\to$ PCA $=40$ $\to$ Harmony per-sample $\to$ Leiden res $=0.7 \to$ 23 top-level clusters; dermal subset re-Leidened into 12 Derm subclusters mapped to Dingwall's Data-S1C Derm0--Derm11 panels by top-50-marker Jaccard). Derm10 recovers at \textbf{206 WT / 30 cKO dermal cells} $=$ $\mathbf{4.32\times}$ WT enrichment, Fisher $p = 2.25\!\times\!10^{-18}$; direction matches Dingwall's paper (their effect is on the 45{,}370-cell superset, ours on the 25{,}800-cell QC subset). Dermal baseline in our replica: $8{,}806$ WT / $5{,}446$ cKO. Artefact: \pathsplit{data/processed/dingwall\_replica/replica\_cluster\_20\_qc.json}. + +\textbf{Line C --- PANDA trained directly on the reproduction's Derm labels.} \panda{}-Marker trained on the reproduction's Derm0--Derm11 labels with a 70/30 genotype-stratified split ($9{,}965$ train / $4{,}287$ test) reaches $\mathbf{81.4\%}$ test accuracy across 10 Derm classes. On held-out cells, predicted Derm10 shows $\mathbf{5.20\times}$ WT-enrichment odds-ratio (Fisher $p = 5.8\!\times\!10^{-9}$); scored against true reproduction labels on the same cells, Derm10 shows $4.34\times$ WT enrichment ($p = 2.9\!\times\!10^{-6}$). The EDEN depletion phenotype is reproduced at inference time from cells the base pan-skin corpus never saw as EDEN. Artefact: \pathsplit{discovery/pan\_skin/marker/104\_dingwall\_derm\_summary.json}. + +Lines B and C together show that Derm10's En1-cKO depletion is a real, learnable, reproducible property; Line A shows the base pan-tissue fibroblast prototype is too broad to resolve it on its own. + +\subsection{Primary EDEN candidate: Derm2} +\label{sec:dingwall-primary-eden} + +We next asked whether \panda{} could extend Dingwall's En1-dependence phenotype up the paper's own developmental lineage. Dingwall's Fig 5 Slingshot pseudotime places Derm10 (Secondary EDEN, cKO-depleted, statistically tested) downstream of a lineage running Derm3 $\to$ Derm6 $\to$ Derm9 $\to$ Derm2 $\to$ Derm10, but the upstream clusters were not statistically tested for cKO depletion in the paper. + +Grouping the reproduction's 14{,}252 dermal cells by their replica-assigned Data-S1C Derm identities (top-50-marker Jaccard mapping) recovers the Derm2 compartment; an independent Leiden sub-clustering (resolution $=1.5$, 21 subclusters) of the PANDA-predicted dermal compartment yields a cluster whose top Wilcoxon markers match \textbf{Derm2 at 7/10 top-marker overlap} (Enpp2, Bmp5, Hpse2, Cacna2d3, Sned1, Rspo2, Slc24a3). The replica-assigned Derm2 compartment contains \textbf{1{,}586 cells (1{,}051 WT / 535 cKO)}, a cKO fraction of $0.337$ vs.\ the dermal baseline $0.382$ --- \textbf{$1.25\times$ WT enrichment, Fisher $p = 9.84 \times 10^{-5}$}. + +\begin{center} +\small +\setlength{\tabcolsep}{4pt} +\begin{tabular}{p{2.0cm}p{3.0cm}p{4.9cm}p{2.3cm}p{2.3cm}} +\toprule +Leiden sub-cluster & Dingwall panel & Marker overlap & Fisher $p$ & Direction \\ +\midrule +Derm2 match & \textbf{Derm2} (Primary EDEN candidate) & \textbf{7/10} (Enpp2, Bmp5, Hpse2, Cacna2d3, Sned1, Rspo2, Slc24a3) & $\mathbf{9.84 \times 10^{-5}}$ & $\downarrow$ cKO ($1.25\times$) \\ +\bottomrule +\end{tabular} +\end{center} + +\textbf{Novel biology.} Derm2 sits immediately upstream of Derm10 in Dingwall's Fig 5 Slingshot pseudotime, and Dingwall's CellChat analysis (Data S2) shows Derm2 sends Bmp5 and Rspo2 signals to epidermal placodes (Epi0/3/5), consistent with a signalling-competent Primary-EDEN precursor. Dingwall hypothesises this direction via Slingshot but does not statistically test Derm2 for En1-cKO depletion; our analysis extends En1-dependence to Derm2 at a modest but well-powered effect size ($1.25\times$, Fisher $p = 9.84\!\times\!10^{-5}$ at $n=1{,}586$). {\sloppy Artefacts: \pathsplit{discovery/pan\_skin/marker/105\_primary\_eden\_full\_dermal.json} (reported Derm2 cell counts and Fisher statistics), \pathsplit{100\_primary\_eden\_discovery.csv}, \pathsplit{100\_primary\_eden\_summary.json}, \pathsplit{101\_derm\_identity\_summary.json}, \pathsplit{101\_derm\_subcluster\_scores.csv}.\par} + +\subsection{Melanoblast Eda-derepression: MITF axis vs neural-crest} +\label{sec:dingwall-melanoblast-mitf} + +The melanoblast Eda-derepression phenotype is MITF-axis-driven rather than a neural-crest reversion. The pan-skin melanoblast prediction on Dingwall totals 5{,}143 cells (2{,}108 cKO / 3{,}035 WT). At baseline, cKO cells in this class show an Eda\_ectodysplasin shift relative to WT ($\Delta = +0.044$, $p = 2.4 \times 10^{-6}$). (The companion Sweat\_gland shift reported in earlier revisions is withdrawn: that module contains \emph{En1} itself, \S\ref{sec:dingwall-pathway}. The class-assignment caveat in \S\ref{sec:dingwall-zshot-class} applies to this entire subsection.) This raises a mechanistic question: are the derepressing cKO melanoblasts \emph{reverting toward a neural-crest / bipotent progenitor state}, or is the melanocyte identity being \emph{dismantled at the MITF master-regulator axis} while the cell keeps its melanoblast identity? + +To distinguish, we split cKO melanoblasts into the top and bottom quartile of within-cKO Eda-derepression score ($n = 527$ each) and computed the delta for three orthogonal modules: Neural\_crest, Melanogenesis\_late, and MITF\_regulon. The top-derepression quartile shows Neural\_crest $\Delta = +0.016$ ($p = 0.50$, ns), Melanogenesis\_late $\Delta = -0.017$ ($p = 0.47$, ns), and \textbf{MITF\_regulon $\mathbf{\Delta = -0.052}$} \textbf{($\mathbf{p = 1.2 \times 10^{-4}}$)} --- a significant but small decrement of the MITF regulon inside the derepressing subset, with no compensatory gain of neural-crest identity. + +\textbf{Novel biology.} cKO cells in this class that derepress the Eda program show a small but significant MITF-regulon decrement with no compensatory neural-crest gain, consistent with the melanoblast identity being partially dismantled at the MITF axis rather than reverting to a bipotent neural-crest state. {\sloppy Artefacts: \pathsplit{discovery/pan\_skin/marker/106\_melanoblast\_nc\_summary.json}, \pathsplit{106\_melanoblast\_nc\_scores.csv}; script \pathsplit{scripts/analysis/106\_melanoblast\_neural\_crest.py}.\par} + +\subsection{Per-class DEGs under En1-cKO} +\label{sec:dingwall-hf-placode-degs} + +HF-placode is the most transcriptionally perturbed pan-skin class under En1-cKO. Per-PANDA-class Wilcoxon DE (cKO vs.\ WT) across all 12 non-trivial predicted skin classes on Dingwall, thresholded at $|\text{LFC}| > 1$ and $p_\text{adj} < 0.05$, ranks the classes by count of significantly differentially expressed genes: + +\begin{center} +\small +\setlength{\tabcolsep}{4pt} +\begin{tabular}{p{4.6cm}p{2.2cm}p{8.0cm}} +\toprule +Predicted class & DEGs & Top genes \\ +\midrule +\textbf{HF-placode} & \textbf{12} (all up) & \emph{complete set:} 2610027F03Rik, Mybpc1, Ccdc3, Kcnh7, Cntn5, Bmpr1b, Tnc, Esr1, Pcdh9, Meis2, Grip1, Runx1t1 \\ +melanoblast & 8 (7 up, 1 down) & Mybpc1, Bmpr1b, \emph{En1}, Kcnh7, Ttn, Esr1 \\ +fibroblast-reticular & 4 (all up) & Mybpc1, Ttn, Krt5 \\ +spinous & 3 (2 up, 1 down) & Slc1a3; \emph{Acer3} down \\ +HF-ORS & 2 & Mybpc1 \\ +basal-IFE & 2 & Meis2 \\ +melanocyte & 2 & Bmpr1b \\ +endothelial, immune, granular, mel-precursor, sebaceous & $\leq 1$ & --- \\ +\bottomrule +\end{tabular} +\end{center} + +\textbf{Novel biology.} The most transcriptionally perturbed pan-skin class under \emph{En1}-cKO is \textbf{HF-placode}, not the sweat-gland-fated compartments the canonical eccrine-specification story would predict. All 12 HF-placode DEGs are up in cKO (consistent with derepression) and include Bmpr1b (placode-induction BMP receptor), Tnc (placode ECM), the transcriptional regulators Meis2 and Runx1t1, the nuclear-receptor/scaffold genes Esr1 and Grip1, and a set of neural-type adhesion and channel genes (Cntn5, Pcdh9, Kcnh7). The largest molecular footprint of \emph{En1} loss lands on placode-fated ectoderm rather than the sweat-gland-committed lineage, consistent with Dingwall's result that hair-follicle placodes populate the cKO interfootpad space --- i.e.\ En1 loss acts at least as strongly on appendage \emph{induction} as on commitment. {\sloppy Artefacts: \pathsplit{discovery/pan\_skin/marker/107\_dingwall\_class\_deg\_count.csv} and \pathsplit{107\_dingwall\_class\_deg\_count.json}; script \pathsplit{scripts/analysis/107\_dingwall\_class\_deg\_count.py}.\par} + +\subsection{Synthesis and testable predictions} + +Taken together, our data are consistent with Dingwall's model in which En1 suppresses the hair-follicle program in volar skin, and add the molecular identity and spatial breadth of the derepressed signal: Eda derepression is spatially distributed across HF-placode, dermis, and melanoblast compartments (\S\ref{sec:dingwall-pathway}); the same class additionally shows a BMP shift; reticular dermis acquires a Basal\_keratinocyte signal; class-composition shifts (\S\ref{sec:dingwall-fisher}) show strong spinous depletion with placode and melanocyte enrichment; and the Primary-EDEN candidate Derm2 (\S\ref{sec:dingwall-primary-eden}) extends the phenotype upstream in the paper's Fig 5 lineage. This model yields the following wet-lab predictions: + +\begin{enumerate} +\item \textbf{En1 as spatial repressor}: ISH for Foxi3/Muc5b/Krt8 will show broad-weak expression across cKO volar epidermis vs.\ discrete-strong expression at WT placode sites. +\item \textbf{Dermal Eda upregulation}: ISH on sorted cKO reticular fibroblasts will show elevated \emph{Eda}. +\item \textbf{Melanocyte-lineage re-annotation (prerequisite)}: before any melanoblast-derived claim is tested, Dct/Mitf/Sox10 immunostaining on P2.5 volar skin should establish what fraction of this compartment is genuinely melanocytic (\S\ref{sec:dingwall-zshot-class}). +\item \textbf{Spinous-layer differentiation defect}: Krt10 IHC on cKO volar skin will show reduced spinous-layer thickness relative to WT. +\item \textbf{Derm2 depletion}: FISH combining Enpp2 $+$ Bmp5 $+$ Rspo2 co-expression on P2.5 dermal sections will show reduced signal in cKO relative to WT. +\end{enumerate} + +\section{Discovery target: Dahlin Kit-mutant hematopoiesis} +\label{sec:dahlin} + +\subsection{Zero-shot classification recovers Dahlin's compositional shifts} + +Held-out target: Dahlin \emph{et al.}\ 2018 \cite{dahlin2018kit} (Wilson lab), GSE107727. After per-sample QC (min\_counts $>1000$, min\_genes $>500$, mt\_frac $<0.10$; ENSMUSG$\to$symbol retains 27{,}044/27{,}998 features), 61{,}122 LSK/Kit$^+$ HSPCs recovered: WT $n=46{,}447$, c-Kit W41/W41 $n=14{,}675$. Evaluated on Dahlin with no cell-level overlap; 80.4\% HVG overlap after symbol conversion. + +\textbf{Sorting-gate control (essential).} Dahlin sequenced two FACS gates: LSK (Lin$^-$Sca-1$^+$Kit$^+$, enriched for immature progenitors) and the broader LK gate. Per GEO \texttt{source\_name}, SIGAB1/C1/D1 are WT \emph{LSK}; SIGAF1/G1/H1 are WT \emph{LK}; both mutant samples (SIGAG8/H8) are \emph{LK}. The two gates give radically different compositions in WT alone (WT-LSK is $83.8\%$ MPP / $1.3\%$ erythroid; WT-LK is $36.2\%$ / $34.9\%$), so pooling all six WT samples against LK-only mutants makes the gate, not the genotype, the dominant contrast. A within-WT negative control (LK vs LSK, \emph{genotype held constant}) yields $\log_2\!\text{fc}$ $+4.78$ erythroid, $+4.21$ myeloid, $-1.21$ MPP --- \emph{larger} than any genotype effect below. Earlier revisions of this section reported the pooled contrast and are withdrawn. We therefore restrict the genotype comparison to the matched LK gate, as Dahlin do. + +Fisher exact class enrichment, gate-matched (Kit-W41 LK, $n=14{,}675$ vs.\ WT LK, $n=23{,}101$): +\begin{center} +\begin{tabular}{lrrrrl} +\toprule +Class & Kit\_W41 \% & WT \% & $\log_2$ fc & Fisher $p$ & Dahlin's report \\ +\midrule +erythroid & 48.0 & 34.9 & $+0.46$ & $1.6 \times 10^{-141}$ & Expanded erythroid $\checkmark$ \\ +myeloid & 11.6 & 7.2 & $+0.69$ & $6.7 \times 10^{-48}$ & Expanded neutrophil prog.\ $\checkmark$ \\ +MPP & 26.1 & 36.2 & $-0.48$ & $3.8 \times 10^{-96}$ & Reduced immature pool $\checkmark$ \\ +megakaryocyte & 10.8 & 17.2 & $-0.67$ & $2.3 \times 10^{-67}$ & --- (not reported) \\ +\textbf{basophil-mast} & \textbf{1.67} & \textbf{2.46} & $\mathbf{-0.56}$ & $\mathbf{1.6 \times 10^{-7}}$ & \textbf{Mast/basophil loss} $\checkmark$ \\ +lymphoid & 0.38 & 0.60 & $-0.67$ & $2.4 \times 10^{-3}$ & --- (not reported) \\ +LT-HSC & 0.17 & 0.14 & $+0.31$ & $0.50$ (ns) & --- \\ +\bottomrule +\end{tabular} +\end{center} + +Gate-matched, \panda{} recovers the three compositional shifts Dahlin report --- expanded erythroid, expanded neutrophil progenitors, and, most importantly, \textbf{loss of the mast-cell/basophil branch} ($-0.56$, $p = 1.6\times10^{-7}$), which the paper emphasises as agreeing with the severe mast-cell deficiency of W41/W41 mice. The pooled analysis \emph{missed} this signal, so gate-matching both shrinks the spurious effects and recovers the real one. Megakaryocyte and lymphoid depletion are additional to Dahlin's reported set and are offered as observations, not reproductions. Artefacts: \pathsplit{discovery/hematopoiesis/marker/66\_dahlin\_enrichment.csv} (gate-matched) and \pathsplit{66\_dahlin\_gate\_negative\_control.csv} (within-WT gate control). + +\textbf{Design limitations.} Two caveats bound every Kit-W41 contrast in this section. (i) Replication is $n=2$ mutant animals vs.\ $n=3$ gate-matched WT samples; all $p$-values below are computed cell-wise over tens of thousands of cells and therefore reflect pseudoreplication --- they index effect consistency within this dataset, not between-animal significance. (ii) Sequencing depth is confounded with genotype (median counts/cell: mutant $\approx 15{,}100$, WT $6{,}100$--$8{,}600$), and depth normalisation does not remove detection bias. Effect sizes, not $p$-values, should be read as the evidence. + +\subsection{Within-class pathway module analysis} +\label{sec:dahlin-pathway} + +Pathway module analysis extends the panel to 30 modules per system (canonical marker programs + LT\_HSC\_quiescence, Erythroid\_dev, Granulopoiesis, Lymphopoiesis\_B/T, OXPHOS, Glycolysis, Redox\_glutathione, Ribosomal, Autophagy, Apoptosis\_pro/anti, ISR, Kit\_signaling, and more). Top Kit-W41 vs.\ WT within-class hits on Dahlin: + +\begin{center} +\small +\setlength{\tabcolsep}{4pt} +\begin{tabular}{p{2.3cm}p{3.45cm}p{1.5cm}p{2.2cm}p{5.15cm}} +\toprule +Class & Module & $\Delta$ (Kit $-$ WT) & MannU $p_\text{adj}$ & Interpretation \\ +\midrule +\textbf{MPP} & \textbf{Redox\_glutathione} & $\mathbf{+0.059}$ & $\mathbf{1.2 \times 10^{-127}}$ & Redox stress lead finding \\ +MPP & Apoptosis\_pro & $+0.067$ & $\mathbf{5.2 \times 10^{-125}}$ & Compensatory pro-apoptosis \\ +megakaryocyte & LT\_HSC\_quiescence & $-0.085$ & $\mathbf{3.2 \times 10^{-104}}$ & Loss of quiescence signature \\ +erythroid & Glycolysis & $+0.051$ & $\mathbf{4.0 \times 10^{-95}}$ & Metabolic reprogramming \\ +MPP & Integrated\_stress & $+0.021$ & $\mathbf{1.4 \times 10^{-28}}$ & ISR upregulation \\ +MPP & Kit\_signaling & $-0.271$ & $\approx 0$ & Positive control (expected from Kit-W41 mutation) \\ +erythroid & Apoptosis\_pro & $-0.017$ & $\mathbf{2.3 \times 10^{-13}}$ & Erythroid apoptosis suppression (Dahlin's central molecular claim) \\ +\bottomrule +\end{tabular} +\end{center} + +\textbf{Interpretation, and what is \emph{not} novel here.} Earlier revisions framed the module hits below as novel Kit-W41 phenotypes. That framing is withdrawn: each is either already reported by Dahlin or already established in the Kit/HSC literature, and each is additionally vulnerable to the gate confound described above. +\begin{itemize}\itemsep2pt +\item \textbf{Integrated stress response.} Dahlin report this themselves --- ``system-wide induction of the transcription factor \emph{Atf4}, which is involved in the integrated stress response''. Our module score reproduces their finding; it does not extend it, and \emph{Atf4} alone is the cleaner readout. +\item \textbf{Redox/glutathione.} SCF/Kit signalling limiting ROS via the glutathione system in Kit-expressing progenitors is established \cite{lennartsson2012kit,ludin2014ros}. Our score is also not MPP-specific (it rises in every class) and is not the largest effect in our own table. +\item \textbf{Erythroid glycolysis and LT-HSC OXPHOS.} Glycolysis/OXPHOS rebalancing across erythroid maturation, and the quiescence$\to$activation OXPHOS switch in HSCs, are textbook \cite{filippi2022metabolism}. Because W41 shifts the erythroid compartment toward more mature cells, a glycolysis shift is expected from composition alone. The LT-HSC row rests on 25 mutant cells and is not statistically supported. +\item \textbf{Megakaryocyte quiescence loss.} The same module falls in every class, and more steeply in LT-HSC and MPP than in megakaryocytes, so it is a global rather than lineage-specific effect. The module also contains \emph{Mpl} (a megakaryocyte gene) and \emph{Egr1} (a dissociation-stress immediate-early gene), so it is not a clean quiescence readout. +\item \textbf{Kit\_signaling collapse} is a sanity check only, not a pathway-activity readout: the module is dominated by the \emph{Kit} transcript itself, and W41 is a kinase-impairing missense allele. +\end{itemize} +Dahlin's erythroid pro-survival signature (\emph{Casp3}/\emph{Casp6}/\emph{Bid} down in late erythroid clusters) is reproduced in direction ($\Delta = -0.017$, $p_\text{adj} = 2.3 \times 10^{-13}$). Because the module contrast was not recomputed gate-matched, its magnitude should be read as provisional pending the same LK-vs-LK restriction applied to the composition analysis. Note also that Dahlin's own central claim is the absence of the mast-cell branch and system-wide pleiotropy, not erythroid apoptosis. + +Artefacts: \pathsplit{discovery/hematopoiesis/marker/57\_pathway\_class\_by\_module\_padj.tsv} and \pathsplit{57\_pathway\_class\_by\_module\_delta.tsv}. The Kit\_signaling module collapse is visible in every predicted class (Figure~\ref{fig:dahlin}). + +\begin{figure}[!htb] +\centering +\includegraphics[width=\textwidth]{figures/fig3_dahlin_heatmap.pdf} +\caption[Dahlin Kit-W41 vs.\ WT pathway heat-map]{Dahlin Kit-W41 vs.\ WT within-class pathway module scores. $\Delta = $ Kit\_W41 mean $-$ WT mean of per-cell module scores.\\ +\textbf{Kit\_signaling:} collapses across every class ($\Delta = -0.27$ in MPP, $p_\text{adj} \approx 0$; expected positive control given the c-Kit W41 loss-of-function mutation).\\ +\textbf{Erythroid Apoptosis\_pro $\downarrow$:} Dahlin's central claim ($p_\text{adj} = 2.3 \times 10^{-13}$; opposite direction from MPP $\Delta = +0.067$).\\ +\textbf{MPP Redox\_glutathione $\uparrow$:} $p_\text{adj} = 1.2 \times 10^{-127}$.\\ +\textbf{MK LT\_HSC\_quiescence $\downarrow$:} $p_\text{adj} = 3.2 \times 10^{-104}$.\\ +\textbf{ISR upregulation:} $p_\text{adj} = 1.4 \times 10^{-28}$ in MPP.} +\label{fig:dahlin} +\end{figure} + +\subsection{Dahlin marker deep-dive per predicted class} +\label{sec:dahlin-markers} + +Full-class Wilcoxon per PANDA-predicted class on raw Dahlin counts (\pathsplit{discovery/hematopoiesis/marker/92\_dahlin\_marker\_deep\_dive.csv}). Distribution of the 10 largest predicted classes (13 classes receive predictions; T-cell $n{=}14$, endothelial $n{=}6$, naive-B $n{=}5$ omitted): + +\begin{center} +\scriptsize +\setlength{\tabcolsep}{3pt} +\begin{tabular}{p{2.1cm}rrrrp{3.1cm}p{2.5cm}} +\toprule +Class & $n$ & WT & Kit\_W41 & frac\_WT & Top markers & Canonical recovery \\ +\midrule +MPP & 31{,}754 & 27{,}931 & 3{,}823 & 0.880 & Cd34, Adgrl4, Sox4, Pim1 & 1/3 Kit-sig (Sox4) \\ +erythroid & 15{,}397 & 8{,}352 & 7{,}045 & 0.542 & Ca1, Blvrb, Klf1, Aqp1 & 2/7 (Klf1, Blvrb) \\ +megakaryocyte & 8{,}180 & 6{,}596 & 1{,}584 & 0.806 & Itga2b, Pbx1, Apoe, Gata2 & 2/4 (Itga2b, Pf4-adj) \\ +myeloid & 3{,}463 & 1{,}755 & 1{,}708 & 0.507 & Elane, Mpo, Prtn3, Ctsg & 4/8 (Elane, Mpo, Prtn3, Ctsg) \\ +basophil-mast & 848 & 603 & 245 & 0.711 & Cpa3, Ms4a2, Csrp3, Hdc & 4/5 (Cpa3, Ms4a2, Gata2, Hdc) \\ +lymphoid & 694 & 639 & 55 & 0.921 & Gimap1, Gimap6, B2m & --- \\ +monocyte & 472 & 315 & 157 & 0.667 & Irf8, Ms4a6c, Tyrobp & 2/8 (Mpo, Ctsg) \\ +\textbf{LT-HSC} & \textbf{160} & \textbf{135} & \textbf{25} & \textbf{0.844} & \textbf{Hlf, Mpl, Meis1} & \textbf{3/7 (Hlf, Meis1, Mecom)} \\ +pro-B & 66 & 60 & 6 & 0.909 & Ebf1, Cd79a, Vpreb3 & 1/4 (Il7r-adj) \\ +macrophage & 63 & 42 & 21 & 0.667 & Irf8, Lgals1, Tyrobp & 1/8 (Ctsg) \\ +\bottomrule +\end{tabular} +\end{center} + +Basophil-mast, megakaryocyte, erythroid, and myeloid predictions recover canonical panels at LFC $3$--$8$. Erythroid and myeloid are the classes with the lowest WT fraction (frac\_WT $0.542$ and $0.507$), consistent with the gate-matched expansion of both compartments in \S\ref{sec:dahlin}. The LT-HSC class recovers 160 cells (84.4\% WT) with canonical Hlf/Mpl/Meis1/Mecom markers --- a Kit-W41-depleted stem population predicted as a first-class label. Note this class rests on 160 cells (25 mutant) and is anchored entirely by one external atlas (\S\ref{sec:dahlin-per-lineage-metabolism}). + +\subsection{Per-lineage metabolic reprogramming under Kit-W41} +\label{sec:dahlin-per-lineage-metabolism} + +LT-HSCs and lymphoid progenitors show the strongest OXPHOS upregulation under Kit-W41. We scored all 61{,}122 Dahlin cells on four canonical metabolic modules --- OXPHOS\_ETC, Glycolysis, Fatty\_acid\_oxidation, and Redox\_glutathione --- and computed per-lineage Kit-W41-minus-WT deltas. Ranking lineages by $\sum_m |\Delta_m|$ (total mobilised metabolic signal): + +\begin{center} +\begin{tabular}{lrrrrr} +\toprule +Predicted class & $\Delta$ OXPHOS & $\Delta$ Glycolysis & $\Delta$ FAO & $\Delta$ Redox & $\sum |\Delta|$ \\ +\midrule +macrophage & $+0.044$ & $+0.190$ & $-0.042$ & $+0.098$ & $0.374$ \\ +\textbf{LT-HSC} & $\mathbf{+0.219}$ & $+0.024$ & $-0.035$ & $+0.058$ & $\mathbf{0.335}$ \\ +\textbf{lymphoid} & $\mathbf{+0.196}$ & $+0.039$ & $-0.002$ & $+0.095$ & $\mathbf{0.332}$ \\ +erythroid & $+0.115$ & $+0.051$ & $-0.060$ & $+0.072$ & $0.299$ \\ +MPP & $+0.138$ & $+0.063$ & $-0.009$ & $+0.059$ & $0.269$ \\ +megakaryocyte & $+0.100$ & $+0.016$ & $-0.012$ & $+0.057$ & $0.185$ \\ +basophil-mast & $-0.036$ & $+0.058$ & $+0.016$ & $+0.037$ & $0.147$ \\ +myeloid & $-0.008$ & $+0.073$ & $+0.014$ & $+0.026$ & $0.122$ \\ +monocyte & $+0.023$ & $+0.013$ & $+0.019$ & $+0.027$ & $0.082$ \\ +\bottomrule +\end{tabular} +\end{center} + +\textbf{Observation, not a novelty claim.} Setting aside macrophage (small $n$), the largest OXPHOS deltas fall on \textbf{LT-HSC} ($+0.219$, from 25 mutant cells) and \textbf{lymphoid} ($+0.196$, from 55): both show large OXPHOS gains with modest Redox/Glycolysis gains and a slight FAO drop --- the signature of a shift from a quiescent, FAO-supported state to an ETC-active state. This is consistent with hypofunctional Kit forcing canonically-quiescent HSCs into an energetically active (and stressed) state, with the effect largest per cell in the LT-HSC compartment. This per-lineage magnitude ranking is not reported in Dahlin. {\sloppy Artefacts: \pathsplit{discovery/hematopoiesis/marker/108\_dahlin\_lineage\_metabolism\_ranked.csv} and \pathsplit{108\_dahlin\_lineage\_metabolism.json}; script \pathsplit{scripts/analysis/108\_dahlin\_lineage\_metabolism.py}.\par} + +\subsection{Testable wet-lab predictions on Kit-mutant mice} + +\begin{enumerate} +\item \textbf{Erythroid-restricted anti-apoptotic switch}: Bcl2 and Bcl2l1 (Bcl-xL) protein levels should be elevated in Kit-W41 erythroid progenitors (Ter119$^+$-gated) but not in Kit-W41 MPP (Lin$^-$Kit$^+$Sca1$^+$-gated), reflecting the direction-flip in Apoptosis\_pro module score ($\Delta = -0.017$ in erythroid vs.\ $\Delta = +0.067$ in MPP). +\item \textbf{Integrated stress response as a druggable node}: ISRIB or GADD34 hyperactivator treatment of Kit-W41 mice should partially rescue the proliferation defect. ATF4$^+$ nuclei should be elevated in Kit-W41 bone-marrow sections across MPP, erythroid, myeloid, and megakaryocyte compartments. +\item \textbf{Redox rescue}: N-acetylcysteine or glutathione supplementation should attenuate the progenitor redox shift. Note this is a test of established Kit--ROS biology \cite{lennartsson2012kit,ludin2014ros} in the W41 background, not a newly implicated node. +\item \textbf{Compensatory Kit-independent proliferation pathway}: because Kit\_signaling collapses across every class ($p \approx 0$) while some proliferation persists, an alternative RTK (Flt3, Csf1r) or non-RTK pathway must be compensating. +\end{enumerate} + +\section{Discovery target: Veres SC-$\beta$ protocol imperfection} +\label{sec:veres} + +\subsection{Zero-shot classification on Veres hPSC differentiation} + +\textbf{Scope note.} The pan-pancreatic corpus is built by mapping Veres's own published cell-type calls into the corpus vocabulary (\texttt{sc\_alpha}$\to$\texttt{alpha\_progenitor}, \texttt{sc\_beta}$\to$\texttt{beta\_progenitor}, \texttt{sc\_ec}$\to$\texttt{epsilon}), and 44{,}711 of these labelled cells are in training. Recovering those class identities and their canonical markers on the held-out slice is therefore label recovery under a supervised mapping, not independent inference, and is reported here as a transfer benchmark rather than as a discovery. Note also that the staged samples analysed here are Veres's \emph{experimental} protocols x1/x2, not the v8 production protocol on which their characterisation is based, so protocol-level efficiency statements do not generalise. + +Veres \emph{et al.}\ 2019 \cite{veres2019scbeta}, \emph{Nature} 569:368--373 (Melton lab), GSE114412: hPSC-directed pancreatic differentiation across Stages 3--6 (foregut endoderm $\to$ stem-cell $\beta$, SC-$\beta$), 44{,}711 cells in training, 12{,}297 held out (\S\ref{sec:zeroshot-veres}). Held-out macro F1 $= 0.900$ (Marker) on 20 canonical Veres classes; 82.8\% HVG overlap after human$\to$mouse case-fold. + +Predicted class distribution across Veres stages (n=9{,}080 cells with parseable stage annotation; the remaining 3{,}217 cells carry no stage tag and are dominated by the primary-islet control sample): +\begin{center} +\small +\setlength{\tabcolsep}{5pt} +\resizebox{\textwidth}{!}{% +\begin{tabular}{p{4.4cm}rrrr} +\toprule +Predicted class & Stage 3 (n=2{,}899) & Stage 4 (n=2{,}577) & Stage 5 (n=1{,}589) & \textbf{Stage 6 (n=2{,}015)} \\ +\midrule +pancreatic-progenitor & 72\% & 0\% & 0\% & 0\% \\ +proliferating & 28\% & 14\% & 5\% & 2\% \\ +endocrine-progenitor-primed & 0\% & 41\% & 0\% & 0\% \\ +alpha\_progenitor & 0\% & 26\% & 49\% & \textbf{59\%} \\ +beta\_progenitor & 0\% & 0\% & 11\% & \textbf{10\%} \\ +epsilon & 0\% & 0\% & 13\% & 11\% \\ +exocrine & 0\% & 0\% & 14\% & 10\% \\ +ductal, delta & 0\% & 0\% & 3\% & 4\% each \\ +\bottomrule +\end{tabular}} +\end{center} + +\texttt{pancreatic-progenitor} dominates Stage 3 (72\%) and collapses across Stages 4--5 as endocrine-progenitor-primed and alpha\_progenitor identities emerge; by Stage 6, alpha\_progenitor is the largest class (59\%) and beta\_progenitor is 10\%. Adult-$\alpha$/adult-$\beta$ sub-prototypes never activate; SC-$\beta$-protocol cells occupy immature progenitor identities in the corpus manifold. Without any stage information supplied as input, \panda{} reproduces Veres's central finding that the differentiation is imperfect --- alpha-lineage outnumbers beta-lineage $\sim\!6\times$ at the SC-$\beta$ target stage (Figure~\ref{fig:veres}). + +\begin{figure}[!htb] +\centering +\includegraphics[width=\textwidth]{figures/fig4_veres_stage_stack.pdf} +\caption[Veres 2019 stage-stack composition]{Veres 2019 iPSC-directed pancreatic differentiation. Stacked bars show PANDA-Marker predicted class fraction across the four Veres protocol stages (Stages 3--6, pooling two batch samples each; $n=9{,}080$ cells with parseable stage annotation).\\ +\textbf{Stage 3:} \texttt{pancreatic-progenitor} dominates (72\% of $n=2{,}899$) and monotonically declines through the protocol.\\ +\textbf{Stage 6:} \texttt{alpha\_progenitor} rises from 0\% at Stage 3 to 59\% at the SC-$\beta$ target stage.\\ +\textbf{Sub-prototypes:} Adult-$\beta$ and adult-$\alpha$ sub-prototypes do not activate on Veres (see \S\ref{sec:zeroshot-veres}).} +\label{fig:veres} +\end{figure} + +\subsection{Stage-6 alpha\_progenitor vs beta\_progenitor mechanism} +\label{sec:veres-stage6-mech} + +Restricting to Stage 6: $1{,}192$ alpha\_progenitor vs $199$ beta\_progenitor predictions. The marker deep-dive (\S\ref{sec:veres-marker-deep-dive}) places the identity boundary on the master-TF axis: alpha\_progenitor recovers 4/4 canonical adult-$\alpha$ markers (GCG, ARX, IRX2, MAFB) at LFC $+3$ to $+7.5$; beta\_progenitor top-3 are ACVR1C ($+4.5$), CALB2 ($+5.9$), INS ($+4.8$). The residual $6\times$ alpha/beta asymmetry at Stage 6 matches Veres's own report of SC-$\beta$-protocol inefficiency. Because the corpus adult-$\beta$ prototype does not activate on Veres (\S\ref{sec:veres-adult-beta}), Stage-6 secretory cells route to \texttt{beta\_progenitor} rather than \texttt{beta}, and immature-$\beta$ cannot be distinguished from committed adult-$\beta$ on this checkpoint. + +\subsection{Veres marker deep-dive per predicted class} +\label{sec:veres-marker-deep-dive} + +Wilcoxon per PANDA-predicted class on raw Veres counts (\pathsplit{discovery/pancreas/marker/91\_veres\_marker\_deep\_dive.csv}). Top-9 classes by cell count shown; distribution reflects the current 18-class predicted vocabulary on the held-out slice: +\begin{center} +\scriptsize +\setlength{\tabcolsep}{3pt} +\begin{tabular}{p{3.0cm}rp{4.8cm}p{2.6cm}r} +\toprule +Predicted class & $n$ & Top-3 markers (LFC) & Canonical panel recovery & \% of class at Stage 6 \\ +\midrule +alpha\_progenitor & 2{,}864 & GCG ($+7.5$), TTR ($+4.9$), CHGA ($+4.7$) & 4/4 alpha (GCG, ARX, IRX2, MAFB) & 42\% \\ +pancreatic-progenitor & 2{,}099 & MDK ($+3.5$), SOX11 ($+3.9$), FN1 ($+4.5$) & --- & 0\% \\ +proliferating & 1{,}273 & TUBA1B ($+2.7$), TUBB ($+2.5$), HMGB1 ($+2.2$) & --- & 3\% \\ +endocrine-progenitor-primed & 1{,}057 & DLK1 ($+6.0$), LDHB ($+3.0$), PDX1 ($+2.9$) & 1/5 beta (PDX1) & 0\% \\ +epsilon$^{\dagger}$ & 701 & DDC ($+5.7$), FEV ($+5.1$), TPH1 ($+7.4$) & --- (see $\dagger$) & 31\% \\ +beta & 560 & INS ($+7.4$), IAPP ($+9.6$), ADCYAP1 ($+7.5$) & 1/5 beta (INS) $+$ MAFA/UCN3 in top-15 & 0\% \\ +beta\_progenitor & 539 & ACVR1C ($+4.5$), CALB2 ($+5.9$), INS ($+4.8$) & 1/2 EP-Fev (INSM1) & 37\% \\ +delta & 533 & SST ($+7.0$), ISL1 ($+2.6$), PCP4 ($+2.7$) & 1/2 EP-Fev (INSM1) & 16\% \\ +acinar & 426 & REG1A ($+11.8$), CTRB1 ($+12.2$), CTRB2 ($+11.8$) & 4/4 acinar (PRSS1, PRSS2, CEL, CTRB1) & 0\% \\ +\bottomrule +\end{tabular} +\end{center} + +$^{\dagger}$\textbf{The \texttt{epsilon} class is a misnomer and is really Veres's SC-EC population.} Its top marker is TPH1 (LFC $+7.4$), the defining enterochromaffin enzyme, with DDC and FEV also high, while GHRL --- the marker that defines a true $\epsilon$/ghrelin cell --- shows no enrichment over background. 89\% of the class is Veres's \texttt{sc\_ec} annotation. This is Veres's own headline discovery (``a previously unreported population that resembles enterochromaffin cells''), and it is being reported here under a mouse-atlas label that has no SC-EC entry. Veres anticipated precisely this: ``these cells may be misclassified as either progenitors or bona fide $\beta$ cells when analysed using methods that are based on preselected groups of genes.'' The class should be read as SC-EC throughout. + +\textbf{Interpretation.} On the current pan-pancreatic checkpoint, \panda{}-Marker does not separate adult-$\alpha$ from juvenile $\alpha$ or adult-$\beta$ from juvenile $\beta$ on the Veres held-out slice. The \texttt{alpha\_progenitor} cluster recovers all four canonical adult-$\alpha$ markers (GCG, ARX, IRX2, MAFB) and dominates Veres Stage 6 ($59\%$ of staged Stage-6 cells). Acinar recovers all four PRSS1/PRSS2/CTRB1/CEL exocrine markers at LFC $\sim 10$. + +\paragraph{Adult-$\beta$ signal within the \texttt{beta} class.}\label{sec:veres-adult-beta} +\panda{}-Marker predicts 0 Veres cells as \texttt{adult-beta} on the current checkpoint; the adult sub-prototype does not activate. \emph{Correction:} earlier revisions reported that the adult-maturity signature nonetheless survived at the gene level inside the \texttt{beta} class (MAFA $16\times$, UCN3 $8\times$, IAPP $12\times$). That claim is withdrawn. All 560 cells in that class carry no protocol-stage tag: they come from the primary-islet control sample bundled in the GEO extract (adult cadaveric human islets, GSE84133), not from the differentiation, and \texttt{stage6\_frac} for the class is $0$. The enrichment was therefore measured in adult islet $\beta$ cells and says nothing about SC-$\beta$ cells. It also runs against Veres, who state explicitly that ``some markers of maturity or age (UCN3, MAFA and SIX3) were not expressed'' in SC-$\beta$ cells. Two consistent readings: (i) Veres SC-$\beta$ cells occupy the corpus \texttt{beta} manifold too broadly for the adult sub-prototype's angular margin to fire; (ii) the corpus \texttt{adult-beta} prototype was learned on primary human islet data and the SC-$\beta$ transcriptome is close-but-not-close-enough. Artefact \texttt{95\_adult\_beta\_validation.json} is marked \texttt{vacuous: true} for the adult-vs-juvenile ratio. + +\subsection{Polyhormonal SC-$\alpha$ sub-clustering} +\label{sec:veres-polyhormonal-subcluster} + +Polyhormonal signal within the SC-$\alpha$ pool is concentrated rather than uniform, but --- as the correction below sets out --- this neither resolves an open question nor supports a bipotent-arrest reading. Veres 2019 characterise the SC-$\alpha$ output as polyhormonal (Gcg + Ins + Sst co-expression) and conclude in their Discussion that these are $\alpha$-like cells transiently mis-expressing insulin. To test this we pooled Veres alpha-lineage predictions (\emph{alpha\_progenitor} + \emph{alpha}, $n = 3{,}473$), defined \emph{polyhormonal} as $\geq 2$ of \{Ins, Gcg, Sst\} above their respective 75\textsuperscript{th} percentiles within the alpha pool, and Leiden-clustered at resolution $= 0.5$ into 8 sub-clusters: + +\begin{center} +\begin{tabular}{lrrrrrr} +\toprule +Sub-cluster & $n$ & \% pool & \% polyhormonal & \% GCG-hi & \% SST-hi & Enrichment \\ +\midrule +\textbf{3} & \textbf{522} & \textbf{15\%} & \textbf{46\%} & \textbf{69\%} & \textbf{43\%} & \textbf{$\mathbf{2.53\times}$} \\ +0 & 868 & 25\% & 30\% & 43\% & 29\% & $1.65\times$ \\ +2 & 585 & 17\% & 15\% & 14\% & 27\% & $0.82\times$ \\ +6, 4, 1, 5, 7 & 1{,}498 & 43\% & $\leq 9\%$ & --- & --- & $\leq 0.48\times$ \\ +\midrule +baseline & 3{,}473 & 100\% & 18.1\% & --- & --- & $1.00\times$ \\ +\bottomrule +\end{tabular} +\end{center} + +\textbf{Correction and attribution.} Earlier revisions presented this as novel biology resolving an open question, and as support for a ``trapped bipotent progenitor'' reading. Both claims are withdrawn. (i) Veres resolve the question explicitly in their Discussion: ``the identity of poly-hormonal cells has previously been controversial: we conclude that they represent $\alpha$-like (that is, SC-$\alpha$) cells that only transiently mis-express insulin.'' Their annotation names the population ``SC-alpha cells (Poly-hormonal cells)'', so an alpha pool defined from those labels is polyhormonal by construction. (ii) The discreteness of polyhormonal cells was established well before 2019 \cite{kelly2011,rezania2011,riedel2012,nostro2015,petersen2017}, and the field's conclusion --- a committed $\alpha$-lineage precursor (ARX$^+$, PDX1$^-$, NKX6.1$^-$) transiently mis-expressing insulin --- is incompatible with a bipotent-arrest reading. (iii) The sub-clustering statistics do not independently support ``discrete'': 609 of the 3{,}473 pooled cells are primary-islet control cells rather than SC-islet cells, and the polyhormonal rate sits only $\sim\!1.1\times$ above a permutation null that shuffles the three hormones independently, so most of the nominal baseline is an artefact of percentile thresholding. What remains genuinely open --- and what this analysis does \emph{not} settle --- is a formal statistical test of discreteness (mixture modelling with doublet control) on the SC-$\alpha$ pool alone. {\sloppy Artefacts: \pathsplit{discovery/pancreas/marker/110\_veres\_polyhormonal\_alpha\_per\_cluster.csv} and \pathsplit{110\_veres\_polyhormonal\_alpha\_summary.json}; script \pathsplit{scripts/analysis/110\_veres\_polyhormonal\_alpha.py}.\par} + +\subsection{Testable wet-lab predictions on Veres SC-$\beta$ differentiation} + +\begin{enumerate} +\item \textbf{Terminal TF-switching (already established in mouse; open in late-stage hPSC)}: NKX6-1 is necessary and sufficient to respecify non-$\beta$ endocrine precursors toward $\beta$ fate, and does so partly by directly repressing \emph{ARX} \cite{schaffer2013}; MNX1 acts similarly. The untested part is only the implementation --- inducible NKX6-1 at the Stage 5$\to$6 transition of an SC-islet protocol, with lineage tracing to show conversion rather than selective survival. We do not claim the mechanism as novel. +\item \textbf{ARX knockdown --- and why the human prediction is not the mouse one}: in mouse, \emph{Arx} inactivation converts $\alpha$-cells to $\beta$-like cells \cite{courtney2013}. In human hPSC differentiation the published result is the opposite direction of usefulness: ARX-null hESC-derived endocrine cells were $94\%$ somatostatin-positive with insulin-positive cells reduced by $\sim\!65\%$ \cite{gage2015}. A well-posed experiment must therefore state a \emph{single} outcome in advance (SC-$\beta$ \emph{or} SC-$\delta$, not either), and the human literature predicts $\delta$. Our data do not adjudicate this. +\item \textbf{Polyhormonal sorting}: single-cell FISH for SST on Stage-6 SC-$\alpha$ output would establish whether GCG$^+$SST$^+$ cells form a separable population; paired with a formal discreteness test (mixture model, doublet-controlled) this is the experiment that would settle the question left open in \S\ref{sec:veres-polyhormonal-subcluster}. (Statistics quoted for this prediction in earlier revisions had no backing artefact in the repository and have been removed.) +\end{enumerate} + +\section{Prototype geometry} +\label{sec:proto-geom} + +For each system, \panda{}'s $K$ class-mean prototype vectors in the 128-d projection space have a participation-ratio effective dimensionality +\[ +\text{eff\_dim}(P) \;=\; \frac{\big(\sum_i \lambda_i\big)^2}{\sum_i \lambda_i^2}, +\] +where $\{\lambda_i\}$ are the eigenvalues of the centred prototype covariance. This is a smooth ``how many independent directions is the prototype set using?'' summary that equals $K$ when prototypes are orthonormal and 1 when they are collinear: + +\begin{center} +\begin{tabular}{lrrr} +\toprule +System & $K$ & Effective dim & Fraction \\ +\midrule +Pan-skin & 13 & \textbf{11.08} & 85\% \\ +Pan-hematopoietic & 15 & \textbf{11.33} & 76\% \\ +Pan-pancreatic & 20 & \textbf{10.85} & 54\% \\ +\bottomrule +\end{tabular} +\end{center} + +\textbf{Interpretation.} The pan-skin prototype set uses $85\%$ of the available dimensions, indicating cleanly-separated identities that occupy nearly-orthogonal directions in the projection space. The pan-hematopoietic prototype set uses $76\%$: still healthy, but the shared myeloid/basophil/megakaryocyte Gata1/Gata2$^+$ progenitor program links several lineages onto a common axis. The pan-pancreatic prototype set uses only $54\%$ of its 20 dimensions --- the largest prototype-set compression in the paper --- consistent with the well-established secondary transition dogma of pancreatic endocrinogenesis, in which $\beta$/$\delta$/endocrine-progenitor/other cell fates share a strong Nkx6-1$^+$/Neurod1$^+$ transcriptional program until terminal Ins1$^\text{hi}$/Mafa$^\text{hi}$ maturation. This is a direct empirical readout of biological lineage compression in the training corpus: the prototype geometry \panda{} learns is determined by transcriptomic distinguishability between labels, so labels whose training cells are close in expression space end up with lower-dimensional prototype sets even if the canonical ontology says they are distinct cell types. + +\section{Cross-system synthesis} + +Three tissue systems, three discovery targets, three paper-central findings independently reproduced by \panda{}: + +\begin{center} +\small +\begin{tabular}{p{2.4cm}p{3.9cm}p{8.6cm}} +\toprule +System & Held-out target & Paper claim reproduced \\ +\midrule +Skin & Dingwall En1-cKO & Unified spatial-repressor model $+$ EDEN lineage extension \\ +Hematopoiesis & Dahlin Kit-W41 & Three reported compositional shifts, gate-matched (incl.\ mast/basophil loss) \\ +Pancreas & Veres hPSC & Stage-wise identity progression and SC-$\alpha$-dominant Stage 6 (label recovery) \\ +\bottomrule +\end{tabular} +\end{center} + +\textbf{Common threads.} Across all three systems: (i) zero-shot classification followed by within-class Wilcoxon and pathway module scoring recovers canonical marker sets never supplied to the model (Cd34/Myc for MPP; Cpa3/Gata2/Hdc/Mcpt8 for basophil-mast; Sox10/Dct/Tyr/Pmel for melanocyte; Nkx6-1/Mnx1/Neurod1 for SC-$\beta$; Arx/Irx2 for SC-$\alpha$); (ii) low prototype-cosine populations mark biology the training corpus does not represent, examined post-hoc rather than by an inference-time gate (\S\ref{sec:inference}); (iii) between-condition module-score contrasts reproduce the source paper's mechanistic interpretations at Wilcoxon $p$-values exceeding textbook thresholds by many orders of magnitude. + +\textbf{Divergences.} The three systems separate cleanly by prototype geometry (\S\ref{sec:proto-geom}): skin at $85\%$ effective-dim utilisation shows fully-resolved terminal identities; hematopoiesis at $76\%$ has moderate lineage-shared compression on the Gata2$^+$ multipotent progenitor axis; pancreas at $54\%$ has strong compression on the shared endocrine-progenitor program. This ordering is itself a testable biological statement: the transcriptomic distinguishability of terminal identities in a training corpus is directly readable off the prototype set's effective dimensionality. Applying this framework prospectively lets an atlas-builder decide, before training, whether adult-anchor cells (or additional developmental stages) are needed to unlock terminal-cell-type separability. + +Figure~\ref{fig:multiumap} shows the projection geometry for the three discovery targets; the condition of interest induces local density shifts within a cluster-preserving embedding, consistent with the within-class contrasts reported above. + +\begin{figure}[!htb] +\centering +\includegraphics[width=\textwidth]{figures/fig6_multi_umap.pdf} +\caption[Three-panel discovery-target UMAP]{Three discovery-target UMAPs in \panda{}'s 128-d projection space (all cells shown: $n=25{,}800$ / $61{,}122$ / $12{,}297$).\\ +\textbf{(a)} Dingwall skin coloured by En1 genotype.\\ +\textbf{(b)} Dahlin hematopoiesis coloured by Kit genotype.\\ +\textbf{(c)} Veres held-out slice ($n=12{,}297$) coloured by protocol stage (3--6); the 3{,}217 primary-islet control cells (no stage tag) are shown in grey.\\ +\textbf{Observation:} between-condition/between-stage differences manifest as local density shifts within a shared, biologically-structured embedding.} +\label{fig:multiumap} +\end{figure} + +\section{Discussion} + +\panda{} delivers a compact prototype-anchored classifier (PCA-only and marker-augmented variants) trained on 100\%-paper-labeled corpora that validates within-corpus, transfers to labeled hold-outs across a spectrum of strictness (strict zero-shot, same-study split, anchor design, held-out slice --- each labeled as such in \S 5), and supports mechanistic discovery via within-class Wilcoxon DE and pathway module scoring. + +Three points about how the checkpoint is used in discovery. First, the Dingwall EDEN validation illustrates the intended workflow of a prototype-anchored classifier as a hypothesis-generation tool: the pan-tissue fibroblast prototype is calibrated broader than any narrow tissue-specific compartment, so an independent clustering (Line B) or target-native training pass (Line C) is required to resolve EDEN as a single class; once either is provided, the depletion phenotype is a learnable and reproducible property of the labelling rather than a clustering artefact, and sub-cluster panel-matching extends the phenotype to the untested Derm2 precursor. Second, Dahlin is a cautionary case about experimental design rather than a source of new phenotypes: the dataset contains two FACS gates, and a within-genotype gate contrast produces larger compositional shifts than the mutation itself (\S\ref{sec:dahlin}). Restricting to the matched gate both removes the inflated effects and recovers the mast/basophil loss that Dahlin emphasise and that the pooled analysis missed. The pathway-module hits we initially read as novel are, on checking, either reported by Dahlin themselves (the ATF4/integrated-stress response) or established Kit and HSC-metabolism biology \cite{lennartsson2012kit,ludin2014ros,filippi2022metabolism}. Third, on Veres the adult-$\beta$ sub-prototype does not activate on SC-$\beta$ cells --- consistent with Veres's own report that MAFA, UCN3 and SIX3 are not expressed in SC-$\beta$ --- and the class we label \texttt{epsilon} is in fact their SC-EC enterochromaffin population, a vocabulary limit of the corpus rather than a discovery. Both are the kind of failure the prototype-geometry compression (54\% effective dim, \S\ref{sec:proto-geom}) makes quantitatively concrete: a pan-tissue corpus cannot resolve identities it has no prototype for, and will route them to the nearest available label. + +Future work will extend the interpretability toolkit --- prototype--gene attribution, counterfactual single-gene knockout, and Hessian gene-gene interaction probes --- using the analytic PCA back-projection that \panda{}'s architecture makes exact. + +\section{Limitations} + +\begin{itemize} +\item Pan-tissue prototypes are calibrated for cross-dataset generality; narrow tissue-specific compartments like Dingwall's EDEN require either an independent clustering step (Line B) or Dingwall-native training (Line C) to be recovered as a single class. +\item Zero-shot transfer to mature terminal cell types requires developmentally-mature anchors in the training corpus (documented via the Baron test-half zero-shot performance and the pancreatic prototype-geometry compression). +\item Sparse-class predictions (HF-placode, HF-DP, gamma/immune on Baron test) are underpowered when the training-corpus support is small. +\item Novel-population claims derived from low-confidence cells are hypotheses examined post-hoc; every mechanistic claim in the discovery reports requires wet-lab confirmation. +\item Fibroblast-papillary predictions in Dingwall show smooth-muscle contamination (Cald1, Myh11, Acta2), consistent with the class being conflated with myofibroblast identity in adult skin. +\item Cross-platform generalisation to Smart-seq2 (Nestorowa) remains a hard benchmark; a Smart-seq2 LT-HSC training anchor would be the natural next step. +\end{itemize} + +\section{Reproducibility} + +\textbf{Multiple-testing scope.} Pathway-module $p$-values reported in the class $\times$ module tables (\S\ref{sec:dingwall-pathway}, \S\ref{sec:dahlin-pathway}) are Bonferroni-corrected within-system across the full $n_\text{classes} \times n_\text{modules}$ family scanned per system (Dingwall: $12 \times 15+$ modules; Dahlin: $10 \times 30$ modules). Per-class Wilcoxon DEG $p$-values (\S\ref{sec:dingwall-hf-placode-degs}, \S\ref{sec:dahlin-markers}, and the Veres per-class marker deep-dive) use Benjamini--Hochberg per class. Fisher exact class enrichments (\S\ref{sec:dingwall-fisher}) use Benjamini--Hochberg across the classes tested per system. + +All quantitative claims in this paper trace to a specific artefact: +\begingroup +\RaggedRight +\sloppy +\begin{itemize}\itemsep2pt +\item Model checkpoints: \pathsplit{checkpoints/\{system\}/\{variant\}/panda\_final.pt} where variant $\in \{$pca, marker$\}$ +\item Marker channel gene lists: \pathsplit{panda/markers.yaml} +\item Corpus builders: \pathsplit{scripts/\{pan\_skin,hematopoiesis,pancreas\}/}; unified trainer: \pathsplit{scripts/common/train\_panda.py}; unified 5-fold CV driver: \pathsplit{scripts/common/cv\_holdout.py}; zero-shot inference: \pathsplit{scripts/common/run\_all\_zero\_shot.py} +\item External-label supplements: \pathsplit{data/external\_labels/\{dingwall\_supp,haensel,joost2016,mca,mia,byrnes,yu,baccin,melanocyte\_anchor\}/} +\item Per-configuration 5-fold CV (three seeds per system): \pathsplit{discovery/\{system\}/\{variant\}/cv\_5fold.json}, \pathsplit{cv\_5fold\_seed1.json}, \pathsplit{cv\_5fold\_seed2.json} +\item Zero-shot labeled targets: + \begin{itemize}\itemsep2pt + \item Baron test-half: \pathsplit{discovery/pancreas/\{pca,marker\}/baron\_summary.json} (current 20-class vocabulary; the paper body's Baron numbers are read from this file) + \item Veres held-out 12{,}297: \pathsplit{discovery/pancreas/\{pca,marker\}/veres\_summary.json} + \item Nestorowa: \pathsplit{discovery/hematopoiesis/\{pca,marker\}/nestorowa\_summary.json} + \item Sulic (anchor design, \S\ref{sec:zeroshot-sulic}): \pathsplit{discovery/pan\_skin/\{pca,marker\}/97\_sulic\_anchor\_zero\_shot.json}. The older \pathsplit{sulic\_summary.json} holds the withdrawn training-set numbers and must not be cited. + \item Belote: \pathsplit{discovery/pan\_skin/\{pca,marker\}/belote\_summary.json} + \end{itemize} +\item Adult-$\beta$ marker-panel validation on Veres: \pathsplit{discovery/pancreas/marker/95\_adult\_beta\_validation.json} +\item Expanded pathway module analysis (all systems): \pathsplit{scripts/analysis/57\_pathway\_analysis.py}, outputs at \pathsplit{discovery/\{system\}/marker/57\_pathway\_class\_by\_module\_\{padj,delta\}.tsv} +\item \textbf{Dingwall EDEN validation (three lines of evidence)}: + \begin{itemize}\itemsep2pt + \item Line A post-hoc scoring: \pathsplit{scripts/analysis/98\_eden\_posthoc\_detection.py}, output \pathsplit{discovery/pan\_skin/marker/98\_eden\_summary.json} + \item Line B scanpy reproduction: \pathsplit{scripts/analysis/103\_replicate\_dingwall\_seurat\_pipeline.py}, output \pathsplit{data/processed/dingwall\_replica/dingwall\_replica.h5ad}, \pathsplit{data/processed/dingwall\_replica/replica\_cluster\_20\_qc.json}, \pathsplit{replica\_marker\_matches.csv} + \item Line C PANDA on Derm labels: \pathsplit{scripts/analysis/104\_train\_on\_dingwall\_derm\_labels.py}, output \pathsplit{discovery/pan\_skin/marker/104\_dingwall\_derm\_summary.json} $+$ prediction/depletion CSVs + \end{itemize} +\item Primary EDEN Derm2 discovery: \pathsplit{scripts/analysis/100\_primary\_eden\_discovery.py}, \pathsplit{101\_primary\_eden\_derm\_scoring.py}; outputs \pathsplit{discovery/pan\_skin/marker/100\_primary\_eden\_discovery.csv}, \pathsplit{100\_primary\_eden\_summary.json}, \pathsplit{101\_derm\_identity\_summary.json}, \pathsplit{101\_derm\_subcluster\_scores.csv} +\item Dingwall other mechanistic outputs: \pathsplit{discovery/pan\_skin/marker/57\_pathway\_analysis.csv, 90\_dingwall\_marker\_deep\_dive.csv} +\item Dahlin Kit-mutant: \pathsplit{discovery/hematopoiesis/marker/92\_dahlin\_marker\_deep\_dive.csv, dahlin\_summary.json} +\item Veres marker deep-dive: \pathsplit{discovery/pancreas/marker/91\_veres\_marker\_deep\_dive.csv} +\end{itemize} +\endgroup + +\begin{thebibliography}{99} +\bibitem{dingwall2024en1cko} Dingwall HL, Tomizawa RR, Aharoni A, Hu P, Qiu Q, Kokalari B, Martinez SM, Donahue JC, Aldea D, Mendoza M, Glass IA, Wu H, Kamberov YG (2024). ``Sweat gland development requires an eccrine dermal niche and couples two epidermal programs.'' \emph{Developmental Cell} 59(1):20--32.e6 \url{https://doi.org/10.1016/j.devcel.2023.11.015}. GSE220977. +\bibitem{lennartsson2012kit} Lennartsson J, R\"onnstrand L (2012). ``Stem cell factor receptor/c-Kit: from basic science to clinical implications.'' \emph{Physiological Reviews} 92(4):1619--1649 \url{https://doi.org/10.1152/physrev.00046.2011}. +\bibitem{ludin2014ros} Ludin A, \emph{et al.} (2014). ``Reactive oxygen species regulate hematopoietic stem cell self-renewal, migration and development, as well as their bone marrow microenvironment.'' \emph{Antioxidants \& Redox Signaling} 21(11):1605--1619 \url{https://doi.org/10.1089/ars.2014.5941}. +\bibitem{filippi2022metabolism} Filippi MD, Ghaffari S (2022). ``Mitochondria in the maintenance of hematopoietic stem cells: new perspectives and opportunities.'' \emph{HemaSphere} 6:e740 \url{https://doi.org/10.1097/HS9.0000000000000740}. +\bibitem{kelly2011} Kelly OG, \emph{et al.} (2011). ``Cell-surface markers for the isolation of pancreatic cell types derived from human embryonic stem cells.'' \emph{Nature Biotechnology} 29(8):750--756 \url{https://doi.org/10.1038/nbt.1931}. +\bibitem{rezania2011} Rezania A, \emph{et al.} (2011). ``Production of functional glucagon-secreting alpha-cells from human embryonic stem cells.'' \emph{Diabetes} 60(1):239--247 \url{https://doi.org/10.2337/db10-0573}. +\bibitem{riedel2012} Riedel MJ, \emph{et al.} (2012). ``Immunohistochemical characterisation of cells co-producing insulin and glucagon in the developing human pancreas.'' \emph{Diabetologia} 55(2):372--381 \url{https://doi.org/10.1007/s00125-011-2344-9}. +\bibitem{nostro2015} Nostro MC, \emph{et al.} (2015). ``Efficient generation of NKX6-1+ pancreatic progenitors from multiple human pluripotent stem cell lines.'' \emph{Stem Cell Reports} 4(4):591--604 \url{https://doi.org/10.1016/j.stemcr.2015.02.017}. +\bibitem{petersen2017} Petersen MBK, \emph{et al.} (2017). ``Single-cell gene expression analysis of a human ESC model of pancreatic endocrine development reveals different paths to beta-cell differentiation.'' \emph{Stem Cell Reports} 9(4):1246--1261 \url{https://doi.org/10.1016/j.stemcr.2017.08.009}. +\bibitem{schaffer2013} Schaffer AE, \emph{et al.} (2013). ``Nkx6.1 controls a gene regulatory network required for establishing and maintaining pancreatic beta cell identity.'' \emph{PLoS Genetics} 9(1):e1003274 \url{https://doi.org/10.1371/journal.pgen.1003274}. +\bibitem{courtney2013} Courtney M, \emph{et al.} (2013). ``The inactivation of Arx in pancreatic alpha-cells triggers their neogenesis and conversion into functional beta-like cells.'' \emph{PLoS Genetics} 9(10):e1003934 \url{https://doi.org/10.1371/journal.pgen.1003934}. +\bibitem{gage2015} Gage BK, \emph{et al.} (2015). ``The role of ARX in human pancreatic endocrine specification.'' \emph{PLoS ONE} 10(12):e0144100 \url{https://doi.org/10.1371/journal.pone.0144100}. +\bibitem{veres2019scbeta} Veres A, \emph{et al.}\ (Melton DA \emph{corresp.}) (2019). ``Charting cellular identity during human in vitro $\beta$-cell differentiation.'' \emph{Nature} 569:368--373 \url{https://doi.org/10.1038/s41586-019-1168-5}. GSE114412. +\bibitem{dahlin2018kit} Dahlin JS, \emph{et al.}\ (Wilson NK \emph{corresp.}) (2018). ``A single-cell hematopoietic landscape resolves 8 lineage trajectories and defects in Kit mutant mice.'' \emph{Blood} 131(21):e1--e11 \url{https://doi.org/10.1182/blood-2017-12-821413}. GSE107727. +\bibitem{haensel2020skin} Haensel D, \emph{et al.}\ (Annusver K \emph{et al.}) (2020). ``Defining epidermal basal cell states during skin homeostasis and wound healing using single-cell transcriptomics.'' \emph{Cell Reports} 30(11):3932--3947.e6 \url{https://doi.org/10.1016/j.celrep.2020.02.091}. GSE142471. +\bibitem{joost2016} Joost S, \emph{et al.}\ (2016). ``Single-cell transcriptomics reveals that differentiation and spatial signatures shape epidermal and hair follicle heterogeneity.'' \emph{Cell Systems} 3(3):221--237.e9 \url{https://doi.org/10.1016/j.cels.2016.08.010}. GSE67602. +\bibitem{belote2021} Belote RL, \emph{et al.}\ (2021). ``Human melanocyte development and melanoma dedifferentiation at single-cell resolution.'' \emph{Nature Cell Biology} 23(9):1035--1047 \url{https://doi.org/10.1038/s41556-021-00740-8}. GSE151091. +\bibitem{han2018mca} Han X, \emph{et al.}\ (Guo G \emph{corresp.}) (2018). ``Mapping the mouse cell atlas by Microwell-Seq.'' \emph{Cell} 172(5):1091--1107.e17 \url{https://doi.org/10.1016/j.cell.2018.02.001}. GSE108097. +\bibitem{tms2020} Tabula Muris Consortium (2020). ``A single-cell transcriptomic atlas characterizes ageing tissues in the mouse.'' \emph{Nature} 583:590--595 \url{https://doi.org/10.1038/s41586-020-2496-1}. GSE132042. +\bibitem{baccin2020} Baccin C, \emph{et al.}\ (2020). ``Combined single-cell and spatial transcriptomics reveal the molecular, cellular and spatial bone marrow niche organization.'' \emph{Nature Cell Biology} 22(1):38--48 \url{https://doi.org/10.1038/s41556-019-0439-6}. GSE122465. +\bibitem{bastidas2019} Bastidas-Ponce A, \emph{et al.}\ (Bakhti M, Lickert H) (2019). ``Comprehensive single cell mRNA profiling reveals a detailed roadmap for pancreatic endocrinogenesis.'' \emph{Development} 146(12):dev173849 \url{https://doi.org/10.1242/dev.173849}. GSE132188. +\bibitem{byrnes2018} Byrnes LE, \emph{et al.}\ (Sneddon JB \emph{corresp.}) (2018). ``Lineage dynamics of murine pancreatic development at single-cell resolution.'' \emph{Nature Communications} 9(1):3922 \url{https://doi.org/10.1038/s41467-018-06176-3}. GSE101099. +\bibitem{yu2021} Yu X, \emph{et al.}\ (2021). ``Single-cell RNA-seq of the developing pancreas identifies novel endocrine progenitors and endocrine subtypes.'' \emph{Cell Research} 31:669--686 \url{https://doi.org/10.1038/s41422-021-00509-6}. GSE139627. +\bibitem{hrovatin2023mia} Hrovatin K, \emph{et al.}\ (2023). ``Delineating mouse $\beta$-cell identity during lifetime and in diabetes with a single cell atlas.'' \emph{Nature Metabolism} 5:1615--1637 \url{https://doi.org/10.1038/s42255-023-00876-x}. GSE211796. +\end{thebibliography} + +\clearpage +\subsection*{Supplement figure index} +\label{sec:supp-index} + +Supporting figures are provided as PDFs under \pathsplit{figures/supplement/} (S\emph{n}) and \pathsplit{figures/biology/} (B\emph{n}). Each entry lists filename, one-line description, and the paper section that motivates it. + +\begingroup +\RaggedRight +\small +\begin{itemize}\itemsep1pt +\item \textbf{S1} \pathsplit{01\_cv\_summary.pdf} --- Held-out 5-fold CV summary across the three systems (\S\ref{sec:multiseed}; Fig~\ref{fig:cv}). +\item \textbf{S2} \pathsplit{02\_per\_class\_f1.pdf} --- Per-class F1 bars for skin/HSC/pancreas (Fig~\ref{fig:cv}). +\item \textbf{S3} \pathsplit{03\_prototype\_cosine.pdf} --- Prototype-prototype cosine similarity matrices per system (\S\ref{sec:proto-geom}). +% S4 removed --- legacy training-trajectory figure contained stale class vocab and wrong K counts. +\item \textbf{S5} \pathsplit{05\_adversary\_purification.pdf} --- Dataset/depth-adversary AUROC vs.\ ramp step (curriculum diagnostics for \S\ref{sec:multiseed}). +\item \textbf{S6} \pathsplit{06\_cross\_system\_prototypes.pdf} --- Cross-system prototype-geometry comparison (\S\ref{sec:proto-geom}). +% S7 removed --- attribution-heatmap contained deprecated class labels and leaked cell-barcode strings as gene columns. +% S8 removed --- TF-enrichment used stale skin/pancreas class rosters. +% S9 removed --- KO-essentials contained deprecated class labels and leaked cell-barcode strings. +% S10 removed --- Hessian gene-pair figure contained stale classes and illegible pair labels. +\item \textbf{S11} \pathsplit{11\_novel\_populations.pdf} --- \emph{Legacy figure.} Low-cosine clusters, rendered by a superseded abstain-gated script; the canonical pipeline has no abstention step (\S\ref{sec:inference}). Retained for provenance only. +% S12 removed --- Co-attention modules referenced deprecated classes (nascent-eccrine-gland, UNK, unassigned, mesenchymal). +\item \textbf{S13} \pathsplit{13\_dingwall\_umap.pdf} --- Full Dingwall UMAP by predicted class and genotype (\S\ref{sec:dingwall}; Fig~\ref{fig:umap}). +\item \textbf{S14} \pathsplit{14\_dahlin\_umap.pdf} --- Full Dahlin UMAP by predicted class and Kit genotype (\S\ref{sec:dahlin}). +\item \textbf{S15} \pathsplit{15\_veres\_umap.pdf} --- Full Veres UMAP by predicted class and stage (\S\ref{sec:veres}). +\item \textbf{S16} \pathsplit{16\_dingwall\_discovery.pdf} --- Dingwall mechanistic panel: En1-cKO class-enrichment bars and melanocyte pathway-module deltas (\S\ref{sec:dingwall-fisher}, \S\ref{sec:dingwall-melanoblast-mitf}). +\item \textbf{S17} \pathsplit{17\_dahlin\_discovery.pdf} --- Dahlin discovery panel: MPP redox, MK quiescence, per-lineage metabolism (\S\ref{sec:dahlin-pathway}, \S\ref{sec:dahlin-per-lineage-metabolism}). +\item \textbf{S18} \pathsplit{18\_veres\_discovery.pdf} --- Veres discovery panel: Stage-6 alpha/beta, polyhormonal sub-cluster (\S\ref{sec:veres-polyhormonal-subcluster}). +\item \textbf{S19} \pathsplit{19\_myeloid\_network.pdf} --- Myeloid marker co-expression network (context: Dahlin \S\ref{sec:dahlin-markers}). +\item \textbf{S20} \pathsplit{20\_placode\_wnt\_module.pdf} --- Pan-skin HF-placode gene co-attribution module network (from \pathsplit{82\_pan\_skin\_coatt\_modules.csv}; filename is historical). +% S21--S22 not included (no corresponding supplement figures). +\item \textbf{S23} \pathsplit{23\_anchor\_delta\_recall.pdf} --- Anchor-vs-holdout recall $\Delta$ across the four held-out labeled targets. +\item \textbf{S24} \pathsplit{24\_pca\_vs\_marker\_umaps\_dingwall\_by\_genotype.pdf} --- PCA vs.\ Marker UMAP of Dingwall coloured by En1 genotype (Marker vs.\ PCA ablation; \S\ref{sec:dingwall}). +\item \textbf{S24b} \pathsplit{24b\_pca\_vs\_marker\_umaps\_dingwall\_by\_class.pdf} --- Same Dingwall PCA vs.\ Marker UMAP coloured by predicted class. +\item \textbf{S25} \pathsplit{25\_pca\_vs\_marker\_umaps\_dahlin\_by\_genotype.pdf} --- PCA vs.\ Marker UMAP of Dahlin coloured by Kit genotype (\S\ref{sec:dahlin}). +\item \textbf{S25b} \pathsplit{25b\_pca\_vs\_marker\_umaps\_dahlin\_by\_class.pdf} --- Same Dahlin PCA vs.\ Marker UMAP coloured by predicted class. +\item \textbf{S26} \pathsplit{26\_pca\_vs\_marker\_umaps\_veres\_by\_stage.pdf} --- PCA vs.\ Marker UMAP of Veres coloured by protocol stage (\S\ref{sec:veres}). +\item \textbf{S26b} \pathsplit{26b\_pca\_vs\_marker\_umaps\_veres\_by\_class.pdf} --- Same Veres PCA vs.\ Marker UMAP coloured by predicted class. +\item \textbf{S27} \pathsplit{27\_dingwall\_en1\_enrichment.pdf} --- En1-cKO enrichment per PANDA class (\S\ref{sec:dingwall-fisher}). +\item \textbf{S28} \pathsplit{28\_melanocyte\_pathway\_modules.pdf} --- Melanocyte-lineage pathway module scores in Dingwall (\S\ref{sec:dingwall-melanoblast-mitf}). +\item \textbf{B1} \pathsplit{biology\_01\_dingwall\_umap.pdf} --- Publication-style Dingwall UMAP by predicted class (\S\ref{sec:dingwall}). +\item \textbf{B2} \pathsplit{biology\_02\_primary\_eden.pdf} --- Primary-EDEN Derm2 sub-cluster and marker overlap (\S\ref{sec:dingwall-primary-eden}). +\item \textbf{B3} \pathsplit{biology\_03\_melanoblast\_mitf.pdf} --- Melanoblast MITF-regulon vs.\ neural-crest scoring (\S\ref{sec:dingwall-melanoblast-mitf}). +\item \textbf{B4} \pathsplit{biology\_04\_dahlin\_metabolism.pdf} --- Per-lineage metabolic reprogramming heat-map on Dahlin (\S\ref{sec:dahlin-per-lineage-metabolism}). +\item \textbf{B5} \pathsplit{biology\_05\_dahlin\_composition.pdf} --- Compositional shifts under Kit-W41 (\S\ref{sec:dahlin}, Fisher table). +\item \textbf{B6} \pathsplit{biology\_06\_veres\_beta\_quadrant.pdf} --- INS $\times$ MAFA/UCN3 quadrants for the \texttt{beta}-predicted cells. \emph{Note:} those cells are primary-islet controls, not differentiation cells (\S\ref{sec:veres-adult-beta}), so this panel describes adult islet $\beta$ cells. +\item \textbf{B7} \pathsplit{biology\_07\_veres\_polyhormonal.pdf} --- Veres polyhormonal SC-$\alpha$ sub-cluster 3 (\S\ref{sec:veres-polyhormonal-subcluster}). +\item \textbf{B8} \pathsplit{biology\_08\_prototype\_geometry.pdf} --- Cross-system prototype effective-dimensionality plot (\S\ref{sec:proto-geom}). +\end{itemize} +\endgroup + +\end{document}