| \section{Introduction} |
|
|
| \emph{Compositional data}---nonnegative vectors (or components) whose summed total carries no intrinsic meaning---arise across scientific and statistical settings, including microbiome profiles, cell type proportions, ecological counts, and probabilistic model outputs \citep{gloor2017microbiome, billheimer2001statistical, buettner2021sccoda}. Such data live on the simplex, where inference is based on \emph{relative} not absolute values. Often, the components are not exchangeable and are organized by domain hierarchies reflecting evolutionary, functional, or semantic relationships \citep{silverman2017phylogenetic, harmon2019phylogenetic}. These trees describe which comparisons among components are meaningful and at what resolution. |
|
|
| The standard tool in compositional data analysis (CoDA) is \emph{Aitchison geometry} \citep{aitchison1982statistical}, which formalizes that only ratios are informative. This is achieved via an isometric log‑ratio (ILR) embedding of compositional data into Euclidean space \citep{egozcue2003isometric}. Here, one treats components \emph{symmetrically}, partly reflecting origins in domains where no external structure was assumed \citep{egozcue2005groups, mandal2015analysis}. But in many applications, domain-specific hierarchies are known: trees specify which comparisons are meaningful and how they are related. To address this gap, the literature provides strategies to incorporate tree structure, such as tree‑based balances and phylogeny‑aware coordinates. Doing so improves interpretability and downstream analysis. Yet, existing approaches typically address {specific regimes (e.g., binary trees as in phylogenetic reconstruction, selected contrasts)} \citep{silverman2017phylogenetic, washburne2017phylogenetic, morton2017balance} or rely on construction choices {\em external} to the geometry \citep{lozupone2005unifrac, mao2022dirichlet}. {For instance, a polytomous node admits no canonical binary resolution, yet binary-tree methods force an arbitrary choice (Figure~\ref{fig:binary-vs-poly}).} Hence, the key question we study is whether hierarchies can be incorporated in a way that is general, geometrically principled, and canonical for arbitrary trees. |
|
|
| \textbf{Main difficulty.} The core issue is structural incompatibility. Aitchison geometry identifies compositions up to a \emph{global} scaling: a $(d-1)$-dimensional Euclidean tangent space where ILR coordinates live. But trees impose \emph{local} and \emph{nested} constraints: distinctions are meaningful within clades (i.e., subtrees) and comparisons happen at multiple resolutions. Aligning these requires {\em decomposing} the Aitchison tangent space to respect branching structure at every internal node. Binary trees sidestep this issue by reducing each node to a single contrast (Figure~\ref{fig:binary-vs-poly}). For multi‑branching (i.e., \emph{polytomous}) hierarchies, no canonical construction exists. |
|
|
| \textbf{This paper.} We ask if an orthonormal decomposition of the Aitchison tangent space can be compatible with arbitrary tree structure and the simplex invariances---without sacrificing isometry or introducing arbitrary choices. Such a decomposition would give each internal node a geometrically motivated signature and provide a multiscale coordinate system for simplex-valued data on the same footing as standard ILR methods \cite{pawlowsky2001geometric}, while grounding the analysis firmly in tree structure. |
|
|
| The absence of such a coordinate system extends beyond traditional CoDA to probabilistic model representations. Softmax outputs are simplex-valued but often represented in flat probability coordinates, while structure over outcomes is increasingly explicit: class taxonomies, semantic hierarchies, grammars, tree-structured search spaces \citep{silla2011survey}. The absence of a canonical tree-aligned coordinate system beyond heuristics limits analysis of where probability mass, errors, or learning signals concentrate. |
|
|
| \textbf{Contributions.} We (1) construct \textbf{PolyILR} (Polytomous ILR), a canonical orthonormal decomposition of the Aitchison tangent space aligned with arbitrary trees, answering affirmatively the question raised above, (2) demonstrate its utility in CoDA; stable feature selection and tree-level inference in standard microbiome and single-cell datasets, (3) establish a novel theoretical connection to softmax classifiers via shared invariance structure. {\color{black}Our goal is interpretability of the representation itself, not downstream performance: each coordinate corresponds to a specific tree location (i.e., a node-contrast pair) that yields consistent, tree-grounded analysis and inference (see Table~\ref{table:tree-subparts}).} {\color{black} Code is available at \url{https://github.com/vsingh-group/polyilr}.} |
|
|
| {\color{black} |
| \textbf{Conflict of Interest Disclosure.} The authors declare no financial conflicts of interest related to this work. |
| } |
|
|