hv-manifold
The shape of a corpus in style space.
Given a corpus of texts, embed each text into a style vector, then compute the geometry of the resulting point cloud: its intrinsic dimension, its curvature, its connectivity, and its boundary.
The claim in one sentence
Every other hv- model reads a text. hv-manifold reads a corpus:
a single geometric signature that describes the shape of an entire
collection in style space.
What it produces
A ManifoldReport containing:
- pr_dim β participation-ratio dimension (continuous intrinsic dim)
- pca_dim β eigenvalues > 1 (Kaiser rule)
- curvature β clustered (positive) or sprawling (negative)
- beta_1_at_connection β cycle rank when the corpus first connects
- n_clusters β connected components at the median edge length
- cluster_sizes β size of each cluster
- per_cluster_dim β per-cluster intrinsic dimension
- boundary_points β outliers on the edge of the manifold
- log_volume β log-determinant spread of the covariance
Plus per-text detail: curvature, local density, 2D projection, and component label.
Plus a global style label: bimodal, clustered, connected,
uniform-clustered, core-periphery, fragmented, or
loosely-clustered.
Install
pip install numpy
Actually β no dependencies. Pure stdlib.
## Usage
### Demo
```bash
python hv_manifold.py
Runs a 30-text synthetic corpus in five groups β casual, academic,
technical, poetic, repetitive β and prints the full geometric report.
### Analyze a directory
```bash
python hv_manifold.py --dir ./corpus
python hv_manifold.py --dir ./corpus --glob "*.txt" --limit 100
python hv_manifold.py --dir ./corpus --json
python hv_manifold.py --dir ./corpus --save-to ./model
### Python
```python
from hv_manifold import HVManifold
m = HVManifold()
report = m.analyze(texts, labels=labels)
print(report.pr_dim) # 8.474
print(report.pca_dim) # 5
print(report.curvature_mean) # -0.400
print(report.n_clusters) # 2
print(report.style) # "bimodal"
print(report.summary)
The ten style axes
Each text is embedded into a 10-dimensional style vector, then standardized across the corpus.
| axis | what it measures |
|---|---|
lexical_diversity |
types / tokens |
rare_rate |
fraction of long, uncommon words |
hedge_rate |
may, might, perhaps, probably |
certainty_rate |
definitely, certainly, proven |
negation_rate |
not, never, without |
passive_rate |
passive constructions per sentence |
clause_rate |
commas + subordinators per word |
mean_sentence_length |
words / sentence |
drift |
1 β Jaccard between first and last sentence |
composition_depth |
max nesting depth (subordinators, brackets) |
The manifold is the shape of these 10 vectors after standardization.
The geometry
pr_dim β participation-ratio dimension
Continuous intrinsic dimension. Reflects how many eigenvalues carry non-trivial variance:
pr_dim = (Σλ)² / Σλ²
A corpus spread evenly across 10 axes has pr_dim β 10. A corpus
collapsing onto 2 axes has pr_dim β 2. The demo corpus reports 8.47.
pca_dim β Kaiser rule
Count of eigenvalues greater than 1 after standardization. The demo reports 5.
curvature β clustered or sprawling
Per-text local density (mean k-NN distance) versus the median:
curvature_i = (median β density_i) / median
- Positive mean β clustered. Texts have close neighbours; the corpus has centres of gravity.
- Negative mean β sprawling. Texts sit on the edge of the manifold; the corpus has no centre.
- High std β mixed. Some regions dense, others sparse.
The demo reports mean = β0.400, std = 1.045.
beta_1_at_connection β cycle rank
Cycle rank of the graph at the moment it first becomes connected, built by union-find over edges sorted by distance. Zero means the corpus connects as a tree. Positive means there are redundant paths.
n_clusters and cluster_sizes
Connected components at the median edge length. A robust threshold that does not require a bandwidth parameter.
The demo reports 2 clusters (sizes 29, 1) β one giant component and
one outlier.
per_cluster_dim
Intrinsic dimension computed inside each cluster. Cluster 0 has 7.31; cluster 1 has 0.00 (a singleton has no dimension).
boundary_points
Texts whose local density exceeds boundary_factor Γ median. These are
the texts on the edge of the manifold β the outliers, the experiments,
the ones that do not belong to any cluster.
log_volume
Log-determinant of the covariance eigenvalues. A monotone measure of spread. The demo reports β0.879.
The style label
A single string summarizing the manifold's shape:
| label | condition |
|---|---|
connected |
one component, near-zero curvature |
uniform-clustered |
one component, curvature > 0.10 |
core-periphery |
one component, curvature < β0.10 |
bimodal |
exactly two components |
clustered |
three to five components |
loosely-clustered |
six to n/4 components |
fragmented |
more than max(6, n/4) components |
The demo corpus is bimodal: 29 texts in one mass, one text alone.
Benchmarks
The demo corpus
30 synthetic texts across five groups:
| group | count | character |
|---|---|---|
| C β casual | 6 | short, hedged, everyday |
| A β academic | 6 | long, subordinated, hedged |
| T β technical | 6 | numeric, dense, clause-heavy |
| P β poetic | 6 | long sentences, high drift, low clause rate |
| R β repetitive | 6 | short sentences, high lexical repetition |
Reported geometry:
| metric | value |
|---|---|
| PR dim | 8.474 |
| PCA dim | 5 |
| curvature mean | β0.400 |
| curvature std | 1.045 |
| clusters @ median | 2 |
| cluster sizes | 29, 1 |
| local dims | 7.31, 0.00 |
| Ξ²β @ connect | 0 |
| connectivity Ξ΅ | 5.620 |
| boundary points | 8 |
| log-volume | β0.879 |
| style | bimodal |
Reading the report:
- PR dim 8.47 β the corpus spreads across nearly the full 10-dimensional style space. It is not low-dimensional.
- PCA dim 5 β but only 5 axes carry above-unit variance. Five style directions explain the corpus.
- Curvature β0.40 β sprawling. The corpus does not have a centre.
- 2 clusters (29, 1) β one text is geometrically isolated at the
median edge length. That text is
Pβ the poetic outlier the classifier flagged. - 8 boundary points β the two groups at the edges (
CandA) contribute most of them. Casual and academic texts are the extremes of this corpus in style space. - Style: bimodal β the honest label. This corpus is not one thing.
Why the geometry is not the sum of its parts
A corpus can be high-dimensional (pr_dim large) but low-rank
(pca_dim small). A corpus can be clustered (positive curvature) but
one single component (n_clusters = 1). The manifold is the joint
description of dimension, curvature, connectivity, and boundary. Every
single-axis summary collapses one of these. The report is the artifact.
When to use it
- Corpus diagnostics. "Is my training corpus one thing or five?"
- Outlier detection.
boundary_pointsnames the texts on the edge. - Cluster triage. Before running k-means, ask how many clusters the geometry actually supports.
- Style drift monitoring. Track
pr_dimand curvature over time as a corpus grows. - Dataset cards. Report the manifold as a fingerprint of the corpus.
- Retrieval sanity checks. If a corpus is
bimodal, one retrieval index may not serve it.
When not to use it
- For fewer than ~15 texts. The Jacobi eigensolver and k-NN density are unstable on tiny corpora.
- As a ground-truth clustering. The clusters are at the median edge length β a robust default, not the answer.
- For non-English corpora. The style lexicons are English.
- For very long texts (whole books). Style vectors saturate; the manifold flattens.
- For corpora with fewer than 3 features. PR dim and curvature both degenerate.
- As a replacement for reading. The manifold tells you the shape. It does not tell you what the corpus says.
Honest limitations
- The style axes are hand-designed. Not learned. They are one reasonable basis, not the only one.
- The manifold is basis-dependent. Swap the 10 features and the
geometry changes.
hv-manifoldreports the shape in this basis. - PCA dim depends on standardization. The Kaiser rule (eigen > 1) assumes unit-variance features. Correlated features can hide here.
- Curvature is a k-NN density proxy. It is not Riemannian curvature. It is a signal of clustered vs. sprawling, nothing more.
- Ξ²β at connection uses union-find only. It is a cycle-rank proxy, not a full persistent-homology computation. Loops that close after the first connection are not counted.
- The style label is thresholded.
bimodalandclusteredsit on either side of an integer. Small corpora will flip between them. - Boundary points are density outliers, not semantic ones. A text can be a boundary point and still be completely on-topic.
- No calibration against human judgment. The demo is synthetic. The labels have not been validated against expert corpus annotation.
- The demo corpus is designed to be separable. Its geometry looks
clean because the five groups are genuinely different. Real corpora
will report
connectedandloosely-clusteredmore often.
Reference
Part of the corpus-model series. Sits above the text-level models:
| model | reads |
|---|---|
hv-tempo |
one text's pace |
hv-forget |
one text's memory |
hv-fold |
one text's passes |
hv-slip |
one text's slip |
hv-wall |
one text's wall |
hv-reader |
one text's reading profile |
hv-manifold |
a whole corpus's shape |
hv-manifold is orthogonal to the reading models. It does not score
any single text. It describes how a collection of texts sits together
in style space.
License
Apache-2.0
- Downloads last month
- 8