excel-chunker: spreadsheet cell-role GNNs
Graph neural networks that classify every cell of a spreadsheet into one of 13 structural roles (data value, column/row headers at three nesting depths, aggregation, metadata, comment, junk, empty). The predicted roles drive structure-aware chunking of arbitrary spreadsheets for retrieval-augmented question answering.
Models from the paper “Q&A on Any Spreadsheet Requires Interpreting Its Grid Structure” (Zofia Smoleń, 2026, arXiv:2609.20732).
Models in this repository
| Folder | Architecture | Params | This checkpoint (pooled macro-F1) | Architecture mean ± SD (15 runs) | Notes |
|---|---|---|---|---|---|
gat/ |
GAT (edge-featured graph attention) | 2.18 M | 0.815 | 0.672 ± 0.091 | The architecture used end-to-end in the RAG pipeline; best end-to-end RAG quality in the paper’s open-panel evaluation |
dual_modality_gnn/ |
DualModalityGNN (structural + content streams) | 1.93 M | 0.830 | 0.749 ± 0.057 | Best cell-classification macro-F1 among all tested architectures |
Metric: 13-class macro-F1 with per-sheet confusion matrices pooled within a run (test cells of the held-out fold), the paper’s canonical metric. Each model was trained 15× (3 seeds × 5 cross-validation folds); the shipped checkpoint is the run with the highest pooled macro-F1 (GAT: seed 10042010 / fold 4; Dual: seed 10042010 / fold 1), i.e. a fold model trained on 4/5 of the corpus. Full per-run and per-class numbers are in each folder’s config.json under provenance.
Baselines from the same protocol, for context: SpatialEdgeTransformer 0.726 ± 0.049, AdjTransformer 0.718 ± 0.064, GCN 0.629 ± 0.081, MLP (no graph) 0.584 ± 0.075.
Files per model
model.safetensors— weights (safe to load, no pickle)config.json— architecture name, model/graph hyperparameters, the 13 cell labels, parameter count, and provenance (selection metric, run id, per-class F1)pytorch_checkpoint.pt— the original training checkpoint ({model_state_dict, config, cell_labels, arch}); drop-in for the excel-chunker codebase. Loading it requirestorch.load(..., weights_only=False)(pickle), hence the safetensors copy.
Label space
value, aggregation, header, metadata, comment, empty, junk, col_header_1, col_header_2, col_header_3, row_header_1, row_header_2, row_header_3
How to use
These are not standalone text models: the input is a cell graph (857-dim node features, 23-dim edge features) built from an .xlsx sheet by the excel-chunker feature pipeline. You need the accompanying code to construct the graph.
Drop-in with the excel-chunker codebase (simplest):
from pathlib import Path
from predict import load_model_from_path # excel-chunker repo
model = load_model_from_path(Path("gat/pytorch_checkpoint.pt"))
From safetensors:
import json
from safetensors.torch import load_file
from models import build_model # excel-chunker repo
cfg = json.load(open("gat/config.json"))
mc = cfg["model_config"]
model = build_model(cfg["arch"], in_dim=mc["in_dim"], hidden_dim=mc["hidden_dim"],
num_classes=mc["num_classes"], edge_dim=mc["edge_type_dim"])
model.enable_two_stage_head(mc["hidden_dim"]) # use_two_stage: true
model.load_state_dict(load_file("gat/model.safetensors"))
model.eval()
Intended use & limitations
- Research artifact accompanying the paper; intended for reproducing its results and for structure-aware spreadsheet chunking / RAG experiments.
- Each checkpoint is a cross-validation fold model (trained on 4/5 of the annotated corpus), not a model retrained on the full corpus.
- Tail classes (deeply nested headers such as
col_header_3,row_header_3) are rare in the corpus and their F1 varies strongly across folds; seeprovenance.per_class_f1_this_runin eachconfig.json. - Inputs must be produced by the same graph-construction configuration recorded in
config.json → model_config.graph_config.
Citation
@misc{smolen2026spreadsheet,
title = {Q\&A on Any Spreadsheet Requires Interpreting Its Grid Structure},
author = {Smole{\'n}, Zofia},
year = {2026},
eprint = {2609.20732},
archivePrefix = {arXiv},
primaryClass = {cs.AI},
url = {https://arxiv.org/abs/2609.20732}
}