datacheck-concepts v1
This model tells what a column of a research data file measures. It reads the column's name, its
file, its neighbouring columns, summary statistics and up to 200 of its values, and returns one of
12 concepts. pytacheck's data_check uses it
offline to fill in the concepts that
metacheck's rules leave blank. metacheck asks an LLM
for those columns; without an LLM they stay blank.
pip install "pytacheck[concepts]"
With that extra installed, data_check downloads this model once (tag v1, ~840 MB) and uses it
by default. concepts="cascade" sends the columns it is least sure of (confidence below 0.92) to
the configured LLM, and concepts="llm" is metacheck's behaviour.
Labels
reaction_time, accuracy, age, gender, race, likert, condition, id, date,
timestamp, measure, other, in the order of concept_model.json. measure (some other
measured variable) and other mean "no concept": pytacheck leaves those columns blank, as it does
when the LLM gives these answers.
Input
The model was trained on text in exactly this form (pytacheck's concept_text() builds it):
column: {name} | file: {basename} | neighbours: {3 before, 2 after} | n=.. missing=.. unique=.. min=.. max=.. mean=.. sd=.. type={numeric|text|...} | values: {k} distinct in sample: {v1} ; {v2} ; ...
- The values are up to 200 evenly spaced non-missing values, each cut to 200 characters. The first 40 distinct ones are listed, cut to 40 characters.
- Statistics are printed to 4 significant digits.
- The input is truncated at 256 tokens.
Files
model.onnx: inputs areinput_idsandattention_mask(int64, batch × tokens). The outputprobsholds the softmax probabilities (float32, batch × 12).- The fp32 export (opset 18) has its matmuls quantised weight-only: MatMulNBits, 8 bits, symmetric
blocks of 32,
accuracy_level1, so the arithmetic stays fp32. - The embedding tables are stored as fp16 and cast back to fp32 after the lookup.
- Against the PyTorch fp32 model it agrees on the top label for 320/320 reference columns (mean absolute difference of the top probability 0.0023). Its confidences are therefore usable for the cascade threshold.
- onnxruntime's dynamic int8 quantisation was not used: XLM-R's outlier activations break it, and it cost 0.023 F1 in the cascade.
- The fp32 export (opset 18) has its matmuls quantised weight-only: MatMulNBits, 8 bits, symmetric
blocks of 32,
tokenizer.json: XLM-RoBERTa's tokenizer, for thetokenizerslibrary.concept_model.json: labels,max_len,pad_id, the default cascade threshold, the quantisation and the training settings.
On CPU it takes about 0.1 s per column with 8 threads and about 2 GB of memory.
Training
- Base model: FacebookAI/xlm-roberta-large, fine-tuned for sequence classification.
- Settings: 2 epochs, learning rate 1.5e-5, batch size 32, max length 256, seed 0, bf16, on one GPU (10 minutes).
- Data: 107,643 columns of public research data files. The sources are ResearchBox boxes and OSF,
GitHub and Zenodo repositories linked from psychology papers. They were collected by running
metacheck's
data_checkwith its LLM calls recorded instead of sent. - Labels: Muse Spark 1.3 Contributor, at reasoning effort xhigh, with the labelling prompt. For the 2,145 benchmark columns below, the benchmark's gold labels replace Muse's.
Evaluation
The same recipe was scored without the evaluation repositories:
- RB test: 98 ResearchBox boxes, 961 columns, 5-fold cross-fit by box.
- OOD: 86 OSF/GitHub/Zenodo repositories, 812 columns, none of them in training.
The F1 combines the precision of the concepts written with the recall of the columns that have
one. measure and other count as blank, and each repository is weighted equally.
| system | RB test | OOD |
|---|---|---|
| this model (the recipe, trained without the evaluation repositories) | 0.795 | 0.736-0.761 (two runs) |
| Muse Spark 1.3 Contributor, metacheck's prompt | 0.784 | 0.740 |
| Gemini 3.5 Flash-Lite, metacheck's prompt | 0.762 | 0.694 |
| gpt-oss-20b, metacheck's prompt | 0.613 | 0.515 |
| this model, then Muse Spark below confidence 0.92 (~37% of columns sent) | 0.852 | 0.786-0.801 |
On its own the model is level with Muse Spark: the differences (+0.011 RB, +0.021 OOD) are within their repository-clustered 95% confidence intervals. What it adds is that it runs offline, costs nothing to run, and gives the same answer every time.
Limitations
- Most training columns come from English-language psychology data. The base model is multilingual, but other languages and fields are less tested.
- The model leans on column names. For example, a lowercase
rtcolumn next tosubject, block, trialcan come out asid. metacheck's rules label such columns before the model sees them. - The labels are an LLM's (Muse Spark), not people's, and the model inherits its mistakes.
- The input contains values from the data file. The model runs locally, and nothing is sent anywhere.
License
AGPL-3.0-or-later, as pytacheck and metacheck. The base model, XLM-RoBERTa-large, is MIT-licensed.
Model tree for scienceverse/datacheck-concepts
Base model
FacebookAI/xlm-roberta-large