datacheck-concepts v1

This model tells what a column of a research data file measures. It reads the column's name, its file, its neighbouring columns, summary statistics and up to 200 of its values, and returns one of 12 concepts. pytacheck's data_check uses it offline to fill in the concepts that metacheck's rules leave blank. metacheck asks an LLM for those columns; without an LLM they stay blank.

pip install "pytacheck[concepts]"

With that extra installed, data_check downloads this model once (tag v1, ~840 MB) and uses it by default. concepts="cascade" sends the columns it is least sure of (confidence below 0.92) to the configured LLM, and concepts="llm" is metacheck's behaviour.

Labels

reaction_time, accuracy, age, gender, race, likert, condition, id, date, timestamp, measure, other, in the order of concept_model.json. measure (some other measured variable) and other mean "no concept": pytacheck leaves those columns blank, as it does when the LLM gives these answers.

Input

The model was trained on text in exactly this form (pytacheck's concept_text() builds it):

column: {name} | file: {basename} | neighbours: {3 before, 2 after} | n=.. missing=.. unique=.. min=.. max=.. mean=.. sd=.. type={numeric|text|...} | values: {k} distinct in sample: {v1} ; {v2} ; ...
  • The values are up to 200 evenly spaced non-missing values, each cut to 200 characters. The first 40 distinct ones are listed, cut to 40 characters.
  • Statistics are printed to 4 significant digits.
  • The input is truncated at 256 tokens.

Files

  • model.onnx: inputs are input_ids and attention_mask (int64, batch × tokens). The output probs holds the softmax probabilities (float32, batch × 12).
    • The fp32 export (opset 18) has its matmuls quantised weight-only: MatMulNBits, 8 bits, symmetric blocks of 32, accuracy_level 1, so the arithmetic stays fp32.
    • The embedding tables are stored as fp16 and cast back to fp32 after the lookup.
    • Against the PyTorch fp32 model it agrees on the top label for 320/320 reference columns (mean absolute difference of the top probability 0.0023). Its confidences are therefore usable for the cascade threshold.
    • onnxruntime's dynamic int8 quantisation was not used: XLM-R's outlier activations break it, and it cost 0.023 F1 in the cascade.
  • tokenizer.json: XLM-RoBERTa's tokenizer, for the tokenizers library.
  • concept_model.json: labels, max_len, pad_id, the default cascade threshold, the quantisation and the training settings.

On CPU it takes about 0.1 s per column with 8 threads and about 2 GB of memory.

Training

  • Base model: FacebookAI/xlm-roberta-large, fine-tuned for sequence classification.
  • Settings: 2 epochs, learning rate 1.5e-5, batch size 32, max length 256, seed 0, bf16, on one GPU (10 minutes).
  • Data: 107,643 columns of public research data files. The sources are ResearchBox boxes and OSF, GitHub and Zenodo repositories linked from psychology papers. They were collected by running metacheck's data_check with its LLM calls recorded instead of sent.
  • Labels: Muse Spark 1.3 Contributor, at reasoning effort xhigh, with the labelling prompt. For the 2,145 benchmark columns below, the benchmark's gold labels replace Muse's.

Evaluation

The same recipe was scored without the evaluation repositories:

  • RB test: 98 ResearchBox boxes, 961 columns, 5-fold cross-fit by box.
  • OOD: 86 OSF/GitHub/Zenodo repositories, 812 columns, none of them in training.

The F1 combines the precision of the concepts written with the recall of the columns that have one. measure and other count as blank, and each repository is weighted equally.

system RB test OOD
this model (the recipe, trained without the evaluation repositories) 0.795 0.736-0.761 (two runs)
Muse Spark 1.3 Contributor, metacheck's prompt 0.784 0.740
Gemini 3.5 Flash-Lite, metacheck's prompt 0.762 0.694
gpt-oss-20b, metacheck's prompt 0.613 0.515
this model, then Muse Spark below confidence 0.92 (~37% of columns sent) 0.852 0.786-0.801

On its own the model is level with Muse Spark: the differences (+0.011 RB, +0.021 OOD) are within their repository-clustered 95% confidence intervals. What it adds is that it runs offline, costs nothing to run, and gives the same answer every time.

Limitations

  • Most training columns come from English-language psychology data. The base model is multilingual, but other languages and fields are less tested.
  • The model leans on column names. For example, a lowercase rt column next to subject, block, trial can come out as id. metacheck's rules label such columns before the model sees them.
  • The labels are an LLM's (Muse Spark), not people's, and the model inherits its mistakes.
  • The input contains values from the data file. The model runs locally, and nothing is sent anywhere.

License

AGPL-3.0-or-later, as pytacheck and metacheck. The base model, XLM-RoBERTa-large, is MIT-licensed.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for scienceverse/datacheck-concepts

Quantized
(11)
this model