System One fraud detector for FDB ccfraud: Credit Card Fraud Detection (ULB / Worldline, 2013)

This model detects fraudulent credit card transactions. It is a two-stage stack: a gradient-boosted tree model scores each transaction from engineered features, and System One, a typed-decision model (Laya ModernBERT-large encoder with a LoRA adapter and a decision head), reads that score, a few readable engineered fields and the original fields (as many as fit in its token window) as JSON evidence, and returns a calibrated probability that the transaction is fraudulent. The released output is a logistic blend of System One's probability and the tree score, fitted on System One's holdout split. On this set the blend gives System One a negative weight, so the released output mostly follows the tree and partly undoes System One's correction.

On FDB's official test split (56,962 transactions, 75 fraudulent) it scores 0.9867 AUROC and catches 86.7% of fraud at a 1% false-positive rate, against the best published 0.992 AUROC (H2O) and 88.0% recall (a three-way tie of AFD OFI, AFD TFI and AutoGluon). It trails the published AUROC and recall, but the published values lie inside this model's 95% confidence intervals.

It was trained for the Fraud Dataset Benchmark (FDB) and is a benchmark model, not a production fraud system; read Limitations first.

Results

Official FDB test split, scored once per model. AUROC and recall at 1% FPR follow FDB's own evaluation (recall is np.interp(0.01, fpr, tpr) on the test ROC). 95% confidence intervals: class-stratified bootstrap, 1,000 resamples.

Model AUROC [95% CI] Recall at 1% FPR [95% CI] Average precision Recall at 0.1% FPR
This model: blend of System One and the tree 0.9867 [0.9748, 0.9961] 86.7% [78.7, 93.3] 0.8079 80.0%
System One alone 0.9855 [0.9717, 0.9961] 85.3% [77.3, 93.3] 0.8075 80.0%
Tree alone (System One's input) 0.9866 [0.9747, 0.9961] 86.7% [77.3, 93.3] 0.8070 80.0%
Best published AUROC: H2O 0.992
Best published recall at 1% FPR: a three-way tie of AFD OFI, AFD TFI and AutoGluon 88.0%

AUROC leaders are from the FDB paper (arXiv 2208.14417 v3). The paper reports no recall, so recall leaders are from the results table in FDB's README at commit 54cdefa211; later versions of the README keep that table inside an HTML comment, so it no longer renders. A later, commented-out version of the table also lists Auto-sklearn at 88.0%. With 75 fraudulent transactions in the test set, each one is worth 1.3 points of recall.

How the released configuration was chosen. All models were trained and calibrated without test data, and each test set was scored once per candidate. After the test results were known, one model per benchmark was released: System One on every set (for a uniform, calibrated decision interface), and for this set the blend of System One and the tree as the default output, because it scored highest on the test set; its margin over System One alone (one more fraud caught, +0.001 AUROC) is within noise (paired 95% intervals: AUROC -0.0003 to +0.0031, recall 0 to +4.0 points). Choosing among candidates after seeing test results can flatter the headline slightly, so the table shows every candidate.

Model details

  • Developed by: STEAV (steav.io), trained with CID Model Studio.
  • Model type: stacked binary classifier: gradient-boosted trees, then System One with offset stacking (System One adds a learned correction to the tree's log-odds), then a logistic blend of the two.
  • System One base: convaiinnovations/laya at revision 7b928d828b7b (Apache-2.0), a ModernBERT-large encoder (about 395M parameters) with Laya's typed decision head. The base weights are downloaded from the Hub at that pinned revision and checked against their sha256; they are not redistributed here.
  • Trained parameters: a LoRA adapter (rank 16, alpha 32, on the attention Wqkv/Wo and MLP Wo projections of all 28 layers; 4.39M parameters) and the decision head (26.2M parameters, initialised from Laya's head; its 0.26M-parameter action sub-head stays frozen).
  • Question: "Is this card transaction fraudulent, meaning it was not made by the legitimate cardholder?" (a yes/no decision; the output is P(yes)).
  • License: Apache-2.0 for the weights and code; some data files carry their own terms (see License and attribution).
  • Version: 1.0.0.
  • Contact: questions and issues go to this repository's Community tab.

Uses

Intended: benchmark research and comparison on FDB; a reference implementation of stacking a typed-decision model on a tree model's score; a starting point for fine-tuning on your own data.

Out of scope: declining real cardholders' transactions; any automated decision about a person without human review; use on data from a different distribution without retraining and validation.

How to use

Python 3.12 with the pinned versions in requirements.txt. On the reference GPU (NVIDIA GB10 in a DGX Spark, torch 2.11.0, bfloat16) the code reproduced the evaluated scores bit for bit; other GPUs and CPUs give slightly different scores (see Reproducibility). On CPU (Apple M4 Max, macOS arm64, float32) System One takes about 0.12 s per transaction, about 1.9 h for the 56,962-row test set.

hf download getSTEAV/system-one-fdb-ccfraud --local-dir system-one-fdb-ccfraud
cd system-one-fdb-ccfraud && pip install -r requirements.txt
# run from the repository folder
from system_one_fraud import FraudPipeline
from system_one_fraud.fdb import load_split

pipe = FraudPipeline.from_pretrained(".")      # the tree, the adapter and the pinned Laya base
train, test, labels = load_split("ccfraud")       # FDB's loader (needs Kaggle API credentials)
p_fraud = pipe.predict_proba(test)              # one probability per row

predict_proba returns, per row, the probability that the transaction is fraudulent. The default output is the blend; output="system_one" or output="tree" returns one stage alone. System One's own output is shifted from its rebalanced training sample to the natural fraud rate (0.183%). states, tree_out = pipe.states(rows) returns the exact JSON evidence System One reads and the tree model's outputs. examples/quickstart.py runs the same steps on a few test rows.

The adapter does not load with PeftModel.from_pretrained alone: Laya's root config.json is a stub that transformers cannot instantiate, and the decision head and evidence format live in the bundled system_one_fraud package. On macOS the tree model runs in a helper Python process, because PyTorch and the gradient-boosting libraries load two OpenMP runtimes that crash in one process; results are identical. System One's bfloat16 GPU scores depend slightly on how rows are batched (padding); the reference scores used batch size 128 in file order.

Input contract. System One reads the FDB columns v1-v28 and amount; the tree also reads EVENT_TIMESTAMP (FDB encodes the original Time in seconds as 2021-09-01T00:00Z plus that many minutes; the hour of day is recovered from it). Every other column is ignored, including FDB's randomly generated EVENT_ID and ENTITY_ID, its constant ENTITY_TYPE and its load-time LABEL_TIMESTAMP; a missing field raises an error.

Training data

  • Source: 284,807 card transactions made by European cardholders over two days in September 2013, 492 of them fraudulent, with 28 PCA-anonymised features (V1-V28), the amount and the time; published by the Machine Learning Group of the Universite Libre de Bruxelles with Worldline (Kaggle mlg-ulb/creditcardfraud, version 3).
  • Benchmark split: out of time: FDB trains on the first 80% of transactions by time and tests on the last 20%; the test fraud rate (0.13%) is lower than in training (0.18%). Training: 227,845 transactions (417 fraudulent, 0.18%). Test: 56,962 transactions (75 fraudulent).
  • Licence: The database is under the Open Database License (ODbL) v1.0 and its contents under the Database Contents License (DbCL) v1.0; commercial use is permitted. The trained weights are a Produced Work under the ODbL, so the weights and code here are Apache-2.0. Four sorted arrays extracted from the database (tree/oof_margin_ref.npy, tree/amount_ref.npy, tree/pca_norm_ref.npy, system_one/percentile_reference.npy) and the row lists in derivation/ are made available under the ODbL v1.0. As ODbL section 4.6 asks, derivation/ describes how the training data was derived (with FDB's loader and the feature code in this repository). No transaction rows are included.

Training procedure

Stage 1, tree model. A bag of 15 XGBoost models (3 seeds x 5 cross-validation folds; depth 5, learning rate 0.03, fraud weight 10, 1,200-1,300 trees chosen on pooled out-of-fold predictions) over V1-V28, the amount and the hour of day recovered from the time. The averaged margin is converted to a probability by its rank among the 227,845 out-of-fold training margins followed by an isotonic calibration map. Because the official test is out of time, configurations were compared on out-of-time blocks of the training data under pre-declared rules: fit on the first 80% of training rows by time and validate on the last 20% (52 frauds), with a second round that also used the 60-80% block. About 20 single configurations and 40 blends were scored on these blocks, so no untouched validation data remains. On the last-20% block the released recipe averaged 0.9853 AUROC and 84.0% recall at 1% FPR over three seeds.

Training rows carry out-of-fold tree scores: the models that scored a training row never saw it, so System One never learns from a score fit on its own row. Only the final calibration map (monotone, 26 knots) was fitted on all training rows' out-of-fold scores and labels.

Stage 2, System One. Each training row is rendered as a JSON state: the tree's out-of-fold score (Tree_fraud_score), its percentile among all out-of-fold training scores (Tree_percentile), readable engineered fields (Context_*) and the original fields. Raw values are rounded to 4 significant digits in the evidence. System One trained on a 100,000-row sample: all 417 training frauds plus 99,583 random legitimate transactions (the rows are listed in derivation/); CID Model Studio removed 57 rows whose evidence duplicated another row, leaving 99,943 rows with 413 frauds. System One's own output is shifted from the sample's fraud rate (0.417%) to the natural 0.183%. The rows were split 79,954 / 9,994 / 9,995 (train / validation / holdout). System One reads at most 256 tokens and cuts the end of the evidence (keys are sorted) when it does not fit: on the official test set this happens on every row. On this set v3-v9 and v25-v28 never reach the model, and v22, v23 and v24 reach it on only 96%, 28% and 0.4% of rows; the tree's own fields, computed from every input, always fit.

  • LoRA r = 16, alpha = 32, dropout 0.05; learning rate 1e-4 (head 5e-5), weight decay 0.01, 5% warm-up, batch 32; bfloat16; maximum sequence length 256 tokens; seed 42.
  • one supervised epoch with soft cross-entropy, then one RLCD epoch (reinforcement learning on the decision distribution with proper-scoring-rule rewards, following Laya's method).
  • 4,998 optimiser steps, 40.9M training tokens, 2.4 h on one NVIDIA GB10 (DGX Spark).
  • Calibration: one temperature fitted on the validation split (L-BFGS).
  • Validation: accuracy 0.9993, expected calibration error 0.00021.

Stage 3, blend. A logistic regression fitted on System One's holdout split (9,995 rows) over the log-odds of System One's probability and of the tree score as System One reads it: weight -0.237 on System One, +1.279 on the tree, intercept +0.513. Its output is used as is, without a prior shift.

Evaluation

  • Protocol. FDB's official split; the test set was never used for training, feature selection or calibration, and labels were read only to compute the metrics.
  • Calibration on test (10 equal-width bins, released output): expected calibration error 0.00015, Brier score 0.000376.
  • Cost of high detection (released output): catching 95% of fraud requires a false-positive rate of 9.84%, and catching 99% requires 29.63%.
  • Reproduce: python eval/reproduce_fdb.py rebuilds the split with FDB's own loader, checks it against the frames used here, scores it with the released pipeline and compares the metrics with eval/test_metrics.json.

Limitations

  • Anonymised features. V1-V28 are PCA components with no published meaning, and the data is two days of 2013 European transactions. The model cannot be applied to other card data without retraining on the same PCA projection, which is not public.
  • Out-of-time drift. The test period follows training directly and has a lower fraud rate; longer gaps will cost more.
  • System One adds nothing here. Its evidence is mostly numbers it cannot interpret, and its token window never shows it 11 of the 28 PCA fields and shows three more only sometimes (see Training procedure). The blend fitted on its holdout split gives it a negative weight, so the released output follows the tree; System One stays in the pipeline for a uniform decision interface.

Bias, risks and ethical considerations

Fraud models make errors in both directions: false positives block or delay legitimate customers, and their burden can fall unevenly across groups. No fairness evaluation was done for this model; the benchmark data carries no protected attributes we could audit. Use the score as one input to a reviewed decision, monitor error rates on your own population, and retrain when the data drifts.

Reproducibility

These checks passed before release:

  • Tree model. Refit from scratch, the tree model in tree/ reproduces the evaluated tree scores bit for bit on all 56,962 test rows (Apple M4 Max, macOS arm64, the pinned versions). On other platforms (other BLAS or OpenMP builds) the last bits can differ, which can occasionally change the Tree_fraud_score text System One reads.
  • Evidence. The state builder reproduces the evaluated System One inputs exactly on all 56,962 test rows.
  • System One. The inference code in system_one_fraud/ reproduces the evaluated scores bit for bit on all 56,962 test rows on the reference GPU (NVIDIA GB10 in a DGX Spark, torch 2.11.0, bfloat16, batch size 128 in file order, base downloaded from the Hub), run in the training environment with the same pinned versions rather than a fresh install.
  • Fresh install. A new virtual environment built from requirements.txt (Apple M4 Max, macOS arm64, pandas 3.0.6, torch 2.11.0) reproduces the tree scores bit for bit and the System One inputs exactly.
  • Fresh download. FDB's loader, run fresh with the pinned pandas, rebuilds the evaluated frames exactly (content hashes in eval/fdb_split_hashes.json) apart from the identifier and timestamp columns FDB randomises on every load, which the models never read.
  • End to end on CPU. eval/reproduce_fdb.py (fresh download, released pipeline, Apple M4 Max, macOS arm64, float32) gives 0.9867 AUROC and 86.7% recall at 1% FPR, against the reported 0.9867 and 86.7% (identical to 4 decimals).

Measured, not gated: CPU drift. On CPU in float32, System One's scores differ slightly from the bfloat16 GPU reference. On 287 holdout transactions (31 fraudulent) the mean absolute difference in probability is 4.8e-06 and the largest 3.0e-04; decisions at 0.5 agree on 100.0% of rows, and AUROC on that sample is 0.9971 on the GPU and 0.9972 on CPU.

Environmental impact

System One training used one NVIDIA GB10 (DGX Spark, on premises) for 2.4 h, roughly 0.6 kWh at the system's 240 W rating (an upper-bound estimate). The tree models trained in minutes on a laptop CPU.

Files

Path Contents
adapter_model.safetensors, adapter_config.json LoRA adapter
system_one/head.safetensors decision head
system_one/config.json, system_one/calibration.json, release.json head configuration, temperature, release manifest (question, pinned base, priors, blend weights)
system_one/percentile_reference.npy sorted out-of-fold training tree scores, for the Tree_percentile evidence field
tree/ 15 XGBoost boosters (JSON, sliced to the trees used), the calibration map and feature constants (JSON), three sorted training reference distributions (.npy) and manifest.json with every file's sha256 (checked on load)
system_one_fraud/ inference code: pipeline.py (end to end), state.py (JSON evidence), system_one.py (System One), tree_ccfraud.py (tree model), tree_runner.py (macOS helper process), fdb.py (FDB loader)
eval/ reproduce_fdb.py, the split hashes and the test metrics
examples/quickstart.py load, score and print a few test rows
derivation/ how the training data was derived from the source database (ODbL section 4.6)
LICENSE, NOTICE, LICENSES/ licence texts and attributions
requirements.txt pinned dependencies
checksums.sha256 sha256 of every file except this README

Embedded data. tree/ holds model statistics derived from the training data: XGBoost split thresholds and node covers, the sorted out-of-fold training margins (for the rank calibration), the sorted training amounts and a sorted one-number summary of each training row's V1-V28 vector (for two readable percentile fields), and a 26-knot isotonic map. system_one/percentile_reference.npy holds the sorted out-of-fold training tree scores. None of these contains rows, identifiers or labels. Because the four sorted arrays could be read as extractions from the source database, they are offered under the ODbL v1.0 (see License and attribution).

Citation

@misc{steav2026systemonefdb,
  title        = {System One fraud detectors for the Fraud Dataset Benchmark},
  author       = {{STEAV}},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/getSTEAV/system-one-fdb-ccfraud}}
}

@misc{grover2023fraud,
  title         = {Fraud Dataset Benchmark and Applications},
  author        = {Prince Grover and Julia Xu and Justin Tittelfitz and Anqi Cheng and Zheng Li and Jakub Zablocki and Jianbo Liu and Hao Zhou},
  year          = {2023},
  eprint        = {2208.14417},
  archivePrefix = {arXiv},
  primaryClass  = {cs.LG}
}

@misc{convai2026laya,
  title        = {Laya: a non-autoregressive System 1 decision model},
  author       = {{Convai Innovations}},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/convaiinnovations/laya}},
  note         = {Revision 7b928d828b7b0e022f929d9bd2e44165aa270148}
}

@misc{warner2024modernbert,
  title         = {Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference},
  author        = {Benjamin Warner and Antoine Chaffin and Benjamin Clavi{\'e} and Orion Weller and Oskar Hallstr{\"o}m and Said Taghadouini and Alexis Gallagher and Raja Biswas and Faisal Ladhak and Tom Aarsen and Nathan Cooper and Griffin Adams and Jeremy Howard and Iacopo Poli},
  year          = {2024},
  eprint        = {2412.13663},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL}
}

@inproceedings{dalpozzolo2015calibrating,
  title     = {Calibrating Probability with Undersampling for Unbalanced Classification},
  author    = {Dal Pozzolo, Andrea and Caelen, Olivier and Johnson, Reid A. and Bontempi, Gianluca},
  booktitle = {2015 IEEE Symposium Series on Computational Intelligence},
  pages     = {159--166},
  year      = {2015}
}

License and attribution

Files Terms
tree/oof_margin_ref.npy, tree/amount_ref.npy, tree/pca_norm_ref.npy, system_one/percentile_reference.npy, derivation/*.csv Open Database License v1.0 (LICENSES/ODbL-1.0.txt): extracted from the Credit Card Fraud Detection database
everything else (weights, code, configuration, model files) Apache-2.0 (LICENSE)

Attributions are collected in NOTICE.

  • Data: Contains information from the Credit Card Fraud Detection dataset (Machine Learning Group, Universite Libre de Bruxelles, and Worldline; kaggle.com/datasets/mlg-ulb/creditcardfraud), which is made available here under the Open Database License (ODbL) v1.0. Please cite: Andrea Dal Pozzolo, Olivier Caelen, Reid A. Johnson and Gianluca Bontempi, "Calibrating Probability with Undersampling for Unbalanced Classification", IEEE Symposium Series on Computational Intelligence and Data Mining (CIDM), 2015. The files tree/oof_margin_ref.npy, tree/amount_ref.npy, tree/pca_norm_ref.npy, system_one/percentile_reference.npy and derivation/*.csv are derived from that database and are made available under the Open Database License v1.0 (LICENSES/ODbL-1.0.txt; opendatacommons.org/licenses/odbl/1-0/); derivation/README.md describes how the training data was derived. Every other file in this repository is Apache-2.0.
  • Base model: Laya by Convai Innovations (Apache-2.0), built on ModernBERT-large by Answer.AI and LightOn (Apache-2.0). The sequence layout and the decision-head architecture follow Laya's reference code.
  • Benchmark: Fraud Dataset Benchmark, Grover et al. 2023 (MIT; Copyright (c) 2021-2022 Prince Grover and Zheng Li, and (c) 2022 Jianbo Liu, Jakub Zablocki, Hao Zhou, Julia Xu and Anqi Cheng).
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for getSTEAV/system-one-fdb-ccfraud

Adapter
(23)
this model

Collection including getSTEAV/system-one-fdb-ccfraud

Papers for getSTEAV/system-one-fdb-ccfraud