Scope warning. This is a research model for synthetic causal inference. It is measurably unreliable on real observational studies — on IHDP it reports an ATE of
+0.05 +/- 0.12where the truth is+1.84, an error of roughly fifteen sigma — and it loses to plain correlation for structure discovery on real biological data. Read Known failures below before using it for anything consequential.
causal-1
- Version: 1.0.0
- Architecture: amortized neural structural causal model
- Parameters: 26,119,098
- Variable capacity: up to 32 variables per graph
- Inference: explicit structural simulation; no language model is involved
What it does
One shared set of weights answers observational, interventional and counterfactual queries for any DAG within its variable capacity. Graph structure is an input, not an architectural constant, so DAGs that never appeared in training are ordinary inputs.
Intended use
Research use on synthetic or well-specified systems of at most twenty
variables, where calibrated uncertainty matters: estimating
P(Y | do(X=x)), average and individual treatment effects, and unit-level
counterfactuals from a supplied or discovered DAG.
Out of scope
This model should not be used to estimate causal effects from real observational studies, and should not be used for structure discovery on real biological data. It is measurably unreliable on both, in a way that its own uncertainty does not warn you about. See Known failures below before using it for anything consequential.
Known failures
These are measured, not hypothetical, and they are the reason for the scope limits above.
- Confidently wrong on IHDP. On this standard benchmark (747 rows, 26
covariates, binary treatment) the model reports an ATE of
+0.053with an uncertainty of0.121against a true ATE of+1.842. That is an error of roughly fifteen sigma, and the two-sigma interval excludes the truth by a factor of six. A t-learner reaches an ATE error of 0.026 on the same data. Predicting exactly zero would score a PEHE of 1.884; the model scores 1.844, so it contributes almost nothing. - Beaten by correlation on Sachs. Edge AUROC 0.654 against 0.738 for ranking pairs by absolute correlation. Skeleton recovery is 0.700 against 0.788, and orientation is at chance (7 of 17 edges). Feeding more rows does not help: 0.627 at 256 rows, 0.654 at 853.
- A parent's influence is diluted by its siblings. The mechanism
aggregates parents with a softmax, so influence decays as roughly
0.26 / n_parents. A twenty-six-covariate outcome dilutes the treatment twenty-six fold, which is the leading explanation for the IHDP result. The model's predictions correlate 0.92 with the truth while having a regression slope of 0.10: it recovers the direction and relative size of an effect and then reports about a tenth of its magnitude. Three attempted fixes all measured no better. - Calibration holds in distribution only. On synthetic worlds where the graph is silent about a real effect, uncertainty is near-perfectly calibrated (error/sigma 0.96). On real data it is not, as the IHDP figure above shows. Do not rely on the uncertainty estimate as a safety margin outside the training distribution.
Other limitations
- Effects that are not identified from the supplied graph cannot be
recovered. The model reports elevated uncertainty on such queries and the
API marks them
not_identified; it does not invent a point estimate it can justify. - Discovery returns a DAG consistent with the data, which in general identifies a Markov equivalence class rather than a unique orientation.
- Accuracy degrades outside the treatment range seen in the observational
sample; the
outside_rangebenchmark quantifies this. - Effect estimation extrapolates poorly to graphs larger than training: first of seven estimators on held-out worlds, but fourth of seven on twenty-one to twenty-seven variable graphs. Discovery extrapolates fine.
- Training worlds are homogeneous single SCMs with 256 observational rows. The model was never trained to sharpen an estimate given more data, and measurably does not.
Why it fails on real data
The model is amortized: it never fits the dataset in front of it, it applies a function learned from synthetic structural causal models. A t-learner fits IHDP directly and does well. This model instead asks what effect a world of that shape would typically have, and "typically" is defined entirely by the generator. Two properties of that generator are measured and matter: training tasks whose outcome has eight or more parents average a normalised ATE of 0.639 against 1.543 overall, so the prior actively teaches that wide graphs have modest effects; and every node is calibrated to unit variance, so a treatment cannot dominate a long covariate list the way it does in a real study. Strong synthetic scores and weak real scores are therefore not in tension - nothing in the training objective ever rewarded being right off-distribution.
Held-out benchmark
| metric | held-out worlds | unseen structures |
|---|---|---|
| normalised ATE error | 0.4231 | 0.5004 |
| normalised ITE error | 0.5149 | 0.5848 |
| normalised counterfactual error | 0.2890 | 0.3433 |
| edge F1 | 0.7057 | 0.6931 |
| edge AUROC | 0.9193 | 0.9375 |
| SHD | 11.3086 | 43.6406 |
| 1-sigma coverage | 0.7070 | 0.7812 |
Baseline comparison (normalised ATE error, lower is better)
| estimator | error |
|---|---|
| proposed_model | 0.4231 |
| g_computation | 0.4304 |
| backdoor_adjustment | 0.4930 |
| naive_correlation | 0.5140 |
| propensity_weighting | 0.5330 |
| linear_regression | 0.5332 |
| neural_regression | 1.0974 |
Reproduction
python -m causal.cli train --steps 12000 --export model
python -m causal.cli benchmark --artifact model
Full benchmark record
{
"adversarial": {
"scenarios": {
"collider_bias": {
"ate_error": 0.0535808801651001,
"ate_sigma": 0.24587032198905945,
"baseline_backdoor_adjustment_error": 0.07940781116485596,
"baseline_g_computation_error": 0.0839529037475586,
"baseline_linear_regression_error": 1.0498626232147217,
"baseline_naive_correlation_error": 1.0376276969909668,
"baseline_neural_regression_error": 2.398400902748108,
"baseline_propensity_weighting_error": 0.4118959903717041,
"causal_slope": 1.1000001448268764,
"counterfactual_error": 0.07461073994636536,
"edge_auroc": 0.20000000298023224,
"edge_f1": 0.0,
"edge_precision": 0.0,
"edge_recall": 0.0,
"error_to_sigma_ratio": 0.2179233334533328,
"identified": 1.0,
"ite_error": 0.10452093183994293,
"marginal_slope": 1.6864241361618042,
"n_predicted_edges": 4.0,
"n_true_edges": 5.0,
"n_variables": 4.0,
"normalised_ate_error": 0.02874200515882636,
"normalised_ate_sigma": 0.13189044377839107,
"normalised_counterfactual_error": 0.04002290118853031,
"normalised_ite_error": 0.05606741085648537,
"reversed_edges": 4.0,
"shd": 5.0,
"true_ate": 1.9693048000335693
},
"confounding": {
"ate_error": 2.2281384468078613,
"ate_sigma": 0.1943705826997757,
"baseline_backdoor_adjustment_error": 2.565906524658203,
"baseline_g_computation_error": 2.878518581390381,
"baseline_linear_regression_error": 2.565906524658203,
"baseline_naive_correlation_error": 2.6589088439941406,
"baseline_neural_regression_error": 0.04033350944519043,
"baseline_propensity_weighting_error": 2.565906524658203,
"causal_slope": 0.8999999460761252,
"counterfactual_error": 1.5746076107025146,
"edge_auroc": 0.75,
"edge_f1": 0.4,
"edge_precision": 0.3333333333333333,
"edge_recall": 0.5,
"error_to_sigma_ratio": 11.463352200005689,
"identified": 0.0,
"ite_error": 2.229747772216797,
"marginal_slope": 1.8474929332733154,
"n_predicted_edges": 3.0,
"n_true_edges": 2.0,
"n_variables": 3.0,
"normalised_ate_error": 0.8004391092629032,
"normalised_ate_sigma": 0.06982591961734422,
"normalised_counterfactual_error": 0.5656639133690219,
"normalised_ite_error": 0.8010172247886658,
"reversed_edges": 1.0,
"shd": 2.0,
"true_ate": 2.4372920989990234
},
"distribution_shift": {
"ate_error": 4.67472767829895,
"ate_sigma": 3.688892126083374,
"baseline_backdoor_adjustment_error": 17.955249071121216,
"baseline_g_computation_error": 13.376930952072144,
"baseline_linear_regression_error": 28.42704129219055,
"baseline_naive_correlation_error": 30.10206913948059,
"baseline_neural_regression_error": 3.2862675189971924,
"baseline_propensity_weighting_error": 23.800013780593872,
"causal_slope": 1.3033428528651472,
"counterfactual_error": 1.8358383178710938,
"edge_auroc": 0.5555555820465088,
"edge_f1": 0.6666666666666666,
"edge_precision": 0.6666666666666666,
"edge_recall": 0.6666666666666666,
"error_to_sigma_ratio": 1.2672443428868365,
"identified": 1.0,
"ite_error": 5.052611351013184,
"marginal_slope": 16.117734909057617,
"n_predicted_edges": 3.0,
"n_true_edges": 3.0,
"n_variables": 3.0,
"normalised_ate_error": 1.0795293952085245,
"normalised_ate_sigma": 0.8518715441643328,
"normalised_counterfactual_error": 0.4239479955575024,
"normalised_ite_error": 1.1667935848236084,
"reversed_edges": 1.0,
"shd": 1.0,
"true_ate": 2.500958204269409
},
"hidden_variables": {
"ate_error": 2.5442872047424316,
"ate_sigma": 0.5252423882484436,
"baseline_backdoor_adjustment_error": 2.57559871673584,
"baseline_g_computation_error": 2.9756970405578613,
"baseline_linear_regression_error": 2.57559871673584,
"baseline_naive_correlation_error": 2.7625694274902344,
"baseline_neural_regression_error": 2.9572925567626953,
"baseline_propensity_weighting_error": 2.57559871673584,
"causal_slope": 0.9974367782446439,
"counterfactual_error": 1.7296648025512695,
"edge_auroc": NaN,
"edge_f1": 0.0,
"edge_precision": 0.0,
"edge_recall": 0.0,
"error_to_sigma_ratio": 4.844024895300272,
"identified": 0.0,
"ite_error": 2.5377490520477295,
"marginal_slope": 2.007148265838623,
"n_predicted_edges": 1.0,
"n_true_edges": 0.0,
"n_variables": 2.0,
"normalised_ate_error": 0.9220334736203645,
"normalised_ate_sigma": 0.19034449523885227,
"normalised_counterfactual_error": 0.6268195049374066,
"normalised_ite_error": 0.9196640849113464,
"reversed_edges": 0.0,
"shd": 1.0,
"true_ate": 2.5442872047424316
},
"nonlinear": {
"ate_error": 0.173872709274292,
"ate_sigma": 0.09913791716098785,
"baseline_backdoor_adjustment_error": 0.6513547897338867,
"baseline_g_computation_error": 0.9216175079345703,
"baseline_linear_regression_error": 0.6513547897338867,
"baseline_naive_correlation_error": 1.2258453369140625,
"baseline_neural_regression_error": 2.005765378475189,
"baseline_propensity_weighting_error": 0.6513547897338867,
"causal_slope": 0.9333865277512824,
"counterfactual_error": 0.20367702841758728,
"edge_auroc": 0.875,
"edge_f1": 0.5,
"edge_precision": 0.5,
"edge_recall": 0.5,
"error_to_sigma_ratio": 1.753846704202429,
"identified": 1.0,
"ite_error": 0.2109552025794983,
"marginal_slope": 1.1372480392456055,
"n_predicted_edges": 2.0,
"n_true_edges": 2.0,
"n_variables": 3.0,
"normalised_ate_error": 0.12467042747883068,
"normalised_ate_sigma": 0.07108399336162348,
"normalised_counterfactual_error": 0.1460407576693404,
"normalised_ite_error": 0.15125936269760132,
"reversed_edges": 1.0,
"shd": 1.0,
"true_ate": 2.9822487831115723
},
"outside_range": {
"ate_error": 0.2608070373535156,
"ate_sigma": 5.753205299377441,
"baseline_backdoor_adjustment_error": 29.526872634887695,
"baseline_g_computation_error": 133.46777725219727,
"baseline_linear_regression_error": 14.822834014892578,
"baseline_naive_correlation_error": 24.001174926757812,
"baseline_neural_regression_error": 1.5242118835449219,
"baseline_propensity_weighting_error": 43.52757549285889,
"causal_slope": 6.159307069261119,
"counterfactual_error": 7.563058853149414,
"edge_auroc": 0.8888888955116272,
"edge_f1": 0.8,
"edge_precision": 1.0,
"edge_recall": 0.6666666666666666,
"error_to_sigma_ratio": 0.04533247533887548,
"identified": 1.0,
"ite_error": 4.7379150390625,
"marginal_slope": 4.536205291748047,
"n_predicted_edges": 2.0,
"n_true_edges": 3.0,
"n_variables": 3.0,
"normalised_ate_error": 0.045549851192149216,
"normalised_ate_sigma": 1.0047951463415306,
"normalised_counterfactual_error": 1.3208853902644846,
"normalised_ite_error": 0.8274750709533691,
"reversed_edges": 0.0,
"shd": 1.0,
"true_ate": 56.24929428100586
},
"selection_bias": {
"ate_error": 0.13124418258666992,
"ate_sigma": 0.30816394090652466,
"baseline_backdoor_adjustment_error": 0.32472920417785645,
"baseline_g_computation_error": 0.5110855102539062,
"baseline_linear_regression_error": 0.32472920417785645,
"baseline_naive_correlation_error": 0.28928041458129883,
"baseline_neural_regression_error": 0.6534314155578613,
"baseline_propensity_weighting_error": 0.32472920417785645,
"causal_slope": 1.200000014507461,
"counterfactual_error": 0.06795314699411392,
"edge_auroc": 0.2222222238779068,
"edge_f1": 0.0,
"edge_precision": 0.0,
"edge_recall": 0.0,
"error_to_sigma_ratio": 0.4258907846277842,
"identified": 0.0,
"ite_error": 0.11888636648654938,
"marginal_slope": 1.068271279335022,
"n_predicted_edges": 1.0,
"n_true_edges": 3.0,
"n_variables": 3.0,
"normalised_ate_error": 0.14681131025663335,
"normalised_ate_sigma": 0.3447158650895488,
"normalised_counterfactual_error": 0.07601320187795306,
"normalised_ite_error": 0.13298770785331726,
"reversed_edges": 1.0,
"shd": 3.0,
"true_ate": 2.9581568241119385
},
"simpsons_paradox": {
"ate_error": 0.7794182300567627,
"ate_sigma": 0.27058735489845276,
"baseline_backdoor_adjustment_error": 0.01840519905090332,
"baseline_g_computation_error": 0.07550394535064697,
"baseline_linear_regression_error": 2.463747024536133,
"baseline_naive_correlation_error": 2.397135704755783,
"baseline_neural_regression_error": 0.013787031173706055,
"baseline_propensity_weighting_error": 1.3367897272109985,
"causal_slope": -1.0000001452680656,
"counterfactual_error": 0.47993287444114685,
"edge_auroc": 1.0,
"edge_f1": 1.0,
"edge_precision": 1.0,
"edge_recall": 1.0,
"error_to_sigma_ratio": 2.88046804829171,
"identified": 1.0,
"ite_error": 1.018937110900879,
"marginal_slope": 0.2009255588054657,
"n_predicted_edges": 3.0,
"n_true_edges": 3.0,
"n_variables": 3.0,
"normalised_ate_error": 1.0139955220160821,
"normalised_ate_sigma": 0.3520245685826795,
"normalised_counterfactual_error": 0.6243756776335451,
"normalised_ite_error": 1.3256011009216309,
"reversed_edges": 0.0,
"shd": 0.0,
"true_ate": -2.0515401363372803
}
},
"summary": {
"mean_normalised_ate_error": 0.5202214121818542,
"mean_sigma_identified": 0.48233312368392944,
"mean_sigma_not_identified": 0.20162875950336456,
"overconfident_scenarios": [
"confounding",
"hidden_variables"
],
"uncertainty_ordering_correct": false,
"worst_error_to_sigma_ratio": 11.463352200005689
}
},
"baselines": {
"backdoor_adjustment": {
"normalised_ate_error": 0.4930418133735657
},
"g_computation": {
"normalised_ate_error": 0.43038827180862427
},
"linear_regression": {
"normalised_ate_error": 0.5331624746322632
},
"naive_correlation": {
"normalised_ate_error": 0.5140460729598999
},
"neural_regression": {
"normalised_ate_error": 1.0974092483520508
},
"propensity_weighting": {
"normalised_ate_error": 0.5330392122268677
},
"proposed_model": {
"normalised_ate_error": 0.423109233379364
}
},
"baselines_unseen": {
"backdoor_adjustment": {
"normalised_ate_error": 0.4284541606903076
},
"g_computation": {
"normalised_ate_error": 0.4658964276313782
},
"linear_regression": {
"normalised_ate_error": 0.46599188446998596
},
"naive_correlation": {
"normalised_ate_error": 0.503264307975769
},
"neural_regression": {
"normalised_ate_error": 1.071964979171753
},
"propensity_weighting": {
"normalised_ate_error": 0.47844383120536804
},
"proposed_model": {
"normalised_ate_error": 0.500436007976532
}
},
"configuration": {
"counterfactual_rows": 64,
"edge_threshold": 0.5,
"held_out_tasks": 256,
"parameters": 26119098,
"replicas": 64,
"seed": 0,
"unseen_structure_tasks": 128
},
"held_out": {
"ate_bias": -0.0419006422162056,
"ate_error": 0.4225336015224457,
"ate_mae": 0.4225336015224457,
"ate_normalised_mae": 0.423109233379364,
"ate_normalised_rmse": 0.733408510684967,
"ate_rmse": 0.7471932172775269,
"ate_sigma": 0.5646735429763794,
"calibration_coverage_1sigma": 0.70703125,
"calibration_coverage_2sigma": 0.91015625,
"calibration_mean_abs_z": 0.8785681128501892,
"calibration_mean_sigma": 0.5783447623252869,
"calibration_uncertainty_error_correlation": 0.4662990868091583,
"counterfactual_error": 0.28894758224487305,
"edge_auroc": 0.9193468689918518,
"edge_f1": 0.7056852579116821,
"edge_precision": 0.6512736082077026,
"edge_recall": 0.8075383901596069,
"identified": 0.765625,
"ite_error": 0.5081815719604492,
"n_predicted_edges": 21.75390625,
"n_true_edges": 14.80078125,
"n_variables": 11.65625,
"normalised_ate_error": 0.423109233379364,
"normalised_ate_sigma": 0.5783447623252869,
"normalised_counterfactual_error": 0.28902265429496765,
"normalised_ite_error": 0.5149435997009277,
"reversed_edges": 1.86328125,
"shd": 11.30859375,
"sigma_identification_gap": 0.3878597021102905,
"sigma_identified": 0.4874401092529297,
"sigma_not_identified": 0.8752998113632202
},
"identification": {
"held_out": {
"fraction_backdoor": 0.265625,
"fraction_identified": 0.765625,
"n_tasks": 256.0
},
"unseen_structures": {
"fraction_backdoor": 0.3046875,
"fraction_identified": 0.796875,
"n_tasks": 128.0
}
},
"real_data": {
"ihdp": {
"ate_error": 1.7996239438652992,
"baseline_linear_adjustment_error": 0.04008984565734863,
"baseline_naive_difference_error": 0.0023185014724731445,
"baseline_t_learner_ate_error": 0.025522828102111816,
"baseline_t_learner_pehe": 0.2680015563964844,
"description": "747 infants, 25 covariates, known response surfaces",
"ite_mae": 1.805983543395996,
"n_samples": 747.0,
"n_variables": 27.0,
"pehe": 1.8435747623443604,
"predicted_ate": 0.042645715177059174,
"true_ate": 1.8422696590423584
},
"sachs": {
"baseline_correlation_auroc": 0.7375078797340393,
"description": "853 single-cell measurements of 11 phosphoproteins (log-standardised)",
"edge_auroc": 0.6540164351463318,
"edge_f1": 0.2790697674418605,
"edge_precision": 0.23076923076923078,
"edge_recall": 0.35294117647058826,
"n_predicted_edges": 26.0,
"n_samples": 853.0,
"n_true_edges": 17.0,
"reversed_edges": 4.0,
"shd": 27.0,
"variants": {
"log_standardised": {
"baseline_correlation_auroc": 0.7375078797340393,
"description": "853 single-cell measurements of 11 phosphoproteins (log-standardised)",
"edge_auroc": 0.6540164351463318,
"edge_f1": 0.2790697674418605,
"edge_precision": 0.23076923076923078,
"edge_recall": 0.35294117647058826,
"n_predicted_edges": 26.0,
"n_samples": 853.0,
"n_true_edges": 17.0,
"reversed_edges": 4.0,
"shd": 27.0
},
"standardised": {
"baseline_correlation_auroc": 0.7305502891540527,
"description": "853 single-cell measurements of 11 phosphoproteins (standardised)",
"edge_auroc": 0.6521189212799072,
"edge_f1": 0.0,
"edge_precision": 0.0,
"edge_recall": 0.0,
"n_predicted_edges": 5.0,
"n_samples": 853.0,
"n_true_edges": 17.0,
"reversed_edges": 4.0,
"shd": 18.0
}
}
}
},
"training": {
"promoted_stage": "unseen_structures",
"promoted_step": 22000,
"validation_score": 1.0260702967643738
},
"unseen_structures": {
"ate_bias": -0.13809093832969666,
"ate_error": 0.490360826253891,
"ate_mae": 0.490360826253891,
"ate_normalised_mae": 0.500436007976532,
"ate_normalised_rmse": 1.0647010803222656,
"ate_rmse": 1.0703402757644653,
"ate_sigma": 0.6500532627105713,
"calibration_coverage_1sigma": 0.78125,
"calibration_coverage_2sigma": 0.9375,
"calibration_mean_abs_z": 0.6925615072250366,
"calibration_mean_sigma": 0.6671552062034607,
"calibration_uncertainty_error_correlation": 0.6199653148651123,
"counterfactual_error": 0.3348301649093628,
"edge_auroc": 0.9375064969062805,
"edge_f1": 0.6931359171867371,
"edge_precision": 0.6278237104415894,
"edge_recall": 0.8394230008125305,
"identified": 0.796875,
"ite_error": 0.5735531449317932,
"n_predicted_edges": 68.078125,
"n_true_edges": 38.578125,
"n_variables": 24.0390625,
"normalised_ate_error": 0.500436007976532,
"normalised_ate_sigma": 0.6671552062034607,
"normalised_counterfactual_error": 0.3433118462562561,
"normalised_ite_error": 0.5847880244255066,
"reversed_edges": 5.734375,
"shd": 43.640625,
"sigma_identification_gap": 0.8151648640632629,
"sigma_identified": 0.5015749335289001,
"sigma_not_identified": 1.316739797592163
}
}
Training data
frontal-labs/causal-1-training-data — synthetic
SCMs with exact ground-truth counterfactuals. That card documents measured biases
in the generator which explain the real-data failures above; they are worth
reading alongside this one.
Usage
from causal.api import CausalAPI
api = CausalAPI.from_artifact("causal-1")
result = api.effect("treatment", "outcome", 0.0, 1.0)
print(result.effect, result.uncertainty)
- Downloads last month
- 21
Dataset used to train frontal-labs/causal-1
Evaluation results
- Normalised ATE error on Causal-1 held-out worldsvalidation set self-reported0.423
- Edge AUROC on Causal-1 held-out worldsvalidation set self-reported0.919