| --- |
| license: mit |
| library_name: leap-kt |
| tags: |
| - knowledge-tracing |
| - education |
| - educational-data-mining |
| - pfa |
| datasets: |
| - algebra2005 |
| - assist2009 |
| - assist2015 |
| - dbe_kt22 |
| - ednet500 |
| - statics2011 |
| metrics: |
| - auc |
| - accuracy |
| - f1 |
| --- |
| |
| # leap-kt · PFA |
|
|
| **PFA** — |
|
|
| Part of [leap-kt-toolkit](https://github.com/LEAP-LAB-KUS/leap-kt-toolkit), a systematic re-implementation of published Knowledge Tracing models under one protocol. This repository holds **every fold of every dataset** this model has been run on, with the per-epoch training logs and the exact user split alongside the checkpoints. |
|
|
| ## Protocol |
|
|
| User-level 80/20 train/test split · **5-fold** cross-validation over the training portion · held-out fold as validation · early stopping patience 10 on validation AUC · max 200 epochs. |
|
|
| Every cell in the project runs under identical settings; a cell that cannot is recorded as a documented failure rather than re-run under bespoke settings. |
|
|
| ## Results |
|
|
| | dataset | AUC | ACC | F1 | published reference | delta | |
| |---|---|---|---|---|---| |
| | `algebra2005` | **0.7590** ± 0.0003 | 0.7907 | 0.8738 | — | — | |
| | `assist2009` | **0.7029** ± 0.0001 | 0.6990 | 0.8016 | — | — | |
| | `assist2015` | **0.6873** ± 0.0002 | 0.7416 | 0.8466 | — | — | |
| | `dbe_kt22` | **0.7300** ± 0.0002 | 0.7761 | 0.8671 | — | — | |
| | `ednet500` | **0.6113** ± 0.0016 | 0.6540 | 0.7816 | — | — | |
| | `statics2011` | **0.7862** ± 0.0007 | 0.7926 | 0.8755 | — | — | |
|
|
| Per-fold values are in each dataset's `summary.json`. The mean is never reported without the spread — 0.75 ± 0.001 and 0.75 ± 0.09 are different claims. |
|
|
| ## Why these numbers may differ from other reproductions |
|
|
| Multi-concept questions are **not expanded into multiple rows**. Toolkits that do expand them place consecutive test positions carrying the same question and the same response, so a model is shown the answer one step before predicting it; on ASSIST2009 that is around 37% of positions and lifts DKT from a published ~0.75 to ~0.89 AUC. Here concepts are an extra axis on the interaction rather than extra rows, so the leak is not expressible and every interaction is scored exactly once. |
|
|
| Each published cell passed a leak audit before being recorded: train/test user disjointness, no window crossing the split boundary, exactly-once scoring, and a label-shuffle control that must collapse AUC to chance. |
|
|
| ## Files |
|
|
| ``` |
| <dataset>/summary.json mean ± std and per-fold AUC |
| <dataset>/split.json the exact user partition, with a checksum |
| <dataset>/fold<k>/checkpoint/ config.json + weights |
| <dataset>/fold<k>/epochs.jsonl every epoch's train loss and validation metrics; |
| each row carries its own model/dataset/fold |
| <dataset>/fold<k>/run.json protocol and package version for that run |
| ``` |
|
|
| ## Provenance |
|
|
| Produced by `leap-kt` at commit(s) `92091dd`. |
|
|