Unggi's picture
leap-kt: 6 dataset(s), 138 files
7b9179c verified
|
Raw
History Blame Contribute Delete
3.17 kB
---
license: mit
library_name: leap-kt
tags:
- knowledge-tracing
- education
- educational-data-mining
- addressingtwoproblemsdee
datasets:
- algebra2005
- assist2009
- assist2012
- bridge2algebra2006
- dbe_kt22
- junyi
metrics:
- auc
- accuracy
- f1
---
# leap-kt · ADDRESSINGTWOPROBLEMSDEE
**ADDRESSINGTWOPROBLEMSDEE**
Part of [leap-kt-toolkit](https://github.com/LEAP-LAB-KUS/leap-kt-toolkit), a systematic re-implementation of published Knowledge Tracing models under one protocol. This repository holds **every fold of every dataset** this model has been run on, with the per-epoch training logs and the exact user split alongside the checkpoints.
## Protocol
User-level 80/20 train/test split · **5-fold** cross-validation over the training portion · held-out fold as validation · early stopping patience 10 on validation AUC · max 200 epochs.
Every cell in the project runs under identical settings; a cell that cannot is recorded as a documented failure rather than re-run under bespoke settings.
## Results
| dataset | AUC | ACC | F1 | published reference | delta |
|---|---|---|---|---|---|
| `algebra2005` | **0.8196** ± 0.0013 | 0.8144 | 0.8847 | — | — |
| `assist2009` | **0.7618** ± 0.0014 | 0.7374 | 0.8162 | — | — |
| `assist2012` | **0.7323** ± 0.0002 | 0.7337 | 0.8269 | — | — |
| `bridge2algebra2006` | **0.7930** ± 0.0006 | 0.8492 | 0.9148 | — | — |
| `dbe_kt22` | **0.7947** ± 0.0007 | 0.7937 | 0.8737 | — | — |
| `junyi` | **0.7606** ± 0.0002 | 0.7487 | 0.8352 | — | — |
Per-fold values are in each dataset's `summary.json`. The mean is never reported without the spread — 0.75 ± 0.001 and 0.75 ± 0.09 are different claims.
## Why these numbers may differ from other reproductions
Multi-concept questions are **not expanded into multiple rows**. Toolkits that do expand them place consecutive test positions carrying the same question and the same response, so a model is shown the answer one step before predicting it; on ASSIST2009 that is around 37% of positions and lifts DKT from a published ~0.75 to ~0.89 AUC. Here concepts are an extra axis on the interaction rather than extra rows, so the leak is not expressible and every interaction is scored exactly once.
Cells carrying a published reference value are additionally leak-audited before release: train/test user disjointness, no window crossing the split boundary, exactly-once scoring, and a label-shuffle control that must collapse AUC to chance. Cells with no comparable published number rely on the structural guarantee above rather than on that audit.
## Files
```
<dataset>/summary.json mean ± std and per-fold AUC
<dataset>/split.json the exact user partition, with a checksum
<dataset>/fold<k>/checkpoint/ config.json + weights
<dataset>/fold<k>/epochs.jsonl every epoch's train loss and validation metrics;
each row carries its own model/dataset/fold
<dataset>/fold<k>/run.json protocol and package version for that run
```
## Provenance
Produced by `leap-kt` at commit(s) `2c1ccef`, `a782d79`, `bb08bf2`, `e3a8dc3`.