xdfdet: explainable deepfake detection with EfficientNet-B4 and Grad-CAM

xdfdet detects manipulated face videos and shows which facial regions influence each prediction. The released aug-cutout-black detector scores 0.8981 AUC; the baseline detector scores 0.8684 on FaceForensics++. Each checkpoint used its own random test split, so read these as separate run results.

Detect a video

pip install git+https://github.com/mertkayacs/xdfdet && xdfdet predict video.mp4 --gradcam cam.png

The command downloads the default aug-cutout-black model, prints the probability that the video is real and saves a Grad-CAM heatmap. Select another detector with --model baseline. Use Python 3.10 to 3.12.

import xdfdet

model = xdfdet.load_model("aug-cutout-black")

Use the package for face detection, alignment and normalization. Each Keras model takes 12 RGB face crops, shaped (batch, 12, 224, 224, 3), and outputs real probabilities shaped (batch, 12, 1). Average the frames for a video score; below 0.5 is classified as fake.

Released detectors

File Augmentation and cutout AUC
aug-cutout-black.keras standard augmentation, black-fill cutout 0.8981
aug-cutout-random.keras standard augmentation, random-fill cutout 0.8820
aug-cutout-white.keras standard augmentation, white-fill cutout 0.8734
cutout-white.keras white-fill cutout 0.8700
baseline.keras none 0.8684
cutout-black.keras black-fill cutout 0.8669
cutout-random.keras random-fill cutout 0.8642
aug-standard.keras standard augmentation 0.8616

AUC measures how well a detector separates real and fake examples: 0.5 is chance and 1.0 is perfect. The eight weights study how augmentation and masking parts of a face during training change detection and attention. Grad-CAM shows where the model looks.

Four detectors attend to different regions of the same real face

Research and limits

The UBMK 2026 paper by Mert Kaya and Venera Adanova is accepted, with no DOI yet. Its results average three runs per setting; the table above reports the released checkpoints' single runs. A ninth setting, aug-intense, is in the paper but its checkpoint was lost. Research code and full metrics | MSc thesis.

Training uses 1,000 real/fake FaceForensics++ pairs. On 398 unseen DFDC videos, AUC drops to 0.60 to 0.66, and most fakes pass as real. DFDC evaluation. Performance across age, sex and skin tone was not measured. A score alone cannot establish whether a video is authentic or justify a decision about a person.

The released weights are float32 copies of mixed-float16 training runs. For evaluation in the original GPU precision, use xdfdet.load_model(name, mixed_precision=True).

License and citation

Weights: CC BY-NC 4.0, for non-commercial research and education under the FaceForensics++ terms. Code: MIT.

Cite the paper and thesis
@inproceedings{kaya2026augmentation,
  title     = {Augmentation and Cutout in Deepfake Detection: A Comparative Study of
               Accuracy, Calibration, and Attention},
  author    = {Kaya, Mert and Adanova, Venera},
  booktitle = {11th International Conference on Computer Science and Engineering (UBMK 2026)},
  year      = {2026},
  note      = {To appear}
}

@mastersthesis{kaya2025xdfdet,
  title  = {Explainable Deepfake Detection Using Frame Level CNN Models:
            A Comparative Study of Augmentation and Cutout Techniques},
  author = {Kaya, Mert},
  school = {TED University},
  year   = {2025},
  doi    = {10.5281/zenodo.18998566}
}

Project and demonstrations | Kaggle weights and quickstart.

An Eschatia Labs project. Mert Kaya.

Downloads last month
634
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Space using mertkayacs/xdfdet 1

Collections including mertkayacs/xdfdet