JunXueTech commited on
Commit
cd728f1
·
verified ·
1 Parent(s): 649716f

Release original LaST-Net checkpoint and usage instructions

Browse files
Files changed (2) hide show
  1. README.md +44 -0
  2. best.pt +3 -0
README.md ADDED
@@ -0,0 +1,44 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ tags:
3
+ - audio
4
+ - speech-deepfake-detection
5
+ - pytorch
6
+ - fairseq
7
+ base_model: facebook/wav2vec2-xls-r-300m
8
+ ---
9
+
10
+ # LaST-Net
11
+
12
+ Model weights for **LaST-Net: Length-Aware Layer and Scale-Adaptive Temporal Network for Speech Deepfake Detection**.
13
+
14
+ **Training and inference code:** [JunXue-tech/LaST-Net v1.0.0](https://github.com/JunXue-tech/LaST-Net/tree/v1.0.0).
15
+
16
+ Release **v1.0.0** pairs this original checkpoint with the same version tag in the code repository.
17
+
18
+ ## Checkpoint
19
+
20
+ `best.pt` is the original epoch-52 checkpoint selected by the lowest mean development EER across 1, 2, 4 and 6 seconds on ASVspoof 2019 LA. It includes the fine-tuned XLS-R 300M frontend, LaST-Net backend, optimizer state and original training metadata. The original checkpoint is distributed without conversion.
21
+
22
+ The matching code uses strict state-dict loading. Model settings are router temperature 0.8, uniform floor 0.05, and layer dropout 0.1. Paired Prediction Consistency is used only during training.
23
+
24
+ ## Usage
25
+
26
+ Follow the environment setup in the [code repository](https://github.com/JunXue-tech/LaST-Net). From that repository:
27
+
28
+ ```bash
29
+ python download_model.py
30
+ python infer.py example.wav --checkpoint checkpoints/best.pt \
31
+ --ssl-path /path/to/xlsr2_300m.pt --seconds 6
32
+ ```
33
+
34
+ The model constructor requires the original fairseq-format XLS-R 300M checkpoint, available from the [official XLS-R repository](https://github.com/facebookresearch/fairseq/tree/main/examples/wav2vec/xlsr), before loading the fine-tuned parameters from `best.pt`.
35
+
36
+ Input audio must be mono at 16 kHz. The supplied inference code evaluates 1–6 second inputs using prefix cropping and repetition of shorter recordings. Higher `bonafide_log_score` values favor bona fide speech. Scores are not calibrated probabilities.
37
+
38
+ ## Evaluation
39
+
40
+ Duration-averaged EER (%) across 1–6 second inputs: 19LA **1.29**, 21LA **4.98**, 21DF **3.62**, and In-the-Wild **7.58**. Per-duration results and evaluation commands are provided in the code repository. Results depend on the evaluation protocol and preprocessing.
41
+
42
+ ## Authors
43
+
44
+ Jun Xue, Qichang Li, Yan Li, Daixian Li, Yanzhen Ren, Yihuan Huang, and Zhuolin Yi. School of Cyber Science and Engineering, Wuhan University. Corresponding author: Yanzhen Ren.
best.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:8fee6178d19b398fac9e1b1bfcf13720c4eb5ff722b5b52e4bc698e0de0fa28b
3
+ size 3811985991