Heterogeneous Semantic and Spatial-Frequency Learning for Cross-Domain Deepfake Detection
Official implementation of the paper HSSF: Heterogeneous Semantic and Spatial-Frequency Learning for Cross-Domain Deepfake Detection.
HSSF is a heterogeneous three-member framework for cross-dataset deepfake detection. Its premise: a spatial-frequency forensic branch (low-level manipulation artifacts) and CLIP pre-trained visual-representation branches (broad natural-image priors) make weakly correlated errors, so fusing their prediction scores — without any target-domain supervision — improves cross-domain robustness where either family alone struggles.
The two families' performance ordering inverts across domains — forensic wins in-domain (FF++ video AUC 0.9939 vs. CLIP v2's 0.9558), CLIP wins on Celeb-DF (0.8399 vs. 0.7894) — which is precisely the complementarity the fusion exploits. The framework report is candid about boundaries: the fusion gain is dataset-dependent (significant on Celeb-DF, not on DFD / FaceShifter), and this is a core finding, not a caveat buried in a footnote.
Three members, trained completely independently (no shared weights, no feature-level fusion); fusion happens once, at inference time, on prediction scores:
- SFCF-GSAF (forensic branch) — ResNet-50 spatial features fused with stationary-wavelet-transform (SWT) high-frequency features (db4, L=3, 27 channels) by GSAF: grouped (G=8) dual bottleneck-gated branches (channel attention Eq. 11 / spatial attention Eq. 12), cross-group shuffle, global gate (Eq. 13), scaled identity residual (α=0.3). 27.9M parameters / ≈4.2 GFLOPs.
- CLIP v2 — ViT-L/14 with the last four visual blocks + ln_post fine-tuned (6 epochs), MLP head.
- CLIP v1-MLP — ViT-L/14 fully frozen, MLP probe on cached features.
Fusion. Equal-weight rank average of the members' scores
(transductive; label-free, no trainable parameters, ranks over the
evaluation set). For online deployment, an inductive probability
average (per-sample function of member scores) preserves the core
gain. See docs/method.md for every equation mapped to code.
HSSF/
├── src/hssf/ # installable package (refactor of the paper's training code)
│ ├── models/ # SFCF-GSAF, SWT, GSAF, CLIP branches
│ ├── training/ # trainer, focal loss / MixUp, GPU augmentation
│ ├── data/ # FF++ video-split dataset, eval datasets
│ ├── fusion/ # rank average (transductive) / prob average (inductive)
│ └── evaluation/ # frame/video AUC, paired bootstrap
├── scripts/ # CLI entry points (train / evaluate / fuse / bootstrap / figures)
│ ├── original/ # UNMODIFIED scripts that produced the paper's numbers
│ ├── baselines/ # Xception (same protocol), attention baselines, Wavelet-CLIP re-impl.
│ └── figures/ # paper figure scripts + shared style
├── configs/ # declarative protocol snapshots (common/5m, members, fusion)
├── docs/ # protocol, method, results, limitations, data, reproducibility
├── results/tables/ # frozen statistical outputs (bootstrap10k.json)
├── tests/ # smoke tests (pure-function units)
├── paper/HSSF_manuscript.docx
└── TODO.md # open items for the authors
Python ≥ 3.9, PyTorch ≥ 2.0 with CUDA.
git clone <repo-url> && cd HSSF
pip install -r requirements.txt
pip install -e . # installs the `hssf` package (or use scripts/ with src/ on path)Datasets are not included — see docs/data.md for the expected
layout and the environment variables (HSSF_FFPP_ROOT,
HSSF_CDF_ROOT, HSSF_DFD_ROOT, HSSF_FS_ROOT, …) that point the
loaders at them.
# Train the three members (single A100 80GB; independent processes)
python scripts/train_sfcf_gsaf.py --variant full \
--data_root $FFPP_ROOT --save_dir checkpoints/sfcf_gsaf_full5m # 60 epochs
python scripts/train_clip_v2.py --data_root $FFPP_ROOT \
--clip_weights checkpoints/ViT-L-14.pt --save_dir checkpoints/clip_v2
python scripts/train_clip_v1_mlp.py --data_root $FFPP_ROOT \
--clip_weights checkpoints/ViT-L-14.pt --save_dir checkpoints/clip_v1
# Evaluate a member on all domains (dual caliber; writes eval.json + probs.npz)
python scripts/evaluate_sfcf_gsaf.py --variant full \
--ckpt checkpoints/sfcf_gsaf_full5m/best.pth \
--out_dir checkpoints/sfcf_gsaf_full5m
# Fuse member score dumps (transductive rank average + inductive prob average)
python scripts/fuse_rank_average.py \
--member v2=checkpoints/clip_v2/clipv2_probs.npz:p \
--member v1mlp=checkpoints/clip_v1/clip_probs.npz:mlp \
--member full5m=checkpoints/sfcf_gsaf_full5m/probs.npz:p \
--domains celebdf dfd faceshifter --out results/fusion_3way.json
# Paired bootstrap (Table 10; asserts the paper anchors first)
python scripts/paired_bootstrap.py --results_dir results/per_sample \
--out results/tables/bootstrap10k.json
# Ablations (19 variants) and baselines
python scripts/train_sfcf_gsaf.py --variant no_swt ... # 40-epoch budget
python scripts/baselines/train_baseline_xception.py ...All experiments follow one frozen common/5m protocol — see
docs/protocol.md:
- Source domain: FaceForensics++ (c23), all five manipulation methods; video-level split 720/140/139 (train/val/test; frames of a video never span splits).
- Model selection: FF++ validation split only; target domains never touched for tuning.
- Evaluation: frame-level and video-level (video-mean) AUC on FF++ test, Celeb-DF v2 (6,529 videos), DFD, and an independent FaceShifter set; plus AP / EER / TPR@low-FPR / F1 (Table 11).
- Statistics: paired bootstrap, B = 10,000, videos as resampling units, ranks recomputed inside each resample (transductive modeling), two-sided p with +1 correction.
Frame / video AUC under common/5m (paper Tables 2, 7, 8; full tables in
docs/results.md):
| Model | FF++ (F/V) | Celeb-DF (F/V) | DFD (F/V) | FaceShifter (F/V) |
|---|---|---|---|---|
| SFCF-GSAF (full5m) | 0.9626 / 0.9939 | 0.7192 / 0.7894 | 0.7870 / 0.8420 | 0.7439 / 0.7716 |
| CLIP v1 + MLP | 0.8495 / 0.9416 | 0.7278 / 0.7661 | 0.8354 / 0.8799 | 0.7455 / 0.7770 |
| CLIP v2 | 0.9313 / 0.9558 | 0.8065 / 0.8399 | 0.8831 / 0.9191 | 0.8476 / 0.8708 |
| HSSF three-way (rank fusion) | — | 0.8244 / 0.8801 | — / 0.9224 | — / 0.8672 |
| Three-way (inductive prob. avg.) | — | — / 0.8700 | — / 0.9267 | — / 0.8726 |
Statistical testing and negative controls (Tables 9–10):
- Three-way − CLIP v2 on Celeb-DF: +0.0311, 95% CI [+0.0217, +0.0406], p < 0.001 (B = 10,000).
- DFD (+0.0033, p = 0.399) and FaceShifter (−0.0035, p = 0.635): not significant — the fusion gain is dataset-dependent.
- Adding a 4th, weaker member (four-way) or a source-domain learnable gate (LSFF) hurts (Table 9): gains come from heterogeneous complementarity, not extra members or learned weights.
- Member error correlations on Celeb-DF are only φ = 0.07–0.20 (on DFD the two CLIP members correlate at 0.57 — matching the null fusion result there).
- Detection-quality: fusion lifts DFD TPR@1%FPR 0.3592 → 0.5368 (Table 11).
- Frozen statistics ship with the repo
(
results/tables/bootstrap10k.json);scripts/paired_bootstrap.pyre-derives them and asserts the paper anchors before running. scripts/original/preserves the unmodified scripts that produced the paper's numbers, as a provenance layer next to the refactoredsrc/hssfpackage.- Training is not bit-deterministic (GPU); multi-seed options quantify
variance. Epoch budgets follow the paper: full model 60 epochs,
ablations 40. Details:
docs/reproducibility.md.
If you use this code, please cite the paper. (CITATION.cff is a
placeholder until the manuscript is accepted — see TODO.md.)
@article{hssf2026,
title = {HSSF: Heterogeneous Semantic and Spatial-Frequency Learning
for Cross-Domain Deepfake Detection},
author = {[Author]},
journal = {[Journal]},
year = {2026},
note = {placeholder -- update on acceptance}
}MIT (see LICENSE). Datasets and pre-trained CLIP weights remain under
their own licenses; this repository distributes neither. Checkpoint
redistribution is pending author decision (TODO.md).