← Back
Luna-Xin

Luna-Xin/HSSF

Heterogeneous Semantic and Spatial-Frequency Learning for Cross-Domain Deepfake Detection

View on GitHub ↗
Stars
3
Forks
0
Watchers
3
Open issues
0
Contributors
1
Language
Python
License
MIT License
Default branch
main
Created Oct 1, 2026Updated Oct 1, 2026

Star growth

Today—
This week—
This month—

Star history will appear here once this repo has been tracked for a couple of days.

README

HSSF

Heterogeneous Semantic and Spatial-Frequency Learning for Cross-Domain Deepfake Detection

Official implementation of the paper HSSF: Heterogeneous Semantic and Spatial-Frequency Learning for Cross-Domain Deepfake Detection.

Overview

HSSF is a heterogeneous three-member framework for cross-dataset deepfake detection. Its premise: a spatial-frequency forensic branch (low-level manipulation artifacts) and CLIP pre-trained visual-representation branches (broad natural-image priors) make weakly correlated errors, so fusing their prediction scores — without any target-domain supervision — improves cross-domain robustness where either family alone struggles.

The two families' performance ordering inverts across domains — forensic wins in-domain (FF++ video AUC 0.9939 vs. CLIP v2's 0.9558), CLIP wins on Celeb-DF (0.8399 vs. 0.7894) — which is precisely the complementarity the fusion exploits. The framework report is candid about boundaries: the fusion gain is dataset-dependent (significant on Celeb-DF, not on DFD / FaceShifter), and this is a core finding, not a caveat buried in a footnote.

Method Summary

Three members, trained completely independently (no shared weights, no feature-level fusion); fusion happens once, at inference time, on prediction scores:

  1. SFCF-GSAF (forensic branch) — ResNet-50 spatial features fused with stationary-wavelet-transform (SWT) high-frequency features (db4, L=3, 27 channels) by GSAF: grouped (G=8) dual bottleneck-gated branches (channel attention Eq. 11 / spatial attention Eq. 12), cross-group shuffle, global gate (Eq. 13), scaled identity residual (α=0.3). 27.9M parameters / ≈4.2 GFLOPs.
  2. CLIP v2 — ViT-L/14 with the last four visual blocks + ln_post fine-tuned (6 epochs), MLP head.
  3. CLIP v1-MLP — ViT-L/14 fully frozen, MLP probe on cached features.

Fusion. Equal-weight rank average of the members' scores (transductive; label-free, no trainable parameters, ranks over the evaluation set). For online deployment, an inductive probability average (per-sample function of member scores) preserves the core gain. See docs/method.md for every equation mapped to code.

Repository Structure

HSSF/
├── src/hssf/                 # installable package (refactor of the paper's training code)
│   ├── models/               #   SFCF-GSAF, SWT, GSAF, CLIP branches
│   ├── training/             #   trainer, focal loss / MixUp, GPU augmentation
│   ├── data/                 #   FF++ video-split dataset, eval datasets
│   ├── fusion/               #   rank average (transductive) / prob average (inductive)
│   └── evaluation/           #   frame/video AUC, paired bootstrap
├── scripts/                  # CLI entry points (train / evaluate / fuse / bootstrap / figures)
│   ├── original/             # UNMODIFIED scripts that produced the paper's numbers
│   ├── baselines/            # Xception (same protocol), attention baselines, Wavelet-CLIP re-impl.
│   └── figures/              # paper figure scripts + shared style
├── configs/                  # declarative protocol snapshots (common/5m, members, fusion)
├── docs/                     # protocol, method, results, limitations, data, reproducibility
├── results/tables/           # frozen statistical outputs (bootstrap10k.json)
├── tests/                    # smoke tests (pure-function units)
├── paper/HSSF_manuscript.docx
└── TODO.md                   # open items for the authors

Installation

Python ≥ 3.9, PyTorch ≥ 2.0 with CUDA.

git clone <repo-url> && cd HSSF
pip install -r requirements.txt
pip install -e .            # installs the `hssf` package (or use scripts/ with src/ on path)

Datasets are not included — see docs/data.md for the expected layout and the environment variables (HSSF_FFPP_ROOT, HSSF_CDF_ROOT, HSSF_DFD_ROOT, HSSF_FS_ROOT, …) that point the loaders at them.

Quick Start

# Train the three members (single A100 80GB; independent processes)
python scripts/train_sfcf_gsaf.py --variant full \
    --data_root $FFPP_ROOT --save_dir checkpoints/sfcf_gsaf_full5m     # 60 epochs
python scripts/train_clip_v2.py --data_root $FFPP_ROOT \
    --clip_weights checkpoints/ViT-L-14.pt --save_dir checkpoints/clip_v2
python scripts/train_clip_v1_mlp.py --data_root $FFPP_ROOT \
    --clip_weights checkpoints/ViT-L-14.pt --save_dir checkpoints/clip_v1

# Evaluate a member on all domains (dual caliber; writes eval.json + probs.npz)
python scripts/evaluate_sfcf_gsaf.py --variant full \
    --ckpt checkpoints/sfcf_gsaf_full5m/best.pth \
    --out_dir checkpoints/sfcf_gsaf_full5m

# Fuse member score dumps (transductive rank average + inductive prob average)
python scripts/fuse_rank_average.py \
    --member v2=checkpoints/clip_v2/clipv2_probs.npz:p \
    --member v1mlp=checkpoints/clip_v1/clip_probs.npz:mlp \
    --member full5m=checkpoints/sfcf_gsaf_full5m/probs.npz:p \
    --domains celebdf dfd faceshifter --out results/fusion_3way.json

# Paired bootstrap (Table 10; asserts the paper anchors first)
python scripts/paired_bootstrap.py --results_dir results/per_sample \
    --out results/tables/bootstrap10k.json

# Ablations (19 variants) and baselines
python scripts/train_sfcf_gsaf.py --variant no_swt ...   # 40-epoch budget
python scripts/baselines/train_baseline_xception.py ...

Data and Protocol

All experiments follow one frozen common/5m protocol — see docs/protocol.md:

  • Source domain: FaceForensics++ (c23), all five manipulation methods; video-level split 720/140/139 (train/val/test; frames of a video never span splits).
  • Model selection: FF++ validation split only; target domains never touched for tuning.
  • Evaluation: frame-level and video-level (video-mean) AUC on FF++ test, Celeb-DF v2 (6,529 videos), DFD, and an independent FaceShifter set; plus AP / EER / TPR@low-FPR / F1 (Table 11).
  • Statistics: paired bootstrap, B = 10,000, videos as resampling units, ranks recomputed inside each resample (transductive modeling), two-sided p with +1 correction.

Main Results

Frame / video AUC under common/5m (paper Tables 2, 7, 8; full tables in docs/results.md):

Model FF++ (F/V) Celeb-DF (F/V) DFD (F/V) FaceShifter (F/V)
SFCF-GSAF (full5m) 0.9626 / 0.9939 0.7192 / 0.7894 0.7870 / 0.8420 0.7439 / 0.7716
CLIP v1 + MLP 0.8495 / 0.9416 0.7278 / 0.7661 0.8354 / 0.8799 0.7455 / 0.7770
CLIP v2 0.9313 / 0.9558 0.8065 / 0.8399 0.8831 / 0.9191 0.8476 / 0.8708
HSSF three-way (rank fusion) — 0.8244 / 0.8801 — / 0.9224 — / 0.8672
Three-way (inductive prob. avg.) — — / 0.8700 — / 0.9267 — / 0.8726

Statistical testing and negative controls (Tables 9–10):

  • Three-way − CLIP v2 on Celeb-DF: +0.0311, 95% CI [+0.0217, +0.0406], p < 0.001 (B = 10,000).
  • DFD (+0.0033, p = 0.399) and FaceShifter (−0.0035, p = 0.635): not significant — the fusion gain is dataset-dependent.
  • Adding a 4th, weaker member (four-way) or a source-domain learnable gate (LSFF) hurts (Table 9): gains come from heterogeneous complementarity, not extra members or learned weights.
  • Member error correlations on Celeb-DF are only φ = 0.07–0.20 (on DFD the two CLIP members correlate at 0.57 — matching the null fusion result there).
  • Detection-quality: fusion lifts DFD TPR@1%FPR 0.3592 → 0.5368 (Table 11).

Reproducibility

  • Frozen statistics ship with the repo (results/tables/bootstrap10k.json); scripts/paired_bootstrap.py re-derives them and asserts the paper anchors before running.
  • scripts/original/ preserves the unmodified scripts that produced the paper's numbers, as a provenance layer next to the refactored src/hssf package.
  • Training is not bit-deterministic (GPU); multi-seed options quantify variance. Epoch budgets follow the paper: full model 60 epochs, ablations 40. Details: docs/reproducibility.md.

Citation

If you use this code, please cite the paper. (CITATION.cff is a placeholder until the manuscript is accepted — see TODO.md.)

@article{hssf2026,
  title   = {HSSF: Heterogeneous Semantic and Spatial-Frequency Learning
             for Cross-Domain Deepfake Detection},
  author  = {[Author]},
  journal = {[Journal]},
  year    = {2026},
  note    = {placeholder -- update on acceptance}
}

License

MIT (see LICENSE). Datasets and pre-trained CLIP weights remain under their own licenses; this repository distributes neither. Checkpoint redistribution is pending author decision (TODO.md).