Ilmu Komputer & AI editorial
Complementary rPPG-Derived and Lip-Region Frequency Cues for Talking-Face Deepfake Detection
The core problem
Innovation
In-domain evaluation shows that lip-region DCT matches or exceeds the rPPG-derived 1D ResNet on every TF method except SadTalker. The Concat fusion achieves an AUC of 0.891, compared to 0.824 for the rPPG-only baseline and 0.827 for the lip-region DCT-only baseline. This indicates that the two cues are complementary in-domain. Under leave-one-generator-out (LOGO) evaluation, the cues exhibit a split in transferability: each cue transfers clearly better to three held-out methods, while IP-LAP is near chance for both. The Concat fusion averages an AUC of 0.798 across LOGO scenarios but falls below rPPG alone when DCT transfers poorly. This suggests that static fusion only partly exploits the complementarity. Additionally, lip-region DCT outperforms full-face DCT on six of the seven methods, highlighting the importance of focusing on the mouth region for frequency-based deepfake detection. The results are summarized in the following table (values approximated from the abstract):
| Method | rPPG AUC | Lip-DCT AUC | Concat AUC |
|-----------------|----------|-------------|------------|
| In-domain (avg) | 0.824 | 0.827 | 0.891 |
| LOGO (avg) | - | -
Why it matters
Who should read this
Opening member content…