Jadwal Sholat

Memuat jadwal sholat…

Ilmu Komputer & AI editorial

Open AccessOA2026

Complementary rPPG-Derived and Lip-Region Frequency Cues for Talking-Face Deepfake Detection

A subject-independent study on Celeb-DF++ reveals that lip-region DCT and rPPG-derived waveforms offer complementary strengths across seven talking-face generators, with static fusion only partially exploiting their synergy.
Othmane Harraq; Tamer Aldwairi· 2026· DOI 10.48550/arXiv.2609.22284

The core problem

Talking-face (TF) deepfakes pose a growing threat to digital media integrity, and detection methods based on remote photoplethysmography (rPPG) have shown promise but exhibit uneven performance across different generators. This study investigates two lightweight visual-only cues for TF deepfake detection: rPPG-derived waveforms extracted by RhythmFormer and lip-region discrete cosine transform (DCT) coefficients. The authors evaluate these cues on the seven TF methods of the Celeb-DF++ dataset under a subject-independent protocol, ensuring that identity-specific artifacts do not inflate performance. The central hypothesis is that these cues capture complementary information—rPPG signals reflecting subtle temporal color variations and lip-region DCT capturing frequency-domain artifacts in the mouth area—and that their fusion can improve detection robustness. The paper explicitly treats the rPPG-derived signal as an empirical cue and does not claim it is cardiac in origin, avoiding overinterpretation of the physiological basis. The study aims to quantify the individual and combined performance of these cues, assess their generalization across generators, and analyze the limitations o

Innovation

In-domain evaluation shows that lip-region DCT matches or exceeds the rPPG-derived 1D ResNet on every TF method except SadTalker. The Concat fusion achieves an AUC of 0.891, compared to 0.824 for the rPPG-only baseline and 0.827 for the lip-region DCT-only baseline. This indicates that the two cues are complementary in-domain. Under leave-one-generator-out (LOGO) evaluation, the cues exhibit a split in transferability: each cue transfers clearly better to three held-out methods, while IP-LAP is near chance for both. The Concat fusion averages an AUC of 0.798 across LOGO scenarios but falls below rPPG alone when DCT transfers poorly. This suggests that static fusion only partly exploits the complementarity. Additionally, lip-region DCT outperforms full-face DCT on six of the seven methods, highlighting the importance of focusing on the mouth region for frequency-based deepfake detection. The results are summarized in the following table (values approximated from the abstract):

| Method | rPPG AUC | Lip-DCT AUC | Concat AUC |
|-----------------|----------|-------------|------------|
| In-domain (avg) | 0.824 | 0.827 | 0.891 |
| LOGO (avg) | - | -

Talking-face (TF) deepfakes pose a growing threat to digital media integrity, and detection methods based on remote photoplethysmography (rPPG) have shown promise but exhibit uneven performance across different generators. This study investigates two lightweight visual-only cues for TF deepfake detection: rPPG-derived waveforms extracted by RhythmFormer and lip-region discrete cosine transform (DCT) coefficients. The authors evaluate these cues on the seven TF methods of the Celeb-DF++ dataset under a subject-independent protocol, ensuring that identity-specific artifacts do not inflate performance. The central hypothesis is that these cues capture complementary information—rPPG signals reflecting subtle temporal color variations and lip-region DCT capturing frequency-domain artifacts in the mouth area—and that their fusion can improve detection robustness. The paper explicitly treats the rPPG-derived signal as an empirical cue and does not claim it is cardiac in origin, avoiding overinterpretation of the physiological basis. The study aims to quantify the individual and combined performance of these cues, assess their generalization across generators, and analyze the limitations of static fusion strategies.
The authors employ a subject-independent evaluation protocol on Celeb-DF++, which contains seven talking-face generation methods. Two feature extraction pipelines are used: (1) rPPG-derived waveforms extracted using RhythmFormer, a state-of-the-art rPPG estimator, resulting in 1D signals processed by a 1D ResNet classifier; and (2) lip-region DCT coefficients computed from the mouth area, capturing frequency-domain statistics. For fusion, the features are concatenated (Concat) before classification. The study compares these cues against unimodal baselines and also evaluates full-face DCT as an alternative. The evaluation includes in-domain testing (training and testing on the same generator) and leave-one-generator-out (LOGO) cross-generator testing to assess generalization. Performance is measured using the area under the ROC curve (AUC). The architecture can be summarized as follows:

Why it matters

The study demonstrates that rPPG-derived waveforms and lip-region DCT coefficients provide complementary information for talking-face deepfake detection. The in-domain results show that fusion improves performance, but the LOGO evaluation reveals that the cues generalize differently across generators. This suggests that a static concatenation fusion is suboptimal; adaptive or attention-based fusion mechanisms could better leverage the strengths of each cue depending on the generator characteristics. The near-chance performance on IP-LAP for both cues indicates that this generator may produce artifacts that are not captured by either modality, or that the subject-independent protocol is particularly challenging for IP-LAP. The finding that lip-region DCT outperforms full-face DCT on most methods underscores the importance of localized analysis, as deepfake artifacts are often concentrated in the mouth area during speech. The authors caution that the rPPG-derived signal is treated as an empirical cue and not necessarily cardiac in origin, which is an important caveat for interpretability. Future work could explore dynamic fusion strategies, incorporate additional cues, and extend the evaluation to more diverse datasets. The study contributes to the understanding of how different visual cues complement each other and highlights the need for generator-aware fusion in deepfake detection.

Who should read this

CS practitioners and researchers

Opening member content…