Computer Science editorial
Open AccessOA2026
AffectAI-Capture: A Reproducible Multimodal Protocol for Small-Group Meeting Research
A synchronized data collection architecture for affective, behavioral, and meeting-analytics research
Meisam Jamshidi Seikavandi; Alice Modica; Anna Obara; Fabricio Batista Narcizo; Tanya Ignatenko; Ted Vucurevich; Jesper Bรผnsow Boldt; Paolo Burelli; Andrew Burke Dittbernerยท 2026ยท DOI 10.48550/arXiv.2605.19794
The core problem
Understanding affective and behavioral dynamics in small-group meetings requires rich, multimodal data captured in controlled yet naturalistic settings. Existing datasets often lack synchronization across modalities or standardized packaging, hindering reproducibility and cross-study comparison. AffectAI-Capture addresses this gap by proposing a protocol for four-person meeting-like interactions that combines eye tracking, wearable physiology, close-talk and room audio, multi-view video, event logging, and structured self-report. The protocol is grounded in established group-interaction paradigms, with fixed task blocks designed to elicit relevant social and cognitive processes. The authors emphasize a synchronization philosophy centered on a single authoritative event timeline, ensuring temporal alignment across all data streams. This digest outlines the experimental rationale, synchronization approach, data organization, and practical trade-offs, while noting that pilot-level validation of audio quality and video synchronization has been conducted, with full sessions with participants remaining ongoing work.
Innovation
As the protocol is in the pilot validation stage, results are primarily technical. Bench tests assessed audio quality and video synchronization. Audio quality was evaluated by recording known signals and measuring signal-to-noise ratio (SNR) and frequency response. Video synchronization was tested by capturing a clapperboard event across multiple cameras and measuring the temporal offset between frames. Preliminary findings indicate that the synchronization approach achieves sub-frame accuracy for video and sample-accurate alignment for audio when using the central event timeline. However, full validation with human participants is ongoing, and no empirical results from actual meetings are reported yet. The authors note that the protocol's reproducibility stems from its standardized outputs and detailed documentation, which will enable other researchers to replicate the setup and compare findings across studies.
Understanding affective and behavioral dynamics in small-group meetings requires rich, multimodal data captured in controlled yet naturalistic settings. Existing datasets often lack synchronization across modalities or standardized packaging, hindering reproducibility and cross-study comparison. AffectAI-Capture addresses this gap by proposing a protocol for four-person meeting-like interactions that combines eye tracking, wearable physiology, close-talk and room audio, multi-view video, event logging, and structured self-report. The protocol is grounded in established group-interaction paradigms, with fixed task blocks designed to elicit relevant social and cognitive processes. The authors emphasize a synchronization philosophy centered on a single authoritative event timeline, ensuring temporal alignment across all data streams. This digest outlines the experimental rationale, synchronization approach, data organization, and practical trade-offs, while noting that pilot-level validation of audio quality and video synchronization has been conducted, with full sessions with participants remaining ongoing work.
The AffectAI-Capture protocol orchestrates data collection through a modular instrumentation setup. Four participants engage in structured task blocks derived from established group-interaction paradigms, such as collaborative problem-solving or decision-making scenarios. Data streams include:
Why it matters
The AffectAI-Capture protocol represents a significant step toward reproducible multimodal research in small-group settings. By centralizing synchronization around a single event timeline, it mitigates common issues of data misalignment and facilitates integration across diverse data types. The inclusion of structured self-report adds a subjective dimension that complements objective measures. However, challenges remain: the complexity of the setup may limit scalability, and the pilot validation has not yet been extended to full participant sessions. The authors discuss trade-offs between naturalistic interaction and controlled tasks, noting that fixed task blocks enhance comparability but may reduce spontaneity. Future work will involve deploying the protocol in real meetings, assessing data quality in less controlled environments, and refining the synchronization architecture. The protocol's emphasis on standardized packaging and timing provenance aligns with open science principles, potentially enabling large-scale sharing and meta-analyses. As affective computing and meeting analytics advance, such protocols will be crucial for building robust, generalizable models of group behavior.
Who should read this
CS practitioners and researchers
Opening member contentโฆ