VLDB 2026 Research / reviewers in the wild / expert
Maksim Siniukov
dblp:297/3461
· DBLP profile ↗
7ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0001-8529-2281ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LibreFace 2.0: A Generalizable Facial Expression Analysis Toolkit Leveraging Synthetic DataabstractFacial expression analysis is central to social AI and human-computer interaction. However, existing toolkits often struggle to generalize across diverse demographics, largely due to the limited diversity of training data for tasks such as action unit (AU) detection, which typically require costly per-frame annotations. In this work, we introduce LibreFace 2.0, a toolkit that leverages recent advances in face generation and motion retargeting to enrich AU datasets with broader demographic coverage. Specifically, we employ stable diffusion to synthesize a wide range of identities spanning age, gender, race and facial attributes and retarget AU motions from annotated datasets onto these generated identities. Training on this large-scale, demographically diverse dataset yields consistent improvements in benchmark performance and enhances fairness across demographic groups. Beyond AU detection and intensity estimation, LibreFace 2.0 also supports facial expression recognition and gaze estimation through lightweight models that achieve competitive accuracy with substantially fewer parameters, enabling efficient inference. Our work provides a scalable approach to achieving fairer face analysis in real-world applications. The code and the synthetic data will be released publicly at https://github.com/ihp-lab/LibreFace. Xulang Guan, Ashutosh Chaubey, Maksim Siniukov, Annabelle Hsieh, Zongjian Li, Mohammad Soleymani 0001 |
FG | 3 |
| 2026 | Discrete Facial Encoding: A Framework for Data-driven Facial Display DiscoveryabstractFacial expression analysis is central to understanding human behavior, yet existing coding systems such as the Facial Action Coding System (FACS) are constrained by limited coverage and costly manual annotation. In this work, we introduce Discrete Facial Encoding (DFE), an unsupervised, data-driven alternative of compact and interpretable dictionary of facial expressions from 3D mesh sequences learned through a Residual Vector Quantized Variational Autoencoder (RVQ-VAE). Our approach first extracts identity-invariant expression features from images using a 3D Morphable Model (3DMM), effectively disentangling factors such as head pose and facial geometry. We then encode these features using an RVQ-VAE, producing a sequence of discrete tokens from a shared codebook, where each token captures a specific, reusable facial deformation pattern that contributes to the overall expression. Through extensive experiments, we demonstrate that Discrete Facial Encoding captures more precise facial behaviors than FACS and other facial encoding alternatives. We evaluate the utility of our representation across three high-level psychological tasks: stress detection, personality prediction, and depression detection. Using a simple Bag-of-Words model built on top of the learned tokens, our system consistently outperforms both FACS-based pipelines and strong image and video representation learning models such as Masked Autoencoders. Further analysis reveals that our representation covers a wider variety of facial displays, highlighting its potential as a scalable and effective alternative to FACS for psychological and affective computing applications. We released the source code and model weights at https://github.com/ihp-lab/vqface. Minh Tran 0004, Maksim Siniukov, Zhangyu Jin, Mohammad Soleymani 0001 |
WACV | 2 |
| 2025 | Ditailistener: Controllable High Fidelity Listener Video Generation with DiffusionabstractGenerating naturalistic and nuanced listener motions for extended interactions remains an open problem. Existing methods often rely on low-dimensional motion codes for facial behavior generation followed by photorealistic rendering, limiting both visual fidelity and expressive richness. To address these challenges, we introduce DiTaiListener, powered by a video diffusion model with multimodal conditions. Our approach first generates short segments of listener responses conditioned on the speaker's speech and facial motions with DiTaiListener-Gen. It then refines the transitional frames via DiTaiListener-Edit for a seamless transition. Specifically, DiTaiListener-Gen adapts a Diffusion Transformer (DiT) for the task of listener head portrait generation by introducing a Causal Temporal Multimodal Adapter (CTM-Adapter) to process speakers' auditory and visual cues. CTM-Adapter integrates speakers' input in a causal manner into the video generation process to ensure temporally coherent listener responses. For long-form video generation, we introduce DiTaiListener-Edit, a transition refinement video-to-video diffusion model. The model fuses video segments into smooth and continuous videos, ensuring temporal consistency in facial expressions and image quality when merging short video segments produced by DiTaiListener-Gen. Quantitatively, DiTaiListener achieves the state-of-the-art performance on benchmark datasets in both photorealism (+73.8% in FID on RealTalk) and motion representation (+6.1% in FD metric on VICO) spaces. User studies confirm the superior performance of DiTaiListener, with the model being the clear preference in terms of feedback, diversity, and smoothness, outperforming competitors by a significant margin. Maksim Siniukov, Di Chang, Minh Tran 0004, Hongkun Gong, Ashutosh Chaubey, Mohammad Soleymani 0001 |
ICCV | 1 |
| 2024 | DIM: Dyadic Interaction Modeling for Social Behavior Generation
Minh Tran 0004, Di Chang, Maksim Siniukov, Mohammad Soleymani 0001 |
ECCV (37) | 3 |
| 2024 | SEMPI: A Database for Understanding Social Engagement in Video-Mediated Multiparty InteractionabstractWe present a database for automatic understanding of Social Engagement in MultiParty Interaction (SEMPI). Social engagement is an important social signal characterizing the level of participation of an interlocutor in a conversation. Social engagement involves maintaining attention and establishing connection and rapport. Machine understanding of social engagement can enable an autonomous agent to better understand the state of human participation and involvement to select optimal actions in human-machine social interaction. Recently, video-mediated interaction platforms, e.g., Zoom, have become very popular. The ease of use and increased accessibility of video calls have made them a preferred medium for multiparty conversations, including support groups and group therapy sessions. To create this dataset, we first collected a set of publicly available video calls posted on YouTube. We then segmented the videos by speech turn and cropped the videos to generate single-participant videos. We developed a questionnaire for assessing the level of social engagement by listeners in a conversation probing the relevant nonverbal behaviors for social engagement, including back-channeling, gaze, and expressions. We used Prolific, a crowd-sourcing platform, to annotate 3,505 videos of 76 listeners by three people, reaching a moderate to high inter-rater agreement of 0.693. This resulted in a database with aggregated engagement scores from the annotators. We developed a baseline multimodal pipeline using the state-of-the-art pre-trained models to track the level of engagement achieving the CCC score of 0.454. The results demonstrate the utility of the database for future applications in video-mediated human-machine interaction and human-human social skill assessment. Our dataset and code are available at https://github.com/ihp-lab/SEMPI. Maksim Siniukov, Yufeng Yin 0002, Eli Fast, Yingshan Qi, Aarav Monga, Audrey Kim, Mohammad Soleymani 0001 |
ICMI | 1 |
| 2023 | Applicability limitations of differentiable full-reference image-quality metricsabstractMore and more visual-quality metrics are being developed to assess the quality of images, but little research has considered their limitations. In this paper, we demonstrate that image preprocessing before compression can artificially increase the quality scores provided by the popular metrics DISTS, LPIPS, HaarPSI, and VIF. We propose a series of neural-network preprocessing models that increase DISTS by up to 34.5%, LPIPS by up to 36.8%, VIF by up to 98.0%, and HaarPSI by up to 22.6% in the case of JPEG-compressed images. However, a subjective comparison of these preprocessed images showed that the visual quality either dropped or remained unchanged, indicating the limited applicability of these metrics. We used a ResNet-like lightweight CNN architecture for preprocessing and the differentiable DiffJPEG algorithm for compression. Maksim Siniukov, Dmitriy L. Kulikov, Dmitriy S. Vatolin |
DCC | 1 |
| 2023 | Unveiling the Limitations of Novel Image Quality MetricsabstractSubjective image quality measurement plays a crucial role in the advancement of image-processing technologies. The primary goal of a visual quality metric is to reflect subjective evaluation results. Despite the rapid development of these metrics, their potential limitations have not been sufficiently explored. This research paper aims to address this gap by demonstrating how image preprocessing before compression can artificially inflate the quality scores of widely used metrics such as DISTS, LPIPS, HaarPSI, VIF, STLPIPS, ADISTS, MR-Perceptual, AHIQ, IQT, and CONTRIQUE. We present several CNN-based preprocessing models that significantly increase these metrics when images are JPEG-compressed. However, a subjective assessment (with 1027 participants) of preprocessed images reveals that the visual quality either decreases or remains unchanged, thereby challenging the universal applicability of these metrics. The detection of metric attacks integrated into image processing systems has emerged as a significant issue. Attack detection can be achieved by comparing subjective evaluation results with metric results. If these results are anticorrelated, it suggests that the metric has been attacked. However, the time-consuming nature of subjective evaluations and the need for numerous participants make them impractical for routine attack detection. To address this, we propose using other metrics to determine whether the target metric has been attacked. Our results show that attacking one metric affects the output of other metrics, offering a potential method for attack detection. Maksim Siniukov, Dmitriy L. Kulikov, Dmitriy S. Vatolin |
MMSP | 1 |