VLDB 2026 Research / reviewers in the wild / expert
Ashutosh Chaubey
dblp:258/8999
· DBLP profile ↗
6ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0002-8463-0012ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LibreFace 2.0: A Generalizable Facial Expression Analysis Toolkit Leveraging Synthetic DataabstractFacial expression analysis is central to social AI and human-computer interaction. However, existing toolkits often struggle to generalize across diverse demographics, largely due to the limited diversity of training data for tasks such as action unit (AU) detection, which typically require costly per-frame annotations. In this work, we introduce LibreFace 2.0, a toolkit that leverages recent advances in face generation and motion retargeting to enrich AU datasets with broader demographic coverage. Specifically, we employ stable diffusion to synthesize a wide range of identities spanning age, gender, race and facial attributes and retarget AU motions from annotated datasets onto these generated identities. Training on this large-scale, demographically diverse dataset yields consistent improvements in benchmark performance and enhances fairness across demographic groups. Beyond AU detection and intensity estimation, LibreFace 2.0 also supports facial expression recognition and gaze estimation through lightweight models that achieve competitive accuracy with substantially fewer parameters, enabling efficient inference. Our work provides a scalable approach to achieving fairer face analysis in real-world applications. The code and the synthetic data will be released publicly at https://github.com/ihp-lab/LibreFace. Xulang Guan, Ashutosh Chaubey, Maksim Siniukov, Annabelle Hsieh, Zongjian Li, Mohammad Soleymani 0001 |
FG | 2 |
| 2026 | Face-LLaVA: Facial Expression and Attribute Understanding through Instruction Tuning
Ashutosh Chaubey, Xulang Guan, Mohammad Soleymani 0001 |
WACV | 1 |
| 2025 | Ditailistener: Controllable High Fidelity Listener Video Generation with DiffusionabstractGenerating naturalistic and nuanced listener motions for extended interactions remains an open problem. Existing methods often rely on low-dimensional motion codes for facial behavior generation followed by photorealistic rendering, limiting both visual fidelity and expressive richness. To address these challenges, we introduce DiTaiListener, powered by a video diffusion model with multimodal conditions. Our approach first generates short segments of listener responses conditioned on the speaker's speech and facial motions with DiTaiListener-Gen. It then refines the transitional frames via DiTaiListener-Edit for a seamless transition. Specifically, DiTaiListener-Gen adapts a Diffusion Transformer (DiT) for the task of listener head portrait generation by introducing a Causal Temporal Multimodal Adapter (CTM-Adapter) to process speakers' auditory and visual cues. CTM-Adapter integrates speakers' input in a causal manner into the video generation process to ensure temporally coherent listener responses. For long-form video generation, we introduce DiTaiListener-Edit, a transition refinement video-to-video diffusion model. The model fuses video segments into smooth and continuous videos, ensuring temporal consistency in facial expressions and image quality when merging short video segments produced by DiTaiListener-Gen. Quantitatively, DiTaiListener achieves the state-of-the-art performance on benchmark datasets in both photorealism (+73.8% in FID on RealTalk) and motion representation (+6.1% in FD metric on VICO) spaces. User studies confirm the superior performance of DiTaiListener, with the model being the clear preference in terms of feedback, diversity, and smoothness, outperforming competitors by a significant margin. Maksim Siniukov, Di Chang, Minh Tran 0004, Hongkun Gong, Ashutosh Chaubey, Mohammad Soleymani 0001 |
ICCV | 5 |
| 2025 | ContextIQ: A Multimodal Expert-Based Video Retrieval System for Contextual AdvertisingabstractContextual advertising serves ads that are aligned to the content that the user is viewing. The rapid growth of video content on social platforms and streaming services, along with privacy concerns, has increased the need for contextual advertising. Placing the right ad in the right context creates a seamless and pleasant ad viewing experience, resulting in higher audience engagement and, ultimately, better ad monetization. From a technology standpoint, effective contextual advertising requires a video retrieval system capable of understanding complex video content at a very granular level. Current text-to-video retrieval models based on joint multimodal training demand large datasets and computational resources, limiting their practicality and lacking the key functionalities required for ad ecosystem integration. We introduce ContextIQ, a multimodal expert-based video retrieval system designed specifically for contextual advertising. ContextIQ utilizes modality-specific experts-video, audio, transcript (captions), and metadata such as objects, actions, emotion, etc.-to create semantically rich video representations. We show that our system, without joint training, achieves better or comparable results to state-of-the-art models and commercial solutions on multiple text-to-video retrieval benchmarks. Our ablation studies highlight the benefits of leveraging multiple modalities for enhanced video retrieval accuracy instead of using a vision-language model alone. Furthermore, we show how video retrieval systems such as ContextIQ can be used for contextual advertising in an ad ecosystem while also addressing concerns related to brand safety and filtering in-appropriate content. Ashutosh Chaubey, Anoubhav Agrawal, Sartaki Sinha Roy, Aayush Agrawal, Susmita Ghose |
WACV | 1 |
| 2023 | Meta-Learning Framework for End-to-End Imposter Identification in Unseen Speaker RecognitionabstractSpeaker identification systems are deployed in diverse environments, often different from the lab conditions on which they are trained and tested. In this paper, first, we show the problem of generalization using fixed thresholds (computed using EER metric) for imposter identification in unseen speaker recognition and then introduce a robust speaker-specific thresholding technique for better performance. Secondly, inspired by the recent use of meta-learning techniques in speaker verification, we propose an end-to-end meta-learning framework for imposter detection which decouples the problem of imposter detection from unseen speaker identification. Thus, unlike most prior works that use some heuristics to detect imposters, the proposed network learns to detect imposters by leveraging the utterances of the enrolled speakers. Furthermore, we show the efficacy of the proposed techniques on VoxCeleb1, VCTK and the FFSVC 2022 datasets, beating the baselines by up to $10 \%$. Ashutosh Chaubey, Sparsh Sinha, Susmita Ghose |
ASRU | 1 |
| 2022 | Improved Relation Networks for End-to-End Speaker Verification and IdentificationabstractSpeaker identification systems in a real-world scenario are tasked to identify a speaker amongst a set of enrolled speakers given just a few samples for each enrolled speaker.This paper demonstrates the effectiveness of meta-learning and relation networks for this use case.We propose improved relation networks for speaker verification and few-shot (unseen) speaker identification.The use of relation networks facilitates joint training of the frontend speaker encoder and the backend model.Inspired by the use of prototypical networks in speaker verification and to increase the discriminability of the speaker embeddings, we train the model to classify samples in the current episode amongst all speakers present in the training set.Furthermore, we propose a new training regime for faster model convergence by extracting more information from a given metalearning episode with negligible extra computation.We evaluate the proposed techniques on VoxCeleb, SITW and VCTK datasets on the tasks of speaker verification and unseen speaker identification.The proposed approach outperforms the existing approaches consistently on both tasks. Ashutosh Chaubey, Sparsh Sinha, Susmita Ghose |
INTERSPEECH | 1 |