VLDB 2026 Research / reviewers in the wild / expert
Minh Tran 0004
dblp:50/2494-4
· DBLP profile ↗
16ranked-venue papers
12as first author
15since 2021 · last 2026
0009-0004-2391-3563ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 10 first-author · 13 since 2021Artificial intelligence and machine learning · 10 · 8 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Discrete Facial Encoding: A Framework for Data-driven Facial Display DiscoveryabstractFacial expression analysis is central to understanding human behavior, yet existing coding systems such as the Facial Action Coding System (FACS) are constrained by limited coverage and costly manual annotation. In this work, we introduce Discrete Facial Encoding (DFE), an unsupervised, data-driven alternative of compact and interpretable dictionary of facial expressions from 3D mesh sequences learned through a Residual Vector Quantized Variational Autoencoder (RVQ-VAE). Our approach first extracts identity-invariant expression features from images using a 3D Morphable Model (3DMM), effectively disentangling factors such as head pose and facial geometry. We then encode these features using an RVQ-VAE, producing a sequence of discrete tokens from a shared codebook, where each token captures a specific, reusable facial deformation pattern that contributes to the overall expression. Through extensive experiments, we demonstrate that Discrete Facial Encoding captures more precise facial behaviors than FACS and other facial encoding alternatives. We evaluate the utility of our representation across three high-level psychological tasks: stress detection, personality prediction, and depression detection. Using a simple Bag-of-Words model built on top of the learned tokens, our system consistently outperforms both FACS-based pipelines and strong image and video representation learning models such as Masked Autoencoders. Further analysis reveals that our representation covers a wider variety of facial displays, highlighting its potential as a scalable and effective alternative to FACS for psychological and affective computing applications. We released the source code and model weights at https://github.com/ihp-lab/vqface. Minh Tran 0004, Maksim Siniukov, Zhangyu Jin, Mohammad Soleymani 0001 |
WACV | 1 |
| 2025 | Ditailistener: Controllable High Fidelity Listener Video Generation with DiffusionabstractGenerating naturalistic and nuanced listener motions for extended interactions remains an open problem. Existing methods often rely on low-dimensional motion codes for facial behavior generation followed by photorealistic rendering, limiting both visual fidelity and expressive richness. To address these challenges, we introduce DiTaiListener, powered by a video diffusion model with multimodal conditions. Our approach first generates short segments of listener responses conditioned on the speaker's speech and facial motions with DiTaiListener-Gen. It then refines the transitional frames via DiTaiListener-Edit for a seamless transition. Specifically, DiTaiListener-Gen adapts a Diffusion Transformer (DiT) for the task of listener head portrait generation by introducing a Causal Temporal Multimodal Adapter (CTM-Adapter) to process speakers' auditory and visual cues. CTM-Adapter integrates speakers' input in a causal manner into the video generation process to ensure temporally coherent listener responses. For long-form video generation, we introduce DiTaiListener-Edit, a transition refinement video-to-video diffusion model. The model fuses video segments into smooth and continuous videos, ensuring temporal consistency in facial expressions and image quality when merging short video segments produced by DiTaiListener-Gen. Quantitatively, DiTaiListener achieves the state-of-the-art performance on benchmark datasets in both photorealism (+73.8% in FID on RealTalk) and motion representation (+6.1% in FD metric on VICO) spaces. User studies confirm the superior performance of DiTaiListener, with the model being the clear preference in terms of feedback, diversity, and smoothness, outperforming competitors by a significant margin. Maksim Siniukov, Di Chang, Minh Tran 0004, Hongkun Gong, Ashutosh Chaubey, Mohammad Soleymani 0001 |
ICCV | 3 |
| 2025 | Head2Body: Body Pose Generation from Multi-Sensory Head-Mounted Inputs
Minh Tran 0004, Hongda Mao, Qingshuang Chen, Yelin Kim |
ICCV | 1 |
| 2025 | SetPeER: Set-Based Personalized Emotion Recognition With Weak SupervisionabstractIndividual variability of expressive behaviors is a major challenge for emotion recognition systems. Personalized emotion recognition strives to adapt machine learning models to individual behaviors, thereby enhancing emotion recognition performance and overcoming the limitations of generalized emotion recognition systems. However, existing datasets for audiovisual emotion recognition either have a very low number of data points per speaker or include a limited number of speakers. The scarcity of data significantly limits the development and assessment of personalized models, hindering their ability to effectively learn and adapt to individual expressive styles. This paper introduces EmoCeleb: a large-scale, weakly labeled emotion dataset generated via cross-modal labeling. EmoCeleb comprises over 150 hours of audiovisual content from approximately 1,500 speakers, with a median of 50 utterances per speaker. This rich dataset provides a rich resource for developing and benchmarking personalized emotion recognition methods, including those requiring substantial data per individual, such as set learning approaches. We also propose SetPeER: a novel personalized emotion recognition architecture employing set learning. SetPeER effectively captures individual expressive styles by learning representative speaker features from limited data, achieving strong performance with as few as eight utterances per speaker. By leveraging set learning, SetPeER overcomes the limitations of previous approaches that struggle to learn effectively from limited data per individual. Through extensive experiments on EmoCeleb and established benchmarks,i.e, MSP-Podcast and MSP-Improv, we demonstrate the effectiveness of our dataset and the superior performance of SetPeER compared to existing methods for emotion recognition. Our work paves the way for more robust and accurate personalized emotion recognition systems. Minh Tran 0004, Yufeng Yin 0002, Mohammad Soleymani 0001 |
IEEE Trans. Affect. Comput. | 1 |
| 2024 | DIM: Dyadic Interaction Modeling for Social Behavior Generation
Minh Tran 0004, Di Chang, Maksim Siniukov, Mohammad Soleymani 0001 |
ECCV (37) | 1 |
| 2024 | Ex2Eg-MAE: A Framework for Adaptation of Exocentric Video Masked Autoencoders for Egocentric Social Role Understanding
Minh Tran 0004, Yelin Kim, Che-Chun Su, Cheng-Hao Kuo, Min Sun 0001, Mohammad Soleymani 0001 |
ECCV (80) | 1 |
| 2024 | Social-MAE: A Transformer-Based Multimodal Autoencoder for Face and VoiceabstractHuman social behaviors are inherently multi-modal necessitating the development of powerful audiovisual models for their perception. In this paper, we present Social-MAE, our pre-trained audiovisual Masked Autoencoder based on an extended version of Contrastive Audio-Visual Masked Auto-Encoder (CAV-MAE), which is pre-trained on audiovisual social data. Specifically, we modify CAV-MAE to receive a larger number of frames as input and pre-train it on a large dataset of human social interaction (VoxCeleb2) in a self-supervised manner. We demonstrate the effectiveness of this model by fine-tuning and evaluating the model on different social and affective downstream tasks, namely, emotion recognition, laughter detection and apparent personality estimation. The model achieves state-of-the-art results on multimodal emotion recognition and laughter recognition and competitive results for apparent personality estimation, demonstrating the effectiveness of in-domain self-supervised pre-training. Code and model weight are available here https://github.com/HuBohy/SocialMAE. Hugo Bohy, Minh Tran 0004, Kevin El Haddad, Thierry Dutoit, Mohammad Soleymani 0001 |
FG | 2 |
| 2024 | LibreFace: An Open-Source Toolkit for Deep Facial Expression AnalysisabstractFacial expression analysis is an important tool for human-computer interaction. In this paper, we introduce LibreFace, an open-source toolkit for facial expression analysis. This open-source toolbox offers real-time and offline analysis of facial behavior through deep learning models, including facial action unit (AU) detection, AU intensity estimation, and facial expression recognition. To accomplish this, we employ several techniques, including the utilization of a large-scale pre-trained network, feature-wise knowledge distillation, and task-specific fine-tuning. These approaches are designed to effectively and accurately analyze facial expressions by leveraging visual information, thereby facilitating the implementation of real-time interactive applications. In terms of Action Unit (AU) intensity estimation, we achieve a Pearson Correlation Coefficient (PCC) of 0.63 on DISFA, which is 7% higher than the performance of OpenFace 2.0 [4] while maintaining highly-efficient inference that runs two times faster than OpenFace 2.0 [4]. Despite being compact, our model also demonstrates competitive performance to state-of-the-art facial expression analysis methods on AffecNet, FFHQ, and RAF-DB. Our code will be released at https://github.com/ihp-lab/LibreFace Di Chang, Yufeng Yin 0002, Zongjian Li, Minh Tran 0004, Mohammad Soleymani 0001 |
WACV | 4 |
| 2023 | A Speech Representation Anonymization Framework via Selective Noise PerturbationabstractPrivacy and security are major concerns when communicating speech signals to cloud services such as automatic speech recognition (ASR) and speech emotion recognition (SER). Existing solutions for speech anonymization mainly focus on voice conversion or voice modification to convert a raw utterance into another one with similar content but different, or no, identity-related information. However, an alternative approach to share speech data under the form of privacy-preserving representation has been largely under-explored. In this paper, we propose a speech anonymization framework that achieves privacy via noise perturbation to a selected subset of the high-utility representations extracted using a pre-trained speech encoder. The subset is chosen with a Transformer-based privacy-risk saliency estimator. We validate our framework on four tasks, namely, Automatic Speaker Verification (ASV), ASR, SER and Intent Classification (IC) for privacy and utility assessment. Experimental results show that our approach is able to achieve a competitive, or even superior, utility compared to the speech anonymization baselines from the VoicePrivacy2022 Challenges, while maintaining the same level of privacy. Moreover, the easily-controlled amount of perturbation allows our framework to have a flexible range of privacy-utility trade-offs without re-training any component. Minh Tran 0004, Mohammad Soleymani 0001 |
ICASSP | 1 |
| 2023 | Personalized Adaptation with Pre-trained Speech Encoders for Continuous Emotion Recognition
Minh Tran 0004, Yufeng Yin 0002, Mohammad Soleymani 0001 |
INTERSPEECH | 1 |
| 2023 | Privacy-preserving Representation Learning for Speech Understanding
Minh Tran 0004, Mohammad Soleymani 0001 |
INTERSPEECH | 1 |
| 2023 | SAAML: A Framework for Semi-supervised Affective Adaptation via Metric LearningabstractSocially intelligent systems such as home robots should be able to perceive emotions and social behaviors. Affect recognition datasets have limited labeled data, and existing large unlabeled datasets, e.g., VoxCeleb2, suitable for pre-training, mostly contain neutral expressions, limiting their application to affective downstream tasks. We introduce a novel Semi-supervised Affective Adaptation framework via Metric Learning (SAAML) to adapt pre-trained audiovisual models (e.g., AV-HuBERT) to expressive behaviors associated with emotions and social communication. The proposed framework automatically retrieves a large number of emotional excerpts (>100 hours) from the VoxCeleb2 dataset via metric learning from two emotion recognition datasets (MSP-IMPROV and CREMA-D), and learns domain-invariant emotion-aware representations. Experimental results show that fine-tuning the proposed affect-aware AV-HuBERT (AW-HuBERT) improves the emotion recognition accuracy by 3-6% compared to fine-tuning the original pre-trained models. We further validate the effectiveness of the AW-HuBERT on human-centered visual understanding tasks, namely, facial expression recognition, video highlight detection, and continuous emotion recognition. The proposed approach consistently outperforms AV-HuBERT and delivers competitive performance compared to the existing methods. With this work, we demonstrate the effectiveness of adaptive pre-training for existing models on domain-specific data to enhance their performance for human-centered tasks. Minh Tran 0004, Yelin Kim, Che-Chun Su, Cheng-Hao Kuo, Mohammad Soleymani 0001 |
ACM Multimedia | 1 |
| 2022 | A Pre-Trained Audio-Visual Transformer for Emotion RecognitionabstractIn this paper, we introduce a pretrained audio-visual Transformer trained on more than 500k utterances from nearly 4000 celebrities from the VoxCeleb2 dataset for human behavior understanding. The model aims to capture and extract useful information from the interactions between human facial and auditory behaviors, with application in emotion recognition. We evaluate the model performance on two datasets, namely CREMAD-D (emotion classification) and MSP-IMPROV (continuous emotion regression). Experimental results show that fine-tuning the pre-trained model helps improving emotion classification accuracy by 5-7% and Concordance Correlation Coefficients (CCC) in continuous emotion recognition by 0.03-0.09 compared to the same model trained from scratch. We also demonstrate the robustness of finetuning the pre-trained model in a low-resource setting. With only 10% of the original training set provided, finetuning the pre-trained model can lead to at least 10% better emotion recognition accuracy and a CCC score improvement by at least 0.1 for continuous emotion recognition. Minh Tran 0004, Mohammad Soleymani 0001 |
ICASSP | 1 |
| 2021 | Modeling Dynamics of Facial Behavior for Mental Health AssessmentabstractFacial action unit (FAU) intensities are popular descriptors for the analysis of facial behavior. However, FAUs are sparsely represented when only a few are activated at a time. In this study, we explore the possibility of representing the dynamics of facial expressions by adopting algorithms used for word representation in natural language processing. Specifically, we perform clustering on a large dataset of temporal facial expressions with 5.3M frames before applying the Global Vector representation (GloVe) algorithm to learn the embeddings of the facial clusters. We evaluate the usefulness of our learned representations on two downstream tasks: schizophrenia symptom estimation and depression severity regression. These experimental results show the potential of our approach for improving the assessment of mental health symptoms over baseline models that use FAU intensities alone. Minh Tran 0004, Ellen Bradley, Michelle Matvey, Joshua Woolley, Mohammad Soleymani 0001 |
FG | 1 |
| 2021 | A Systematic Cross-Corpus Analysis of Human Reactions to Robot Conversational FailuresabstractIn this paper, we analyze multimodal behavioral responses to robot failures across different tasks. Two multimodal datasets are examined in which humans interact with guided-task robots in task-oriented dialogues. In both datasets, the robots simulated failures of conversational breakdown and miscommunication typically observed in human-robot interactions. We closely examine human reactions to these failures looking at facial and acoustic features. Our analyses identify the significant behavioral features for automatic detection of such failures in interaction. We also examine human responses to different types of robot failures and if failures occurred early or late in the interaction cause variation in the responses. Our findings indicate that several nonverbal behaviors are consistently present in responses to robots’ failures, e.g., gaze and speech prosody, whereas, linguistic features appear to be task-dependent. We discuss how these findings may generalize to other tasks, and how autonomous robots may identify opportunities to detect and recover from failures in interactions with humans. Dimosthenis Kontogiorgos, Minh Tran 0004, Joakim Gustafson, Mohammad Soleymani 0001 |
ICMI | 2 |
| 2020 | Towards A Friendly Online Community: An Unsupervised Style Transfer Framework for Profanity RedactionabstractOffensive and abusive language is a pressing problem on social media platforms.In this work, we propose a method for transforming offensive comments, statements containing profanity or offensive language, into non-offensive ones.We design a RETRIEVE, GENERATE and EDIT unsupervised style transfer pipeline to redact the offensive comments in a word-restricted manner while maintaining a high level of fluency and preserving the content of the original text.We extensively evaluate our method's performance and compare it to previous style transfer models using both automatic metrics and human evaluations.Experimental results show that our method outperforms other models on human evaluations and is the only approach that consistently performs well on all automatic evaluation metrics. Minh Tran 0004, Mohammad Soleymani 0001 |
COLING | 1 |