VLDB 2026 Research / reviewers in the wild / expert
Guanyu Hu 0003
dblp:172/3540-3
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2026
0009-0006-1911-6283ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Fair Domain Generalization: An Information-Theoretic ViewabstractDomain generalization (DG) and algorithmic fairness are two key challenges in machine learning. However, most DG methods focus solely on minimizing expected risk in the unseen target domain, without considering algorithmic fairness. Conversely, fairness methods typically do not account for domain shifts, so the fairness achieved during training may not generalize to unseen test domains. In this work, we bridge these gaps by studying the problem of Fair Domain Generalization (FairDG), which aims to minimize both expected risk and fairness violations in unseen target domains. We derive novel mutual information-based upper bounds for expected risk and fairness violations in multi-class classification tasks with multi-group sensitive attributes. These bounds provide key insights for algorithm design from an information-theoretic perspective. Guided by these insights, we propose a practical method that solves the FairDG problem through Pareto optimization. Experiments on real-world vision and language datasets show that our method achieves superior utility–fairness trade-offs compared to existing approaches. Tangzheng Lian, Guanyu Hu 0003, Dimitris Kollias, Xinyu Yang 0001, Oya Çeliktutan |
AAAI | 2 |
| 2026 | From Cognitive Priors to Instance Semantics: A Unified Framework for Multi-task Affective ComputingabstractUnderstanding human affect via Valence-Arousal, Expressions, and Action Unit is essential for human-machine interaction. While recent multi-task learning (MTL) methods seek to unify these tasks, they overlook three key challenges: (i) the absence of unified modeling all three affective task types: regression, detection, and classification; (ii) reliance on complete annotations for all tasks, leaving disjoint single-task datasets underutilized; and (iii) task conflicts caused by Noisy Gradients, Negative Transfer (NT), and Task-specific Performance Misalignment (TPM). We introduce COIN, a novel two-stage MTL framework that bridges Cognitive Priors and Instance Semantics for robust training. First, we design a cognitively guided cross-task label induction strategy to propagate supervision under sparse annotations and mitigate NT, yielding strong task-specific CogXperts. Second, we introduce two complementary branches to address TPM: (i) Task-Specific Branch: transferring cognitive knowledge from task-optimal CogXperts to jointly optimize objectives under partial supervision, and (ii) Semantic Alignment Branch: enhancing instance-level semantic representations via Class-Conditioned and Instance-Adaptive Prompts. Experiments across six diverse datasets demonstrate COIN’s robustness and generalization. Code is available at https://github.com/imhgy/COIN. Guanyu Hu 0003, Dimitris Kollias, Xinyu Yang 0001 |
WACV | 1 |
| 2025 | Grounding Emotion Recognition with Visual Prototypes: VEGA - Revisiting CLIP in MERCabstractMultimodal Emotion Recognition in Conversations remains a challenging task due to the complex interplay of textual, acoustic and visual signals. While recent models have improved performance via advanced fusion strategies, they often lack psychologically meaningful priors to guide multimodal alignment. In this paper, we revisit the use of CLIP and propose a novel Visual Emotion Guided Anchoring (VEGA) mechanism that introduces class-level visual semantics into the fusion and classification process. Distinct from prior work that primarily utilizes CLIP's textual encoder, our approach leverages its image encoder to construct emotion-specific visual anchors based on facial exemplars. These anchors guide unimodal and multimodal features toward a perceptually grounded and psychologically aligned representation space, drawing inspiration from cognitive theories (prototypical emotion categories and multisensory integration). A stochastic anchor sampling strategy further enhances robustness by balancing semantic stability and intra-class diversity. Integrated into a dual-branch architecture with self-distillation, our VEGA-augmented model achieves sota performance on IEMOCAP and MELD. Code is available at: https://github.com/dkollias/VEGA Guanyu Hu 0003, Dimitris Kollias, Xinyu Yang 0001 |
ACM Multimedia | 1 |
| 2024 | Bridging the Gap: Protocol Towards Fair and Consistent Affect AnalysisabstractThe increasing integration of machine learning algorithms in daily life underscores the critical need for fairness and equity in their deployment. As these technologies play a pivotal role in decision-making, addressing biases across diverse subpopulation groups, including age, gender, and race, becomes paramount. Automatic affect analysis, at the inter-section of physiology, psychology, and machine learning, has seen significant development. However, existing databases and methodologies lack uniformity, leading to biased evaluations. This work addresses these issues by analyzing six affective databases, annotating demographic attributes, and proposing a common protocol for database partitioning. Emphasis is placed on fairness in evaluations. Extensive experiments with baseline and state-of-the-art methods demonstrate the impact of these changes, revealing the inadequacy of prior assessments. The findings underscore the importance of considering demographic attributes in affect analysis research and provide a foundation for more equitable methodologies. Our annotations, code and pre-trained models are available at: https://github.com/dkollias/Fair-Consistent-Affect-Analysis Guanyu Hu 0003, Eleni Papadopoulou, Dimitris Kollias, Paraskevi K. Tzouveli, Xinyu Yang 0001 |
FG | 1 |
| 2024 | Robust Facial Reactions Generation: An Emotion-Aware Framework with Modality CompensationabstractThe objective of the Multiple Appropriate Facial Reaction Generation (MAFRG) task is to produce contextually appropriate and diverse listener facial behavioural responses based on the multimodal behavioural data of the conversational partner (i.e., the speaker). Current methodologies typically assume continuous availability of speech and facial modality data, neglecting real-world scenarios where these data may be intermittently unavailable, which often results in model failures. Furthermore, despite utilising advanced deep learning models to extract information from the speaker’s multimodal inputs, these models fail to adequately leverage the speaker’s emotional context, which is vital for eliciting appropriate facial reactions from human listeners. To address these limitations, we propose an Emotion-aware Modality Compensatory (EMC) framework. This versatile solution can be seamlessly integrated into existing models, thereby preserving their advantages while significantly enhancing performance and robustness in scenarios with missing modalities. Our framework ensures resilience when faced with missing modality data through the Compensatory Modality Alignment (CMA) module. It also generates more appropriate emotion-aware reactions via the Emotion-aware Attention (EA) module, which incorporates the speaker’s emotional information throughout the entire encoding and decoding process. Experimental results demonstrate that our framework improves the appropriateness metric FRCorr by an average of 57.2% compared to the original model structure. In scenarios where speech modality data is missing, the performance of appropriate generation shows an improvement, and when facial data is missing, it only exhibits minimal degradation. Guanyu Hu 0003, Siyang Song, Dimitris Kollias, Xinyu Yang 0001, Zhonglin Sun, Odysseus Kaloidas |
IJCB | 1 |
| 2024 | Learning facial expression and body gesture visual information for video emotion recognition
Guanyu Hu 0003, Xinyu Yang 0001, Anh Tuan Luu, Yizhuo Dong |
Expert Syst. Appl. | 2 |
| 2024 | Artificial intelligence for optimizing recruitment and retention in clinical trials: a scoping reviewabstractOBJECTIVE: The objective of our research is to conduct a comprehensive review that aims to systematically map, describe, and summarize the current utilization of artificial intelligence (AI) in the recruitment and retention of participants in clinical trials. MATERIALS AND METHODS: A comprehensive electronic search was conducted using the search strategy developed by the authors. The search encompassed research published in English, without any time limitations, which utilizes AI in the recruitment process of clinical trials. Data extraction was performed using a data charting table, which included publication details, study design, and specific outcomes/results. RESULTS: The search yielded 5731 articles, of which 51 were included. All the studies were designed specifically for optimizing recruitment in clinical trials and were published between 2004 and 2023. Oncology was the most covered clinical area. Applying AI to recruitment in clinical trials has demonstrated several positive outcomes, such as increasing efficiency, cost savings, improving recruitment, accuracy, patient satisfaction, and creating user-friendly interfaces. It also raises various technical and ethical issues, such as limited quantity and quality of sample size, privacy, data security, transparency, discrimination, and selection bias. DISCUSSION AND CONCLUSION: While AI holds promise for optimizing recruitment in clinical trials, its effectiveness requires further validation. Future research should focus on using valid and standardized outcome measures, methodologically improving the rigor of the research carried out. Xiaoran Lu, Guanyu Hu 0003, Ziyi Zhong, Zihao Jiang 0013 |
J. Am. Medical Informatics Assoc. | 4 |
| 2023 | Multi-Scale Receptive Field Graph Model for Emotion Recognition in ConversationsabstractEmotion recognition in conversations (ERC) has gained more attention, where contextual information modeling and multimodal fusion have been the focus and challenges in recent years. In this paper, we proposed a Multi-Scale Receptive Field Graph model (MSRFG) to tackle the challenges of ERC. Specifically, MSRFG constructs multi-scale perception graphs and learns contextual information via parallel multi-scale receptive field paths. To compensate for the deficiency of temporal information learning by the graph network, MSRFG injects temporal dependencies into the graph network to model the temporal relationships between utterances. Moreover, to achieve the effective fusion of multimodal information, MSRFG converges the multi-scale features of each modality separately and performs the learning of attention weights after the integration of converged features. We carried out experiments on IEMOCAP and MELD datasets to validate the effectiveness of the proposed method, and the results proved the superiority of our model over the existing SOTA methods. The code is available at https://github.com/Janie1996/MSRFG1. Guanyu Hu 0003, Anh Tuan Luu, Xinyu Yang 0001 |
ICASSP | 2 |
| 2022 | Audio-Visual Domain Adaptation Feature Fusion for Speech Emotion Recognition
Guanyu Hu 0003, Xinyu Yang 0001, Anh Tuan Luu, Yizhuo Dong |
INTERSPEECH | 2 |