VLDB 2026 Research / reviewers in the wild / expert
Yan Zhuang 0002
dblp:02/5194-2
· DBLP profile ↗
13ranked-venue papers
7as first author
13since 2021 · last 2026
0000-0001-7444-3275ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TMDC: A Two-Stage Modality Denoising and Complementation Framework for Multimodal Sentiment Analysis with Missing and Noisy ModalitiesabstractMultimodal Sentiment Analysis (MSA) aims to infer human sentiment by integrating information from multiple modalities such as text, audio, and video. In real-world scenarios, however, the presence of missing modalities and noisy signals significantly hinders the robustness and accuracy of existing models. While prior works have made progress on these issues, they are typically addressed in isolation, limiting overall effectiveness in practical settings. To jointly mitigate the challenges posed by missing and noisy modalities, we propose a framework called Two-stage Modality Denoising and Complementation (TMDC). TMDC comprises two sequential training stages. In the Intra-Modality Denoising Stage, denoised modality-specific and modality-shared representations are extracted from complete data using dedicated denoising modules, reducing the impact of noise and enhancing representational robustness. In the Inter-Modality Complementation Stage, these representations are leveraged to compensate for missing modalities, thereby enriching the available information and further improving robustness. Extensive evaluations on MOSI, MOSEI, and IEMOCAP demonstrate that TMDC consistently achieves superior performance compared to existing methods, establishing new state-of-the-art results. Yan Zhuang 0002, Minhao Liu, Yanru Zhang, Jiawen Deng 0006, Fuji Ren |
AAAI | 1 |
| 2026 | Beyond Explicit Refusals: Soft-Failure Attacks on Retrieval-Augmented GenerationabstractExisting jamming attacks on Retrieval-Augmented Generation (RAG) systems typically induce explicit refusals or denial-ofservice behaviors, which are conspicuous and easy to detect.In this work, we formalize a subtler availability threat, termed soft failure, which degrades system utility by inducing fluent and coherent yet non-informative responses rather than overt failures.We propose Deceptive Evolutionary Jamming Attack (DEJA), an automated black-box attack framework that generates adversarial documents to trigger such soft failures by exploiting safety-aligned behaviors of large language models.DEJA employs an evolutionary optimization process guided by a fine-grained Answer Utility Score (AUS), computed via an LLM-based evaluator, to systematically degrade the certainty of answers while maintaining high retrieval success.Extensive experiments across multiple RAG configurations and benchmark datasets show that DEJA consistently drives responses toward low-utility soft failures, achieving SASR above 79% while keeping hard-failure rates below 15%, significantly outperforming prior attacks.The resulting adversarial documents exhibit high stealth, evading perplexity-based detection and resisting query paraphrasing, and transfer across model families to proprietary systems without retargeting. Wentao Zhang 0011, Yan Zhuang 0002, ZhuHang Zheng, Mingfei Zhang, Jiawen Deng 0006, Fuji Ren |
ACL (1) | 2 |
| 2026 | ReNoRD: Learning from Relations under Noisy Pseudo Labels via Relational Distillation for Multimodal Sentiment
Tiantai Zhai, Yan Zhuang 0002, Fuji Ren, Jiawen Deng 0006 |
ICMR | 2 |
| 2026 | Decoupled hypergraph modeling for multimodal sentiment analysis
Yanping Huang, Jiawen Deng 0006, Yan Zhuang 0002, Jiali You 0002, Fuji Ren |
Neurocomputing | 3 |
| 2026 | Intra-Sample and Intra-Modal Enhancement for Multimodal Sentiment Analysis With Missing ModalitiesabstractMultimodal sentiment analysis (MSA) with missing modalities involves understanding the person's sentiment using multimodal data where some modalities are missing. Most existing methods focus on reconstructing the missing modalities using the available modalities from each sample, relying on modality-common information. However, these methods overlook the modality-specific information that other samples can provide. Additionally, these approaches often require the guidance of full modality representations during the reconstruction process, which is impractical in resource-constrained real-world scenarios. To address these challenges, we propose theIntra-sample andIntra-modalEnhancement (IIE) framework. The IIE framework enhances both sample-level and modality-level representations to capture additional modality-common and modality-specific information from existing modalities, without requiring full modalities. Specifically, IIE first learns sample-level representations by distilling modality-common information from the available modalities into learnable latent units. Then, it enhances modality-level representations by leveraging modality-specific information from other samples with the same modality, which is crucial for improving robustness in the presence of missing modalities. Finally, IIE ensures consistency between the enhanced modality-level and sample-level representations, combining the enhanced and initial representations to make predictions. Extensive experiments on three datasets demonstrate that the IIE framework significantly outperforms existing methods in terms of both effectiveness and robustness in handling MSA with missing modalities. Code is available athttps://github.com/YetZzzzzz/IIE. Yan Zhuang 0002, Yanru Zhang, Jiawen Deng 0006, Fuji Ren |
IEEE Trans. Multim. | 1 |
| 2025 | CMAD: Correlation-Aware and Modalities-Aware Distillation for Multimodal Sentiment Analysis with Missing Modalities
Yan Zhuang 0002, Minhao Liu, Yanru Zhang, Jiawen Deng 0006, Fuji Ren |
ICCV | 1 |
| 2025 | FAME: Fusion-Aware Multi-modal Ensemble for Social Media Popularity PredictionabstractAs social media becomes a dominant platform for sharing content, predicting the popularity of user posts has become increasingly important for applications such as content recommendation, trend forecasting, and user engagement. However, this task is challenging due to the diverse and multimodal nature of social media posts, which often include unstructured text, images, and structured metadata. To address this challenge, we propose Fusion-Aware Multi-modal Ensemble (FAME), a framework effectively captures and integrates diverse information sources within social media content. Unlike prior approaches that rely on a single model to process all modalities, FAME leverages four specialized predictors. Three of them-CatBoost, LightGBM, and AutoGluon-are tree-based models that excel at handling structured metadata and its interactions with unstructured features. The fourth is a denoising autoencoder (DAE), which learns robust joint representations from unstructured text and image data. These models are combined through a weighted ensemble strategy, allowing FAME to leverage the complementary strengths of different architectures. Experiments on the Social Media Prediction Dataset demonstrate that FAME significantly outperforms existing baselines, achieving state-of-the-art results and validating its effectiveness in modeling the complex, multimodal nature of social media content. Yan Zhuang 0002, Yanru Zhang, Minhao Liu, Jiawen Deng 0006, Fuji Ren |
ACM Multimedia | 1 |
| 2025 | Hyper-Modality Enhancement for Multimodal Sentiment Analysis with Missing ModalitiesabstractMultimodal Sentiment Analysis (MSA) aims to infer human emotions by integrating complementary signals from diverse modalities. However, in real-world scenarios, missing modalities are common due to data corruption, sensor failure, or privacy concerns, which can significantly degrade model performance. To tackle this challenge, we propose Hyper-Modality Enhancement (HME), a novel framework that avoids explicit modality reconstruction by enriching each observed modality with semantically relevant cues retrieved from other samples. This cross-sample enhancement reduces reliance on fully observed data during training, making the method better suited to scenarios with inherently incomplete inputs. In addition, we introduce an uncertainty-aware fusion mechanism that adaptively balances original and enriched representations to improve robustness. Extensive experiments on three public benchmarks show that HME consistently outperforms state-of-the-art methods under various missing modality conditions, demonstrating its practicality in real-world MSA applications. Yan Zhuang 0002, Minhao Liu, Yanru Zhang, Wei Li 0308, Jiawen Deng 0006, Fuji Ren |
NeurIPS | 1 |
| 2025 | ETS-MM: A Multi-Modal Social Bot Detection Model Based on Enhanced Textual Semantic RepresentationabstractSocial bots are becoming increasingly common in social networks, and their activities affect the security and authenticity of social media platforms. Current state-of-the-art social bot detection methods leverage multimodal approaches that analyze various modalities, such as user metadata, text, and social network relationships. However, these methods may not always extract additional dimensions of semantic feature information that could offer a deeper understanding of users' social patterns. To address this issue, we propose ETS-MM, a multimodal detection framework designed to augment multidimensional information from text and extract the semantic feature representation of user text information. We first analyze the user's tweeting behavior based on topic preference and emotion tendency, integrating them into the textual data. Then, we try to extract enhanced semantic representations that reveal the latent relationship between tweeting behavior and tweet content while identifying potential contextual associations and emotional changes. Additionally, to capture the complex interaction between users, we integrate the user's multimodal information, including metadata, textual features, enhanced semantic features, and social network relationships to propagate and aggregate information across various modalities. Experimental results demonstrate that ETS-MM significantly outperforms existing methods across two widely used social bot detection benchmark datasets, validating its effectiveness and superiority. Wei Li 0308, Jiawen Deng 0006, Jiali You 0002, Yan Zhuang 0002, Fuji Ren |
WWW | 5 |
| 2025 | Enhanced Emotion Recognition in Conversations Through Hybrid Context Encoding and Latent Dependency MiningabstractEmotion recognition in conversations (ERC) is a pivotal component of affective computing, involving a common two-stage paradigm where pre-trained language models first extract context-independent features, followed by the encoding of contextual information and the modeling of emotional dependencies. This paradigm faces two challenges: (1) Existing methods struggle to capture both the intra-dialogue emotional continuity and the inter-dialogue semantic similarity. (2) The complexity of emotional elicitation processes gives rise to entangled dependencies, termed “latent dependencies”, which are difficult for current methods to detect and analyze. To overcome these challenges, we propose a Hybrid-Context Encoder with an Automated Latent Dependency Mining model for ERC. Specifically, we examine the emotional continuity and the semantic similarity from the standpoint of context encoders. We experimentally find that context encoders with different architectures exhibit distinct benefits. Based on these findings, we design a hybrid contextual encoding module that effectively combines the strengths of various encoders. Additionally, we design a lightweight generative module for latent dependency mining that autonomously generates a context mask, enabling the effective discovery of latent dependencies. We conduct extensive experiments on three datasets in the text modality. Our model achieves the best performance, which validates the superiority of our approach. Zheng Hu 0001, Jiawen Deng 0006, Satoshi Nakagawa, Yan Zhuang 0002, Shimin Cai, Fuji Ren |
IEEE Trans. Affect. Comput. | 4 |
| 2025 | Hierarchical Denoising for Robust Social RecommendationabstractSocial recommendations leverage social networks to augment the performance of recommender systems. However, the critical task of denoising social information has not been thoroughly investigated in prior research. In this study, we introduce a hierarchical denoising robust social recommendation model to tackle noise at two levels: 1) intra-domain noise, resulting from user multi-faceted social trust relationships, and 2) inter-domain noise, stemming from the entanglement of the latent factors over heterogeneous relations (e.g., user-item interactions, user-user trust relationships). Specifically, our model advances a preference and social psychology-aware methodology for the fine-grained and multi-perspective estimation of tie strength within social networks. This serves as a precursor to an edge weight-guided edge pruning strategy that refines the model's diversity and robustness by dynamically filtering social ties. Additionally, we propose a user interest-aware cross-domain denoising gate, which not only filters noise during the knowledge transfer process but also captures the high-dimensional, nonlinear information prevalent in social domains. We conduct extensive experiments on three real-world datasets to validate the effectiveness of our proposed model against state-of-the-art baselines. We perform empirical studies on synthetic datasets to validate the strong robustness of our proposed model. Zheng Hu 0001, Satoshi Nakagawa, Yan Zhuang 0002, Jiawen Deng 0006, Shimin Cai, Tao Zhou 0001, Fuji Ren |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2025 | Multi-Level Contrastive Learning for Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis has garnered increasing attention. The bulk of existing work in multimodal sentiment analysis primarily focuses on designing various networks to align and subsequently fuse representations from individual modalities. Contrastive learning, recognized for its intrinsic alignment capabilities, has also been extensively applied in multimodal sentiment analysis. However, current contrastive learning methods are often limited to pairwise modalities and typically perform contrastive learning prior to modality fusion, neglecting the consistency of interactions across multiple modalities. Moreover, they overlook the overall consistency within samples. To address these issues, we introduce a novel Multi-Level Contrastive Learning (MLCL) framework for multimodal sentiment analysis, composed of Uni-Modal Contrastive Learning (UMCL), Bi-Modal Contrastive Learning (BMCL) and Tri-Modal Contrastive Learning (TMCL). UMCL enhances intra-modal representations by creating positive pairs using modality-specific random dropout, while BMCL leverages the asymmetry of attention mechanisms, using two directional attentions as positive samples. TMCL aligns non-overlapping uni-modal and bi-modal representations, underscoring the complementarity of tri-modal information. The effectiveness of MLCL is demonstrated through its performance on multiple datasets. Our comprehensive experiments across multiple datasets demonstrate the superiority of the MLCL framework, which achieves new state-of-the-art performance. Yan Zhuang 0002, Yanru Zhang, Jiawen Deng 0006, Zheng Hu 0001, Fuji Ren |
IEEE Trans. Multim. | 1 |
| 2024 | GLoMo: Global-Local Modal Fusion for Multimodal Sentiment AnalysisabstractMultimodal Sentiment Analysis (MSA) has witnessed remarkable progress and gained increasing attention in recent decade. However, current MSA methodologies primarily rely on global representations extracted from different modalities, such as the mean of all token representations, to construct sophisticated fusion networks. These approaches often overlook the valuable details present in local representations, which consist of fused representations of consecutive several tokens. Additionally, the integration of multiple local representations, and the fusion of local and global information present significant challenges. To address these limitations, we propose the Global-Local Modal (GLoMo) Fusion framework. It comprises two essential components: (i) modality-specific mixture of experts layers that integrate diverse local representations within each modality, and (ii) a global-guided fusion module that effectively combines global and local representations. The former component leverages specialized expert networks to automatically select and integrate crucial local representations from each modality, while the latter ensures the preservation of global information during the fusion process. We evaluate GLoMo on various datasets, encompassing tasks in multimodal sentiment analysis, multimodal humor detection, and multimodal emotion recognition. Extensive experiments demonstrate that GLoMo outperforms existing state-of-the-art models, validating the effectiveness of our proposed framework. Our code is publicly available at https://github.com/YetZzzzzz/GLoMo. Yan Zhuang 0002, Yanru Zhang, Zheng Hu 0001, Jiawen Deng 0006, Fuji Ren |
ACM Multimedia | 1 |