VLDB 2026 Research / reviewers in the wild / expert
Guanghui Ye
dblp:28/1313
· DBLP profile ↗
17ranked-venue papers
7as first author
17since 2021 · last 2026
0009-0001-7713-6550ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PLUM-Net: Prototype-Induced Label Structuring for Disentangled Multimodal Representation NetworkabstractExisting multimodal representation learning approaches often rely on simple feature concatenation or unified transformations, which fail to effectively disentangle and leverage common and private information across different modalities in a progressive manner. Moreover, they typically lack adaptive modeling tailored to specific task requirements. To address these limitations, we propose a Prototype-Induced Label Structuring for Disentangled Multimodal Representation Network (PLUM-Net). It first employs a multilevel semantic alignment module to synchronize global and local semantics across audio, visual and textual streams. On this aligned foundation, a prototype-based single-modal label generation module derives modality-specific hard and soft-labels that subtly steer the network toward a cleaner split between shared and private cues. Guided by these labels, the task-conditioned feature bifurcator module channels information through the most beneficial common or private pathway for the given task, after which a private refinement module polishes and fuses each modality’s idiosyncratic signals. Extensive experiments show that PLUM-Net delivers strong performance on datasets such as CMU-MOSI, CMU-MOSEI and UR-FUNNY, achieving an ACC-2 of 90.3% on CMU-MOSI, representing a 2%–4% improvement over previous SOTA models. Huan Zhao 0003, Xupeng Zha, Guanghui Ye, Zixing Zhang 0001 |
AAAI | 5 |
| 2026 | Making Visual Dialogue More Engaging: A New Task, Method, and MetricabstractLarge language model (LLM)-based visual dialogue (VD) systems have made response generation for image-grounded conversations more correct and coherent. However, user engagement - the extent to which a user is interested, emotionally involved, and willing to continue the conversation - remains a challenge. To fully explore engaging VD, we propose: (i) a new task named Audio-enhanced VD (AVD), which introduces additional audio dialogue contexts that can more vividly convey the speaker's emotions as input, with the aim of generating correct but more engaging dialogue responses. Specifically, we employ a text-to-speech model as the modality translator to generate the paired acoustic utterances from the inputting textual utterances; (ii) an accompanying approach named Visually-grounded and Interleaved Text-Audio Dialogue Modeling (VITA-DM), which utilizes both image-grounded information and interleaved text-audio utterances for visual dialogue modeling, differentiating from previous multi-modal LLM (MLLM)-based methods that normally model text and audio modalities separately. We also present three pre-training tasks to better learn multi-modal interactions across language, vision, and audio; (iii) a novel metric named Multi-Modal Engagement (MME), which fills the gap of engagement estimation in VD and can provide a fine-grained assessment along emotional, attentional, and reply engagement dimensions (EE, AE, RE). We experiment on two popular datasets and provide extensive evaluations (automatic, engagement-specific, and human), supporting the validity of our approach. Furthermore, based on empirical results that reveal that emotions contribute the most to engagement, we justify our emphasis on the emotional aspect throughout the definition, solution, and evaluation of our task. Guanghui Ye, Huan Zhao 0003, Yingxue Gao, Zhixue Zhao, Xupeng Zha, Zhihua Jiang |
AAAI | 1 |
| 2026 | Boosting Adversarial Transferability via Ensemble Non-AttentionabstractEnsemble attacks integrate the outputs of surrogate models with diverse architectures, which can be combined with various gradient-based attacks to improve adversarial transferability. However, previous work shows unsatisfactory attack performance when transferring across heterogeneous model architectures. The main reason is that the gradient update directions of heterogeneous surrogate models differ widely, making it hard to reduce the gradient variance of ensemble models while making the best of individual model. To tackle this challenge, we design a novel ensemble attack, NAMEA, which for the first time integrates the gradients from the non-attention areas of ensemble models into the iterative gradient optimization process. Our design is inspired by the observation that the attention areas of heterogeneous models vary sharply, thus the non-attention areas of ViTs are likely to be the focus of CNNs and vice versa. Therefore, we merge the gradients respectively from the attention and non-attention areas of ensemble models so as to fuse the transfer information of CNNs and ViTs. Specifically, we pioneer a new way of decoupling the gradients of non-attention areas from those of attention areas, while merging gradients by meta-learning. Empirical evaluations on ImageNet dataset indicate that NAMEA outperforms AdaEA and SMER, the state-of-the-art ensemble attacks by an average of 15.0% and 9.6%, respectively. This work is the first attempt to explore the power of ensemble non-attention in boosting cross-architecture transferability, providing new insights into launching ensemble attacks. Yipeng Zou, Qin Liu 0001, Jie Wu 0001, Yu Peng 0003, Guo Chen 0001, Hui Zhou 0014, Guanghui Ye |
AAAI | 7 |
| 2026 | Towards Mitigating Modality Bias in Vision-Language Models for Temporal Action LocalizationabstractJiaqi Li, Guangming Wang, Shuntian Zheng, Minzhe Ni, Xiaoman Lu, Guanghui Ye, Yu Guan. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jiaqi Li 0008, Guangming Wang 0001, Shuntian Zheng, Minzhe Ni, Xiaoman Lu, Guanghui Ye, Yu Guan 0001 |
ACL (1) | 6 |
| 2026 | Automatic and Reliable Evaluation for Academic Caption-to-Figure Generation with LMMsabstractGuanghui Ye, Huan Zhao, Qin Zhu, Fengnan Li, Jiaqi Li, Yixian Shen, Zhonghao Ren, Zhihua Jiang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Guanghui Ye, Huan Zhao 0003, Fengnan Li, Jiaqi Li 0008, Yixian Shen, Zhonghao Ren, Zhihua Jiang |
ACL (1) | 1 |
| 2026 | Generating Multi-Modal Knowledge Clues as an Image: Toward Improving Image-Sequence Reasoning With Assisted Visual InputabstractRecent multi-modal large language models (MLLMs) have exhibited powerful abilities in addressing complex vision-language tasks such as image-sequence reasoning (ISR). However, significant challenges remain, e.g., it is still difficult for the MLLMs to fully capture and represent cross-image visual knowledge such as scene relations, attributes, and entity links between multiple images, which hinders them from better solving ISR. To alleviate these issues, we introduce a novel concept Visualized Knowledge Clue (VizKC) - synthetic images that encode key visual and external knowledge from a sequence of input images and are then used alongside the original input images within a multi-image MLLM to enhance reasoning performance. Accordingly, we propose an accompanying approach named VizKC-ISR, composed of two modules - VizKC generation and VizKC utilization. Specifically, in the generation module, VizKC-ISR follows aSee-Find-Fusepipeline: (i) “See - Scene Perception”, to construct an initial VizKC that incorporates scene relations of key visual entities detected from an original image; (ii) “Find - Knowledge Generation”, to generate enriched image captions with real-world knowledge and fine-grained entity details and then extract structured knowledge tuples from generated captions; (iii) “Fuse - Image Editing”, to introduce relevant knowledge tuples into the VizKC via iterative image editing. In the utilization module, we employ a multi-image MLLM (e.g., mPLUG-Owl3) to solve the VizKC-assisted ISR tasks by reasoning with generated knowledge clues. We evaluate VizKC-ISR on nine ISR benchmarks categorized into three multi-image scenarios. The results show that our VizKC-ISR performs best in all tasks, e.g., obtaining the highest average accuracy of 63.1% and surpassing the mPLUG-Owl3 baseline by 6.4 absolute points, due to the bridge between visually-grounded reasoning and multi-modal knowledge challenges. Guanghui Ye, Huan Zhao 0003, Yixian Shen, Jiaqi Li 0008, Fengnan Li, Zhihua Jiang, Keqin Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Dual-View Learning for Conversational Emotion Recognition Through Context and Emotion-Shift ModelingabstractConversational Emotion Recognition (CER) has recently been explored through conversational context modeling to learn the emotion distribution, i.e., the likelihood over emotion categories associated with each utterance. While these methods have shown promising results in emotion classification, they often focus on the interactions between utterances (utterance-view) and overlook shifts in the speaker's emotions (emotion-view). This emphasis on homogeneous view modeling limits their overall effectiveness. To address this limitation, we propose DVL-CER, a novel Dual-View Learning approach for CER. DVL-CER integrates both the utterance-view and emotion-view using two projection heads, enabling cross-view projection of emotion distributions. Our approach offers several key advantages: (1) We introduce an emotion-view that captures shifts in a speaker's emotions from initial to subsequent states within a conversation. This view enriches the conversation modeling and supports seamless integration with various CER baseline models. (2) Our dual-view projection learning strategy flexibly balances consistency and independence between the two heterogeneous views, promoting view-specific adaptation learning and incorporating the emotion verification capability within CER. We validate DVL-CER through extensive experiments on two widely-used datasets, IEMOCAP and EmoryNLP. The results demonstrate that DVL-CER achieves state-of-the-art performance, delivering robust and high-quality emotion distributions compared with existing CER methods and other dual-view learning strategies. Xupeng Zha, Huan Zhao 0003, Guanghui Ye, Zixing Zhang 0001 |
AAAI | 3 |
| 2025 | Knowledge Image Matters: Improving Knowledge-Based Visual Reasoning with Multi-Image Large Language ModelsabstractWe revisit knowledge-based visual reasoning (KB-VR) in light of modern advances in multimodal large language models (MLLMs), and make the following contributions: (i) We propose Visual Knowledge Card (VKC) -a novel image that incorporates not only internal visual knowledge (e.g., scene-aware information) detected from the raw image, but also external world knowledge (e.g., attribute or object knowledge) produced by a knowledge generator; (ii) We present VKC-enhanced Multi-Image Reasoning (VKC-MIR) -a fourstage pipeline which harnesses a state-of-theart scene perception engine to construct an initial VKC (Stage-1), a powerful LLM to generate relevant domain knowledge (Stage-2), an excellent image editing toolkit to introduce generated knowledge into an iteratively-edited VKC (Stage-3), and finally, an emerging multiimage MLLM to solve the VKC-enhanced task (Stage-4).By performing experiments on three popular KB-VR benchmarks, our approach achieves new state-of-the-art results compared to previous top-performing models.Our code is available at: https://github. com/yyy1103/VKC. Guanghui Ye, Huan Zhao 0003, Zhixue Zhao, Xupeng Zha, Zhihua Jiang |
ACL (1) | 1 |
| 2025 | Enhanced Multimodal Emotion Recognition in Conversations via Contextual Filtering and Multi-Frequency Graph PropagationabstractMultimodal Emotion Recognition in Conversations (ERC) plays a crucial role in understanding human language and behavior in real-world scenarios. However, existing research tends to simply concatenate multimodal representations, failing to capture the complex relationships between modalities. Recent advances have shown that Graph Neural Networks (GNNs) are effective in capturing complex data relationships, offering a promising solution for multimodal ERC. Despite this, current GNN-based methods still face challenges, including weak interactions between modalities, neglecting the information entropy of utterances, and erasure of high-frequency signals that capture key variations and discrepancies between closely related nodes. To address these limitations, we propose a GNNs-based multi-frequency propagation method enhanced by contextual filtering for multimodal ERC. Our approach introduces a context filtering module that combines a similarity matrix and an information entropy matrix, enabling GNNs to effectively capture the inherent relationships among utterances and provide sufficient multimodal and contextual modeling. Additionally, our method explores multivariate relationships by recognizing the varying importance of emotional discrepancies and commonalities through multi-frequency signals. Experimental results on two benchmark datasets, IEMOCAP and MELD, demonstrate that our method outperforms the latest (non-)graph-based works. Our method is available at https://github.com/G22-web/ConFilMER. Huan Zhao 0003, Yingxue Gao, Haijiao Chen, Guanghui Ye, Zixing Zhang 0001 |
ICASSP | 5 |
| 2025 | DSSM: Dual State Space Model For Human Motions GenerationabstractText-driven human motion generation has attracted considerable critical attention in recent years. The task requires generating movements that are diverse, natural, and comfortable in accordance with the text description. However, while generating the human motion, there is a significant gap in the amount of feature information contained in word-based text description modality and joint-based human motion modality. This results in an extreme imbalance of feature information in the latent space across the modalities, which seriously affects the effect of feature fusion. To alleviate the imbalance between modalities, we propose the Dual State Space Model (DSSM), which reconstruction the fused feature from coarse-to-fine. The DSSM contains two unit structures: the Masked State Space Model (MSSM) and the Hierarchical State Space Block (HSSB). At the same time, in order to make better use of timing information and reduce the computational complexity of the model, the DSSM is also the first method to introduce the state space model (SSM) into the text-driven motion sequence generation. We evaluated the DSSM on the HumanML3D and KITML benchmark datasets, and the experimental results show that our approach achieves state-of-the-art performance. The source code is available on GitHub at https://github.com/yimingliu123/DSSM. Huan Zhao 0003, Yaqian Liu, Haijiao Chen, Guanghui Ye, Zixing Zhang 0001 |
ICASSP | 6 |
| 2025 | UniDE: A multi-level and low-resource framework for automatic dialogue evaluation via LLM-based data augmentation and multitask learning
Guanghui Ye, Huan Zhao 0003, Zixing Zhang 0001, Zhihua Jiang |
Inf. Process. Manag. | 1 |
| 2025 | CCDE: A Compact and Competitive Dialogue Evaluation Framework via Knowledge Distillation of Large Language ModelsabstractAutomatic evaluation metrics not only play a vital role in developing dialogue and interactive systems but also have a great impact on social activities in our daily life. However, previous specialized metrics for evaluating dialogues exhibit a relatively low correlation with human judgments. In addition, today’s state-of-the-art (SOTA) evaluators that leverage large language models (LLMs) are challenging to deploy in real-world applications due to their sheer size. To this end, we propose a novel evaluation framework, compact and competitive dialogue evaluation (CCDE), which leverages knowledge distillation of LLMs to generate training data and sequentially learn a multitask evaluator regarding diversified quality dimensions. Specifically, we first employ ChatGPT asteacherto generate a high-quality and rich-annotation corpus, CCDE-data. Then, we implement astudentevaluator CCDE (1.3B) via using InstructGPT as the backbone model that is trained and fine-tuned on CCDE-data. We conduct extensive experiments on three public benchmarks: fine-grained evaluation of dialog (FED), PersonaChat, and TopicalChat. The results demonstrate that our model CCDE can outperform the current SOTA model G-Eval which calls GPT-4 ($\boldsymbol{\geq}$175B) by 4.3 on the FED dataset, 3.5 on the PersonaChat dataset, and 0.3 on the TopicalChat dataset, in terms of the Spearman correlation metric (%). We release the data and code at:https://anonymous.4open.science/r/ccde-3827. Guanghui Ye, Huan Zhao 0003, Haijiao Chen, Zhixue Zhao, Zhihua Jiang, Keqin Li 0001 |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2024 | Leveraging Context-Aware Prompting for Commit Message GenerationabstractWriting comprehensive commit messages is tedious yet important, because these messages describe changes of code, such as fixing bugs or adding new features.However, most existing methods focus on either only the changed lines or nearest context lines, without considering the effectiveness of selecting useful contexts.On the other hand, it is possible that introducing excessive contexts can lead to noise.To this end, we propose a code model COMMIT (Context-aware prOMpting based comMIt-message generaTion) in conjunction with a code dataset CODEC (COntext and metaData Enhanced Code dataset).Leveraging program slicing, CODEC consolidates code changes along with related contexts via property graph analysis.Further, utilizing CodeT5+ as the backbone model, we train COMMIT via context-aware prompt on CODEC.Experiments show that COMMIT can surpass all compared models including pre-trained language models for code (code-PLMs) such as Com-mitBART and large language models for code (code-LLMs) such as Code-LlaMa.Besides, we investigate several research questions (RQs), further verifying the effectiveness of our approach.We release the data and code at: https: //github.com/Jnunlplab/COMMIT.git. Zhihua Jiang, Dongning Rao, Guanghui Ye |
EMNLP | 4 |
| 2024 | LSTDial: Enhancing Dialogue Generation via Long- and Short-Term Measurement FeedbackabstractGuanghui Ye, Huan Zhao, Zixing Zhang, Xupeng Zha, Zhihua Jiang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Guanghui Ye, Huan Zhao 0003, Zixing Zhang 0001, Xupeng Zha, Zhihua Jiang |
NAACL-HLT | 1 |
| 2023 | An effective negative sampling approach for contrastive learning of sentence embedding
Qitao Tan, Guanghui Ye, Chuan Wu 0003 |
Mach. Learn. | 3 |
| 2022 | IM⌃2: an Interpretable and Multi-category Integrated Metric Framework for Automatic Dialogue EvaluationabstractEvaluation metrics shine the light on the best models and thus strongly influence the research directions, such as the recently developed dialogue metrics USR, FED, and GRADE.However, most current metrics evaluate the dialogue data as isolated and static because they only focus on a single quality or several qualities.To mitigate the problem, this paper proposes an interpretable, multi-faceted, and controllable framework IM 2 (Interpretable and M ulti-category Integrated M etric) to combine a large number of metrics which are good at measuring different qualities.The IM 2 framework first divides current popular dialogue qualities into different categories and then applies or proposes dialogue metrics to measure the qualities within each category and finally generates an overall IM 2 score.An initial version of IM 2 was submitted to the AAAI 2022 Track5.1@DSTC10challenge 1 and took the 2 nd place on both of the development and test leaderboard.After the competition, we develop more metrics and improve the performance of our model.We compare IM 2 with other 13 current dialogue metrics and experimental results show that IM 2 correlates more strongly with human judgments than any of them on each evaluated dataset 2 . Zhihua Jiang, Guanghui Ye, Dongning Rao |
EMNLP | 2 |
| 2021 | Multi-faceted Classification for the Identification of Informative Communications during Crises: Case of COVID-19abstractSocial media data are used to enhance crisis management, as people widely adopt social media to share and acquire information to cope with uncertainties in crises. Identification and extraction of informative communications out of large volumes of data is critical for accurate situational awareness and timely response. Existing studies use conditions of geolocations, keywords, and topics separately or jointly to retrieve data that can be crisis related, but are not enough to filter subsets of data for different crisis management tasks. We propose that the crisis communication purposes of users can be detected to enhance data selection and prioritization for different crisis management tasks. A classification framework was built to identify three facets of a message: content type, audience type, and information source. The definitions of these categories are not dependent on a specific type of crises. So the classification framework can be potentially applied to different crisis scenarios. Machine learning models were created for the automatic classification of messages. Results showed the CNN-based model achieved the best accuracy (88.5%) for the classification of content type. The proposed Naive Bayes and logistic repression with predetermined features can best differentiate audience types and information source with an accuracy of 72.7% and 72.2%, respectively. Zhuoli Xie, Ajay Jayanth, Kapil Yadav, Guanghui Ye, Lingzi Hong |
COMPSAC | 4 |