Zhihua Jiang

dblp:164/4279 · DBLP profile ↗
← Back
21ranked-venue papers
5as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 4 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Text-Routed Sparse Mixture-of-Experts Model with Explanation and Temporal Alignment for Multi-Modal Sentiment Analysis
abstract
Human-interaction-involved applications underscore the need for Multi-modal Sentiment Analysis (MSA). Although many approaches have been proposed to address the subtle emotions in different modalities, the power of explanations and temporal alignments is still underexplored. Thus, this paper proposes the Text-routed sparse mixture-of-Experts model with eXplanation and Temporal alignment for MSA (TEXT). TEXT first augments explanations for MSA via Multi-modal Large Language Models (MLLM), and then novelly aligns the representations of audio and video through a temporality-oriented neural network block. TEXT aligns different modalities with explanations and facilitates a new text-routed sparse mixture-of-experts with gate fusion. Our temporal alignment block merges the benefits of Mamba and temporal cross-attention. As a result, TEXT achieves the best performance across four datasets among all tested models, including three recently proposed approaches and three MLLMs. TEXT wins on at least four metrics out of all six metrics. For example, TEXT decreases the mean absolute error to 0.353 on the CH-SIMS dataset, which signifies a 13.5% decrement compared with recently proposed approaches.
Dongning Rao, Yunbiao Zeng, Zhihua Jiang, Jujian Lv
AAAI3
2026 Making Visual Dialogue More Engaging: A New Task, Method, and Metric
abstract
Large language model (LLM)-based visual dialogue (VD) systems have made response generation for image-grounded conversations more correct and coherent. However, user engagement - the extent to which a user is interested, emotionally involved, and willing to continue the conversation - remains a challenge. To fully explore engaging VD, we propose: (i) a new task named Audio-enhanced VD (AVD), which introduces additional audio dialogue contexts that can more vividly convey the speaker's emotions as input, with the aim of generating correct but more engaging dialogue responses. Specifically, we employ a text-to-speech model as the modality translator to generate the paired acoustic utterances from the inputting textual utterances; (ii) an accompanying approach named Visually-grounded and Interleaved Text-Audio Dialogue Modeling (VITA-DM), which utilizes both image-grounded information and interleaved text-audio utterances for visual dialogue modeling, differentiating from previous multi-modal LLM (MLLM)-based methods that normally model text and audio modalities separately. We also present three pre-training tasks to better learn multi-modal interactions across language, vision, and audio; (iii) a novel metric named Multi-Modal Engagement (MME), which fills the gap of engagement estimation in VD and can provide a fine-grained assessment along emotional, attentional, and reply engagement dimensions (EE, AE, RE). We experiment on two popular datasets and provide extensive evaluations (automatic, engagement-specific, and human), supporting the validity of our approach. Furthermore, based on empirical results that reveal that emotions contribute the most to engagement, we justify our emphasis on the emotional aspect throughout the definition, solution, and evaluation of our task.
Guanghui Ye, Huan Zhao 0003, Yingxue Gao, Zhixue Zhao, Xupeng Zha, Zhihua Jiang
AAAI7
2026 Automatic and Reliable Evaluation for Academic Caption-to-Figure Generation with LMMs
abstract
Guanghui Ye, Huan Zhao, Qin Zhu, Fengnan Li, Jiaqi Li, Yixian Shen, Zhonghao Ren, Zhihua Jiang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Guanghui Ye, Huan Zhao 0003, Fengnan Li, Jiaqi Li 0008, Yixian Shen, Zhonghao Ren, Zhihua Jiang
ACL (1)8
2026 Leveraging dynamic few-shot prompting and ensemble method for task-oriented dialogue with subjective knowledge
abstract
Subjective knowledge is key to meeting customer needs. Thus, the Subjective Knowledge-grounded Task-oriented Dialogue (SK-TOD) task tries to accommodate subjective user requests like “Does the restaurant have a good atmosphere?” by choosing relevant subjective knowledge snippets and generating appropriate responses. However, unlike existing methods like retrieval-augmented generation using external objective knowledge, selecting subjective knowledge and summarizing opinions from reviews in a specified scope pose new challenges. Therefore, this paper proposes the DESIGN ( D ynamic f E w- S hot prompt I n G and e N semble) method for SK-TOD. Specifically, DESIGN first adopts Aspect-Based Sentiment Analysis (ABSA) to enhance subjective knowledge snippets and then builds an ensemble composed of diverse base models for knowledge selection (KS). Here, the base models include both classification models and generative models. At last, for response generation (RG), DESIGN employs generative models conditioned on dialogue context and ABSA-enhanced knowledge. Particularly, we devise the sample selection via the similarity-alignment algorithm to choose similar samples dynamically for the few-shot prompting of KS and RG. We experiment on the 11th Dialog System Technology Challenge (DSTC11) SK-TOD benchmark and an extended dataset, ReDial, with 6147 instances. For KS, we beat the winner of DSTC11 and boosted the F1 for 7% regarding the baseline and achieved 86.16%. For RG, DESIGN outperforms baselines and the DSTC11 winner across eight metrics.E.g., DESIGN improves entailment performance by 5% over the DSTC11 winner and 10% over the baseline. 1
Dongning Rao, Jietao Zhuang, Zhihua Jiang
Inf. Process. Manag.3
2026 Generating Multi-Modal Knowledge Clues as an Image: Toward Improving Image-Sequence Reasoning With Assisted Visual Input
abstract
Recent multi-modal large language models (MLLMs) have exhibited powerful abilities in addressing complex vision-language tasks such as image-sequence reasoning (ISR). However, significant challenges remain, e.g., it is still difficult for the MLLMs to fully capture and represent cross-image visual knowledge such as scene relations, attributes, and entity links between multiple images, which hinders them from better solving ISR. To alleviate these issues, we introduce a novel concept Visualized Knowledge Clue (VizKC) - synthetic images that encode key visual and external knowledge from a sequence of input images and are then used alongside the original input images within a multi-image MLLM to enhance reasoning performance. Accordingly, we propose an accompanying approach named VizKC-ISR, composed of two modules - VizKC generation and VizKC utilization. Specifically, in the generation module, VizKC-ISR follows aSee-Find-Fusepipeline: (i) “See - Scene Perception”, to construct an initial VizKC that incorporates scene relations of key visual entities detected from an original image; (ii) “Find - Knowledge Generation”, to generate enriched image captions with real-world knowledge and fine-grained entity details and then extract structured knowledge tuples from generated captions; (iii) “Fuse - Image Editing”, to introduce relevant knowledge tuples into the VizKC via iterative image editing. In the utilization module, we employ a multi-image MLLM (e.g., mPLUG-Owl3) to solve the VizKC-assisted ISR tasks by reasoning with generated knowledge clues. We evaluate VizKC-ISR on nine ISR benchmarks categorized into three multi-image scenarios. The results show that our VizKC-ISR performs best in all tasks, e.g., obtaining the highest average accuracy of 63.1% and surpassing the mPLUG-Owl3 baseline by 6.4 absolute points, due to the bridge between visually-grounded reasoning and multi-modal knowledge challenges.
Guanghui Ye, Huan Zhao 0003, Yixian Shen, Jiaqi Li 0008, Fengnan Li, Zhihua Jiang, Keqin Li 0001
IEEE Trans. Circuits Syst. Video Technol.6
2025 Knowledge Image Matters: Improving Knowledge-Based Visual Reasoning with Multi-Image Large Language Models
abstract
We revisit knowledge-based visual reasoning (KB-VR) in light of modern advances in multimodal large language models (MLLMs), and make the following contributions: (i) We propose Visual Knowledge Card (VKC) -a novel image that incorporates not only internal visual knowledge (e.g., scene-aware information) detected from the raw image, but also external world knowledge (e.g., attribute or object knowledge) produced by a knowledge generator; (ii) We present VKC-enhanced Multi-Image Reasoning (VKC-MIR) -a fourstage pipeline which harnesses a state-of-theart scene perception engine to construct an initial VKC (Stage-1), a powerful LLM to generate relevant domain knowledge (Stage-2), an excellent image editing toolkit to introduce generated knowledge into an iteratively-edited VKC (Stage-3), and finally, an emerging multiimage MLLM to solve the VKC-enhanced task (Stage-4).By performing experiments on three popular KB-VR benchmarks, our approach achieves new state-of-the-art results compared to previous top-performing models.Our code is available at: https://github. com/yyy1103/VKC.
Guanghui Ye, Huan Zhao 0003, Zhixue Zhao, Xupeng Zha, Zhihua Jiang
ACL (1)6
2025 A Comprehensive Literary Chinese Reading Comprehension Dataset with an Evidence Curation Based Solution
abstract
Low-resource language understanding is a challenging task, even for large language models (LLMs).An epitome of this problem is the CompRehensive lIterary chineSe readIng comprehenSion (CRISIS), whose difficulties include limited linguistic data, long input, and insight-required questions.Besides the compelling need to provide a larger dataset for CRISIS, excessive information, order bias, and entangled conundrums still plague the CRISIS solutions.Thus, we present the eVIdence cuRation with opTion shUffling and Abstract meaning representation-based cLauses segmenting (VIRTUAL) procedure for CRISIS, with the most extensive dataset.While the dataset is also named CRISIS, it results from a three-phase construction process, including question selection, data cleaning, and a silver-standard data augmentation step, which augments translations, celebrity profiles, government jobs, reign mottos, and dynasty to CRISIS.The six steps of VIRTUAL include embedding, shuffling, abstract meaning representation-based option segmenting, evidence extraction, solving, and voting.Notably, the evidence extraction algorithm facilitates the extraction of literary Chinese evidence sentences, translated evidence sentences, and annotations of keywords using a similaritybased ranking strategy.While CRISIS compiles understanding-required questions from seven sources, the experiments on CRISIS substantiate the effectiveness of VIRTUAL, with a 7 percent increase in accuracy compared to the baseline.Interestingly, both non-LLMs and LLMs exhibit order bias, and abstract meaning representation-based option segmenting is beneficial for CRISIS.
Dongning Rao, Rongchu Zhou, Zhihua Jiang
EMNLP4
2025 Leveraging meta-data of code for adapting prompt tuning for code summarization
Zhihua Jiang, Dongning Rao
Appl. Intell.1
2025 A robust dialogue evaluation metric exploiting denoising, pre-training and ensembling
Dongning Rao, Lianyong Ling, Zhihua Jiang
Eng. Appl. Artif. Intell.3
2025 UniDE: A multi-level and low-resource framework for automatic dialogue evaluation via LLM-based data augmentation and multitask learning
Guanghui Ye, Huan Zhao 0003, Zixing Zhang 0001, Zhihua Jiang
Inf. Process. Manag.4
2025 CCDE: A Compact and Competitive Dialogue Evaluation Framework via Knowledge Distillation of Large Language Models
abstract
Automatic evaluation metrics not only play a vital role in developing dialogue and interactive systems but also have a great impact on social activities in our daily life. However, previous specialized metrics for evaluating dialogues exhibit a relatively low correlation with human judgments. In addition, today’s state-of-the-art (SOTA) evaluators that leverage large language models (LLMs) are challenging to deploy in real-world applications due to their sheer size. To this end, we propose a novel evaluation framework, compact and competitive dialogue evaluation (CCDE), which leverages knowledge distillation of LLMs to generate training data and sequentially learn a multitask evaluator regarding diversified quality dimensions. Specifically, we first employ ChatGPT asteacherto generate a high-quality and rich-annotation corpus, CCDE-data. Then, we implement astudentevaluator CCDE (1.3B) via using InstructGPT as the backbone model that is trained and fine-tuned on CCDE-data. We conduct extensive experiments on three public benchmarks: fine-grained evaluation of dialog (FED), PersonaChat, and TopicalChat. The results demonstrate that our model CCDE can outperform the current SOTA model G-Eval which calls GPT-4 ($\boldsymbol{\geq}$175B) by 4.3 on the FED dataset, 3.5 on the PersonaChat dataset, and 0.3 on the TopicalChat dataset, in terms of the Spearman correlation metric (%). We release the data and code at:https://anonymous.4open.science/r/ccde-3827.
Guanghui Ye, Huan Zhao 0003, Haijiao Chen, Zhixue Zhao, Zhihua Jiang, Keqin Li 0001
IEEE Trans. Comput. Soc. Syst.6
2024 Leveraging Context-Aware Prompting for Commit Message Generation
abstract
Writing comprehensive commit messages is tedious yet important, because these messages describe changes of code, such as fixing bugs or adding new features.However, most existing methods focus on either only the changed lines or nearest context lines, without considering the effectiveness of selecting useful contexts.On the other hand, it is possible that introducing excessive contexts can lead to noise.To this end, we propose a code model COMMIT (Context-aware prOMpting based comMIt-message generaTion) in conjunction with a code dataset CODEC (COntext and metaData Enhanced Code dataset).Leveraging program slicing, CODEC consolidates code changes along with related contexts via property graph analysis.Further, utilizing CodeT5+ as the backbone model, we train COMMIT via context-aware prompt on CODEC.Experiments show that COMMIT can surpass all compared models including pre-trained language models for code (code-PLMs) such as Com-mitBART and large language models for code (code-LLMs) such as Code-LlaMa.Besides, we investigate several research questions (RQs), further verifying the effectiveness of our approach.We release the data and code at: https: //github.com/Jnunlplab/COMMIT.git.
Zhihua Jiang, Dongning Rao, Guanghui Ye
EMNLP1
2024 LSTDial: Enhancing Dialogue Generation via Long- and Short-Term Measurement Feedback
abstract
Guanghui Ye, Huan Zhao, Zixing Zhang, Xupeng Zha, Zhihua Jiang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Guanghui Ye, Huan Zhao 0003, Zixing Zhang 0001, Xupeng Zha, Zhihua Jiang
NAACL-HLT5
2024 An Empirical Study of Leveraging PLMs and LLMs for Long-Text Summarization
Zhihua Jiang, Junzhan Yang, Dongning Rao
PRICAI (2)1
2023 Ancient Chinese Machine Reading Comprehension Exception Question Dataset with a Non-trivial Model
Dongning Rao, Guanju Huang, Zhihua Jiang
PRICAI (2)3
2022 IM⌃2: an Interpretable and Multi-category Integrated Metric Framework for Automatic Dialogue Evaluation
abstract
Evaluation metrics shine the light on the best models and thus strongly influence the research directions, such as the recently developed dialogue metrics USR, FED, and GRADE.However, most current metrics evaluate the dialogue data as isolated and static because they only focus on a single quality or several qualities.To mitigate the problem, this paper proposes an interpretable, multi-faceted, and controllable framework IM 2 (Interpretable and M ulti-category Integrated M etric) to combine a large number of metrics which are good at measuring different qualities.The IM 2 framework first divides current popular dialogue qualities into different categories and then applies or proposes dialogue metrics to measure the qualities within each category and finally generates an overall IM 2 score.An initial version of IM 2 was submitted to the AAAI 2022 Track5.1@DSTC10challenge 1 and took the 2 nd place on both of the development and test leaderboard.After the competition, we develop more metrics and improve the performance of our model.We compare IM 2 with other 13 current dialogue metrics and experimental results show that IM 2 correlates more strongly with human judgments than any of them on each evaluated dataset 2 .
Zhihua Jiang, Guanghui Ye, Dongning Rao
EMNLP1
2021 STANKER: Stacking Network based on Level-grained Attention-masked BERT for Rumor Detection on Social Media
abstract
Rumor detection on social media puts pretrained language models (LMs), such as BERT, and auxiliary features, such as comments, into use.However, on the one hand, rumor detection datasets in Chinese companies with comments are rare; on the other hand, intensive interaction of attention on Transformer-based models like BERT may hinder performance improvement.To alleviate these problems, we build a new Chinese microblog dataset named Weibo20 1 by collecting posts and associated comments from Sina Weibo and propose a new ensemble named STANKER (Stacking neTwork bAsed-on atteNtion-masKed BERT).STANKER adopts two level-grained attentionmasked BERT (LGAM-BERT) models as base encoders.Unlike the original BERT, our new LGAM-BERT model takes comments as important auxiliary features and masks coattention between posts and comments on lower-layers.Experiments on Weibo20 and three existing social media datasets showed that STANKER outperformed all compared models, especially beating the old state-of-theart on Weibo dataset.
Dongning Rao, Zhihua Jiang
EMNLP (1)3
2021 Syntax and Sentiment Enhanced BERT for Earliest Rumor Detection
Dongning Rao, Zhihua Jiang
NLPCC (1)3
2021 A dual deep neural network with phrase structure and attention mechanism for sentiment analysis
Dongning Rao, Sihong Huang, Zhihua Jiang, Ganesh Gopal Devarajan, Rizwan Patan
Neural Comput. Appl.3
2019 Scalable and optimal planning based on Pregel
abstract
Summary Automated planning generates plans for specific tasks. Optimal planning aims at generating optimal plans under global constraints. As a result, the divide‐and‐conquer method is not applicable for optimal planning. Therefore, engineering applications of optimal planning face the scalability issue. Fortunately, cloud computing tools are on the shelf. For example, the Apache Spark is an engine for big data processing. It supports the Pregel for scalable computing. Therefore, we proposed an optimal Planning method based on the Pregel, called the PbP. Unlike classical planning, the PbP method uses the Pregel as the computation model, instead of the traditional state‐space searching. The core idea is to transform planning problems into graph processing problems. Specifically, actions are mapped into vertices, partial orders between actions are mapped into edges between vertices, and states are mapped into messages. Furthermore, the planning is viewed as message propagating in the graph, and plan traces are stored as attributes of vertices. Experimental results showed the feasibility of the proposed method PbP. Moreover, compared with state‐of‐the‐art optimal planners, our approach is more scalable and faster.
Zhihua Jiang, Dongning Rao
Concurr. Comput. Pract. Exp.1
2016 Cost-Sensitive Action Model Learning
abstract
Action model learning can relieve people from writing planning domain descriptions from scratch. Real-world learners need to be sensitive to all kinds of expenses which it will spend in the learning. However, most of previous studies in this research line only considered the running time as the learning cost. In real-world applications, we will spend extra expense when we carry out actions or get observations, particularly for online learning. The learning algorithm should apply more techniques for saving the total cost when keeping a high rate of accuracy. The cost of carrying out actions and getting observations is the dominated expense in online learning. Therefore, we design a cost-sensitive algorithm to learn action models under partial observability. It combines three techniques to lessen the total cost: constraints, filtering and active learning. These techniques are used in observation reduction in action model learning. First, the algorithm uses constraints to confine the observation space. Second, it removes unnecessary observations by belief state filtering. Third, it actively picks up observations based on the results of the previous two techniques. This paper also designs strategies to reduce the amount of plan steps used in the learning. We performed experiments on some benchmark domains. It shows two results. For one thing, the learning accuracy is high in most cases. For the other, the algorithm dramatically reduces the total cost according to the definition of cost in this paper. Therefore, it is significant for real-world learners, especially, when long plans are unavailable or observations are expensive.
Dongning Rao, Zhihua Jiang
Int. J. Uncertain. Fuzziness Knowl. Based Syst.2