Zhiyue Liu

dblp:227/0986 · DBLP profile ↗
← Back
22ranked-venue papers
11as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 8 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author
YearPublicationVenuePosition
2026 Knowledge-enhanced Chinese multimodal hate speech detection
Qingbao Huang, Pijian Li, Xingmao Zhang, Shizhen Chen, Haonan Cheng, Zhiyue Liu
Expert Syst. Appl.7
2026 ExCap: Entity-aware zero-shot image captioning via faithful synthetic image-text alignment
Qipeng Jiang, Zhiyue Liu, Wenkai Zhou, Qingbao Huang, Jiahai Wang
Knowl. Based Syst.2
2025 Task-Specific Information Decomposition for End-to-End Dense Video Captioning
abstract
Dense video captioning aims to localize events within input videos and generate concise descriptive texts for each event.Advanced endto-end methods require both tasks to share the same intermediate features that serve as event queries, thereby enabling the mutual promotion of two tasks.However, relying on shared queries limits the model's ability to extract taskspecific information, as event semantic perception and localization demand distinct perspectives on video understanding.To address this, we propose a decomposed dense video captioning framework that derives localization and captioning queries from event queries, enabling task-specific representations while maintaining inter-task collaboration.Considering the roles of different queries, we design a contrastive semantic optimization strategy that guides localization queries to focus on event-level visual features and captioning queries to align with textual semantics.Besides, only localization information is considered in existing methods for label assignment, failing to ensure the relevance of the selected queries to descriptions.We jointly consider localization and captioning losses to achieve a semantically balanced assignment process.Extensive experiments on the YouCook2 and ActivityNet Captions datasets demonstrate that our framework achieves state-of-the-art performance.
Zhiyue Liu, Xinru Zhang 0008
ACL (1)1
2025 A Knowledge Noise Mitigation Framework for Knowledge-based Visual Question Answering
abstract
Knowledge-based visual question answering (KB-VQA) requires a model to understand images and utilize external knowledge to provide accurate answers. Existing approaches often directly augment models with retrieved information from knowledge sources while ignoring substantial knowledge redundancy, which introduces noise into the answering process. To address this, we propose a training-free framework with knowledge focusing for KB-VQA, that mitigates the impact of noise by enhancing knowledge relevance and reducing redundancy. First, for knowledge retrieval, our framework concludes essential parts from the image-question pairs, creating low-noise queries that enhance the retrieval of highly relevant knowledge. Considering that redundancy still persists in the retrieved knowledge, we then prompt large models to identify and extract answer-beneficial segments from knowledge. In addition, we introduce a selective knowledge integration strategy, allowing the model to incorporate knowledge only when it lacks confidence in answering the question, thereby mitigating the influence of redundant information. Our framework enables the acquisition of accurate and critical knowledge, and extensive experiments demonstrate that it outperforms state-of-the-art methods.
Zhiyue Liu, Sihang Liu 0010, Xinru Zhang 0008
ICME1
2025 Target-oriented Multimodal Sentiment Classification with Counterfactual-enhanced Debiasing
abstract
Target-oriented multimodal sentiment classification seeks to predict sentiment polarity for specific targets from image-text pairs. While existing works achieve competitive performance, they often over-rely on textual content and fail to consider dataset biases, in particular word-level contextual biases. This leads to spurious correlations between text features and output labels, impairing classification accuracy. In this paper, we introduce a novel counterfactual-enhanced debiasing framework to reduce such spurious correlations. Our framework incorporates a counterfactual data augmentation strategy that minimally alters sentiment-related causal features, generating detail-matched image-text samples to guide the model’s attention toward content tied to sentiment. Furthermore, for learning robust features from counterfactual data and prompting model decisions, we introduce an adaptive debiasing contrastive learning mechanism, which effectively mitigates the influence of biased words. Experimental results on several benchmark datasets show that our proposed method outperforms state-of-the-art baselines.
Zhiyue Liu, Fanrong Ma, Xin Ling
ICME1
2025 Deep Learning-Based Knowledge Injection for Metaphor Detection: A Comprehensive Review
Zhiyue Liu, Xingmao Zhang, Qingbao Huang
NLPCC (3)2
2025 Synthesize then align: Modality alignment augmentation for zero-shot image captioning with synthetic data
Zhiyue Liu, Xin Ling, Qingbao Huang, Jiahai Wang
Knowl. Based Syst.1
2024 Improving Cross-Modal Alignment with Synthetic Pairs for Text-Only Image Captioning
abstract
Although image captioning models have made significant advancements in recent years, the majority of them heavily depend on high-quality datasets containing paired images and texts which are costly to acquire. Previous works leverage the CLIP's cross-modal association ability for image captioning, relying solely on textual information under unsupervised settings. However, not only does a modality gap exist between CLIP text and image features, but a discrepancy also arises between training and inference due to the unavailability of real-world images, which hinders the cross-modal alignment in text-only captioning. This paper proposes a novel method to address these issues by incorporating synthetic image-text pairs. A pre-trained text-to-image model is deployed to obtain images that correspond to textual data, and the pseudo features of generated images are optimized toward the real ones in the CLIP embedding space. Furthermore, textual information is gathered to represent image features, resulting in the image features with various semantics and the bridged modality gap. To unify training and inference, synthetic image features would serve as the training prefix for the language decoder, while real images are used for inference. Additionally, salient objects in images are detected as assistance to enhance the learning of modality alignment. Experimental results demonstrate that our method obtains the state-of-the-art performance on benchmark datasets.
Zhiyue Liu, Jinyuan Liu 0001, Fanrong Ma
AAAI1
2024 Target-Oriented Multimodal Sentiment Classification with Adaptive Modality Weighting
Zhiyue Liu, Fanrong Ma
NLPCC (4)1
2023 RMLM: A Flexible Defense Framework for Proactively Mitigating Word-level Adversarial Attacks
abstract
Adversarial attacks on deep neural networks keep raising security concerns in natural language processing research.Existing defenses focus on improving the robustness of the victim model in the training stage.However, they often neglect to proactively mitigate adversarial attacks during inference.Towards this overlooked aspect, we propose a defense framework that aims to mitigate attacks by confusing attackers and correcting adversarial contexts that are caused by malicious perturbations.Our framework comprises three components: (1) a synonym-based transformation to randomly corrupt adversarial contexts in the word level, (2) a developed BERT defender to correct abnormal contexts in the representation level, and (3) a simple detection method to filter out adversarial examples, any of which can be flexibly combined.Additionally, our framework helps improve the robustness of the victim model during training.Extensive experiments demonstrate the effectiveness of our framework in defending against word-level adversarial attacks.
Zhiyue Liu, Xiaopeng Zheng, Qinliang Su, Jiahai Wang
ACL (1)2
2023 A Two-Stage Chinese Medical Video Retrieval Framework with LLM
Ningjie Lei, Jinxiang Cai, Yixin Qian, Zhilong Zheng, Zhiyue Liu, Qingbao Huang
NLPCC (3)6
2022 UECA-Prompt: Universal Prompt for Emotion Cause Analysis
abstract
Emotion cause analysis (ECA) aims to extract emotion clauses and find the corresponding cause of the emotion. Existing methods adopt fine-tuning paradigm to solve certain types of ECA tasks. These task-specific methods have a deficiency of universality. And the relations among multiple objectives in one task are not explicitly modeled. Moreover, the relative position information introduced in most existing methods may make the model suffer from dataset bias. To address the first two problems, this paper proposes a universal prompt tuning method to solve different ECA tasks in the unified framework. As for the third problem, this paper designs a directional constraint module and a sequential learning module to ease the bias. Considering the commonalities among different tasks, this paper proposes a cross-task training method to further explore the capability of the model. The experimental results show that our method achieves competitive performance on the ECA datasets.
Xiaopeng Zheng, Zhiyue Liu, Zizhen Zhang, Jiahai Wang
COLING2
2022 A Coarse-to-Fine Training Paradigm for Dialogue Summarization
Zhiyue Liu, Jiahai Wang
ICANN (1)1
2022 Multiple userids identification with deep learning
Siyuan Chen 0005, Zhiyue Liu, Jiahai Wang
Expert Syst. Appl.3
2021 A Cooperative Framework with Generative Adversarial Networks and Entropic Auto-Encoders for Text Generation
abstract
Generating text with high quality and sufficient diversity is a fundamental task in natural language generation. Although generative adversarial networks (GANs) achieve promising results in text generation, GAN-based language models suffer from mode collapse, i.e., the generator tends to sacrifice diversity and focus on limited text patterns with high quality. By contrast, maximum likelihood estimation (MLE) based language models could cover various text patterns and generate diversified samples with poor quality. This paper proposes a cooperative framework with GANs and entropic auto-encoders (EAEs), named GAN-EAE, to synthesize their advantages for text generation, where EAEs are powerful MLE-based generative models based on deterministic auto-encoders. By imitating the output distribution of EAEs, the generator shapes its output distribution closer to the real data distribution against mode collapse. Meanwhile, by learning the samples from the generator of GANs, EAEs subtly distribute probability mass on high quality patterns for improving generation quality. The similar samples obtained from the generator may raise mode collapse and should be downplayed during adversarial training. Thus, a sample re-weighting mechanism is adopted to improve diversity by measuring the inner distance of generated samples. Experimental results demonstrate that GAN-EAE could improve both GANs and EAEs to achieve state-of-the-art performance.
Zhiyue Liu, Jiahai Wang
IJCNN1
2021 Multi-agent Deep Reinforcement Learning with Spatio-Temporal Feature Fusion for Traffic Signal Control
Jiahai Wang, Siyuan Chen 0005, Zhiyue Liu
ECML/PKDD (4)4
2021 Topic-to-Essay Generation with Comprehensive Knowledge Enhancement
Zhiyue Liu, Jiahai Wang
ECML/PKDD (5)1
2020 CatGAN: Category-Aware Generative Adversarial Networks with Hierarchical Evolutionary Learning for Category Text Generation
abstract
Generating multiple categories of texts is a challenging task and draws more and more attention. Since generative adversarial nets (GANs) have shown competitive results on general text generation, they are extended for category text generation in some previous works. However, the complicated model structures and learning strategies limit their performance and exacerbate the training instability. This paper proposes a category-aware GAN (CatGAN) which consists of an efficient category-aware model for category text generation and a hierarchical evolutionary learning algorithm for training our model. The category-aware model directly measures the gap between real samples and generated samples on each category, then reducing this gap will guide the model to generate high-quality category samples. The Gumbel-Softmax relaxation further frees our model from complicated learning strategies for updating CatGAN on discrete data. Moreover, only focusing on the sample quality normally leads the mode collapse problem, thus a hierarchical evolutionary learning algorithm is introduced to stabilize the training procedure and obtain the trade-off between quality and diversity while training CatGAN. Experimental results demonstrate that CatGAN outperforms most of the existing state-of-the-art methods.
Zhiyue Liu, Jiahai Wang
AAAI1
2020 Fusion-Extraction Network for Multimodal Sentiment Analysis
Tao Jiang 0059, Jiahai Wang, Zhiyue Liu, Yingbiao Ling
PAKDD (2)3
2019 Targeted Sentiment Classification with Attentional Encoder Network
Youwei Song, Jiahai Wang, Tao Jiang 0059, Zhiyue Liu, Yanghui Rao
ICANN (4)4
2019 Analysis of load imbalance degree using one-dimensional algorithms for cascaded rectifier based on the modified T-type five-level topology
abstract
This paper presents a single-phase rectifier based on the modified T-type five-level topology cascaded. The presented topology eliminates clamping diodes and only eight power switches along with their antiparallel diodes are used, reducing cost and longer life. The one-dimensional algorithms is described for cascaded rectifier based on the modified T-type five-level topology. This paper focuses on the discussion of load imbalance degree and why it exists in the cascaded rectifier. Meanwhile, this paper takes the two-module cascaded as an example to discuss the DC-side load imbalance, and extend it to more modules. Furthermore, the system structure and control architecture are given. Simulation results based on MATLAB/Simulink verify the effectiveness of the topology of the modified T-type applied to cascaded rectifiers and the correctness of the analysis of load imbalance degree using one-dimensional algorithms.
Zhiyue Liu
IECON1
2018 IMA health state evaluation using deep feature learning with quantum neural network
Zehai Gao, Cunbao Ma, Yige Luo, Zhiyue Liu
Eng. Appl. Artif. Intell.4