Ziqiang Cao

dblp:148/4447 · DBLP profile ↗
← Back
45ranked-venue papers
9as first author
34since 2021 · last 2027
0000-0002-1077-9033ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 30 · 9 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 6 first-author · 12 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2027 VotingBench : Evaluating generative capabilities of Large Language Models via voting games
Jiale Mao, Jinyi Zu, Ziqiang Cao
Inf. Process. Manag.3
2026 Interleaved Tool-Call Reasoning for Protein Function Understanding
abstract
Chuanliu Fan, Zicheng Ma, Huanran Meng, Aijia Zhang, Wenjie Du, Jun Zhang, Ziqiang Cao, Guohong Fu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Chuanliu Fan, Zicheng Ma, Huanran Meng, Jun Zhang 0069, Ziqiang Cao, Guohong Fu
ACL (1)7
2026 The Bidirectional Process Reward Model
abstract
Process Reward Models (PRMs), which assign fine-grained scores to intermediate reasoning steps within a solution trajectory, have emerged as a promising approach to enhance the reasoning quality of Large Language Models (LLMs).However, most existing PRMs rely on a unidirectional left-to-right (L2R) evaluation scheme, which restricts their utilization of global context.In light of this challenge, we propose a novel bidirectional evaluation paradigm, named Bidirectional Process Reward Model (BiPRM).BiPRM incorporates a parallel right-to-left (R2L) evaluation stream, implemented via prompt reversal, alongside the conventional L2R flow.Then a gating mechanism is introduced to adaptively fuse the reward scores from both streams to yield a holistic quality assessment.Remarkably, compared to the original PRM, BiPRM introduces only a 0.3% parameter increase for the gating module, and the parallel execution of two streams incurs merely 5% inference time latency.Our extensive empirical evaluations spanning diverse benchmarks, LLM backbones, PRM objectives and sampling policies demonstrate that BiPRM consistently surpasses unidirectional baselines, achieving an average relative gain of 10.6% over 54 solution-level configurations and 37.7% in 12 step-level error detection scenarios.Generally, our results highlight the effectiveness, robustness and general applicability of BiPRM, offering a promising new direction for processbased reward modeling.1
Lingyin Zhang, Xiaoxue Ren, Ziqiang Cao
ACL (1)4
2026 Bidirectional GPT
Chuanliu Fan, Zicheng Ma, Jun Zhang 0071, Yiqin Gao, Ziqiang Cao, Guohong Fu
Inf. Process. Manag.7
2026 Optimizing Prompts With MLLM for Multimodal Product Style Recognition in Industry 5.0
abstract
The transition to Industry 5.0, supported by the high-bandwidth and low-latency capabilities of 6G networks, accelerates the demand for personalized products, making efficient and accurate product style recognition a critical task. However, traditional recognition models, which rely heavily on large-labeled datasets, struggle with generalization and efficiency in such data-rich environments. In contrast, multimodal large language models (MLLMs), such as Qwen-VL, offer new possibilities by integrating visual and textual inputs, requiring less annotated data and demonstrating strong generalization. Nevertheless, most existing MLLMs are general-purpose and lack optimized prompts tailored for domain-specific tasks, limiting their recognition performance. To address this, we propose a prompt optimization framework that combines genetic algorithms (GA) and chain of thought (CoT) reasoning. GA is used to iteratively refine the instruction component of prompts, while CoT encourages step-by-step reasoning, improving the model’s interpretability and recognition accuracy. Validated on a custom-built apparel dataset and the public FashionStyle14 dataset, our method enables Qwen-VL to achieve significant accuracy improvements of 8.7% and 19.6%, respectively, over the baseline. This study makes three contributions: a framework to adapt general MLLMs for fine-grained style recognition; a demonstration of 6G-enabled real-time artificial intelligence reasoning loops; and a high-quality, multidimensional style dataset for future research. This demonstrates the effectiveness and stability of the proposed framework in handling fine-grained style recognition tasks, contributing to the development of intelligent design systems for Industry 5.0.
Mengxue Li, Hujiang Huang, Yan Hong 0002, Ziqiang Cao, Jie Zhang 0076, Song Guo 0001
IEEE Trans. Ind. Informatics4
2025 AIM: Let Any Multimodal Large Language Models Embrace Efficient In-Context Learning
abstract
In-context learning (ICL) advances Large Language Models (LLMs) exhibiting emergent ability on downstream tasks without updating billions of parameters. However, in the area of multimodal Large Language Models (MLLMs), two problems hinder the application of multimodal ICL: (1) Most primary MLLMs are only trained on single-image datasets, making them unable to read extra multimodal demonstrations. (2) With the demonstrations increasing, thousands of visual tokens highly challenge hardware and degrade ICL performance. During preliminary explorations, we discovered that the inner LLM focuses more on the linguistic modality within multimodal demonstrations during generation. Therefore, we propose a general and lightweight framework AIM to tackle the mentioned problems through Aggregating Image information of Multimodal demonstrations to the latent space of the corresponding textual labels. After aggregation, AIM substitutes each demonstration with generated fused virtual tokens whose length is reduced to the same as its texts. Except for shortening input length, AIM further upgrades MLLMs pre-trained on image-text pairs to support multimodal ICL, as images from demonstrations are disregarded. Furthermore, benefiting from aggregating different demonstrations independently, AIM configures Demonstration Bank (DB) to avoid repeated aggregation, which significantly boosts model efficiency. We build AIM upon QWen-VL and LLaVA-Next, and AIM is comprehensively evaluated on image caption, VQA, and hateful speech detection. Outstanding results reveal that AIM provides an efficient and effective solution in upgrading MLLMs for multimodal ICL.
Ziqiang Cao, Wenjie Li 0002
AAAI5
2025 UniICL: An Efficient ICL Framework Unifying Compression, Selection, and Generation
abstract
In-context learning (ICL) enhances the reasoning abilities of Large Language Models (LLMs) by prepending a few demonstrations.It motivates researchers to introduce more examples to provide additional contextual information for the generation.However, existing methods show a significant limitation due to the problem of excessive growth in context length, which causes a large hardware burden.In addition, shallow-relevant examples selected by off-the-shelf tools hinder LLMs from capturing useful contextual information for generation.In this paper, we propose UniICL, a novel Unified ICL framework that unifies demonstration compression, demonstration selection, and final response generation.Furthermore, to boost inference efficiency, we design a tailored compression strategy that allows UniICL to cache compression results into Demonstration Bank (DB), which avoids repeated compression of the same demonstration.Extensive outof-domain evaluations prove the advantages of UniICL in both effectiveness and efficiency.
Qi Lv 0001, Ziqiang Cao, Wenjie Li 0002
ACL (1)5
2025 LlaMol: A Unified Molecule Designer via Preference Ranking and Numerical Enhancement
abstract
Goal-oriented de novo molecule design, namely generating molecules with specific property or substructure constraints from scratch, is a crucial yet challenging task in drug discovery. Existing research often relies on separate predictors for distinct properties and struggles with integrating substructure constraints due to the complexities involved in modeling structural information via multitask learning. This separation necessitates a dedicated prediction model for each constraint, limiting the flexibility and posing challenges for realworld applications. To address these limitations, we propose a unified framework for molecular design that incorporates multiple property and substructure constraints, leveraging LLMs to handle diverse constraint settings within a single model. We first integrate feedback learning derived from preference ranking to eliminate the need for separate property predictors. Then, we enhance the model's ability to follow numerical instructions by introducing a unified numerical encoding into the prompt. We conduct extensive experiments across single-property, substructureproperty, and multi-property constrained tasks. Experimental results demonstrate that LlaMol consistently outperforms state-of-the-art baselines across various constraint settings. Notably, in the multi-objective binding affinity maximization task, LlaMol achieves a significantly lower$\mathrm{K}_{\mathrm{D}}$value of 0.25 for the protein target ESR1, while maintaining the highest overall performance, surpassing previous methods by 4.76 %. These results underscore the effectiveness and versatility of LLM-based frameworks for molecule generation under complex constraints.
Chuanliu Fan, Zicheng Ma, Jun Zhang 0069, Ziqiang Cao, Yiqin Gao, Guohong Fu
BIBM6
2025 Personalized Large Language Model Assistant with Evolving Conditional Memory
abstract
With the rapid development of large language models, AI assistants like ChatGPT have become increasingly integrated into people’s works and lives but are limited in personalized services. In this paper, we present a plug-and-play framework that could facilitate personalized large language model assistants with evolving conditional memory. The personalized assistant focuses on intelligently preserving the knowledge and experience from the history dialogue with the user, which can be applied to future tailored responses that better align with the user’s preferences. Generally, the assistant generates a set of records from the dialogue, stores them in a memory bank, and retrieves related memory to improve the quality of the response. For the crucial memory design, we explore different ways of constructing the memory and propose a new memorizing mechanism named conditional memory to enhance the memory management of the framework. We also investigate the retrieval and usage of memory in the generation process. To better evaluate the personalized assistants’ abilities, we build the first evaluation benchmark from three critical aspects: continuing previous dialogue, learning personalized knowledge and learning from user feedback. The experimental results illustrate the effectiveness of our method.
Ruifeng Yuan, Shichao Sun, Yongqi Li 0001, Ziqiang Cao, Wenjie Li 0002
COLING5
2025 Interleaved-Modal Chain-of-Thought
abstract
Chain-of-Thought (CoT) prompting elicits large language models (LLMs) to produce a series of intermediate reasoning steps before arriving at the final answer. However, when transitioning to vision-language models (VLMs), their text-only rationales struggle to express the fine-grained associations with the original image. In this paper, we propose an image-incorporated multimodal Chain-of-Thought, named Interleaved-modal Chain-of-Thought (ICoT), which generates sequential reasoning steps consisting of paired visual and textual rationales to infer the final answer. Intuitively, the novel ICoT requires VLMs to enable the generation of fine-grained interleaved-modal content, which is hard for current VLMs to fulfill. Considering that the required visual information is usually part of the input image, we propose Attention-driven Selection (ADS) to realize ICoT over existing VLMs. ADS intelligently inserts regions of the input image to generate the interleaved-modal reasoning steps with ignorable additional latency. ADS relies solely on the attention map of VLMs without the need for parameterization, and therefore it is a plug-and-play strategy that can be generalized to a spectrum of VLMs. We apply ADS to realize ICoT on two popular VLMs of different architectures. Extensive evaluations of three benchmarks have shown that ICoT prompting achieves substantial performance (up to 14%) and interpretability improvements compared to existing multimodal CoT prompting methods.
Yongqi Li 0001, Ziqiang Cao, Wenjie Li 0002
CVPR3
2025 Select and Order: Optimizing Few-Shot Image Classification with In-Context Learning
Hujiang Huang, Yu Xie 0001, Chuanliu Fan, Ziqiang Cao
MMM (3)5
2025 Prot2Chat: protein large language model with early fusion of text, sequence, and structure
abstract
MOTIVATION: Proteins are of great significance in living organisms. However, understanding their functions encounters numerous challenges, such as insufficient integration of multimodal information, a large number of training parameters, limited flexibility of classification-based methods, and the lack of systematic evaluation metrics for protein question answering systems. To tackle these issues, we propose the Prot2Chat framework. RESULTS: We modified ProteinMPNN to encode protein sequence and structural information in a unified way. We used a large language model (LLM) to encode questions into vectors and developed a protein-text adapter to compress protein information into virtual tokens based on these vectors, achieving the early fusion of text and protein information. Finally, the same LLM reads the virtual tokens and the questions to generate answers. To optimize training efficiency, we froze the encoder and employed low-rank adaptation (LoRA) techniques for the LLM. Experiments on two datasets show that both automated metrics and expert evaluations demonstrate the superior performance of our model, and zero-shot prediction results highlight its generalization ability. We have developed an easy-to-use web interactive platform and a rapid installation option, allowing users to swiftly engage with Prot2Chat. AVAILABILITY AND IMPLEMENTATION: The models and codes are available at https://github.com/wangzc1233/Prot2Chat.
Zhicong Wang, Zicheng Ma, Ziqiang Cao, Changlong Zhou, Jun Zhang 0071, Yiqin Gao
Bioinform.3
2025 Semi-supervised multi-label feature selection via partial label correlation and feature self-representation
Yao Zhang 0017, Jun Tang 0007, Ziqiang Cao
Knowl. Based Syst.3
2024 Improving Copy-oriented Text Generation via EDU Copy Mechanism
abstract
Many text generation tasks are copy-oriented. For instance, nearly 30% content of news summaries is copied. The copy rate is even higher in Grammatical Error Correction (GEC). However, existing generative models generate texts through word-by-word decoding, which may lead to factual inconsistencies and slow inference. While Elementary Discourse Units (EDUs) are outstanding extraction units, EDU-based extractive methods can alleviate the aforementioned problems. As a consequence, we propose EDUCopy, a framework that integrates the behavior of copying EDUs into generative models. The main idea of EDUCopy is to use special index tags to represent the copied EDUs during generation. Specifically, we extract important EDUs from input sequences, finetune generative models to generate sequences with special index tags, and restore the generated special index tags into corresponding text spans. By doing so, EDUCopy reduces the number of generated tokens significantly. To verify the effectiveness of EDUCopy, we conduct experiments on the news summarization datasets CNNDM, NYT and the GEC datasets FCE, WI-LOCNESS. While achieving notable ROUGE and M2 scores, GPT-4 evaluation validates the strength of our models in terms of factual consistency, fluency, and overall performance. Moreover, compared to baseline models, EDUCopy achieves a significant acceleration of 1.65x.
Luozheng Qin, Ziqiang Cao, Chunhui Ai
LREC/COLING4
2024 TabMedBERT: A Tabular Knowledge Enhanced Biomedical Pretrained Language Model
abstract
Most existing biomedical language models are trained on plain text with general learning goals such as random word infilling, failing to capture the knowledge in the biomedical corpus sufficiently. Since biomedical articles usually contain many tables summarising the main entities and their relations, in the paper, we propose a Tabular knowledge enhanced bioMedical pretrained language model, called TabMedBERT. Specifically, we align entities between table cells, and article text spans with pre-defined rules. Then we add two table-related self-supervised tasks to integrate tabular knowledge into the language model: Entity Infilling (EI) and Table Cloze Test (TCT). While EI masks tokens within aligned entities in the article, TCT converts aligned entities in the table layout into a cloze text by erasing one entity and prompts the model to extract the appropriate span to fill in the blank. Experimental results demonstrate that TabMedBERT surpasses all competing language models without adding additional parameters, establishing a new state-of-the-art performance of 85.59% (+1.29%) on the BLURB biomedical NLP benchmark and 7 additional information extraction datasets. Moreover, the model architecture for TCT provides a straightforward solution to revise information extraction with paired entities.
Lei Geng, Ziqiang Cao, Juntao Li 0005, Wenjie Li 0002, Sujian Li, Yang Yang 0074, Jun Zhang 0069
ECAI3
2024 BVRCC: Bootstrapping Video Retrieval via Cross-Matching Correction
Luozheng Qin, Shaoyao Huang, Ziqiang Cao
ICANN (6)5
2024 Contrastive Learning with High-Quality and Low-Quality Augmented Data for Query-Focused Summarization
abstract
Unlike general text summarization, Query-focused summarization (QFS) is severely limited by insufficient datasets, forcing previous research to transform datasets from other tasks into QFS format for data augmentation. However, this approach has resulted in two problems: the task and traintest gaps. To alleviate these gaps, we propose QFS-CL, a novel in-place data augmentation framework equipped with contrastive learning. Firstly, we design diverse prompts for ChatGPT to paraphrase the original QFS data into high-quality/low-quality document-summary pairs, filling the task gap. Then, instead of directly incorporating the augmented data into the training set, we train the QFS baseline model in a contrastive learning scheme. Specifically, our approach encourages the model to imitate high-quality pairs and distinguish itself from low-quality pairs, enabling the model to learn how to acquire reliable information and avoid extracting invalid information. Our method achieves state-of-the-art performance on Debatepedia and DUC datasets in ROUGE scores, GPT-4, and human evaluations.
Shaoyao Huang, Ziqiang Cao, Luozheng Qin, Jun Zhang 0069
ICASSP2
2024 STRA: A Simple Token Replacement Strategy Alleviating Exposure Bias in Text Generation
abstract
In general text generation, models are typically trained on ground truth tokens, where erroneous tokens are not considered during generation. However, during the inference process, erroneous tokens are inevitably generated, and such errors accumulate in the autoregressive generation, leading to exposure bias. To address this problem directly, we propose a token replacement strategy with a sequential decay rate to simulate inference scenarios during training. Firstly, during the training process, we randomly replace tokens in the target text with a certain probability, simulating situations where erroneous tokens may be generated during inference. Then, we simulate inference scenarios by performing probability-decayed replacements from start to end. Since earlier generated tokens can impact the generation of subsequent tokens, the preceding tokens exert a stronger influence. Our method improves the base models’ performance on image captioning, text summarization, and dialogue datasets, achieving state-of-the-art performance on the query-focused summarization dataset.
Shaoyao Huang, Luozheng Qin, Ziqiang Cao
ICME3
2024 Guiding ChatGPT to Generate Salient Domain Summaries
abstract
ChatGPT is instruct-tuned to generate general and human-expected content to align with human preference through Reinforcement Learning from Human Feedback (RLHF), meanwhile resulting in generated responses not salient enough. Therefore, in this case, ChatGPT may fail to satisfy domain requirements in zero-shot settings, leading to poor ROUGE scores. Inspired by the In-Context Learning (ICL) and retelling ability of ChatGPT, this paper proposes PADS, a Pipeline for Assisting ChatGPT in Domain Summarization. PADS consists of a retriever to retrieve similar examples from corpora and a rank model to rerank the multiple candidate summaries generated by ChatGPT. Specifically, given an inference document, we first retrieve an in-context demonstration via the retriever. Then, we require ChatGPT to generate k candidate summaries for the inference document at a time under the guidance of the retrieved demonstration. Finally, the rank model independently scores the k candidate summaries according to their quality and selects the optimal one. We extensively explore dense and sparse retrieval methods to select effective demonstrations for reference and efficiently train the rank model to reflect the quality of candidate summaries for each given summarized document. Additionally, PADS contains merely 400M trainable parameters originating from the rank model and we merely collect 2.5k data to train it. We evaluate PADS on five datasets from different domains, and the result indicates that each module in PADS is committed to effectively guiding ChatGPT to generate salient summaries fitting different domain requirements. Specifically, in the popular summarization dataset Gigaword, PADS achieves over +8 gain on ROUGE-L, compared with the naive ChatGPT in the zero-shot setting.1
Ziqiang Cao, Shaoyao Huang, Luozheng Qin, Chunhui Ai
IJCNN2
2024 DNTextSpotter: Arbitrary-Shaped Scene Text Spotting via Improved Denoising Training
abstract
More and more end-to-end text spotting methods based on Transformer architecture have demonstrated superior performance. These methods utilize a bipartite graph matching algorithm to perform one-to-one optimal matching between predicted objects and actual objects. However, the instability of bipartite graph matching can lead to inconsistent optimization targets, thereby affecting the training performance of the model. Existing literature applies denoising training to solve the problem of bipartite graph matching instability in object detection tasks. Unfortunately, this denoising training method cannot be directly applied to text spotting tasks, as these tasks need to perform irregular shape detection tasks and more complex text recognition tasks than classification. To address this issue, we propose a novel denoising training method (DNTextSpotter) for arbitrary-shaped text spotting. Specifically, we decompose the queries of the denoising part into noised positional and noised content queries. We use the four Bezier control points of the Bezier center curve to generate the noised positional queries. For the noised content queries, considering that the output of the text in a fixed positional order is not conducive to aligning position with content, we employ a masked character sliding method to initialize noised content queries, thereby assisting in the alignment of text content and position. Additionally, to improve the model's perception of the background, we further utilize an additional loss function for background characters classification in the denoising training part. DNTextSpotter outperforms state-of-the-art methods on four benchmarks-Total-Text, SCUT-CTW1500, ICDAR15, and Inverse-Text-most notably achieving an 11.3% improvement over the best approach on Inverse-Text.
Yu Xie 0001, Shaoyao Huang, Jiaqing Fan, Ziqiang Cao, Yue Zhang 0011
ACM Multimedia7
2024 CoUDA: Coherence Evaluation via Unified Data Augmentation
abstract
Dawei Zhu, Wenhao Wu, Yifan Song, Fangwei Zhu, Ziqiang Cao, Sujian Li. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Yifan Song 0002, Fangwei Zhu, Ziqiang Cao, Sujian Li
NAACL-HLT5
2024 SelfCP: Compressing over-limit prompt via the frozen large language model itself
Ziqiang Cao, Wenjie Li 0002
Inf. Process. Manag.2
2024 Dialogue acts enhanced extract-abstract framework for meeting summarization
Shichao Sun, Ruifeng Yuan, Wenjie Li 0002, Ziqiang Cao, Sujian Li
Inf. Process. Manag.4
2024 Efficient Image-Text Retrieval via Keyword-Guided Pre-Screening
abstract
Image-text retrieval is a fundamental task to model a connection between images and natural language. Under its flourishing development in performance, most current methods suffer fromN-related time complexity, which hinders their application in practice to a certain extent. Targeting efficiency improvement, we propose a simple and effective keyword-guided pre-screening framework for image-text retrieval. Specifically, we convert the image and text data into keywords and perform keyword matching across the modalities to exclude a large number of irrelevant gallery samples prior to the retrieval network. For the keyword prediction, we transfer it into a multi-label classification problem and propose a multi-task learning scheme by appending the multi-label classifiers to the image-text retrieval network to achieve a lightweight and high-performance keyword prediction. For keyword matching, we introduce the inverted index from the search engine and thus create a win-win situation on both time and space complexities for the pre-screening. Extensive experiments on the two widely-used datasets,i.e., Flickr30K and MS-COCO, verify the effectiveness of the proposed framework. The proposed framework equipped with only two embedding layers achievesO(1) querying time complexity, while improving the retrieval efficiency and maintaining performance, when applied prior to the common image-text retrieval methods.
Min Cao 0005, Ziqiang Cao, Liqiang Nie, Min Zhang 0005
IEEE Trans. Circuits Syst. Video Technol.3
2023 Preserve Context Information for Extract-Generate Long-Input Summarization Framework
abstract
The Extract-generate framework has been a classic approach for text summarization. As pretrained language models struggling with long-input summarization for their high memory cost, extract-generate framework regains researchers' interests. However, the cost of its effectiveness in dealing with long-input summarization is the loss of context information. In this paper, we present a context-aware extract-generate framework (CAEG) for long-input text summarization. It focuses on preserving both local and global context information in an extract-generate framework with little cost, and can be applied to most of existing extract-generate summarization models. CAEG generates a set of context-related text spans called context prompts for each text snippet and use them to transfer the context information from the extractor and generator. To find such context prompts, we propose to capture the context information based on the interpretation of the extractor, where the text spans having the highest contribution to the extraction decision are considered as containing the richest context information. We evaluate our approach on both long-document and long-dialogue summarization datasets: arXiv and QMSum. The experiment results show that CAEG achieves the-state-of-art result on QMSum and outperforms other extract-generate based models in arXiv.
Ruifeng Yuan, Ziqiang Cao, Wenjie Li 0002
AAAI3
2023 Dynamic and Efficient Inference for Text Generation via BERT Family
abstract
Despite the excellent performance of Pretrained Language Models on many text generation tasks, they suffer from inefficient inference on computation and memory due to their largescale parameters and the universal autoregressive decoding paradigm.In this work, we propose a novel fine-tuning method DEER, which can make a single pre-trained model support Dynamic and Efficient infERence and achieve an adaptive trade-off between model performance and latency.In particular, our critical insight is to jointly utilize the non-autoregressive (NAR) generation and dynamic parameter pruning techniques, which can flexibly control the decoding iteration steps and model sizes according to memory and latency limitations.Besides, we also explore the effectiveness of the pre-trained MLMs (i.e., the BERT family) for text generation tasks since their bidirectional attention nature is more suitable for the NAR training objective.Extensive experiments on both monolingual and multilingual pre-trained MLMs demonstrate the effectiveness of our proposed DEER method by consistently achieving (1) higher BLEU scores than the strong autoregressive Transformer model on three neural machine translation tasks with 3 → 12 times speedup, (2) competitive performance (but with much faster inference speed) compared with the BART model on four GLGE benchmark tasks.Our code will be publicly available at GitHub 1 .
Xiaobo Liang, Juntao Li 0005, Lijun Wu 0003, Ziqiang Cao, Min Zhang 0005
ACL (1)4
2023 RaSa: Relation and Sensitivity Aware Representation Learning for Text-based Person Search
abstract
Text-based person search aims to retrieve the specified person images given a textual description. The key to tackling such a challenging task is to learn powerful multi-modal representations. Towards this, we propose a Relation and Sensitivity aware representation learning method (RaSa), including two novel tasks: Relation-Aware learning (RA) and Sensitivity-Aware learning (SA). For one thing, existing methods cluster representations of all positive pairs without distinction and overlook the noise problem caused by the weak positive pairs where the text and the paired image have noise correspondences, thus leading to overfitting learning. RA offsets the overfitting risk by introducing a novel positive relation detection task (i.e., learning to distinguish strong and weak positive pairs). For another thing, learning invariant representation under data augmentation (i.e., being insensitive to some transformations) is a general practice for improving representation's robustness in existing methods. Beyond that, we encourage the representation to perceive the sensitive transformation by SA (i.e., learning to detect the replaced words), thus promoting the representation's robustness. Experiments demonstrate that RaSa outperforms existing state-of-the-art methods by 6.94%, 4.45% and 15.35% in terms of Rank@1 on CUHK-PEDES, ICFG-PEDES and RSTPReid datasets, respectively. Code is available at: https://github.com/Flame-Chasers/RaSa.
Min Cao 0005, Daming Gao, Ziqiang Cao, Chen Chen 0036, Zhenfeng Fan, Liqiang Nie, Min Zhang 0005
IJCAI4
2023 Text-based Person Search without Parallel Image-Text Data
abstract
Text-based person search (TBPS) aims to retrieve the images of the target person from a large image gallery based on a given natural language description. Existing methods are dominated by training models with parallel image-text pairs, which are very costly to collect. In this paper, we make the first attempt to explore TBPS without parallel image-text data (μ-TBPS), in which only non-parallel images and texts, or even image-only data, can be adopted. Towards this end, we propose a two-stage framework, generation-then-retrieval (GTR), to first generate the corresponding pseudo text for each image and then perform the retrieval in a supervised manner. In the generation stage, we propose a fine-grained image captioning strategy to obtain an enriched description of the person image, which firstly utilizes a set of instruction prompts to activate the off-the-shelf pretrained vision-language model to capture and generate fine-grained person attributes, and then converts the extracted attributes into a textual description via the finetuned large language model or the hand-crafted template. In the retrieval stage, considering the noise interference of the generated texts for training model, we develop a confidence score-based training scheme by enabling more reliable texts to contribute more during the training. Experimental results on multiple TBPS benchmarks (i.e., CUHK-PEDES, ICFG-PEDES and RSTPReid) show that the proposed GTR can achieve a promising performance without relying on parallel image-text data.
Min Cao 0005, Chen Chen 0036, Ziqiang Cao, Liqiang Nie, Min Zhang 0005
ACM Multimedia5
2023 MCG-MNER: A Multi-Granularity Cross-Modality Generative Framework for Multimodal NER with Instruction
abstract
Multimodal named entity recognition (MNER) is an essential task of vision and language, which aims to locate named entities and classify them to the predefined categories using visual scenarios. However, existing MNER studies often suffer from bias issues with fine-grained visual cue fusion, which may produce noisy coarse-grained visual cues for MNER. To accurately capture text-image relations and better refine multimodal representations, we propose a novel instruction-based Multi-granularity Cross-modality Generative framework for MNER, namely MCG-MNER. Concretely, we introduce a multi-granularity relation propagation to infer visual clues relevant to text. Then, we propose a method to jnject multi-granularity visual information into cross-modality interaction and fusion to learn a unified representation. Finally, we integrate task-specific instructions and answers for MCG-MNER. Comprehensive experimental results on three benchmark datasets, such as Twitter2015, Twitter2017 and WikiDiverse, demonstrate the superiority of our proposed method over several state-of-the-art MNER methods. We will publicly release our codes for future studies.
Junjie Wu 0005, Chen Gong 0004, Ziqiang Cao, Guohong Fu
ACM Multimedia3
2023 RSpell: Retrieval-Augmented Framework for Domain Adaptive Chinese Spelling Check
Siqi Song, Qi Lv 0001, Lei Geng, Ziqiang Cao, Guohong Fu
NLPCC (1)4
2023 General and Domain-adaptive Chinese Spelling Check with Error-consistent Pretraining
abstract
The lack of label data is one of the significant bottlenecks for Chinese Spelling Check. Existing researches use the automatic generation method by exploiting unlabeled data to expand the supervised corpus. However, there is a big gap between the real input scenario and automatically generated corpus. Thus, we develop a competitive general speller ECSpell, which adopts the Error-consistent masking strategy to create data for pretraining. This error-consistency masking strategy is used to specify the error types of automatically generated sentences consistent with the real scene. The experimental result indicates that our model outperforms previous state-of-the-art models on the general benchmark. Moreover, spellers often work within a particular domain in real life. Due to many uncommon domain terms, experiments on our built domain-specific datasets show that general models perform terribly. Inspired by the common practice of input methods, we propose to add an alterable user dictionary to handle the zero-shot domain-adaption problem. Specifically, we attach a User Dictionary guided inference module (UD) to a general token classification-based speller. Our experiments demonstrate that ECSpell UD , namely, ECSpell combined with UD, surpasses all the other baselines broadly, even approaching the performance on the general benchmark. 1
Qi Lv 0001, Ziqiang Cao, Lei Geng, Chunhui Ai, Guohong Fu
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2022 Few-shot Query-Focused Summarization with Prefix-Merging
abstract
Query-focused summarization has been considered as an important extension for text summarization.It aims to generate a concise highlight for a given query.Different from text summarization, query-focused summarization has long been plagued by the problem of lacking high-quality large-scale datasets.In this paper, we investigate the idea that whether we can integrate and transfer the knowledge of text summarization and question answering to assist the few-shot learning in query-focused summarization.Here, we propose prefix-merging, a prefix-based pretraining strategy for fewshot learning in query-focused summarization.Drawn inspiration from prefix-tuning, we are allowed to integrate the task knowledge from text summarization and question answering into a properly designed prefix and apply the merged prefix to query-focused summarization.With only a small amount of trainable parameters, prefix-merging outperforms fine-tuning on query-focused summarization.We further discuss the influence of different prefix designs and propose a visualized explanation for how prefix-merging works.
Ruifeng Yuan, Ziqiang Cao, Wenjie Li 0002
EMNLP3
2022 Hierarchical Prediction and Adversarial Learning For Conditional Response Generation
abstract
There are a variety of underlying factors influencing what and how people communicate in their daily life. The ability to capture and utilize these factors enables the conversational systems to generate favorable responses and set up amicable connections with users. In this work, we investigate two major factors in response generation, i.e., emotion and intention. To explore the dependency between them, we develop a hierarchical variational model that predicts in sequence the emotion and intention to be conveyed in a response. The response can then be generated word-by-word based on the predictions. We also apply a novel adversarial-augmented inference network to facilitate model training. The experimental results demonstrate the effectiveness of the proposed model as well as the novel adversarial objective. The hypothesis that emotion shapes human communication behavior is also validated.
Yanran Li, Ruixiang Zhang, Wenjie Li 0002, Ziqiang Cao
IEEE Trans. Knowl. Data Eng.4
2021 BASS: Boosting Abstractive Summarization with Unified Semantic Graph
abstract
Wenhao Wu, Wei Li, Xinyan Xiao, Jiachen Liu, Ziqiang Cao, Sujian Li, Hua Wu, Haifeng Wang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Wei Li 0176, Xinyan Xiao, Ziqiang Cao, Sujian Li, Hua Wu 0003, Haifeng Wang 0001
ACL/IJCNLP (1)5
2018 Faithful to the Original: Fact Aware Neural Abstractive Summarization
abstract
Unlike extractive summarization, abstractive summarization has to fuse different parts of the source text, which inclines to create fake facts. Our preliminary study reveals nearly 30% of the outputs from a state-of-the-art neural summarization system suffer from this problem. While previous abstractive summarization approaches usually focus on the improvement of informativeness, we argue that faithfulness is also a vital prerequisite for a practical abstractive summarization system. To avoid generating fake facts in a summary, we leverage open information extraction and dependency parse technologies to extract actual fact descriptions from the source text. The dual-attention sequence-to-sequence framework is then proposed to force the generation conditioned on both the source text and the extracted fact descriptions. Experiments on the Gigaword benchmark dataset demonstrate that our model can greatly reduce fake summaries by 80%. Notably, the fact descriptions also bring significant improvement on informativeness since they often condense the meaning of the source text.
Ziqiang Cao, Furu Wei, Wenjie Li 0002, Sujian Li
AAAI1
2018 Retrieve, Rerank and Rewrite: Soft Template Based Neural Summarization
abstract
Most previous seq2seq summarization systems purely depend on the source text to generate summaries, which tends to work unstably.Inspired by the traditional template-based summarization approaches, this paper proposes to use existing summaries as soft templates to guide the seq2seq model.To this end, we use a popular IR platform to Retrieve proper summaries as candidate templates.Then, we extend the seq2seq framework to jointly conduct template Reranking and templateaware summary generation (Rewriting).Experiments show that, in terms of informativeness, our model significantly outperforms the state-of-the-art methods, and even soft templates themselves demonstrate high competitiveness.In addition, the import of high-quality external summaries improves the stability and readability of generated summaries.
Ziqiang Cao, Wenjie Li 0002, Sujian Li, Furu Wei
ACL (1)1
2017 Joint Copying and Restricted Generation for Paraphrase
abstract
Many natural language generation tasks, such as abstractive summarization and text simplification, are paraphrase-orientated. In these tasks, copying and rewriting are two main writing modes. Most previous sequence-to-sequence (Seq2Seq) models use a single decoder and neglect this fact. In this paper, we develop a novel Seq2Seq model to fuse a copying decoder and a restricted generative decoder. The copying decoder finds the position to be copied based on a typical attention model. The generative decoder produces words limited in the source-specific vocabulary. To combine the two decoders and determine the final output, we develop a predictor to predict the mode of copying or rewriting. This predictor can be guided by the actual writing mode in the training data. We conduct extensive experiments on two different paraphrase datasets. The result shows that our model outperforms the state-of-the-art approaches in terms of both informativeness and language quality.
Ziqiang Cao, Chuwei Luo, Wenjie Li 0002, Sujian Li
AAAI1
2017 Improving Multi-Document Summarization via Text Classification
abstract
Developed so far, multi-document summarization has reached its bottleneck due to the lack of sufficient training data and diverse categories of documents. Text classification just makes up for these deficiencies. In this paper, we propose a novel summarization system called TCSum, which leverages plentiful text classification data to improve the performance of multi-document summarization. TCSum projects documents onto distributed representations which act as a bridge between text classification and summarization. It also utilizes the classification results to produce summaries of different styles. Extensive experiments on DUC generic multi-document summarization datasets show that, TCSum can achieve the state-of-the-art performance without using any hand-crafted features and has the capability to catch the variations of summary styles with respect to different text categories.
Ziqiang Cao, Wenjie Li 0002, Sujian Li, Furu Wei
AAAI1
2017 DailyDialog: A Manually Labelled Multi-turn Dialogue Dataset
abstract
We develop a high-quality multi-turn dialog dataset, DailyDialog, which is intriguing in several aspects. The language is human-written and less noisy. The dialogues in the dataset reflect our daily communication way and cover various topics about our daily life. We also manually label the developed dataset with communication intention and emotion information. Then, we evaluate existing approaches on DailyDialog dataset and hope it benefit the research field of dialog systems. The dataset is available on http://yanran.li/dailydialog
Yanran Li, Hui Su, Xiaoyu Shen 0001, Wenjie Li 0002, Ziqiang Cao, Shuzi Niu
IJCNLP(1)5
2016 TGSum: Build Tweet Guided Multi-Document Summarization Dataset
abstract
The development of summarization research has been significantly hampered by the costly acquisition of reference summaries. This paper proposes an effective way to automatically collect large scales of news-related multi-document summaries with reference to social media's reactions. We utilize two types of social labels in tweets, i.e., hashtags and hyper-links. Hashtags are used to cluster documents into different topic sets. Also, a tweet with a hyper-link often highlights certain key points of the corresponding document. We synthesize a linked document cluster to form a reference summary which can cover most key points. To this aim, we adopt the ROUGE metrics to measure the coverage ratio, and develop an Integer Linear Programming solution to discover the sentence set reaching the upper bound of ROUGE. Since we allow summary sentences to be selected from both documents and high-quality tweets, the generated reference summaries could be abstractive. Both informativeness and readability of the collected summaries are verified by manual judgment. In addition, we train a Support Vector Regression summarizer on DUC generic multi-document summarization benchmarks. With the collected data as extra training resource, the performance of the summarizer improves a lot on all the test sets. We release this dataset for further research.
Ziqiang Cao, Chengyao Chen, Wenjie Li 0002, Sujian Li, Furu Wei, Ming Zhou 0001
AAAI1
2016 AttSum: Joint Learning of Focusing and Summarization with Neural Attention
abstract
Query relevance ranking and sentence saliency ranking are the two main tasks in extractive query-focused summarization. Previous supervised summarization systems often perform the two tasks in isolation. However, since reference summaries are the trade-off between relevance and saliency, using them as supervision, neither of the two rankers could be trained well. This paper proposes a novel summarization system called AttSum, which tackles the two tasks jointly. It automatically learns distributed representations for sentences as well as the document cluster. Meanwhile, it applies the attention mechanism to simulate the attentive reading of human behavior when a query is given. Extensive experiments are conducted on DUC query-focused summarization benchmark datasets. Without using any hand-crafted features, AttSum achieves competitive performance. We also observe that the sentences recognized to focus on the query indeed meet the query need.
Ziqiang Cao, Wenjie Li 0002, Sujian Li, Furu Wei, Yanran Li
COLING1
2015 A Novel Neural Topic Model and Its Supervised Extension
abstract
Topic modeling techniques have the benefits of modeling words and documents uniformly under a probabilistic framework. However, they also suffer from the limitations of sensitivity to initialization and unigram topic distribution, which can be remedied by deep learning techniques. To explore the combination of topic modeling and deep learning techniques, we first explain the standard topic modelfrom the perspective of a neural network. Based on this, we propose a novel neural topic model (NTM) where the representation of words and documents are efficiently and naturally combined into a uniform framework. Extending from NTM, we can easily add a label layer and propose the supervised neural topic model (sNTM) to tackle supervised tasks. Experiments show that our models are competitive in both topic discovery and classification/regression tasks.
Ziqiang Cao, Sujian Li, Yang Liu 0124, Wenjie Li 0002, Heng Ji 0001
AAAI1
2015 Ranking with Recursive Neural Networks and Its Application to Multi-Document Summarization
abstract
We develop a Ranking framework upon Recursive Neural Networks (R2N2) to rank sentences for multi-document summarization. It formulates the sentence ranking task as a hierarchical regression process, which simultaneously measures the salience of a sentence and its constituents (e.g., phrases) in the parsing tree. This enables us to draw on word-level to sentence-level supervisions derived from reference summaries.In addition, recursive neural networks are used to automatically learn ranking features over the tree, with hand-crafted feature vectors of words as inputs. Hierarchical regressions are then conducted with learned features concatenating raw features.Ranking scores of sentences and words are utilized to effectively select informative and non-redundant sentences to generate summaries.Experiments on the DUC 2001, 2002 and 2004 multi-document summarization datasets show that R2N2 outperforms state-of-the-art extractive summarization approaches.
Ziqiang Cao, Furu Wei, Li Dong 0004, Sujian Li, Ming Zhou 0001
AAAI1
2014 Text-level Discourse Dependency Parsing
abstract
Previous researches on Text-level discourse parsing mainly made use of constituency structure to parse the whole document into one discourse tree. In this paper, we present the limitations of constituency based dis-course parsing and first propose to use de-pendency structure to directly represent the relations between elementary discourse units (EDUs). The state-of-the-art depend-ency parsing techniques, the Eisner algo-rithm and maximum spanning tree (MST) algorithm, are adopted to parse an optimal discourse dependency tree based on the arc-factored model and the large-margin learn-ing techniques. Experiments show that our discourse dependency parsers achieve a competitive performance on text-level dis-course parsing. 1
Sujian Li, Liang Wang 0046, Ziqiang Cao, Wenjie Li 0002
ACL (1)3
2014 Joint Learning of Chinese Words, Terms and Keywords
abstract
Previous work often used a pipelined framework where Chinese word segmentation is followed by term extraction and keyword extraction.Such framework suffers from error propagation and is unable to leverage information in later modules for prior components.In this paper, we propose a four-level Dirichlet Process based model (DP-4) to jointly learn the word distributions from the corpus, domain and document levels simultaneously.Based on the DP-4 model, a sentence-wise Gibbs sampler is adopted to obtain proper segmentation results.Meanwhile, terms and keywords are acquired in the sampling process.Experimental results have shown the effectiveness of our method.
Ziqiang Cao, Sujian Li, Heng Ji 0001
EMNLP1