Quangang Li

dblp:151/6251 · DBLP profile ↗
← Back
22ranked-venue papers
1as first author
15since 2021 · last 2026
0000-0003-4005-7717ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 9 since 2021Databases, data management, data science and information retrieval · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 since 2021Computer networks · 3 · 1 first-author
YearPublicationVenuePosition
2026 Fashion Microscope: Pixel-Level Attribute Perception via Optimal Transport and Neural Semantic Aggregation
abstract
Attribute-specific fashion retrieval aims to enhance fine-grained image retrieval by emphasizing the similarity of specific attributes. Current methods primarily rely on attention mechanisms to extract attribute-related visual features but face two key challenges: the limitations of coarse-grained localization in achieving fine-grained accuracy, and an imbalance between global and local perception, where excessive focus on local features can undermine overall performance. To address these issues, we propose the fashion microscope ProFashion, which achieves pixel-level attribute awareness through optimal transport and neural semantic aggregation. The framework begins by employing optimal transport to align semantic attributes with visual patterns from a global perspective, generating an attribute-visual value map that highlights distinctive regions while reducing interference. This is followed by simulating the human brain's perception of attribute feature patterns through superpixel generation and aggregation, capturing attribute-related features at the pixel semantic level and forming key semantic clusters that preserve microstructures. Building on this, an attribute graph is constructed to facilitate feature clustering, significantly enhancing the framework's capability to handle overlapping features and cross-scale relationships. Comprehensive experiments on the FashionAI, DeepFashion, and DARN datasets demonstrate the framework's effectiveness, achieving overall MAP improvements of 3.11%, 3.70%, and 3.49%, respectively. Additionally, the framework delivers relative average throughput gains of 26.94%, 22.22%, and 24.78% on the FashionAI, DeepFashion, and DARN datasets, respectively.
Shuili Zhang, Hongzhang Mu, Jiawei Sheng, Qianqian Tong 0001, Wenyuan Zhang 0002, Quangang Li, Tingwen Liu
AAAI6
2025 Zero-Shot Cross-Domain Slot Filling with Retrieval Augmented In-Context Learning
abstract
Zero-shot cross-domain slot filling is becoming increasingly important due to its ability to generalize to new domains without the need for annotating domain-specific data, which aligns well with the requirements of industrial deployments. Recent advanced works deal with this task through question answering framework and make remarkable progress. However, they always rely on human efforts to manually construct question templates or prompts for all slot types, which is not only labor consuming, but also experience context inconsistency issue between the manual example and the specific test instance. To alleviate this problem, we introduce a retriever designed to extract reference samples from the training sets, serving as demonstrations to guide the model in generating the target slot entity through in-context learning. Building upon this retriever, we propose a retrieval-augmented generative framework that automatically constructs and tailors prompts to each specific test instance, eliminating the need for manual efforts. Experiment results verify that our approach attains the state-of-the-art.
Mengxiao Song, Tingwen Liu, Quangang Li, Duohe Ma, Ling Tian
ICASSP3
2025 Language Models Understand Themselves Better: A Zero-Shot AI-Generated Text Detection Method via Reading and Writing
abstract
The rapid development and widespread adoption of large language models (LLMs) in recent years have introduced significant risks, necessitating robust detection methods to distinguish between AI-generated content and human-written text. Traditional training-based approaches often lack flexibility and frequently make predictions without supporting evidence, especially when adapting to new domains, leading to a lack of interpretability. To address this issue, we propose a novel zero-shot detection framework named Reading and Writing detection method. Our approach utilizes an autoregressive model to assess the intrinsic complexity of text, while leveraging an autoencoder model to quantify the difficulty of reconstructing the text. By integrating these two metrics, we effectively highlight the substantial differences between machine-generated and human-written text. We conduct extensive experiments on four large public datasets from state-of-the-art LLMs, including GPT-3.5, GPT-4, and open-source models like LLaMa. The results demonstrate that our detection method shows tremendous potential across various language generation models and text domains.
Jinao Huang, Tingwen Liu, Quangang Li, Taoyu Su
IJCNN4
2025 Dual-perspective Data Augmentation and Curriculum Learning Framework for Low-resource Complex Named Entity Recognition
abstract
Low-resource complex named entity recognition focuses on identifying complex entities such as creative work, product name and so on, in scenarios where annotated training data is limited. Recent advanced works deal with this task through data augmentation and make substantial progress. However, existing methods ignore the influence of different types or levels of augmented data on model optimization in different learning stages. To address it, we propose a dual-perspective data augmentation and curriculum learning framework. Specifically, we first employ the large language model (LLM) to construct two kinds of augmented datasets from context-perspective and entity-perspective, respectively. Then, we present a multi-stage curriculum learning strategy including a novel adaptive curriculum arrangement algorithm to automatically select the most suitable kind of augmented set to optimize the target model at each training epoch, thus using the augmented data more effectively and controllably. Experimental results on the public benchmark across various low-resource settings show that our framework outperforms previous works.
Mengxiao Song, Tianyun Liu, Wenyuan Zhang 0002, Quangang Li, Tingwen Liu
SIGIR4
2024 Empowering LLMs for Multi-Page Layout Generation via Consistency-Oriented In-Context Learning
abstract
Document layout generation, a burgeoning field of document intelligence, entails positioning and sizing various elements within given constraints. While significant strides have been made in single-page layout generation, real-world documents predominantly span multiple pages, and exploring multi-page layout generation methods has also become the key to meeting the contemporary dramatically increased document processing demands. Despite the promise of leveraging large language models (LLMs) like GPT-4 for their powerful in-context learning abilities, the task transition to multi-page layouts, which contains considerably complex data, presents formidable challenges including excessively long prompts and strict consistency between pages. To this end, we propose a novel framework called Multi-Page Layout Generation via Consistency-Oriented modeling (MuLCO) that capitalizes on in-context learning of LLMs without the need for training or fine-tuning. MuLCO employs three key components: serialization based on code blocks maps intricate document layouts to code-style exemplars, self-correcting reasoning hint decomposes the complex generation task into numerous steps to improve reasoning interpretability, and consistency-oriented multi-round generation predicts coherent multi-page layouts in form of a continuous dialogue. To summarize, we contribute by proposing MuLCO and developing a task-specific dataset and evaluation mechanism. Extensive experiments validate the effectiveness of the MuLCO framework for multi-page layout generation.
Xinghua Zhang 0001, Quangang Li, Tingwen Liu
CIKM4
2024 Fine-Grained Features Alignment and Fusion for Text-Video Cross-Modal Retrieval
abstract
Text-video cross-modal retrieval is an increasingly prominent and challenging task that has garnered significant attention. Traditional models typically embed videos and texts into global vectors, aiming to capture the global features of these modalities. While the models often fall short in capturing fine-grained semantic details. Relying solely on global features proves insufficient to address this challenge. Hence, there is a pressing need to bridge the gap between different modalities by incorporating fine-grained features. In light of this, we propose a highly efficient model designed to capture the fine-grained features of videos and texts including question answer semantic alignment, object alignment and text-video feature fusion. For texts, our model includes the incorporation of entity information and part-of-speech information including adjectives, nouns and verbs information, while for videos, the identification of objects plays a crucial role in facilitating text-video retrieval. Our model undergoes extensive training on the WebVid and CC3M datasets, yielding unequivocal evidence of its superior performance over baseline models. It excels particularly in zero-shot text-video cross-modal retrieval tasks, offering substantial reductions in required computational resources.
Shuili Zhang, Hongzhang Mu, Quangang Li, Chenglong Xiao, Tingwen Liu
ICASSP3
2024 Dynamic Multi-Modal Representation Learning For Topic Modeling
abstract
Topic modeling aims to identify and group related topics within collections of text documents. As multi-modal data becomes increasingly effective and prevalent, researchers are delving into the integration of diverse data types (e.g., images) into topic modeling. However, existing methods for topic modeling struggle with unrelated modal features (e.g., visual features) and efficiency. In this paper, we propose a dynamic multi-modal representation learning method (DMMR) that adaptively integrates multi-modal features to enhance the effectiveness and efficiency of modeling multi-modal data. Concretely, a gating network controls the modality-level decision to choose text, image, or both. Based on the sample-wise choice predicted by the gating network, DMMR performs the fusion of multi-modal features encoded by modality-related expert network. With the dynamic modality selection and fusion, the representative modal features can be chosen to exclude irrelevant modality and speed the inference. Extensive experiments on public datasets demonstrate that the proposed method significantly improves the topic quality (e.g., coherence and diversity) and increases the efficiency by 20.37%.
Hongzhang Mu, Shuili Zhang, Quangang Li, Tingwen Liu
ICME3
2024 CDRNP: Cross-Domain Recommendation to Cold-Start Users via Neural Process
abstract
Cross-domain recommendation (CDR) has been proven as a promising way to tackle the user cold-start problem, which aims to make recommendations for users in the target domain by transferring the user preference derived from the source domain. Traditional CDR studies follow the embedding and mapping (EMCDR) paradigm, which transfers user representations from the source to target domain by learning a user-shared mapping function, neglecting the user-specific preference. Recent CDR studies attempt to learn user-specific mapping functions in meta-learning paradigm, which regards each user's CDR as an individual task, but neglects the preference correlations among users, limiting the beneficial information for user representations. Moreover, both of the paradigms neglect the explicit user-item interactions from both domains during the mapping process. To address the above issues, this paper proposes a novel CDR framework with neural process (NP), termed as CDRNP. Particularly, it develops the meta-learning paradigm to leverage user-specific preference, and further introduces a stochastic process by NP to capture the preference correlations among the overlapping and cold-start users, thus generating more powerful mapping functions by mapping the user-specific preference and common preference correlations to a predictive probability distribution. In addition, we also introduce a preference remainer to enhance the common preference from the overlapping users, and finally devises an adaptive conditional decoder with preference modulation to make prediction for cold-start users with items in the target domain. Experimental results demonstrate that CDRNP outperforms previous SOTA methods in three real-world CDR scenarios.
Xiaodong Li 0012, Jiawei Sheng, Jiangxia Cao, Wenyuan Zhang 0002, Quangang Li, Tingwen Liu
WSDM5
2024 Enhancing Multimodal Entity and Relation Extraction With Variational Information Bottleneck
abstract
This paper studies the multimodal named entity recognition (MNER) and multimodal relation extraction (MRE), which are important for content analysis and various applications. The core of MNER and MRE lies in incorporating evident visual information to enhance textual semantics, where two issues inherently demand investigations. The first issue is modality-noise, where the task-irrelevant information in each modality may be noises misleading the task prediction. The second issue is modality-gap, where representations from different modalities are inconsistent, preventing from building the semantic alignment between the text and image. To address these issues, we propose a novel method for MNER and MRE byMultiModal representation learning withInformationBottleneck (MMIB). For the first issue, a refinement-regularizer probes the information-bottleneck principle to balance the predictive evidence and noisy information, yielding expressive representations for prediction. For the second issue, an alignment-regularizer is proposed, where a mutual information-based item works in a contrastive manner to regularize the consistent text-image representations. To our best knowledge, we are the first to explore variational IB estimation for MNER and MRE. Experiments show that MMIB achieves the state-of-the-art performances on three public benchmarks.
Shiyao Cui, Jiangxia Cao, Xin Cong, Jiawei Sheng, Quangang Li, Tingwen Liu, Jinqiao Shi
IEEE ACM Trans. Audio Speech Lang. Process.5
2024 Cross-Domain NER under a Divide-and-Transfer Paradigm
abstract
Cross-domain Named Entity Recognition (NER) transfers knowledge learned from a rich-resource source domain to improve the learning in a low-resource target domain. Most existing works are designed based on the sequence labeling framework, defining entity detection and type prediction as a monolithic process. However, they typically ignore the discrepant transferability of these two sub-tasks: the former locating spans corresponding to entities is largely domain-robust, whereas the latter owns distinct entity types across domains. Combining them into an entangled learning problem may contribute to the complexity of domain transfer. In this work, we propose the novel divide-and-transfer paradigm in which different sub-tasks are learned using separate functional modules for respective cross-domain transfer. To demonstrate the effectiveness of divide-and-transfer, we concretely implement two NER frameworks by applying this paradigm with different cross-domain transfer strategies. Experimental results on 10 different domain pairs show the notable superiority of our proposed frameworks. Experimental analyses indicate that significant advantages of the divide-and-transfer paradigm over prior monolithic ones originate from its better performance on low-resource data and a much greater transferability. It gives us a new insight into cross-domain NER. Our code is available on GitHub. 1
Xinghua Zhang 0001, Bowen Yu 0002, Xin Cong, Taoyu Su, Quangang Li, Tingwen Liu
ACM Trans. Inf. Syst.5
2023 Adapt-to-Learn Policy Network for Abstractive Multi-document Summarization
abstract
Abstractive multi-document summarization (MDS) aims to generate a summary for a set of topic-related documents, in which different documents may contain trivial and redundant information and present complementary or contradictory content. It is essential to extract salient information and detect redundancy across documents for abstractive MDS compared with single-document summarization. However, it is challenging for models to exploit salient information and generate concise summaries. In this paper, we propose an Adapt-to-Learn Policy (ALP) network to seamlessly adapt the key sentence selection and word generation over retrieval and generative agents, guided by the multi-document summarization reward. Moreover, our model learns latent dependencies among textual units and explicitly takes advantage of critical information by focusing on semantic similarity or discourse relations. Extensive experiments on WikiSum and Multi-News datasets confirm that our method is superior to prior competitive baselines, and experimental analyses show that higher-quality summaries and more fluent word order can be generated in our ALP network.
Hongzhang Mu, Shuili Zhang, Quangang Li, Tingwen Liu
IJCNN3
2023 Enhancing Table Retrieval with Dual Graph Representations
Tianyun Liu, Xinghua Zhang 0001, Zhenyu Zhang 0006, Quangang Li, Tingwen Liu
ECML/PKDD (4)5
2023 Representation and Labeling Gap Bridging for Cross-lingual Named Entity Recognition
abstract
Cross-lingual Named Entity Recognition (NER) aims to address the challenge of data scarcity in low-resource languages by leveraging knowledge from high-resource languages. Most current work relies on general multilingual language models to represent text, and then uses classic combined tagging (e.g., B-ORG) to annotate entities; However, this approach neglects the lack of cross-lingual alignment of entity representations in language models, and also ignores the fact that entity spans and types have varying levels of labeling difficulty in terms of transferability. To address these challenges, we propose a novel framework, referred to as DLBri, which addresses the issues of representation and labeling simultaneously. Specifically, the proposed framework utilizes progressive contrastive learning with source-to-target oriented sentence pairs to pre-finetune the language model, resulting in improved cross-lingual entity-aware representations. Additionally, a decomposition-then-combination procedure is proposed, which separately transfers entity span and type, and then combines their information, to reduce the difficulty of cross-lingual entity labeling. Extensive experiments on 13 diverse language pairs confirm the effectiveness of DLBri.
Xinghua Zhang 0001, Bowen Yu 0002, Jiangxia Cao, Quangang Li, Tingwen Liu
SIGIR4
2022 Event Causality Extraction with Event Argument Correlations
abstract
Event Causality Identification (ECI), which aims to detect whether a causality relation exists between two given textual events, is an important task for event causality understanding. However, the ECI task ignores crucial event structure and cause-effect causality component information, making it struggle for downstream applications. In this paper, we introduce a novel task, namely Event Causality Extraction (ECE), aiming to extract the cause-effect event causality pairs with their structured event information from plain texts. The ECE task is more challenging since each event can contain multiple event arguments, posing fine-grained correlations between events to decide the cause-effect event pair. Hence, we propose a method with a dual grid tagging scheme to capture the intra- and inter-event argument correlations for ECE. Further, we devise a event type-enhanced model architecture to realize the dual grid tagging scheme. Experiments demonstrate the effectiveness of our method, and extensive analyses point out several future directions for ECE.
Shiyao Cui, Jiawei Sheng, Xin Cong, Quangang Li, Tingwen Liu, Jinqiao Shi
COLING4
2022 Enhancing Joint Multiple Intent Detection and Slot Filling with Global Intent-Slot Co-occurrence
abstract
Multi-intent detection and slot filling joint model attracts more and more attention since it can handle multi-intent utterances, which is closer to complex real-world scenarios.Most existing joint models rely entirely on the training procedure to obtain the implicit correlation between intents and slots.However, they ignore the fact that leveraging the rich global knowledge in the corpus can determine the intuitive and explicit correlation between intents and slots.In this paper, we aim to make full use of the statistical co-occurrence frequency between intents and slots as prior knowledge to enhance joint multiple intent detection and slot filling.To be specific, an intent-slot cooccurrence graph is constructed based on the entire training corpus to globally discover correlation between intents and slots.Based on the global intent-slot co-occurrence, we propose a novel graph neural network to model the interaction between the two subtasks.Experimental results on two public multi-intent datasets demonstrate that our approach outperforms the state-of-the-art models.
Mengxiao Song, Bowen Yu 0002, Quangang Li, Tingwen Liu
EMNLP3
2020 Distilling Knowledge from Well-Informed Soft Labels for Neural Relation Extraction
abstract
Extracting relations from plain text is an important task with wide application. Most existing methods formulate it as a supervised problem and utilize one-hot hard labels as the sole target in training, neglecting the rich semantic information among relations. In this paper, we aim to explore the supervision with soft labels in relation extraction, which makes it possible to integrate prior knowledge. Specifically, a bipartite graph is first devised to discover type constraints between entities and relations based on the entire corpus. Then, we combine such type constraints with neural networks to achieve a knowledgeable model. Furthermore, this model is regarded as teacher to generate well-informed soft labels and guide the optimization of a student network via knowledge distillation. Besides, a multi-aspect attention mechanism is introduced to help student mine latent information from text. In this way, the enhanced student inherits the dark knowledge (e.g., type constraints and relevance among relations) from teacher, and directly serves the testing scenarios without any extra constraints. We conduct extensive experiments on the TACRED and SemEval datasets, the experimental results justify the effectiveness of our approach.
Zhenyu Zhang 0006, Xiaobo Shu, Bowen Yu 0002, Tingwen Liu, Jiapeng Zhao, Quangang Li, Li Guo 0001
AAAI6
2020 Joint Entity Linking and Relation Extraction with Neural Networks for Knowledge Base Population
abstract
Relation extraction and entity linking are two fundamental procedures to extend knowledge bases. Most existing methods typically treat them separately and ignore the semantic relevance between entities and relations. In this paper, we pioneer a general joint learning framework for relation extraction and entity linking, which allows these two tasks boost each other. Based on the framework, a demonstration model is proposed with neural networks. We conduct experiments on variants of a standard benchmark dataset (NYT-10) to verify the effectiveness of our approach. Experimental results show that our approach significantly outperforms traditional separate methods without reducing efficiency, especially on datasets with many ambiguous entity mentions. Furthermore, various mainstream methods for relation extraction and entity linking can be easily integrated into our loosely-coupled framework due to its flexible architecture.
Zhenyu Zhang 0006, Xiaobo Sind, Tingwen Liu, Zheng Fang 0002, Quangang Li
IJCNN5
2019 Beyond Word Attention: Using Segment Attention in Neural Relation Extraction
abstract
Relation extraction studies the issue of predicting semantic relations between pairs of entities in sentences. Attention mechanisms are often used in this task to alleviate the inner-sentence noise by performing soft selections of words independently. Based on the observation that information pertinent to relations is usually contained within segments (continuous words in a sentence), it is possible to make use of this phenomenon for better extraction. In this paper, we aim to incorporate such segment information into neural relation extractor. Our approach views the attention mechanism as linear-chain conditional random fields over a set of latent variables whose edges encode the desired structure, and regards attention weight as the marginal distribution of each word being selected as a part of the relational expression. Experimental results show that our method can attend to continuous relational expressions without explicit annotations, and achieve the state-of-the-art performance on the large-scale TACRED dataset.
Bowen Yu 0002, Zhenyu Zhang 0006, Tingwen Liu, Bin Wang 0004, Sujian Li, Quangang Li
IJCAI6
2019 iMCircle: Automatic Mining of Indicators of Compromise from the Web
abstract
With the rapidly evolving landscape of cyber threats, Indicators of Compromise (IOCs) are aggressively exchanged as forensic artifacts to help security professionals quickly identify and response cyber threats. Previous related studies mostly focus on extracting and generating IOCs from some fixed-point monitoring data sources, which are passive and time-consuming. In this paper, we present iMCircle, an innovation system that automatically mines IOCs from the Web by checking suspicious indicators with the help of open-source threat information. Based on the initial input of several suspicious indicators, iMCircle first collects their relevant public threat information from the Web and generates IOCs by checking whether those indicators are threat indicators in the target threat field. Second, it actively extracts new indicators from the search results as new inputs and checks them as described above. In that way, the system works in a circle and generates IOCs continuously. Running this system for almost two months in the real world, it has the appreciable performances on the active checking of suspicious indicators and the automatic generation of IOCs.
Jing Ya, Tingwen Liu, Quangang Li, Jinqiao Shi, Zhaojun Gu
ISCC4
2018 Character-based BiLSTM-CRF Incorporating POS and Dictionaries for Chinese Opinion Target Extraction
abstract
Opinion target extraction (OTE) is a fundamental step for sentiment analysis and opinion summarization. We analyze the difference between Chinese and the Indo-European languages family, and reduce Chinese OTE to a character-based sequence tagging task. Then we introduce two novel features for each character by distributing POS differentially and using predefined templates over contexts and dictionaries. We further propose a character-based BiLSTM-CRF model incorporating the two feature sequences aligned with the character sequence. Experimental results on real-world consumer review datasets show that our work significantly outperforms the baseline methods for Chinese OTE.
Yanzeng Li, Tingwen Liu, Diying Li, Quangang Li, Jinqiao Shi, Yanqiu Wang
ACML4
2014 A probabilistic approach towards modeling email network with realistic features
abstract
Email plays a very important role in our daily life. Much work have been put into practice on email network. Those studies mostly require real email network datasets and reliable models to analyze user information and understand the mechanisms of network evolution. However, much research work is constrained by the absence of real large-scale email datasets. Although email communication is ubiquitous, there are very few large-scale available email datasets satisfied different research purposes. Due to privacy policy and restricted permissions, it is arduous to collect a real large-scale email dataset in a short time. Various social network models are usually used to create synthetic email networks. However, these models focus on modeling several structural properties of network without considering user behaviour patterns. They are not appropriate to generate large-scale realistic synthetic email network datasets. Towards this end, we propose a probabilistic model by which we can construct large-scale synthetic email datasets with a small captured email log. What is more important is that the generated synthetic dataset matches real email network properties and individual communication patterns. Moreover, it has linear complexity, and can be paralleled easily. Experimental results on Enron dataset demonstrate the above benefits of our model.
Quangang Li, Jinqiao Shi, Tingwen Liu, Li Guo 0001, Zhiguang Qin
ICCCN1
2014 Towards misdirected email detection for preventing information leakage
abstract
With the widespread usage of emails, information leakage via misdirected emails becomes a practical and disastrous problem, which should be addressed at all costs. Prior methods have two limitations: privacy issue as relying on email contents to work, and high cost as building too many targeted models. In this paper, we reduce the detection of misdirected emails to a binary classification problem, and build only a universal model to detect misdirected emails. We introduce some representative features that can vividly describe the characteristics of misdirected emails while not infringe users' privacy. Then we design novel algorithms to get these features. The random forest classifier is chosen to perform the detecting task. Experimental results show that our work is able to detect misdirected emails with 89% precision rate and 82% recall rate in average.
Tingwen Liu, Yiguo Pu, Jinqiao Shi, Quangang Li, Xiaojun Chen 0004
ISCC4