Peng Fu 0008

dblp:185/6822-8 · DBLP profile ↗
← Back
32ranked-venue papers
2as first author
28since 2021 · last 2025
0000-0001-9899-8566ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 1 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Multimodal Hypothetical Summary for Retrieval-based Multi-image Question Answering
abstract
Retrieval-based multi-image question answering (QA) task involves retrieving multiple question-related images and synthesizing these images to generate an answer. Conventional "retrieve-then-answer" pipelines often suffer from cascading errors because the training objective of QA fails to optimize the retrieval stage. To address this issue, we propose a novel method to effectively introduce and reference retrieved information into the QA. Given the image set to be retrieved, we employ a multimodal large language model (visual perspective) and a large language model (textual perspective) to obtain multimodal hypothetical summary in question-form and description-form. By combining visual and textual perspectives, MHyS captures image content more specifically and replaces real images in retrieval, which eliminates the modality gap by transforming into text-to-text retrieval and helps improve retrieval. To more advantageously introduce retrieval with QA, we employ contrastive learning to align queries (questions) with MHyS. Moreover, we propose a coarse-to-fine strategy for calculating both sentence-level and word-level similarity scores, to further enhance retrieval and filter out irrelevant details. Our approach achieves a 3.7% absolute improvement over state-of-the-art methods on RETVQA and a 14.5% improvement over CLIP. Comprehensive experiments and detailed ablation studies demonstrate the superiority of our method.
Peize Li, Qingyi Si, Peng Fu 0008, Zheng Lin 0001, Yan Wang 0028
AAAI3
2025 DIVE into MoE: Diversity-Enhanced Reconstruction of Large Language Models from Dense into Mixture-of-Experts
abstract
Large language models (LLMs) with the Mixture-of-Experts (MoE) architecture achieve high cost-efficiency by selectively activating a subset of the parameters. Despite the inference efficiency of MoE LLMs, the training of extensive experts from scratch incurs substantial overhead, whereas reconstructing a dense LLM into an MoE LLM significantly reduces the training budget. However, existing reconstruction methods often overlook the diversity among experts, leading to potential redundancy. In this paper, we come up with the observation that a specific LLM exhibits notable diversity after being pruned on different calibration datasets, based on which we present a Diversity-Enhanced reconstruction method named DIVE. The recipe of DIVE includes domain affinity mining, pruning-based expert reconstruction, and efficient retraining. Specifically, the reconstruction includes pruning and reassembly of the feed-forward network (FFN) module. After reconstruction, we efficiently retrain the model on routers, experts and normalization modules. We implement DIVE on Llama-style LLMs with open-source training corpora. Experiments show that DIVE achieves training efficiency with minimal accuracy trade-offs, outperforming existing pruning and MoE reconstruction methods with the same number of activated parameters. Code is available at: https://github.com/yuchenblah/DIVE.
Bowen Shen, Naibin Gu, Jiaxuan Zhao, Peng Fu 0008, Zheng Lin 0001, Weiping Wang 0005
ACL (1)5
2025 BeamLoRA: Beam-Constraint Low-Rank Adaptation
abstract
Due to the demand for efficient fine-tuning of large language models, Low-Rank Adaptation (LoRA) has been widely adopted as one of the most effective parameter-efficient fine-tuning methods. Nevertheless, while LoRA improves efficiency, there remains room for improvement in accuracy. Herein, we adopt a novel perspective to assess the characteristics of LoRA ranks. The results reveal that different ranks within the LoRA modules not only exhibit varying levels of importance but also evolve dynamically throughout the fine-tuning process, which may limit the performance of LoRA. Based on these findings, we propose BeamLoRA, which conceptualizes each LoRA module as a beam where each rank naturally corresponds to a potential sub-solution, and the fine-tuning process becomes a search for the optimal sub-solution combination. BeamLoRA dynamically eliminates underperforming sub-solutions while expanding the parameter space for promising ones, enhancing performance with a fixed rank. Extensive experiments across three base models and 12 datasets spanning math reasoning, code generation, and commonsense reasoning demonstrate that BeamLoRA consistently enhances the performance of LoRA, surpassing the other baseline methods.
Naibin Gu, Zhenyu Zhang 0006, Xiyu Liu 0003, Peng Fu 0008, Zheng Lin 0001, Shuohuan Wang, Hua Wu 0003, Weiping Wang 0005, Haifeng Wang 0001
ACL (1)4
2025 Adapt Once, Thrive with Updates: Transferable Parameter-Efficient Fine-Tuning on Evolving Base Models
abstract
Parameter-efficient fine-tuning (PEFT) has become a common method for fine-tuning large language models, where a base model can serve multiple users through PEFT module switching.To enhance user experience, base models require periodic updates.However, once updated, PEFT modules fine-tuned on previous versions often suffer substantial performance degradation on newer versions.Re-tuning these numerous modules to restore performance would incur significant computational costs.Through a comprehensive analysis of the changes that occur during base model updates, we uncover an interesting phenomenon: continual training primarily affects task-specific knowledge stored in Feed-Forward Networks (FFN), while having less impact on the task-specific pattern in the Attention mechanism.Based on these findings, we introduce Trans-PEFT, a novel approach that enhances the PEFT module by focusing on the task-specific pattern while reducing its dependence on certain knowledge in the base model.Further theoretical analysis supports our approach.Extensive experiments across 7 base models and 12 datasets demonstrate that Trans-PEFT trained modules can maintain performance on updated base models without re-tuning, significantly reducing maintenance overhead in real-world applications 1 .
Naibin Gu, Peng Fu 0008, Xiyu Liu 0003, Zheng Lin 0001, Weiping Wang 0005
ACL (1)2
2025 KA-CDRE: Knowledge-Augmented Cross-Document Relation Extraction
Peize Li, Jingzi Gu, Peng Fu 0008, Zheng Lin 0001, Weiping Wang 0005
ADMA (4)4
2025 LEAP: An LLM-Based Evidence Augmented Pipeline for Table-Based Fact Verification
Hanwen Zhang 0010, Qingyi Si, Peng Fu 0008, Zheng Lin 0001, Zhigang Lu 0001, Weiping Wang 0005
ADMA (1)3
2025 CBP-Tuning: Efficient Local Customization for Black-box Large Language Models
abstract
The high costs of customizing large language models (LLMs) fundamentally limit their adaptability to user-specific needs.Consequently, LLMs are increasingly offered as cloud-based services, a paradigm that introduces critical limitations: providers struggle to support personalized customization at scale, while users face privacy risks when exposing sensitive data.To address this dual challenge, we propose Customized Black-box Prompt Tuning (CBP-Tuning), a novel framework that facilitates efficient local customization while preserving bidirectional privacy.Specifically, we design a two-stage framework: (1) a prompt generator trained on the server-side to capture domainspecific and task-agnostic capabilities, and (2) user-side gradient-free optimization that tailors soft prompts for individual tasks.This approach eliminates the need for users to access model weights or upload private data, requiring only a single customized vector per task while achieving effective adaptation.Furthermore, the evaluation of CBP-Tuning in the commonsense reasoning, medical and financial domain settings demonstrates superior performance compared to baselines, showcasing its advantages in task-agnostic processing and privacy preservation.
Jiaxuan Zhao, Naibin Gu, Xiyu Liu 0003, Peng Fu 0008, Zheng Lin 0001, Weiping Wang 0005
EMNLP5
2025 Reconstruction of Differentially Private Text Sanitization via Large Language Models
abstract
Differential privacy (DP) is the de facto privacy standard against privacy leakage attacks, including many recently discovered ones against large language models (LLMs). However, we discovered that LLMs could reconstruct the altered/removed privacy from given DP-sanitized prompts. We propose two attacks (black-box and white-box) based on the accessibility to LLMs and show that LLMs could connect the pair of DPsanitized text and the corresponding private training data of LLMs by giving sample text pairs as instructions (in the blackbox attacks) or fine-tuning data (in the white-box attacks). To illustrate our findings, we conduct comprehensive experiments on modern LLMs (e.g., LLaMA-2, LLaMA-3, ChatGPT-3.5, ChatGPT-4, ChatGPT-4o, Claude-3, Claude-3.5, OPT, GPT-Neo, GPT-J, Gemma-2, and Pythia) using commonly used datasets (such as WikiMIA, Pile-CC, and Pile-Wiki) against both wordlevel and sentence-level DP. The experimental results show promising recovery rates, e.g., the black-box attacks against the word-level DP over WikiMIA dataset gave 72.18% on LLaMA2 (70B), 82.39% on LLaMA-3 (70B), 75.35% on Gemma-2, 91.2% on ChatGPT-4o, and 94.01% on Claude-3.5 (Sonnet). More urgently, this study indicates that these well-known LLMs have emerged as a new security risk for existing DP text sanitization approaches in the current environment.
Shuchao Pang, Zhigang Lu 0001, Haichen Wang, Peng Fu 0008, Yongbin Zhou, Minhui Xue 0001
RAID4
2024 Object Attribute Matters in Visual Question Answering
abstract
Visual question answering is a multimodal task that requires the joint comprehension of visual and textual information. However, integrating visual and textual semantics solely through attention layers is insufficient to comprehensively understand and align information from both modalities. Intuitively, object attributes can naturally serve as a bridge to unify them, which has been overlooked in previous research. In this paper, we propose a novel VQA approach from the perspective of utilizing object attribute, aiming to achieve better object-level visual-language alignment and multimodal scene understanding. Specifically, we design an attribute fusion module and a contrastive knowledge distillation module. The attribute fusion module constructs a multimodal graph neural network to fuse attributes and visual features through message passing. The enhanced object-level visual features contribute to solving fine-grained problem like counting-question. The better object-level visual-language alignment aids in understanding multimodal scenes, thereby improving the model's robustness. Furthermore, to augment scene understanding and the out-of-distribution performance, the contrastive knowledge distillation module introduces a series of implicit knowledge. We distill knowledge into attributes through contrastive loss, which further strengthens the representation learning of attribute features and facilitates visual-linguistic alignment. Intensive experiments on six datasets, COCO-QA, VQAv2, VQA-CPv2, VQA-CPv1, VQAvs and TDIUC, show the superiority of the proposed method.
Peize Li, Qingyi Si, Peng Fu 0008, Zheng Lin 0001, Yan Wang 0028
AAAI3
2024 Are Large Language Models Table-based Fact-Checkers?
abstract
Table-based Fact Verification (TFV) aims to extract the entailment relationship between statements and structured tables. Existing TFV methods based on small-scale models suffer from insufficient labeled data and weak zero-shot ability. Recently, the appearance of Large Language Models (LLMs) has gained lots of attraction in research fields. They have shown strong zero-shot and in-context learning capabilities on several NLP tasks, but their potential on TFV is still unknown. In this work, we implement a preliminary study on whether LLMs are table-based fact-checkers. In detail, we design various prompts to explore how in-context learning can help LLMs in TFV, i.e., zero-shot and few-shot TFV capability. Besides, we carefully design and construct TFV instructions to investigate the performance gain brought by the instruction tuning of LLMs. Experimental results demonstrate that LLMs can achieve acceptable results on zero-shot and few-shot TFV with prompt engineering, while instruction-tuning can stimulate the TFV capability significantly. We also make some valuable findings about the format of zero-shot prompts and the number of in-context examples. Finally, we analyze some possible directions to promote the accuracy of TFV via LLMs, which will benefit further research on table reasoning.
Hanwen Zhang 0010, Qingyi Si, Peng Fu 0008, Zheng Lin 0001, Weiping Wang 0005
CSCWD3
2024 AnchorMine: An Efficient Graph Pattern Matching System for Specific Vertex Matching
abstract
As data scales continue to expand, graph structures are widely applied across multiple domains due to their effective organization of complex data. Graph Pattern Matching (GPM) is a fundamental task in graph analysis to identify all user-interesting subgraphs in a graph. Current GPM systems achieve this goal by generating efficient traversal path strategies. However, when matching patterns that include a specific vertex (S-GPM), current GPM systems often traverse paths without the specific vertex or duplicate traverse some paths. These redundant traversals lead to decreased execution efficiency. In this paper, we introduce AnchorMine, a GPM system designed for S-GPM tasks, aiming to significantly reduce redundant path traversal by identifying and reusing paths that include specific vertex. Specifically, AnchorMine first analyzes the pattern to identify vertices in different positions within the pattern, named Anchors (ACs). Then AnchorMine extracts features of reusable paths based on each Anchor (AC). These features enable the system to identify paths that can be reused during matching. Using these features, it further generates the parameters required for matching based on path reuse, achieving efficient matching for S-GPM tasks. In experiments on 8 real-world graph datasets, AnchorMine significantly outperformed GraphPi, SandSlash, and Peregrine on 6 datasets used for performance testing, with matching performance improvements of 3249.22 ×, 2018.73 × and 7573.27 ×, respectively. On the remaining 2 datasets used for scalability testing, AnchorMine scales well.
Jianhuan Zhuo, Mingzhe Xing, Yinliang Yue, Peng Fu 0008, Weiping Wang 0005
MSN5
2024 Cross-modality Multiple Relations Learning for Knowledge-based Visual Question Answering
abstract
Knowledge-based visual question answering not only needs to answer the questions based on images but also incorporates external knowledge to study reasoning in the joint space of vision and language. To bridge the gap between visual content and semantic cues, it is important to capture the question-related and semantics-rich vision-language connections. Most existing solutions model simple intra-modality relation or represent cross-modality relation using a single vector, which makes it difficult to effectively model complex connections between visual features and question features. Thus, we propose a cross-modality multiple relations learning model, aiming to better enrich cross-modality representations and construct advanced multi-modality knowledge triplets. First, we design a simple yet effective method to generate multiple relations that represent the rich cross-modality relations. The various cross-modality relations link the textual question to the related visual objects. These multi-modality triplets efficiently align the visual objects and corresponding textual answers. Second, to encourage multiple relations to better align with different semantic relations, we further formulate a novel global-local loss. The global loss enables the visual objects and corresponding textual answers close to each other through cross-modality relations in the vision-language space, and the local loss better preserves semantic diversity among multiple relations. Experimental results on the Outside Knowledge VQA and Knowledge-Routed Visual Question Reasoning datasets demonstrate that our model outperforms the state-of-the-art methods.
Yan Wang 0028, Peize Li, Qingyi Si, Hanwen Zhang 0010, Wenyu Zang, Zheng Lin 0001, Peng Fu 0008
ACM Trans. Multim. Comput. Commun. Appl.7
2023 A Gradient Control Method for Backdoor Attacks on Parameter-Efficient Tuning
abstract
Parameter-Efficient Tuning (PET) has shown remarkable performance by fine-tuning only a small number of parameters of the pre-trained language models (PLMs) for the downstream tasks, while it is also possible to construct backdoor attacks due to the vulnerability of pretrained weights.However, a large reduction in the number of attackable parameters in PET will cause the user's fine-tuning to greatly affect the effectiveness of backdoor attacks, resulting in backdoor forgetting.We find that the backdoor injection process can be regarded as multitask learning, which has a convergence imbalance problem between the training of clean and poisoned data.And this problem might result in forgetting the backdoor.Based on this finding, we propose a gradient control method to consolidate the attack effect, comprising two strategies.One controls the gradient magnitude distribution cross layers within one task and the other prevents the conflict of gradient directions between tasks.Compared with previous backdoor attack methods in the scenario of PET, our method improves the effect of the attack on sentiment classification and spam detection respectively, which shows that our method is widely applicable to different tasks.
Naibin Gu, Peng Fu 0008, Xiyu Liu 0003, Zhengxiao Liu, Zheng Lin 0001, Weiping Wang 0005
ACL (1)2
2023 BMIPN: A Biased Multi-Granularity Interaction Prototype Network for Few-Shot Relation Extraction
abstract
Few-shot relation extraction (FSRE) focuses on detecting new relations through a few annotated instances. Most existing works adopt prototypical network-based models for FSRE. They compute prototype representations for each class in the support set separately, and then use prototype representations for relation prediction. In this way, they learn only the knowledge of each class, regardless of the high-level interactions among these classes. However, these interactions can help the model understand diversity and improve discrimination, which is essential for FSRE, especially for similar relations’ prediction. In this work, we introduce a novel Biased Multi-granularity Interaction Prototype Network (BMIPN). Specifically, we mimic human cognitive processes to model explicit and adaptive interactions from intra- and inter-class aspects. Furthermore, we propose a novel biased contrastive learning method that encourages the model to focus on contrasting similar relations, generating discriminative and robust prototype representations. Experimental results on two benchmark datasets demonstrate that BMIPN outperforms state-of-the-art models and achieves better performance with respect to similar relations.
Yile Li, Yinliang Yue, Xiaoyan Gu 0001, Peng Fu 0008, Weiping Wang 0005
ECAI4
2023 Compressing and Debiasing Vision-Language Pre-Trained Models for Visual Question Answering
abstract
Despite the excellent performance of visionlanguage pre-trained models (VLPs) on conventional VQA task, they still suffer from two problems: First, VLPs tend to rely on language biases in datasets and fail to generalize to outof-distribution (OOD) data.Second, they are inefficient in terms of memory footprint and computation.Although promising progress has been made in both problems, most existing works tackle them independently.To facilitate the application of VLP to VQA tasks, it is imperative to jointly study VLP compression and OOD robustness, which, however, has not yet been explored.This paper investigates whether a VLP can be compressed and debiased simultaneously by searching sparse and robust subnetworks.To this end, we systematically study the design of a training and compression pipeline to search the subnetworks, as well as the assignment of sparsity to different modality-specific modules.Our experiments involve 3 VLPs, 2 compression methods, 4 training methods, 2 datasets and a range of sparsity levels.Our results show that there indeed exist sparse and robust subnetworks, which are competitive with the debiased full VLP and clearly outperform the debiasing SoTAs with fewer parameters on OOD datasets VQA-CP v2 and VQA-VS. 1
Qingyi Si, Yuanxin Liu, Zheng Lin 0001, Peng Fu 0008, Yanan Cao 0001, Weiping Wang 0005
EMNLP4
2023 Schema Item Matters in Knowledge Base Question Answering
abstract
Knowledge base question answering is a challenging task that aims to answer questions by querying knowledge bases. Recently, state-of-the-art methods tend to include an enumerator module and a ranker module. They first enumerate by searching the knowledge base and then rank the candidates to select the target logical form. However, these methods sometimes fail to cover the candidates which involve more complex combinations. A recent solution to this issue is to add a generator module after the ranker to generate the uncovered target logical form. However, the enumerator and ranker always discard partial ground truth schema items. Consequently, the lack of them in the generator input results in the failure to generate the target logical form. To address this problem, we present a novel framework, SIMQA, to reuse the neglected schema items, i.e., classes and relations. Specifically, we adopt a matcher module to select the most related schema items for the given question, and feed them to the generator. On this basis, we propose a novel generation model based on contrastive learning to force the model to focus on the supplemental schema items. Experiment results on GRAILQA and WEBQSP datasets demonstrate the highly competitive performance of the proposed method, and verify that schema item matters in KBQA.
Zhe Wen, Qingyi Si, Zheng Lin 0001, Peng Fu 0008, Weiping Wang 0005
IJCNN5
2022 Target Really Matters: Target-aware Contrastive Learning and Consistency Regularization for Few-shot Stance Detection
abstract
Stance detection aims to identify the attitude from an opinion towards a certain target. Despite the significant progress on this task, it is extremely time-consuming and budget-unfriendly to collect sufficient high-quality labeled data for every new target under fully-supervised learning, whereas unlabeled data can be collected easier. Therefore, this paper is devoted to few-shot stance detection and investigating how to achieve satisfactory results in semi-supervised settings. As a target-oriented task, the core idea of semi-supervised few-shot stance detection is to make better use of target-relevant information from labeled and unlabeled data. Therefore, we develop a novel target-aware semi-supervised framework. Specifically, we propose a target-aware contrastive learning objective to learn more distinguishable representations for different targets. Such an objective can be easily applied with or without unlabeled data. Furthermore, to thoroughly exploit the unlabeled data and facilitate the model to learn target-relevant stance features in the opinion content, we explore a simple but effective target-aware consistency regularization combined with a self-training strategy. The experimental results demonstrate that our approach can achieve state-of-the-art performance on multiple benchmark datasets in the few-shot setting.
Rui Liu 0032, Zheng Lin 0001, Huishan Ji, Peng Fu 0008, Weiping Wang 0005
COLING5
2022 Deep Piecewise Hashing for Efficient Hamming Space Retrieval
abstract
Hamming space retrieval can achieve constant-time image search, which is more efficient than linear scan. In Hamming space retrieval, the data points inside the Hamming ball imply retrievable while the data points outside are irretrievable. Therefore, it is crucial to explicitly characterize the Hamming ball. However, for the existing Hamming space retrieval methods, many similar points are found close to the outside of the Hamming ball while many dissimilar points are found close to the query point, leading to the decline of both retrieval accuracy and recall. In this paper, we present a novel method named Deep Piecewise Hashing (DPH), for Efficient Hamming Space Retrieval. A piecewise loss is elaborately designed to guide the learning of hash codes. Meanwhile, a piecewise probability distribution is introduced in the proposed loss function. The piecewise probability distribution pays more attention to the learning of those "marginal" similar points. It considers both discrimination and robustness for the dissimilar points inside the Hamming ball. Comprehensive experiments on two datasets, MS-COCO and NUS-WIDE, demonstrate that DPH can yield state-of-the-art Hamming space retrieval performance.
Jingzi Gu, Dayan Wu, Peng Fu 0008, Bo Li 0063, Weiping Wang 0005
ICASSP3
2022 Cross-Target Stance Detection Via Refined Meta-Learning
abstract
Cross-target stance detection (CTSD) aims to identify the stance of the text towards a target, where stance annotations are available for (though related but) different targets. Recently, models based on external semantic and emotion knowledge have been proposed for CTSD, achieving promising performance. However, such solutions rely on much external resources and harness only one source target, which is a waste of other available targets. To address the problem above, we propose a many-to-one CTSD model based on meta-learning. To make the most of meta-learning, we further refine it with a balanced and easy-to-hard learning pattern. Specifically, for multiple-target training, we feed the model according to the similarity among targets, and utilize two kinds of re-balanced strategies to deal with the imbalance in data. We conduct experiments on SemEval 2016 task 6, and results demonstrate that our method is effective and establishes a new state-of-the-art macro-f1 score for CTSD.
Huishan Ji, Zheng Lin 0001, Peng Fu 0008, Weiping Wang 0005
ICASSP3
2022 Connecting Targets via Latent Topics And Contrastive Learning: A Unified Framework For Robust Zero-Shot and Few-Shot Stance Detection
abstract
Zero-shot and few-shot stance detection (ZFSD) aims to automatically identify the users’ stance toward a wide range of continuously emerging targets without or with limited labeled data. Previous works on in-target and cross-target stance detection typically focus on extremely limited targets, which is not applicable to the zero-shot and few-shot scenarios. Additionally, existing ZFSD models are not good at modeling the relationship between seen and unseen targets. In this paper, we propose a unified end-to-end framework with a discrete latent topic variable that implicitly establishes the connections between targets. Moreover, we apply supervised contrastive learning to enhance the generalization ability of the model. Comprehensive experiments on the ZFSD task verify the effectiveness and superiority of our proposed method.
Rui Liu 0032, Zheng Lin 0001, Peng Fu 0008, Yuanxin Liu, Weiping Wang 0005
ICASSP3
2022 Neutral Utterances are Also Causes: Enhancing Conversational Causal Emotion Entailment with Social Commonsense Knowledge
abstract
Conversational Causal Emotion Entailment aims to detect causal utterances for a non-neutral targeted utterance from a conversation. In this work, we build conversations as graphs to overcome implicit contextual modelling of the original entailment style. Following the previous work, we further introduce the emotion information into graphs. Emotion information can markedly promote the detection of causal utterances whose emotion is the same as the targeted utterance. However, it is still hard to detect causal utterances with different emotions, especially neutral ones. The reason is that models are limited in reasoning causal clues and passing them between utterances. To alleviate this problem, we introduce social commonsense knowledge (CSK) and propose a Knowledge Enhanced Conversation graph (KEC). KEC propagates the CSK between two utterances. As not all CSK is emotionally suitable for utterances, we therefore propose a sentiment-realized knowledge selecting strategy to filter CSK. To process KEC, we further construct the Knowledge Enhanced Directed Acyclic Graph networks. Experimental results show that our method outperforms baselines and infers more causes with different emotions from the targeted utterance.
Fandong Meng, Zheng Lin 0001, Rui Liu 0032, Peng Fu 0008, Yanan Cao 0001, Weiping Wang 0005, Jie Zhou 0016
IJCAI5
2022 Detach and Attach: Stylized Image Captioning without Paired Stylized Dataset
abstract
Stylized Image Captioning aims to generate captions with accurate image content and stylized elements simultaneously. However, large-scaled image and stylized caption pairs cost lots of resources and are usually unavailable. Therefore, it's a challenge to generate stylized captions without paired stylized caption dataset. Previous work on controlling the style of generated captions in an unsupervised way can be divided into two ways: implicitly and explicitly. The former mainly relies on a well-trained language model to capture style knowledge, which is limited to a single style and hard to handle multi-style task. Thus, the latter uses extra style constraints such as outlined style labels or stylized words extracted from stylized sentences to control the style rather than the trained style-specific language model. However, certain styles, such as humorous and romance, are implied in the whole sentence, instead of in some words of a sentence. To address the problems above, we propose a two-step method based on Transformer: firstly detach style representations from large-scaled stylized text-only corpus to provide more holistic style supervision, and secondly attach the style representations to image content to generate stylized captions. We learn a shared image-text space to narrow the gap between the image and the text modality for better attachment. Due to the trade-off between semantics and style, we explore three injection methods of style representations to balance two requirements of image content preservation and stylization. Experiments show that our method outperforms the state-of-the-art systems in overall performance, especially on implied styles.
Yutong Tan, Zheng Lin 0001, Peng Fu 0008, Mingyu Zheng, Lanrui Wang, Yanan Cao 0001, Weiping Wang 0005
ACM Multimedia3
2022 Learning to Win Lottery Tickets in BERT Transfer via Task-agnostic Mask Training
abstract
Yuanxin Liu, Fandong Meng, Zheng Lin, Peng Fu, Yanan Cao, Weiping Wang, Jie Zhou. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Yuanxin Liu, Fandong Meng, Zheng Lin 0001, Peng Fu 0008, Yanan Cao 0001, Weiping Wang 0005, Jie Zhou 0016
NAACL-HLT4
2022 A Win-win Deal: Towards Sparse and Robust Pre-trained Language Models
abstract
Despite the remarkable success of pre-trained language models (PLMs), they still face two challenges: First, large-scale PLMs are inefficient in terms of memory footprint and computation. Second, on the downstream tasks, PLMs tend to rely on the dataset bias and struggle to generalize to out-of-distribution (OOD) data. In response to the efficiency problem, recent studies show that dense PLMs can be replaced with sparse subnetworks without hurting the performance. Such subnetworks can be found in three scenarios: 1) the fine-tuned PLMs, 2) the raw PLMs and then fine-tuned in isolation, and even inside 3) PLMs without any parameter fine-tuning. However, these results are only obtained in the in-distribution (ID) setting. In this paper, we extend the study on PLMs subnetworks to the OOD setting, investigating whether sparsity and robustness to dataset bias can be achieved simultaneously. To this end, we conduct extensive experiments with the pre-trained BERT model on three natural language understanding (NLU) tasks. Our results demonstrate that \textbf{sparse and robust subnetworks (SRNets) can consistently be found in BERT}, across the aforementioned three scenarios, using different training and compression methods. Furthermore, we explore the upper bound of SRNets using the OOD information and show that \textbf{there exist sparse and almost unbiased BERT subnetworks}. Finally, we present 1) an analytical study that provides insights on how to promote the efficiency of SRNets searching process and 2) a solution to improve subnetworks' performance at high sparsity. The code is available at \url{https://github.com/llyx97/sparse-and-robust-PLM}.
Yuanxin Liu, Fandong Meng, Zheng Lin 0001, Peng Fu 0008, Yanan Cao 0001, Weiping Wang 0005, Jie Zhou 0016
NeurIPS5
2022 EmoMix+: An Approach of Depression Detection Based on Emotion Lexicon for Mobile Application
abstract
Emotion lexicon is an important auxiliary resource for text emotion analysis. Previous works mainly focused on positive and negative classification and less on fine-grained emotion classification. Researchers use lexicon-based methods to find that patients with depression express more negative emotions on social media. Emotional characteristics are an effective feature in detecting depression, but the traditional emotion lexicon has limitations in detecting depression and ignores many depression words. Therefore, we build an emotion lexicon for depression to further study the differences between healthy users and patients with depression. The experimental results show that the depression lexicon constructed in this paper is effective and has a better effect of classifying users with depression.
Yuanfei Zhang, Lihua Yin, Zhe Sun 0005, Zheng Lin 0001, Peng Fu 0008, Weiping Wang 0005
Secur. Commun. Networks6
2021 Check It Again: Progressive Visual Question Answering via Visual Entailment
abstract
Qingyi Si, Zheng Lin, Ming yu Zheng, Peng Fu, Weiping Wang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Qingyi Si, Zheng Lin 0001, Mingyu Zheng, Peng Fu 0008, Weiping Wang 0005
ACL/IJCNLP (1)4
2021 Learning Class-Transductive Intent Representations for Zero-shot Intent Detection
abstract
Zero-shot intent detection (ZSID) aims to deal with the continuously emerging intents without annotated training data. However, existing ZSID systems suffer from two limitations: 1) They are not good at modeling the relationship between seen and unseen intents. 2) They cannot effectively recognize unseen intents under the generalized intent detection (GZSID) setting. A critical problem behind these limitations is that the representations of unseen intents cannot be learned in the training stage. To address this problem, we propose a novel framework that utilizes unseen class labels to learn Class-Transductive Intent Representations (CTIR). Specifically, we allow the model to predict unseen intents during training, with the corresponding label names serving as input utterances. On this basis, we introduce a multi-task learning objective, which encourages the model to learn the distinctions among intents, and a similarity scorer, which estimates the connections among intents more accurately. CTIR is easy to implement and can be integrated with existing ZSID and GZSID methods. Experiments on two real-world datasets show that CTIR brings considerable improvement to the baseline systems.
Qingyi Si, Yuanxin Liu, Peng Fu 0008, Zheng Lin 0001, Weiping Wang 0005
IJCAI3
2021 Sarcasm Detection with Commonsense Knowledge
abstract
Sarcasm is commonly used in today's social media platforms such as Twitter and Reddit. Sarcasm detection is necessary for analysing people's real sentiments as people usually use sarcasm to express a flipped emotion against the literal meaning. However, the current works neglect the fact that commonsense knowledge is crucial for sarcasm recognition. In this paper, we propose a novel architecture in deep learning for sarcasm detection by integrating commonsense knowledge. To be specific, we apply the pre-trained COMET model to generate relevant commonsense knowledge. Besides, we compare two kinds of knowledge selection strategies to investigate how commonsense knowledge influences performance. Finally, a knowledge-text integration module is designed to model both text and knowledge. The experimental results demonstrate our model's effectiveness on three datasets, including two Twitter datasets and a Reddit dataset.
Hongliang Pan, Zheng Lin 0001, Peng Fu 0008, Weiping Wang 0005
IEEE ACM Trans. Audio Speech Lang. Process.4
2020 Modeling the Incongruity Between Sentence Snippets for Sarcasm Detection
abstract
Sarcasm is a form of irony used to mock or convey contempt, which occurs when there is an incongruity between the literal meaning of an utterance and its intended meaning. Many studies identify sarcasm by capturing the incongruity in-between the words. However, consider the following example, "I love waking up at 4 am on Saturday", there is no apparent incongruity in-between the words. Intuitively, the word "love" and the snippet "waking up at 4 am on Saturday" form a strong contrast. Thus, capturing the incongruity among the sentence snippets is more reasonable since a sentence snippet usually contains more semantic information than a single word. Additionally, not all snippets are equally important when human beings identify sarcasm. Thus, inspired by the above observations, we propose the Self-Attention of Weighted Snippets (SAWS) model for sarcasm detection, which overcomes the problem that the previous models are inefficient in determining the sarcasm caused by snippet incongruity. The experiment results show that our model achieves state-of-the-art performance on four benchmark datasets, including two short text Twitter datasets and two long text Internet Argument Corpus (IAC) datasets.
Hongliang Pan, Zheng Lin 0001, Peng Fu 0008, Weiping Wang 0005
ECAI3
2019 A Multi-channel Neural Network for Imbalanced Emotion Recognition
abstract
Imbalanced issue becomes one of major bottleneck for further popularizing of emotion recognition in actual applications. Recently, some resampling methods have been proposed to improve performance by balancing the training samples. However, over-sampling methods may lead to overfitting, and undersampling methods would lose useful emotion information. In this paper, we propose a multi-channel deep architecture to improve performance in both samples and features imbalance. Specifically, we design a class correction loss function to overcome the gap between majority and minority emotions. Meanwhile, emotionspecific word embedding and a fine-tuning BERT are used to increase the differentiation of emotion words and sentences. Experimental results on two Chinese micro-blog emotion classification datasets show that our proposed architecture outperforms state-of-the-art in imbalanced emotion recognition.
Qingyi Si, Peng Fu 0008, Zheng Lin 0001, Weiping Wang 0005
ICTAI3
2018 Learning Sentiment-Specific Word Embedding via Global Sentiment Representation
abstract
Context-based word embedding learning approaches can model rich semantic and syntactic information. However, it is problematic for sentiment analysis because the words with similar contexts but opposite sentiment polarities, such as good and bad, are mapped into close word vectors in the embedding space. Recently, some sentiment embedding learning methods have been proposed, but most of them are designed to work well on sentence-level texts. Directly applying those models to document-level texts often leads to unsatisfied results. To address this issue, we present a sentiment-specific word embedding learning architecture that utilizes local context informationas well as global sentiment representation. The architecture is applicable for both sentence-level and document-level texts. We take global sentiment representation as a simple average of word embeddings in the text, and use a corruption strategy as a sentiment-dependent regularization. Extensive experiments conducted on several benchmark datasets demonstrate that the proposed architecture outperforms the state-of-the-art methods for sentiment classification.
Peng Fu 0008, Zheng Lin 0001, Fengcheng Yuan, Weiping Wang 0005, Dan Meng 0002
AAAI1
2016 Quantifying the Effect of Sentiment on Topic Evolution in Chinese Microblog
Peng Fu 0008, Zheng Lin 0001, Hailun Lin, Fengcheng Yuan, Weiping Wang 0005, Dan Meng 0002
APWeb (1)1