EDBT 2026 Demo / reviewers in the wild / expert
Zifeng Cheng
dblp:227/6540
· DBLP profile ↗
21ranked-venue papers
8as first author
20since 2021 · last 2026
0000-0002-8486-2614ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 6 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RegionMarker: A Region-Triggered Semantic Watermarking Framework for Embedding-as-a-Service Copyright ProtectionabstractEmbedding-as-a-Service (EaaS) is an effective and convenient deployment solution for addressing various NLP tasks. Nevertheless, recent research has shown that EaaS is vulnerable to model extraction attacks, which could lead to significant economic losses for model providers. For copyright protection, existing methods inject watermark embeddings into text embeddings and use them to detect copyright infringement. However, current watermarking methods often resist only a subset of attacks and fail to provide comprehensive protection. To this end, we present the region-triggered semantic watermarking framework called RegionMarker, which defines trigger regions within a low-dimensional space and injects watermarks into text embeddings associated with these regions. By utilizing a secret dimensionality reduction matrix to project onto this subspace and randomly selecting trigger regions, RegionMarker makes it difficult for watermark removal attacks to evade detection. Furthermore, by embedding watermarks across the entire trigger region and using the text embedding as the watermark, RegionMarker is resilient to both paraphrasing and dimension-perturbation attacks. Extensive experiments on various datasets show that RegionMarker is effective in resisting different attack methods, thereby protecting the copyright of EaaS. Shufan Yang, Zifeng Cheng, Zhiwei Jiang 0001, Yafeng Yin 0002, Cong Wang 0034, Shiping Ge, Yuchen Fu, Qing Gu 0001 |
AAAI | 2 |
| 2026 | Focusing Condition: Inference-Time Self-Contrastive Steering Elicits Better Conditional Text Embeddings in LLMsabstractZifeng Cheng, Lingyun Qian, Zhiwei Jiang, Cong Wang, Yafeng Yin, Fei Shen, Ao Zhou, Qing Gu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zifeng Cheng, Lingyun Qian, Zhiwei Jiang 0001, Cong Wang 0034, Yafeng Yin 0002, Fei Shen 0004, Qing Gu 0001 |
ACL (1) | 1 |
| 2026 | AEA: Adaptive Expert Allocation Improves Sentence Embeddings from Mixture-of-Experts LLMabstractShufan Yang, Zifeng Cheng, Zhiwei Jiang, Qingfeng Qi, Yafeng Yin, Cong Wang, Ao Zhou, Qing Gu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shufan Yang, Zifeng Cheng, Zhiwei Jiang 0001, Qingfeng Qi, Yafeng Yin 0002, Cong Wang 0034, Qing Gu 0001 |
ACL (1) | 2 |
| 2025 | Contrastive Prompting Enhances Sentence Embeddings in LLMs through Inference-Time SteeringabstractZifeng Cheng, Zhonghui Wang, Yuchen Fu, Zhiwei Jiang, Yafeng Yin, Cong Wang, Qing Gu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zifeng Cheng, Zhonghui Wang, Yuchen Fu, Zhiwei Jiang 0001, Yafeng Yin 0002, Cong Wang 0034, Qing Gu 0001 |
ACL (1) | 1 |
| 2025 | Token Prepending: A Training-Free Approach for Eliciting Better Sentence Embeddings from LLMsabstractYuchen Fu, Zifeng Cheng, Zhiwei Jiang, Zhonghui Wang, Yafeng Yin, Zhengliang Li, Qing Gu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yuchen Fu, Zifeng Cheng, Zhiwei Jiang 0001, Zhonghui Wang, Yafeng Yin 0002, Zhengliang Li, Qing Gu 0001 |
ACL (1) | 2 |
| 2025 | Hierarchical Vision-Language Reasoning for Multimodal Multiple-Choice Question AnsweringabstractMultimodal Large Language Models (MLLMs) have demonstrated remarkable multimodal understanding capabilities in Visual Question Answering (VQA) tasks by integrating visual and textual features. However, under the challenging ten-choice question evaluation paradigm, existing methods still exhibit significant limitations when processing PDF documents with complex layouts and lengthy content. Notably, current mainstream models suffer from a strong bias toward English training data, resulting in suboptimal performance for Japanese and other language scenarios. To address these challenges, this paper proposes a novel Japanese PDF document understanding framework that combines multimodal hierarchical reasoning mechanisms with Colqwen-optimized retrieval methods, while innovatively introducing a semantic verification strategy through sub-question decomposition. Experimental results demonstrate that our framework not only significantly enhances the model's deep semantic parsing capability for complex documents, but also exhibits superior robustness in practical application scenarios. Zebo Gu, Tenghao Sun, Mingsheng Tu, Zifeng Cheng, Yafeng Yin 0002, Zhiwei Jiang 0001, Qing Gu 0001 |
ACM Multimedia | 6 |
| 2025 | Cross-Modal Prototype Augmentation and Dual-Grained Prompt Learning for Social Media Popularity PredictionabstractSocial Media Popularity Prediction is a complex multimodal task that requires effective integration of images, text, and structured information. However, current approaches suffer from inadequate visual-textual alignment and fail to capture the inherent cross-content correlations and hierarchical patterns in social media data. To overcome these limitations, we establish a multi-class framework, introducing hierarchical prototypes for structural enhancement and contrastive learning for improved vision-text alignment. Furthermore, we propose a feature-enhanced framework integrating dual-grained prompt learning and cross-modal attention mechanisms, achieving precise multimodal representation through fine-grained category modeling. Experimental results demonstrate state-of-the-art performance on benchmark metrics, establishing new reference standards for multimodal social media analysis. Mingsheng Tu, Luping Wang 0009, Tenghao Sun, Zifeng Cheng, Yafeng Yin 0002, Zhiwei Jiang 0001, Qing Gu 0001 |
ACM Multimedia | 5 |
| 2025 | Steering When Necessary: Flexible Steering Large Language Models with BacktrackingabstractLarge language models (LLMs) have achieved remarkable performance across many generation tasks.
Nevertheless, effectively aligning them with desired behaviors remains a significant challenge.
Activation steering is an effective and cost-efficient approach that directly modifies the activations of LLMs during the inference stage, aligning their responses with the desired behaviors and avoiding the high cost of fine-tuning.
Existing methods typically indiscriminately intervene to all generations or rely solely on the question to determine intervention, which limits the accurate assessment of the intervention strength.
To this end, we propose the **F**lexible **A**ctivation **S**teering with **B**acktracking (**FASB**) framework, which dynamically determines both the necessity and strength of intervention by tracking the internal states of the LLMs during generation, considering both the question and the generated content.
Since intervening after detecting a deviation from the desired behavior is often too late, we further propose the backtracking mechanism to correct the deviated tokens and steer the LLMs toward the desired behavior.
Extensive experiments on the TruthfulQA dataset and six multiple-choice datasets demonstrate that our method outperforms baselines.
Our code will be released at https://github.com/gjw185/FASB. Zifeng Cheng, Jinwei Gan, Zhiwei Jiang 0001, Cong Wang 0034, Yafeng Yin 0002, Yuchen Fu, Qing Gu 0001 |
NeurIPS | 1 |
| 2025 | Advanced Sign Language Video Generation with Compressed and Quantized Multi-Condition TokenizationabstractSign Language Video Generation (SLVG) seeks to generate identity-preserving sign language videos from spoken language texts. Existing methods primarily rely on the single coarse condition (e.g., skeleton sequences) as the intermediary to bridge the translation model and the video generation model, which limits both the naturalness and expressiveness of the generated videos. To overcome these limitations, we propose SignViP, a novel SLVG framework that incorporate multiple fine-grained conditions for improved generation fidelity. Rather than directly translating error-prone high-dimensional conditions, SignViP adopts a discrete tokenization paradigm to integrate and represent fine-grained conditions (i.e., fine-grained poses and 3D hands). SignViP contains three core components. (1) Sign Video Diffusion Model is jointly trained with a multi-condition encoder to learn continuous embeddings that encapsulate fine-grained motion and appearance. (2) Finite Scalar Quantization (FSQ) Autoencoder is further trained to compress and quantize these embeddings into discrete tokens for compact representation of the conditions. (3) Multi-Condition Token Translator is trained to translate spoken language text to discrete multi-condition tokens. During inference, Multi-Condition Token Translator first translates the spoken language text into discrete multi-condition tokens. These tokens are then decoded to continuous embeddings by FSQ Autoencoder, which are subsequently injected into Sign Video Diffusion Model to guide video generation. Experimental results show that SignViP achieves state-of-the-art performance across metrics, including video quality, temporal coherence, and semantic fidelity. The code is available at https://github.com/umnooob/signvip/. Cong Wang 0034, Zexuan Deng, Zhiwei Jiang 0001, Yafeng Yin 0002, Fei Shen 0004, Zifeng Cheng, Shiping Ge, Shiwei Gan, Qing Gu 0001 |
NeurIPS | 6 |
| 2025 | Fine-Grained Alignment Network for Zero-Shot Cross-Modal RetrievalabstractZero-Shot Cross-Modal Retrieval (ZS-CMR) aims to perform cross-modal retrieval on data of unseen classes, where a key challenge is how to address the modality-gap and domain-shift problems simultaneously. Existing methods tackle this challenge mainly by embracing a sample-label alignment paradigm, which aligns samples of different modalities but of the same class with the word embedding of their class label. However, these methods only focus on the class-level alignment and overlook the alignment of rich fine-grained semantic information in samples, incurring coarse understanding of sample matching and poor generalization on unseen classes. In this article, we propose a novel Fine-Grained Alignment Network, an end-to-end framework that learns representation with two fine-grained alignment strategies, yielding representation space that can be better generalized to unseen classes. Specifically, we extract two kinds of fine-grained representations, region embedding and label distribution, respectively, from aspects of both feature and label. To optimize the region embedding, we propose a Fine-Grained Contrastive Learning (FGCL) strategy to simultaneously conduct class-level alignment and model the intra-class discrepancy. To optimize the label distribution, we propose a Fine-Grained Label Alignment (FGLA) strategy to align diverse fine-grained semantic information among samples, rather than merely label information. Finally, both region embedding and label distribution are utilized together to perform ZS-CMR at a finer granularity. Experimental results on three widely used datasets demonstrate that our method outperforms the state-of-the-art methods by a large margin. Detailed ablation studies have also been carried out, which provably affirm the advantage of each component we propose. Our code will be available at https://github.com/ShipingGe/FGAN . Shiping Ge, Zhiwei Jiang 0001, Yafeng Yin 0002, Cong Wang 0034, Zifeng Cheng, Qing Gu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | Unifying Token- and Span-level Supervisions for Few-shot Sequence LabelingabstractFew-shot sequence labeling aims to identify novel classes based on only a few labeled samples. Existing methods solve the data scarcity problem mainly by designing token-level or span-level labeling models based on metric learning. However, these methods are only trained at a single granularity (i.e., either token-level or span-level) and have some weaknesses of the corresponding granularity. In this article, we first unify token- and span-level supervisions and propose a Consistent Dual Adaptive Prototypical (CDAP) network for few-shot sequence labeling. CDAP contains the token- and span-level networks, jointly trained at different granularities. To align the outputs of two networks, we further propose a consistent loss to enable them to learn from each other. During the inference phase, we propose a consistent greedy inference algorithm that first adjusts the predicted probability and then greedily selects non-overlapping spans with maximum probability. Extensive experiments show that our model achieves new state-of-the-art results on three benchmark datasets. All the code and data of this work will be released at https://github.com/zifengcheng/CDAP . Zifeng Cheng, Qingyu Zhou, Zhiwei Jiang 0001, Xuemin Zhao, Yunbo Cao, Qing Gu 0001 |
ACM Trans. Inf. Syst. | 1 |
| 2023 | Controlling Class Layout for Deep Ordinal Classification via Constrained Proxies LearningabstractFor deep ordinal classification, learning a well-structured feature space specific to ordinal classification is helpful to properly capture the ordinal nature among classes. Intuitively, when Euclidean distance metric is used, an ideal ordinal layout in feature space would be that the sample clusters are arranged in class order along a straight line in space. However, enforcing samples to conform to a specific layout in the feature space is a challenging problem. To address this problem, in this paper, we propose a novel Constrained Proxies Learning (CPL) method, which can learn a proxy for each ordinal class and then adjusts the global layout of classes by constraining these proxies. Specifically, we propose two kinds of strategies: hard layout constraint and soft layout constraint. The hard layout constraint is realized by directly controlling the generation of proxies to force them to be placed in a strict linear layout or semicircular layout (i.e., two instantiations of strict ordinal layout). The soft layout constraint is realized by constraining that the proxy layout should always produce unimodal proxy-to-proxies similarity distribution for each proxy (i.e., to be a relaxed ordinal layout). Experiments show that the proposed CPL method outperforms previous deep ordinal classification methods under the same setting of feature extractor. Cong Wang 0034, Zhiwei Jiang 0001, Yafeng Yin 0002, Zifeng Cheng, Shiping Ge, Qing Gu 0001 |
AAAI | 4 |
| 2023 | Improving Domain Generalization for Prompt-Aware Essay Scoring via Disentangled Representation LearningabstractAutomated Essay Scoring (AES) aims to score essays written in response to specific prompts.Many AES models have been proposed, but most of them are either prompt-specific or prompt-adaptive and cannot generalize well on "unseen" prompts.This work focuses on improving the generalization ability of AES models from the perspective of domain generalization, where the data of target prompts cannot be accessed during training.Specifically, we propose a prompt-aware neural AES model to extract comprehensive representation for essay scoring, including both promptinvariant and prompt-specific features.To improve the generalization of representation, we further propose a novel disentangled representation learning framework.In this framework, a contrastive norm-angular alignment strategy and a counterfactual self-training strategy are designed to disentangle the prompt-invariant information and prompt-specific information in representation.Extensive experimental results on datasets of both ASAP and TOEFL11 demonstrate the effectiveness of our method under the domain generalization setting. Zhiwei Jiang 0001, Yafeng Yin 0002, Zifeng Cheng, Qing Gu 0001 |
ACL (1) | 6 |
| 2023 | Aggregating Multiple Heuristic Signals as Supervision for Unsupervised Automated Essay ScoringabstractAutomated Essay Scoring (AES) aims to evaluate the quality score for input essays.In this work, we propose a novel unsupervised AES approach ULRA, which does not require groundtruth scores of essays for training.The core idea of our ULRA is to use multiple heuristic quality signals as the pseudo-groundtruth, and then train a neural AES model by learning from the aggregation of these quality signals.To aggregate these inconsistent quality signals into a unified supervision, we view the AES task as a ranking problem, and design a special Deep Pairwise Rank Aggregation (DPRA) loss for training.In the DPRA loss, we set a learnable confidence weight for each signal to address the conflicts among signals, and train the neural AES model in a pairwise way to disentangle the cascade effect among partialorder pairs.Experiments on eight prompts of ASPA dataset show that ULRA achieves the state-of-the-art performance compared with previous unsupervised methods in terms of both transductive and inductive settings.Further, our approach achieves comparable performance with many existing domain-adapted supervised models, showing the effectiveness of ULRA.The code is available at https: //github.com/tenvence/ulra. Cong Wang 0034, Zhiwei Jiang 0001, Yafeng Yin 0002, Zifeng Cheng, Shiping Ge, Qing Gu 0001 |
ACL (1) | 4 |
| 2023 | Learning Event-Specific Localization Preferences for Audio-Visual Event LocalizationabstractAudio-Visual Event Localization (AVEL) aims to locate events that are both visible and audible in a video. Existing AVEL methods primarily focus on learning generic localization patterns that are applicable to all events. However, events often exhibit modality biases, such as visual-dominated, audio-dominated, or modality-balanced, which can lead to different localization preferences. These preferences may be overlooked by existing methods, resulting in unsatisfactory localization performance. To address this issue, this paper proposes a novel event-aware localization paradigm, which first identifies the event category and then leverages localization preferences specific to that event for improved event localization. To achieve this, we introduce a memory-assisted metric learning framework, which utilizes historic segments as anchors to adjust the unified representation space for both event classification and event localization. To provide sufficient information for this metric learning, we design a spatial-temporal audio-visual fusion encoder to capture the spatial and temporal interaction between audio and visual modalities. Extensive experiments on the public AVE dataset in both fully-supervised and weakly-supervised settings demonstrate the effectiveness of our approach. Code will be released at https://github.com/ShipingGe/AVEL. Shiping Ge, Zhiwei Jiang 0001, Yafeng Yin 0002, Cong Wang 0034, Zifeng Cheng, Qing Gu 0001 |
ACM Multimedia | 5 |
| 2023 | Learning Robust Multi-Modal Representation for Multi-Label Emotion Recognition via Adversarial Masking and PerturbationabstractRecognizing emotions from multi-modal data is an emotion recognition task that requires strong multi-modal representation ability. The general approach to this task is to naturally train the representation model on training data without intervention. However, such natural training scheme is prone to modality bias of representation (i.e., tending to over-encode some informative modalities while neglecting other modalities) and data bias of training (i.e., tending to overfit training data). These biases may lead to instability (e.g., performing poorly when the neglected modality is dominant for recognition) and weak generalization (e.g., performing poorly when unseen data is inconsistent with overfitted data) of the model on unseen data. To address these problems, this paper presents two adversarial training strategies to learn more robust multi-modal representation for multi-label emotion recognition. Firstly, we propose an adversarial temporal masking strategy, which can enhance the encoding of other modalities by masking the most emotion-related temporal units (e.g., words for text or frames for video) of the informative modality. Secondly, we propose an adversarial parameter perturbation strategy, which can enhance the generalization of the model by adding the adversarial perturbation to the parameters of model. Both strategies boost model performance on the benchmark MMER datasets CMU-MOSEI and NEMu. Experimental results demonstrate the effectiveness of the proposed method compared with the previous state-of-the-art method. Code will be released at https://github.com/ShipingGe/MMER. Shiping Ge, Zhiwei Jiang 0001, Zifeng Cheng, Cong Wang 0034, Yafeng Yin 0002, Qing Gu 0001 |
WWW | 3 |
| 2023 | A Consistent Dual-MRC Framework for Emotion-cause Pair ExtractionabstractEmotion-cause pair extraction (ECPE) is a recently proposed task that aims to extract the potential clause pairs of emotions and its corresponding causes in a document. In this article, we propose a new paradigm for the ECPE task. We cast the task as a two-turn machine reading comprehension (MRC) task, i.e., the extraction of emotions and causes is transformed to the task of identifying answer clauses from the input document specific to a query. This two-turn MRC formalization brings several key advantages: First, the QA manner provides an explicit pairing way to identify causes specific to the target emotion; second, it provides a natural way of jointly modeling the emotion extraction, the cause extraction, and the pairing of emotion and cause; and third, it allows us to exploit the well-developed MRC models. Based on the two-turn MRC formalization, we propose a dual-MRC framework to extract emotion-cause pairs in a dual-direction way, which enables a more comprehensive coverage of all pairing cases. Furthermore, we propose a consistent training strategy for the second-turn query, so the model is able to filter the errors produced by the first turn at inference. Experiments on two benchmark datasets demonstrate that our method outperforms previous methods and achieves state-of-the-art performance. All the code and data of this work can be obtained at https://github.com/zifengcheng/CD-MRC . Zifeng Cheng, Zhiwei Jiang 0001, Yafeng Yin 0002, Cong Wang 0034, Shiping Ge, Qing Gu 0001 |
ACM Trans. Inf. Syst. | 1 |
| 2022 | Learning to Classify Open Intent via Soft Labeling and Manifold MixupabstractOpen intent classification is a practical yet challenging task in dialogue systems. Its objective is to accurately classify samples of known intents while at the same time detecting those of open (unknown) intents. Existing methods usually use outlier detection algorithms combined with K-class classifier to detect open intents, where K represents the class number of known intents. Different from them, in this paper, we consider another way without using outlier detection algorithms. Specifically, we directly train a (K+1)-class classifier for open intent classification, where the (K+1)-th class represents open intents. To address the challenge that training a (K+1)-class classifier with training samples of only K classes, we propose a deep model based on Soft Labeling and Manifold Mixup (SLMM). In our method, soft labeling is used to reshape the label distribution of the known intent samples, aiming at reducing model’s overconfident on known intents. Manifold mixup is used to generate pseudo samples for open intents, aiming at well optimizing the decision boundary of open intents. Experiments on four benchmark datasets demonstrate that our method outperforms previous methods and achieves state-of-the-art performance. All the code and data of this work can be obtained at.1 Zifeng Cheng, Zhiwei Jiang 0001, Yafeng Yin 0002, Cong Wang 0034, Qing Gu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2021 | Learning from Graph Propagation via Ordinal Distillation for One-Shot Automated Essay ScoringabstractOne-shot automated essay scoring (AES) aims to assign scores to a set of essays written specific to a certain prompt, with only one manually scored essay per distinct score. Compared to the previous-studied prompt-specific AES which usually requires a large number of manually scored essays for model training (e.g., about 600 manually scored essays out of totally 1000 essays), one-shot AES can greatly reduce the workload of manual scoring. In this paper, we propose a Transductive Graph-based Ordinal Distillation (TGOD) framework to tackle the task of one-shot AES. Specifically, we design a transductive graph-based model as a teacher model to generate pseudo labels of unlabeled essays based on the one-shot labeled essays. Then, we distill the knowledge in the teacher model into a neural student model by learning from the high confidence pseudo labels. Different from the general knowledge distillation, we propose an ordinal-aware unimodal distillation which makes a unimodal distribution constraint on the output of student model, to tolerate the minor errors existed in pseudo labels. Experimental results on the public dataset ASAP show that TGOD can improve the performance of existing neural AES models under the one-shot AES setting and achieve an acceptable average QWK of 0.69. Zhiwei Jiang 0001, Yafeng Yin 0002, Zifeng Cheng, Qing Gu 0001 |
WWW | 5 |
| 2021 | A Unified Target-Oriented Sequence-to-Sequence Model for Emotion-Cause Pair ExtractionabstractEmotion-cause pair extraction is a recently proposed task that aims at extracting all potential clause-level pairs of emotion and cause in text. To solve this task, researchers first proposed a two-step pipeline method. This method extracts the emotions and causes individually in the first step, then pairs the extracted emotions and causes and filters the invalid emotion-cause pairs in the second step. Due to that the two-step method has the error accumulation problem and is hard to be optimized jointly, several one-step end-to-end models have been proposed. These models share a similar underlying idea, that is, reframing the emotion-cause pair extraction task as a classification problem of candidate clause pairs. Unlike these models, in this paper, we reframe the emotion-cause pair extraction task as a unified sequence labeling problem, which allows to extract emotion-cause pairs through one pass of sequence labeling. This is realized by designing a special set of unified labels. In the unified label, we design a content part for emotion/cause identification and a pairing part for clause pairing. Then the emotion-cause pairs can be implicitly derived from the unified labels. To address this unified sequence labeling problem, we propose a unified target-oriented sequence-to-sequence model, which comprehensively utilizes the information of target clause, global context, and former decoded label, to perform end-to-end unified sequence labeling. The experimental results demonstrate the effectiveness of both our proposed unified sequence labeling scheme and unified target-oriented sequence-to-sequence model. All the code and data of this work can be obtained at https://github.com/zifengcheng/UTOS. Zifeng Cheng, Zhiwei Jiang 0001, Yafeng Yin 0002, Na Li 0018, Qing Gu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2020 | A Symmetric Local Search Network for Emotion-Cause Pair ExtractionabstractEmotion-cause pair extraction (ECPE) is a new task which aims at extracting the potential clause pairs of emotions and corresponding causes in a document.To tackle this task, a two-step method was proposed by previous study which first extracted emotion clauses and cause clauses individually, then paired the emotion and cause clauses, and filtered out the pairs without causality.Different from this method that separated the detection and the matching of emotion and cause into two steps, we propose a Symmetric Local Search Network (SLSN) model to perform the detection and matching simultaneously by local search.SLSN consists of two symmetric subnetworks, namely the emotion subnetwork and the cause subnetwork.Each subnetwork is composed of a clause representation learner and a local pair searcher.The local pair searcher is a specially-designed cross-subnetwork component which can extract the local emotion-cause pairs.Experimental results on the ECPE corpus demonstrate the superiority of our SLSN over existing state-of-the-art methods. Zifeng Cheng, Zhiwei Jiang 0001, Yafeng Yin 0002, Qing Gu 0001 |
COLING | 1 |