EDBT 2026 Demo / reviewers in the wild / expert
Kaiwen Wei
dblp:297/8721
· DBLP profile ↗
33ranked-venue papers
10as first author
33since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 8 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MIRAGE: Scaling Test-Time Inference with Parallel Graph-Retrieval-Augmented Reasoning ChainsabstractLarge reasoning models (LRMs) have shown significant progress in test-time scaling through chain-of-thought prompting. Current approaches like search-o1 integrate retrieval augmented generation (RAG) into multi-step reasoning processes but rely on a single, linear reasoning path while incorporating unstructured textual information in a flat, context-agnostic manner. As a result, these approaches can lead to error accumulation throughout the reasoning chain, which significantly limits its effectiveness in medical question-answering (QA) tasks where both accuracy and traceability are critical requirements. To address these challenges, we propose MIRAGE (Multi-path Inference with Retrieval-Augmented Graph Exploration), a novel test-time scalable reasoning framework that performs dynamic multi-path inference over structured medical knowledge graphs. Specifically, MIRAGE 1) decomposes complex queries into entity-grounded sub-questions, 2) executes parallel inference paths, 3) retrieves evidence adaptively via neighbor expansion and multi-hop traversal, and 4) integrates answers using cross-path verification to resolve contradictions. Experiments on three medical QA benchmarks (GenMedGPT-5k, CMCQA, and ExplainCPE) show that MIRAGE consistently outperforms GPT-4o, Tree-of-Thought variants, and other retrieval-augmented baselines in both automatic and human evaluations. Additionally, MIRAGE improves interpretability by generating explicit reasoning chains that trace each factual claim to concrete paths within the knowledge graph, making it especially suitable for complex medical reasoning scenarios. Kaiwen Wei, Rui Shan, Dongsheng Zou, Jianzhong Yang, Bi Zhao, Junnan Zhu |
AAAI | 1 |
| 2026 | MentalSeek-Dx: Towards Progressive Hypothetico-Deductive Reasoning for Real-world Psychiatric DiagnosisabstractXiao Sun, Ymyang, Xinyi Jiang, Yu Tian, Junnan Zhu, Jiang Zhong, Qin Lei, Jingwang Huang, Haoyang Zeng, Xinyu Zhou, Xin Xiao, Kaiwen Wei. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Junnan Zhu, Qin Lei, Jingwang Huang, Kaiwen Wei |
ACL (1) | 12 |
| 2026 | From Past To Path: Masked History Learning for Next-Item Prediction in Generative RecommendationabstractKaiwen Wei, Kejun he, Xiaomian Kang, Jie Zhang, Ymyang, Li Jin, Zhenyang Li, Jiang Zhong, Richard He Bai, Junnan Zhu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Kaiwen Wei, Kejun He, Xiaomian Kang, Ymyang, He Bai 0002, Junnan Zhu |
ACL (1) | 1 |
| 2026 | Chart-MRAG: Benchmarking Multimodal Retrieval Augmented Generation on Chart-based DocumentsabstractYmyang, Jiang Zhong, Li Jin, Xiao Sun, Jingwang Huang, Gaojinpeng, Qing Liu, Yang Bai, Jingyuan Zhang, Rui Jiang, Qin Lei, Kaiwen Wei. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jingwang Huang, Jinpeng Gao, Qin Lei, Kaiwen Wei |
ACL (1) | 12 |
| 2026 | CFVBench: A Comprehensive Video Benchmark for Fine-grained Multimodal Retrieval-Augmented GenerationabstractMultimodal Retrieval-Augmented Generation (MRAG) enables Multimodal Large Language Models (MLLMs) to generate responses with external multimodal evidence. Recently, numerous video-based MRAG benchmarks have been proposed to evaluate model capabilities across retrieval and generation stages in MRAG. However, existing benchmarks remain limited in modality coverage and format diversity, often focusing on single- or limited-modality tasks, or coarse-grained scene understanding. To address these gaps, we introduce CFVBench, a large-scale, manually verified benchmark constructed from 599 publicly available videos, yielding 5,360 open-ended QA pairs. CFVBench, spans high-density formats and domains such as chart-heavy reports, news broadcasts, and software tutorials, requiring models to retrieve and reason over long temporal video spans while maintaining fine-grained multimodal information. Using CFVBench, we systematically evaluate 7 retrieval methods and 14 widely-used MLLMs, revealing a critical bottleneck: current models (even GPT5 or Gemini) struggle to capture transient yet essential fine-grained multimodal details. To mitigate this, we propose Adaptive Visual Refinement (AVR), a plug-and-paly framework that adaptively increases frame sampling density and selectively invokes external tools when necessary. Experiments show that AVR consistently enhances fine-grained multimodal comprehension and improves performance across all evaluated MLLMs. Kaiwen Wei, Ruida Liu, Changzai Pan, Yidan Zhang 0002, Peijin Wang, Yingchao Feng |
WWW | 1 |
| 2026 | Multi-grained dynamic feature resorter with curriculum bootstrapping for multimodal entity and relation extraction
Zhicong Lu, Li Jin 0001, Linhao Zhang, Kaiwen Wei, Qing Liu 0021, Yanfeng Hu |
Neurocomputing | 5 |
| 2026 | Noise-aware temporal knowledge graph reasoning with query-guided learning and confidence-aware optimization
Longquan Liao, Linjiang Zheng, Jiaxing Shang, Xu Li 0014, Kaiwen Wei |
Knowl. Based Syst. | 6 |
| 2026 | Context-Aware Learning and Pattern Decomposition for Temporal Knowledge Graph ReasoningabstractGraph neural network (GNN)-based approaches have achieved remarkable success in temporal knowledge graph (TKG) reasoning. Despite these advances, two critical challenges remain: 1) inadequate modeling of local contextual dynamics, which limits the adaptability of entity and relation representations to specific queries and 2) inadequate mechanisms for handling emerging patterns, that is, novel interactions absent from historical data, which reduces predictive performance in dynamic environments. To address these limitations, we propose TCDR-PD, a temporal and contextual dynamic representation network with pattern decomposition. TCDR-PD introduces a temporal and contextual dynamic representation learning (TCDR) module to capture both global temporal trends and query-specific contextual dynamics, enabling more precise embeddings. Additionally, the pattern decomposition (PD) prediction module explicitly disentangles the prediction of recurring and emerging patterns, enabling tailored strategies to improve reasoning performance. Experiments on four benchmark datasets demonstrate that TCDR-PD outperforms state-of-the-art methods, effectively supporting stable reasoning over evolving TKGs. Longquan Liao, Linjiang Zheng, Jiaxing Shang, Xu Li 0014, Kaiwen Wei |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | SARA: Salience-Aware Reinforced Adaptive Decoding for Large Language Models in Abstractive SummarizationabstractNayu Liu, Junnan Zhu, Yiming Ma, Zhicong Lu, Wenlei Xu, Yong Yang, Jiang Zhong, Kaiwen Wei. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Nayu Liu, Junnan Zhu, Zhicong Lu, Wenlei Xu, Yong Yang 0001, Kaiwen Wei |
ACL (1) | 8 |
| 2025 | Chain-of-Specificity: Enhancing Task-Specific Constraint Adherence in Large Language ModelsabstractLarge Language Models (LLMs) exhibit remarkable generative capabilities, enabling the generation of valuable information. Despite these advancements, previous research found that LLMs sometimes struggle with adhering to specific constraints, such as being in a specific place or at a specific time, and at times even overlook them, which leads to responses that are either too generic or not fully satisfactory. Existing approaches attempted to address this issue by decomposing and rewriting input instructions or reflecting on prior failings, yet they fall short in adequately emphasizing specific constraints and unlocking the underlying knowledge, such as programming within the context of software development. In response, this paper proposes a simple yet effective method called Chain-of-Specificity (CoS). Specifically, CoS emphasizes the specific constraints in the input instructions, unlocks knowledge within LLMs, and refines responses. Experiments conducted on publicly available and self-built complex datasets demonstrate that CoS outperforms existing methods in enhancing generated content, especially in terms of specificity. Additionally, as the number of specific constraints increases, other baselines falter, while CoS still performs well. Moreover, we show that distilling responses generated by CoS effectively enhances the ability of smaller models to follow constrained instructions. Kaiwen Wei, Li Jin 0001 |
COLING | 1 |
| 2025 | T2R-BENCH: A Benchmark for Real World Table-to-Report TaskabstractJie Zhang, Changzai Pan, Sishi Xiong, Kaiwen Wei, Yu Zhao, Xiangyu Li, Jiaxin Peng, Xiaoyan Gu, Jian Yang, Wenhan Chang, Zhenhe Wu, Jiang Zhong, Shuangyong Song, Xuelong Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Changzai Pan, Sishi Xiong, Kaiwen Wei, Yu Zhao 0007, Jian Yang 0037, Wenhan Chang, Zhenhe Wu, Shuangyong Song, Xuelong Li 0001 |
EMNLP | 4 |
| 2025 | EEG Decoding and Visual Reconstruction via 3D Geometric with Nonstationarity ModellingabstractElectroencephalogram (EEG) signal processing has advanced in revealing the mechanisms of human visual perception, but existing methods often overlook two key EEG properties: (1) 3D geometric relationships between EEG electrodes, which reflects the ability to model the brain in stereoscopic terms; and (2) nonstationarity of EEG signals, which involves capturing the dynamic changes in the frequency spectrum. To address these limitations, we introduce the GeoCap framework in this paper. Specifically, to effectively model the 3D geometry of EEG electrodes, we propose spherical manifold encoding (SME), which represents EEG channels on a 3D spherical manifold. Additionally, drawing inspiration from the capacity of cumulative fractional derivatives to incorporate historical context, we introduce the discrete multiscale Caputo derivative (DMSCD) to more accurately capture the temporal dynamics of EEG signals across multiple scales. Experiments on EEG decoding and visual reconstruction tasks demonstrate that GeoCap outperforms other state-of-the-art methods. Kaiwen Wei, Xuekai Wei, Jielu Yan |
ICASSP | 2 |
| 2025 | U-MERE: Unconstrained Multimodal Entity and Relation Extraction with Collaborative Modeling and Order-Sensitive OptimizationabstractExisting multimodal entity and relation extraction tasks primarily focus on text-to-text or text-to-visual entity relations, overlooking real-world complexities involving visual-to-text and visual-to-visual cases, thus failing to capture the richer semantic structures in complex cross-modal interactions. To address the limitations, we propose a new task, Unconstrained Multimodal Entity and Relation Extraction (U-MERE), which jointly extracts arbitrary visual and textual entities, and their relations from image-text pairs. To accomplish U-MERE, we construct UMERE-Bench, a benchmark with over 9,000 samples that comprehensively covers four cross-modal entity relation directions and three task settings. Given the difficulty of jointly modeling diverse directions of cross-modal entity relations, we introduce Collaborative Modeling and Order-Sensitive (CMOS), which collaboratively guides large vision-language models (LVLMs) to decompose task complexity and mitigates generation order bias from fixed target relation sequences. CMOS employs small models to generate candidate entities, guiding LVLMs to capture key information and jointly optimizes multiple feasible relation orderings to reduce order dependency. Additionally, we design a Multimodal Order-aware Matching (MOM) evaluation method to align predictions with ground truth for precise assessment. Experimental results reveal that current LVLMs show limited performance on U-MERE, underscoring its inherent challenges, while CMOS consistently achieves superior performance across multiple advanced LVLMs, demonstrating its effectiveness and generalization capability. The dataset and code will be available in https://github.com/jiaweidoris/U-MERE. Li Jin 0001, Kaiwen Wei, Yuying Shang, Nayu Liu, Zhicong Lu, Qing Liu 0021, Linhao Zhang, Yanfeng Hu |
ACM Multimedia | 3 |
| 2025 | RAIN: Reconstructed-aware in-context enhancement with graph denoising for session-based recommendation
Xinyi Zeng, Shuchao Li, Zequn Zhang, Li Jin 0001, Zhi Guo, Kaiwen Wei |
Neural Networks | 6 |
| 2024 | Multimodal Event Causality Reasoning with Scene Graph Enhanced Interaction NetworkabstractMultimodal event causality reasoning aims to recognize the causal relations based on the given events and accompanying image pairs, requiring the model to have a comprehensive grasp of visual and textual information. However, existing studies fail to effectively model the relations of the objects within the image and capture the object interactions across the image pair, resulting in an insufficient understanding of visual information by the model. To address these issues, we propose a Scene Graph Enhanced Interaction Network (SEIN) in this paper, which can leverage the interactions of the generated scene graph for multimodal event causality reasoning. Specifically, the proposed method adopts a graph convolutional network to model the objects and their relations derived from the scene graph structure, empowering the model to exploit the rich structural and semantic information in the image adequately. To capture the object interactions between the two images, we design an optimal transport-based alignment strategy to match the objects across the images, which could help the model recognize changes in visual information and facilitate causality reasoning. In addition, we introduce a cross-modal fusion module to combine textual and visual features for causality prediction. Experimental results indicate that the proposed SEIN outperforms state-of-the-art methods on the Vis-Causal dataset. Kaiwen Wei |
AAAI | 2 |
| 2024 | Video Event Extraction with Multi-View Interaction Knowledge DistillationabstractVideo event extraction (VEE) aims to extract key events and generate the event arguments for their semantic roles from the video. Despite promising results have been achieved by existing methods, they still lack an elaborate learning strategy to adequately consider: (1) inter-object interaction, which reflects the relation between objects; (2) inter-modality interaction, which aligns the features from text and video modality. In this paper, we propose a Multi-view Interaction with knowledge Distillation (MID) framework to solve the above problems with the Knowledge Distillation (KD) mechanism. Specifically, we propose the self-Relational KD (self-RKD) to enhance the inter-object interaction, where the relation between objects is measured by distance metric, and the high-level relational knowledge from the deeper layer is taken as the guidance for boosting the shallow layer in the video encoder. Meanwhile, to improve the inter-modality interaction, the Layer-to-layer KD (LKD) is proposed, which integrates additional cross-modal supervisions (i.e., the results of cross-attention) with the textual supervising signal for training each transformer decoder layer. Extensive experiments show that without any additional parameters, MID achieves the state-of-the-art performance compared to other strong methods in VEE. Kaiwen Wei, Runyan Du, Li Jin 0001, Jian Liu 0032, Jianhua Yin 0001, Linhao Zhang, Nayu Liu, Zhi Guo |
AAAI | 1 |
| 2024 | CAMEL: Capturing Metaphorical Alignment with Context Disentangling for Multimodal Emotion RecognitionabstractUnderstanding the emotional polarity of multimodal content with metaphorical characteristics, such as memes, poses a significant challenge in Multimodal Emotion Recognition (MER). Previous MER researches have overlooked the phenomenon of metaphorical alignment in multimedia content, which involves non-literal associations between concepts to convey implicit emotional tones. Metaphor-agnostic MER methods may be misinformed by the isolated unimodal emotions, which are distinct from the real emotions blended in multimodal metaphors. Moreover, contextual semantics can further affect the emotions associated with similar metaphors, leading to the challenge of maintaining contextual compatibility. To address the issue of metaphorical alignment in MER, we propose to leverage a conditional generative approach for capturing metaphorical analogies. Our approach formulates schematic prompts and corresponding references based on theoretical foundations, which allows the model to better grasp metaphorical nuances. In order to maintain contextual sensitivity, we incorporate a disentangled contrastive matching mechanism, which undergoes curricular adjustment to regulate its intensity during the learning process. The automatic and human evaluation experiments on two benchmarks prove that, our model provides considerable and stable improvements in recognizing multimodal emotion with metaphor attributes. Linhao Zhang, Li Jin 0001, Guangluan Xu, Xiaoyu Li 0004, Kaiwen Wei, Nayu Liu |
AAAI | 6 |
| 2024 | TAeKD: Teacher Assistant Enhanced Knowledge Distillation for Closed-Source Multilingual Neural Machine TranslationabstractKnowledge Distillation (KD) serves as an efficient method for transferring language knowledge from open-source large language models (LLMs) to more computationally efficient models. However, challenges arise when attempting to apply vanilla KD methods to transfer knowledge from closed-source Multilingual Neural Machine Translation (MNMT) models based on LLMs. In this scenario, the soft labels and training data are not accessible, making it difficult to achieve effective knowledge transfer. To address this issue, this paper proposes a Teacher Assistant enhanced Knowledge Distillation (TAeKD) method to augment the knowledge transfer capacity from closed-source MNMT models. Specifically, TAeKD designs a fusion model that integrates translation outputs from multiple closed-source models to generate soft labels and training samples. Furthermore, a quality assessment learning mechanism is introduced to enhance the generalization of the fusion model and elevate the quality of the fusion data used to train the student model. To facilitate research on knowledge transfer from MNMT models, we also introduce FuseData, a benchmark consisting of a blend of translations from multiple closed-source systems. The experimental results show that TAeKD outperforms the previous state-of-the-art KD methods on both WMT22 and FLORES-101 test sets. Kaiwen Wei |
LREC/COLING | 3 |
| 2024 | Multi-View Prompt for Fine-Grained Multimodal Named Entity Recognition and GroundingabstractFine-Grained Multimodal Named Entity Recognition and Grounding (FMNERG) aims to extract entity name, fine-grained entity type, and its corresponding object from paired text and image. This task demands fundamental reasoning capability for complex language and multimodal comprehension. Despite encouraging results, existing methods face two critical issues: (1) Insufficient knowledge of the entity poses challenges to fine-grained entity recognition; (2) Limited correlations between entities and objects hinder the visual grounding of entities. To tackle these issues, we propose a Multi-View Prompt (MVP) method for the FMNERG task in this paper, which collaborates with Large Language Models (LLMs) and Visual Grounding Models (VGMs) for reasoning. Concretely, MVP constructs a knowledgeable prompt in a chain-of-thought format, progressively refining possible entity types from coarse-grained to fine-grained levels. It leverages a heuristic method to select demonstration examples, which could provide guiding knowledge about entities from LLMs. To establish correlations between entities and potential objects, MVP introduces a grounded prompt that exploits information from guiding knowledge and image caption, enabling VGMs to detect related objects. Experimental results indicate that MVP achieves state-of-the-art performance on the Twitter dataset. Kaiwen Wei |
ECAI | 3 |
| 2024 | GOME: Grounding-based Metaphor Binding With Conceptual Elaboration For Figurative Language IllustrationabstractThe illustration or visualization of figurative language, such as linguistic metaphors, is an emerging challenge for existing Large Language Models (LLMs) and multimodal models. Linhao Zhang, Li Jin 0001, Kaiwen Wei, Guangluan Xu |
EMNLP | 5 |
| 2024 | S2D: Enhancing Zero-Shot Cross-Lingual Event Argument Extraction with Semantic Knowledge
Zongkai Zhao, Xiuhua Li 0001, Kaiwen Wei |
NLPCC (1) | 3 |
| 2024 | DuaPIN: Auxiliary task enhanced dual path interaction network for civil court view generation
Nayu Liu, Yiquan Wu 0001, Kaiwen Wei, Cunhang Fan |
Knowl. Based Syst. | 4 |
| 2024 | Multimodal Cross-Lingual Summarization for Videos: A Revisit in Knowledge Distillation Induced Triple-Stage Training MethodabstractMultimodal summarization (MS) for videos aims to generate summaries from multi-source information (e.g., video and text transcript), showing promising progress recently. However, existing works are limited to monolingual scenarios, neglecting non-native viewers' needs to understand videos in other languages. It stimulates us to introduce multimodal cross-lingual summarization for videos (MCLS), which aims to generate cross-lingual summaries from multimodal input of videos. Considering the challenge of high annotation cost and resource constraints in MCLS, we propose a knowledge distillation (KD) induced triple-stage training method to assist MCLS by transferring knowledge from abundant monolingual MS data to those data with insufficient volumes. In the triple-stage training method, a video-guided dual fusion network (VDF) is designed as the backbone network to integrate multimodal and cross-lingual information through diverse fusion strategies in the encoder and decoder; What's more, we propose two cross-lingual knowledge distillation strategies: adaptive pooling distillation and language-adaptive warping distillation (LAWD), designed for encoder-level and vocab-level distillation objects to facilitate effective knowledge transfer across cross-lingual sequences of varying lengths between MS and MCLS models. Specifically, to tackle lingual sequences of varying lengths between MS and MCLS models. Specifically, to tackle the challenge of unequal length of parallel cross-language sequences in KD, LAWD can directly conduct cross-language distillation while keeping the language feature shape unchanged to reduce potential information loss. We meticulously annotated the How2-MCLS dataset based on the How2 dataset to simulate MCLS scenarios. Experimental results show that the proposed method achieves competitive performance compared to strong baselines, and can bring substantial performance improvements to MCLS models by transferring knowledge from the MS model. Nayu Liu, Kaiwen Wei, Yong Yang 0001, Jianhua Tao 0001, Xian Sun 0001, Fanglong Yao, Li Jin 0001, Zhao Lv, Cunhang Fan |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | More Than Syntaxes: Investigating Semantics to Zero-shot Cross-lingual Relation Extraction and Event Argument Role LabellingabstractSyntactic dependency structures are commonly utilized as language-agnostic features to solve the word order difference issues in zero-shot cross-lingual relation and event extraction tasks. However, while sentences in multiple forms can be employed to express the same meaning, the syntactic structure may vary considerably in specific scenarios. To fix this problem, we find semantics are rarely considered, which could provide a more consistent semantic analysis of sentences and be served as another bridge between different languages. Therefore, in this article, we introduce Syntax and Semantic Driven Network (SSDN) to equip syntax and semantic knowledge across languages simultaneously. Specifically, predicate–argument structures from semantic role labelling are explicitly incorporated into word representations. Then, a semantic-aware relational graph convolutional network and a transformer-based encoder are utilized to model both semantic dependency and syntactic dependency structures, respectively. Finally, a fusion module is introduced to integrate output representations adaptively. We conduct experiments on the widely used Automatic Content Extraction 2005 English, Chinese, and Arabic datasets. The evaluation results demonstrate that the proposed method achieves the state-of-the-art performance. Further study also indicates SSDN could produce robust representations that facilitate the transfer operations across languages. Kaiwen Wei, Li Jin 0001, Zequn Zhang, Zhi Guo, Xiaoyu Li 0004, Qing Liu 0021, Weimiao Feng |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2023 | Guide the Many-to-One Assignment: Open Information Extraction via IoU-aware Optimal TransportabstractKaiwen Wei, Yiran Yang, Li Jin, Xian Sun, Zequn Zhang, Jingyuan Zhang, Xiao Li, Linhao Zhang, Jintao Liu, Guo Zhi. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Kaiwen Wei, Li Jin 0001, Xian Sun 0001, Zequn Zhang, Linhao Zhang, Zhi Guo |
ACL (1) | 1 |
| 2023 | 1% VS 100%: Parameter-Efficient Low Rank Adapter for Dense PredictionsabstractFine-tuning large-scale pretrained vision models to downstream tasks is a standard technique for achieving state-of-the-art performance on computer vision benchmarks. However, fine-tuning the whole model with millions of parameters is inefficient as it requires storing a same-sized new model copy for each task. In this work, we propose LoRand, a method for fine-tuning large-scale vision models with a better tradeoff between task performance and the number of trainable parameters. LoRand generates tiny adapter structures with low-rank synthesis while keeping the original backbone parameters fixed, resulting in high parameter sharing. To demonstrate LoRand's effectiveness, we implement extensive experiments on object detection, semantic segmentation, and instance segmentation tasks. By only training a small percentage (1% to 3%) of the pretrained backbone parameters, LoRand achieves comparable performance to standard fine-tuning on COCO and ADE20K and outperforms fine-tuning in low-resource PASCAL VOC dataset. Dongshuo Yin, Zhechao Wang, Kaiwen Wei, Xian Sun 0001 |
CVPR | 5 |
| 2023 | Event Causality Extraction via Implicit Cause-Effect InteractionsabstractEvent Causality Extraction (ECE) aims to extract the cause-effect event pairs from the given text, which requires the model to possess a strong reasoning ability to capture event causalities.However, existing works have not adequately exploited the interactions between the cause and effect event that could provide crucial clues for causality reasoning.To this end, we propose an Implicit Cause-Effect interaction (ICE) framework, which formulates ECE as a template-based conditional generation problem.The proposed method captures the implicit intra-and inter-event interactions by incorporating the privileged information (ground truth event types and arguments) for reasoning, and a knowledge distillation mechanism is introduced to alleviate the unavailability of privileged information in the test stage.Furthermore, to facilitate knowledge transfer from teacher to student, we design an event-level alignment strategy named Cause-Effect Optimal Transport (CEOT) to strengthen the semantic interactions of cause-effect event types and arguments.Experimental results indicate that ICE achieves state-of-the-art performance on the ECE-CCKS dataset. Zequn Zhang, Kaiwen Wei, Zhi Guo, Xian Sun 0001, Li Jin 0001, Xiaoyu Li 0004 |
EMNLP | 3 |
| 2023 | Emotion-cause pair extraction with bidirectional multi-label sequence tagging
Zequn Zhang, Zhi Guo, Li Jin 0001, Xiaoyu Li 0004, Kaiwen Wei, Xian Sun 0001 |
Appl. Intell. | 6 |
| 2023 | KEPT: Knowledge Enhanced Prompt Tuning for event causality identification
Zequn Zhang, Zhi Guo, Li Jin 0001, Xiaoyu Li 0004, Kaiwen Wei, Xian Sun 0001 |
Knowl. Based Syst. | 6 |
| 2023 | Implicit Event Argument Extraction With Argument-Argument Relational KnowledgeabstractAs a challenging sub-task of event argument extraction, implicit event argument extraction seeks to identify document-level arguments that play direct or implicit roles in a given event. Prior work mainly focuses on capturing direct relations between arguments and the event trigger; however, the lack of reasoning ability imposes limitations to the extraction of implicit arguments. In this work, we propose anArgument-argumentRelation-enhancedEventArgument extraction (AREA) learning framework to tackle this issue through reasoning in event frame-level scope. The proposed method leverages related arguments of the expected one as clues, and utilizes such argument-argument dependencies to guide the reasoning process. To bridge the distribution gap between oracle knowledge used in the training phase and the imperfect related arguments in the test stage, we introduce a conventional knowledge distillation strategy to drive a final model that can work without extra inputs by mimicking the behaviour of a well-informed teacher model. In addition, considering that conventional knowledge distillation methods transfer knowledge individually, we integrate it with a novel relational knowledge distillation mechanism to explicitly capture the structural mutual argument-argument relation. Moreover, since the training process is not compatible with the real situation, a curriculum learning method is further introduced to make the training process smoother. Experimental results demonstrate that the learning framework obtains state-of-the-art performance on the RAMS and Wikievents datasets. Ablation study and further discussion also show it could handle long-range dependency and implicit argument problems effectively. Kaiwen Wei, Xian Sun 0001, Zequn Zhang, Li Jin 0001, Jianwei Lv, Zhi Guo |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | Assist Non-native Viewers: Multimodal Cross-Lingual Summarization for How2 VideosabstractMultimodal summarization for videos aims to generate summaries from multi-source information (videos, audio transcripts), which has achieved promising progress.However, existing works are restricted to monolingual video scenarios, ignoring the demands of non-native video viewers to understand the cross-language videos in practical applications.It stimulates us to propose a new task, named Multimodal Cross-Lingual Summarization for videos (MCLS), which aims to generate cross-lingual summaries from multimodal inputs of videos.First, to make it applicable to MCLS scenarios, we conduct a Video-guided Dual Fusion network (VDF) that integrates multimodal and cross-lingual information via diverse fusion strategies at both encoder and decoder.Moreover, to alleviate the problem of high annotation costs and limited resources in MCLS, we propose a triple-stage training framework to assist MCLS by transferring the knowledge from monolingual multimodal summarization data, which includes: 1) multimodal summarization on sufficient prevalent language videos with a VDF model; 2) knowledge distillation (KD) guided adjustment on bilingual transcripts; 3) multimodal summarization for cross-lingual videos with a KD induced VDF model.Experiment results on the reorganized How2 dataset show that the VDF model alone outperforms previous methods for multimodal summarization, and the performance further improves by a large margin via the proposed triple-stage training framework. * Equal contribution. † Corresponding author.Portuguese (Pt) Transcript: vamos falar hoje sobre o solo.em primeiro lugar, precisamos de uma grande quan dade de solo bom para transplantes na primavera.ela vai adicionar partes iguais de musgo de turfa e composto de jardinagem que extraímos do nosso sistema interno de compostagem, e então um agregado orgânico, uma pedra chamada perlite, que serve para adicionar volume e aumentar a capacidade de retenção de água e de aeração de sua mistura... English (En) Summary: mix sterile soil for plan ng greens in trays to keep in a hoop house.learn to mix soil for growing greens from an organic farmer in this free gardening video. Nayu Liu, Kaiwen Wei, Xian Sun 0001, Fanglong Yao, Li Jin 0001, Zhi Guo, Guangluan Xu |
EMNLP | 2 |
| 2022 | HEFT: A History-Enhanced Feature Transfer framework for incremental event detection
Kaiwen Wei, Zequn Zhang, Li Jin 0001, Zhi Guo, Shuchao Li, Jianwei Lv |
Knowl. Based Syst. | 1 |
| 2021 | Trigger is Not Sufficient: Exploiting Frame-aware Knowledge for Implicit Event Argument ExtractionabstractKaiwen Wei, Xian Sun, Zequn Zhang, Jingyuan Zhang, Guo Zhi, Li Jin. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Kaiwen Wei, Xian Sun 0001, Zequn Zhang, Zhi Guo, Li Jin 0001 |
ACL/IJCNLP (1) | 1 |