EDBT 2026 Demo / reviewers in the wild / expert
Zhuoran Jin
dblp:320/9888
· DBLP profile ↗
24ranked-venue papers
4as first author
24since 2021 · last 2026
0000-0002-1667-8145ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 4 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Explainable Diagnosis: A Self-learned Explanatory Knowledge Base ApproachabstractDongqi Huang, Tong Zhou, Zhuoran Jin, Shenghui Shi, Maoyujiao, Kang Liu, Jun Zhao, Yubo Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Dongqi Huang, Tong Zhou 0014, Zhuoran Jin, Shenghui Shi, Yujiao Mao, Kang Liu 0001, Jun Zhao 0001, Yubo Chen 0001 |
ACL (1) | 3 |
| 2026 | Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot DoabstractZhuoran Jin, Kejian Zhu, Hongbang Yuan, Yupu Hao, Pengfei Cao, Yubo Chen, Kang Liu, Jun Zhao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhuoran Jin, Kejian Zhu, Hongbang Yuan, Yupu Hao, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001 |
ACL (1) | 1 |
| 2026 | Empowering GUI Agents via Autonomous Experience Exploration and Hindsight Experience Utilization for Task PlanningabstractMultimodal web agents can assist humans in operating repetitive GUI tasks, where effective task planning is essential for decomposing complex tasks into executable actions.While small open-source MLLMs are cost-efficient and privacy-preserving compared with commercial large models, they suffer from weak planning and limited cross-website generalization.To address these limitations, we introduce the planning experience exploration and utilization (PEEU) method, which autonomously explores environments to discover experiences and utilizes hindsight experience to synthesize strictly aligned, high-level training data.To quantitatively analyze the generalization behaviors driving this performance, we propose the task decomposition hierarchical analysis framework (TDHAF) to systematically study compositional generalization across three task granularities: low, middle and high levels.Our analysis reveals that mastering low-level atomic skills does not guarantee high-level planning competence, while high-level task training yields stronger OOD generalization.Experiments on real-world benchmarks demonstrate PEEU's superior effectiveness: our 7B model achieves 30.6% accuracy, outperforming the much larger Qwen2.5-VL-32Bmodel.These demonstrate constructing hindsight high-level tasks and leveraging experiences is crucial for OOD planning abilities of small MLLMs. Tianyi Men, Zhuoran Jin, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001 |
ACL (1) | 2 |
| 2026 | One mind, many tongues: A deep dive into language-agnostic knowledge neurons in large language models
Zhuoran Jin, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001 |
Artif. Intell. | 3 |
| 2025 | CITI: Enhancing Tool Utilizing Ability in Large Language Models Without Sacrificing General PerformanceabstractTool learning enables Large Language Models (LLMs) to interact with the external environment by invoking tools, enriching the accuracy and capability scope of LLMs. However, previous works predominantly focus on improving the model's tool-utilizing accuracy and the ability to generalize to new, unseen tools, excessively forcing LLMs to adjust specific tool-invoking pattern without considering the harm to the model's general performance. This deviates from the actual applications and original intention of integrating tools to enhance the model. To tackle this problem, we dissect the capability trade-offs by examining the hidden representation changes and the gradient-based importance score of the model's components. Based on the analysis result, we propose a Component Importance-based Tool-utilizing ability Injection method (CITI). According to the gradient-based importance score of different components, it alleviates the capability conflicts caused by the fine-tuning process by applying distinct training strategies to different components. CITI applies Mixture-Of-LoRA (MOLoRA) for important components. Meanwhile, it fine-tunes the parameters of a few components deemed less important in the backbone of the LLM, while keeping other parameters frozen. CITI can effectively enhance the model's tool-utilizing capability without excessively compromising its general performance. Experimental results demonstrate that our approach achieves outstanding performance across a range of evaluation metrics. Yupu Hao, Zhuoran Jin, Huanxuan Liao, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001 |
AAAI | 3 |
| 2025 | Towards Robust Knowledge Unlearning: An Adversarial Framework for Assessing and Improving Unlearning Robustness in Large Language ModelsabstractLLM have achieved success in many fields but still troubled by problematic content in the training corpora. LLM unlearning aims at reducing their influence and avoid undesirable behaviours. However, existing unlearning methods remain vulnerable to adversarial queries and the unlearned knowledge resurfaces after the manually designed attack queries. As part of a red-team effort to proactively assess the vulnerabilities of unlearned models, we design Dynamic Unlearning Attack (DUA), a dynamic and automated framework to attack these models and evaluate their robustness. It optimizes adversarial suffixes to reintroduce the unlearned knowledge in various scenarios. We find that unlearned knowledge can be recovered in 55.2% of the questions, even without revealing the unlearned model's parameters. In response to this vulnerability, we propose Latent Adversarial Unlearning (LAU), a universal framework that effectively enhances the robustness of the unlearned process. It formulates the unlearning process as a min-max optimization problem and resolves it through two stages: an attack stage, where perturbation vectors are trained and added to the latent space of LLMs to recover the unlearned knowledge, and a defense stage, where previously trained perturbation vectors are used to enhance unlearned model's robustness. With our LAU framework, we obtain two robust unlearning methods, AdvGA and AdvNPO. We conduct extensive experiments across multiple unlearning benchmarks and various models, and demonstrate that they improve the unlearning effectiveness by over 53.5%, cause only less than a 11.6% reduction in neighboring knowledge, and have almost no impact on the model's general capabilities. Hongbang Yuan, Zhuoran Jin, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001 |
AAAI | 2 |
| 2025 | Evaluating Personalized Tool-Augmented LLMs from the Perspectives of Personalization and ProactivityabstractPersonalized tool utilization is essential for aligning large language models (LLMs) with user preference in interaction scenarios with various tools. However, most of the current benchmarks primarily focus on either personalization of text generation or direct tool-utilizing, without considering both. In this work, we introduce a novel benchmark ETAPP for evaluating personalized tool invocation, establishing a sandbox environment, and a comprehensive dataset of 800 testing cases covering diverse user profiles. To improve the accuracy of our evaluation, we propose a key-point-based LLM evaluation method, mitigating biases in the LLM-as-a-judge system by manually annotating key points for each test case and providing them to LLM as the reference. Additionally, we evaluate the excellent LLMs and provide an in-depth analysis. Furthermore, we investigate the impact of different tool-invoking strategies on LLMs’ personalization performance and the effects of fine-tuning in our task. The effectiveness of our preference-setting and key-point-based evaluation method is also validated. Our findings offer insights into improving personalized LLM agents. Our code is available at https://github.com/hypasd-art/ETAPP. Yupu Hao, Zhuoran Jin, Huanxuan Liao, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001 |
ACL (1) | 3 |
| 2025 | A Troublemaker with Contagious Jailbreak Makes Chaos in Honest TownsabstractWith the development of large language models, they are widely used as agents in various fields. A key component of agents is memory, which stores vital information but is susceptible to jailbreak attacks. Existing research mainly focuses on single-agent attacks and shared memory attacks. However, real-world scenarios often involve independent memory. In this paper, we propose the Troublemaker Makes Chaos in Honest Town (TMCHT) task, a large-scale, multi-agent, multi-topology text-based attack evaluation framework. TMCHT involves one attacker agent attempting to mislead an entire society of agents. We identify two major challenges in multi-agent attacks: (1) Non-complete graph structure, (2) Large-scale systems. We attribute these challenges to a phenomenon we term toxicity disappearing. To address these issues, we propose an Adversarial Replication Contagious Jailbreak (ARCJ) method, which optimizes the retrieval suffix to make poisoned samples more easily retrieved and optimizes the replication suffix to make poisoned samples have contagious ability. We demonstrate the superiority of our approach in TMCHT, with 23.51%, 18.95%, and 52.93% improvements in line, star topologies, and 100-agent settings. It reveals potential contagion risks in widely used multi-agent architectures. Tianyi Men, Zhuoran Jin, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001 |
ACL (1) | 3 |
| 2025 | Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal AgentsabstractAs Multimodal Large Language Models (MLLMs) advance, multimodal agents show promise in real-world tasks like web navigation and embodied intelligence.However, due to limitations in a lack of external feedback, these agents struggle with self-correction and generalization.A promising approach is to use reward models as external feedback, but there is no clear on how to select reward models for agents.Thus, there is an urgent need to build a reward bench targeted at agents.To address these challenges, we propose Agent-RewardBench, a benchmark designed to evaluate reward modeling ability in MLLMs.The benchmark is characterized by three key features: (1) Multiple dimensions and real-world agent scenarios evaluation.It covers perception, planning, and safety with 7 scenarios; (2) Step-level reward evaluation.It allows for the assessment of agent capabilities at the individual steps of a task, providing a more granular view of performance during the planning process; and (3) Appropriately difficulty and highquality.We carefully sample from 10 diverse models, difficulty control to maintain task challenges, and manual verification to ensure the integrity of the data.Experiments demonstrate that even state-of-the-art multimodal models show limited performance, highlighting the need for specialized training in agent reward modeling.Code is available at github. Tianyi Men, Zhuoran Jin, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001 |
ACL (1) | 2 |
| 2025 | Establishing Trustworthy LLM Evaluation via Shortcut Neuron AnalysisabstractThe development of large language models (LLMs) depends on trustworthy evaluation.However, most current evaluations rely on public benchmarks, which are prone to data contamination issues that significantly compromise fairness.Previous researches have focused on constructing dynamic benchmarks to address contamination.However, continuously building new benchmarks is costly and cyclical.In this work, we aim to tackle contamination by analyzing the mechanisms of contaminated models themselves.Through our experiments, we discover that the overestimation of contaminated models is likely due to parameters acquiring shortcut solutions in training.We further propose a novel method for identifying shortcut neurons through comparative and causal analysis.Building on this, we introduce an evaluation method called shortcut neuron patching to suppress shortcut neurons.Experiments validate the effectiveness of our approach in mitigating contamination.Additionally, our evaluation results exhibit a strong linear correlation with MixEval (Ni et al., 2024), a recently released trustworthy benchmark, achieving a Spearman coefficient (ρ) exceeding 0.95.This high correlation indicates that our method closely reveals true capabilities of the models and is trustworthy.We conduct further experiments to demonstrate the generalizability of our method across various benchmarks and hyperparameter settings.Code: Kejian Zhu, Shangqing Tu, Zhuoran Jin, Lei Hou 0001, Juan-Zi Li, Jun Zhao 0001 |
ACL (1) | 3 |
| 2025 | FinRAGBench-V: A Benchmark for Multimodal RAG with Visual Citation in the Financial DomainabstractRetrieval-Augmented Generation (RAG) plays a vital role in the financial domain, powering applications such as real-time market analysis, trend forecasting, and interest rate computation. However, most existing RAG research in finance focuses predominantly on textual data, overlooking the rich visual content in financial documents, resulting in the loss of key analytical insights. To bridge this gap, we present FinRAGBench-V, a comprehensive visual RAG benchmark tailored for finance. This benchmark effectively integrates multimodal data and provides visual citation to ensure traceability. It includes a bilingual retrieval corpus with 60,780 Chinese and 51,219 English pages, along with a high-quality, human-annotated question-answering (QA) dataset spanning heterogeneous data types and seven question categories. Moreover, we introduce RGenCite, an RAG baseline that seamlessly integrates visual citation with generation. Furthermore, we propose an automatic citation evaluation method to systematically assess the visual citation capabilities of Multimodal Large Language Models (MLLMs). Extensive experiments on RGenCite underscore the challenging nature of FinRAGBench-V, providing valuable insights for the development of multimodal RAG systems in finance. Suifeng Zhao, Zhuoran Jin, Sujian Li, Jun Gao 0003 |
EMNLP | 2 |
| 2025 | MIRAGE: Evaluating and Explaining Inductive Reasoning Process in Language ModelsabstractInductive reasoning is an essential capability for large language models (LLMs) to achieve higher intelligence, which requires the model to generalize rules from observed facts and then apply them to unseen examples. We present {\scshape Mirage}, a synthetic dataset that addresses the limitations of previous work, specifically the lack of comprehensive evaluation and flexible test data. In it, we evaluate LLMs' capabilities in both the inductive and deductive stages, allowing for flexible variation in input distribution, task scenario, and task difficulty to analyze the factors influencing LLMs' inductive reasoning. Based on these multi-faceted evaluations, we demonstrate that the LLM is a poor rule-based reasoner. In many cases, when conducting inductive reasoning, they do not rely on a correct rule to answer the unseen case. From the perspectives of different prompting methods, observation numbers, and task forms, models tend to consistently conduct correct deduction without correct inductive rules. Besides, we find that LLMs are good neighbor-based reasoners. In the inductive reasoning process, the model tends to focus on observed facts that are close to the current test example in feature space. By leveraging these similar examples, the model maintains strong inductive capabilities within a localized region, significantly improving its deductive performance. Zhuoran Jin, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001 |
ICLR | 3 |
| 2025 | DTELS: Towards Dynamic Granularity of Timeline SummarizationabstractChenlong Zhang, Tong Zhou, Pengfei Cao, Zhuoran Jin, Yubo Chen, Kang Liu, Jun Zhao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Tong Zhou 0014, Zhuoran Jin, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001 |
NAACL (Long Papers) | 4 |
| 2025 | RULE: Reinforcement UnLEarning Achieves Forget-retain Pareto OptimalityabstractThe widespread deployment of Large Language Models (LLMs) trained on massive, uncurated corpora has raised growing concerns about the inclusion of sensitive, copyrighted, or illegal content. This has led to increasing interest in LLM unlearning: the task of selectively removing specific information from a model without retraining from scratch or degrading overall utility.
However, existing methods often rely on large-scale forget and retain datasets, and suffer from unnatural responses, poor generalization, or catastrophic utility loss.
In this work, we propose $\textbf{R}$einforcement $\textbf{U}$n$\textbf{LE}$arning ($\textbf{RULE}$), an efficient framework that formulates unlearning as a refusal boundary optimization problem. RULE is trained with a small portion of forget set and synthesized boundary queries, using a verifiable reward function that encourages safe refusal on forget-related queries while preserving helpful responses on permissible inputs.
We provide both theoretical and empirical evidence demonstrating the effectiveness of RULE in achieving targeted unlearning without compromising model utility. Experimental results show that, with only 12\% forget set and 8\% synthesized boundary data, RULE outperforms existing baselines by up to $17.4\%$ forget quality and $16.3\%$ naturalness response while maintaining general utility, achieving $\textit{forget-retain Pareto Optimality}$. Remarkably, we further observe that RULE improves the $\textit{naturalness}$ of model outputs, enhances training $\textit{efficiency}$, and exhibits strong $\textit{generalization ability}$, generalizing refusal behavior to semantically related but unseen queries. Zhuoran Jin, Hongbang Yuan, Jiaheng Wei, Tong Zhou 0014, Kang Liu 0001, Jun Zhao 0001, Yubo Chen 0001 |
NeurIPS | 2 |
| 2025 | Prompt robust large language model for Chinese medical named entity recognition
Yubo Chen 0001, Baoli Zhang, Zhuoran Jin, Zhengyuan Cai, Yingzheng Wang, Delai Qiu, Shengping Liu, Jun Zhao 0001 |
Inf. Process. Manag. | 4 |
| 2024 | Focus on Your Question! Interpreting and Mitigating Toxic CoT Problems in Commonsense ReasoningabstractJiachun Li, Pengfei Cao, Chenhao Wang, Zhuoran Jin, Yubo Chen, Daojian Zeng, Kang Liu, Jun Zhao. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Chenhao Wang 0004, Zhuoran Jin, Yubo Chen 0001, Daojian Zeng, Kang Liu 0001, Jun Zhao 0001 |
ACL (1) | 4 |
| 2024 | MULFE: A Multi-Level Benchmark for Free Text Model EditingabstractChenhao Wang, Pengfei Cao, Zhuoran Jin, Yubo Chen, Daojian Zeng, Kang Liu, Jun Zhao. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Chenhao Wang 0004, Zhuoran Jin, Yubo Chen 0001, Daojian Zeng, Kang Liu 0001, Jun Zhao 0001 |
ACL (1) | 3 |
| 2024 | Zero-Shot Cross-Lingual Document-Level Event Causality Identification with Heterogeneous Graph Contrastive Transfer LearningabstractEvent Causality Identification (ECI) refers to the detection of causal relations between events in texts. However, most existing studies focus on sentence-level ECI with high-resource languages, leaving more challenging document-level ECI (DECI) with low-resource languages under-explored. In this paper, we propose a Heterogeneous Graph Interaction Model with Multi-granularity Contrastive Transfer Learning (GIMC) for zero-shot cross-lingual document-level ECI. Specifically, we introduce a heterogeneous graph interaction network to model the long-distance dependencies between events that are scattered over a document. Then, to improve cross-lingual transferability of causal knowledge learned from the source language, we propose a multi-granularity contrastive transfer learning module to align the causal representations across languages. Extensive experiments show our framework outperforms the previous state-of-the-art model by 9.4% and 8.2% of average F1 score on monolingual and multilingual scenarios respectively. Notably, in the multilingual scenario, our zero-shot framework even exceeds GPT-3.5 with few-shot learning by 24.3% in overall performance. Zhitao He 0001, Zhuoran Jin, Yubo Chen 0001, Kang Liu 0001, Mengshu Sun, Jun Zhao 0001 |
LREC/COLING | 3 |
| 2024 | Tug-of-War between Knowledge: Exploring and Resolving Knowledge Conflicts in Retrieval-Augmented Language ModelsabstractRetrieval-augmented language models (RALMs) have demonstrated significant potential in refining and expanding their internal memory by retrieving evidence from external sources. However, RALMs will inevitably encounter knowledge conflicts when integrating their internal memory with external sources. Knowledge conflicts can ensnare RALMs in a tug-of-war between knowledge, limiting their practical applicability. In this paper, we focus on exploring and resolving knowledge conflicts in RALMs. First, we present an evaluation framework for assessing knowledge conflicts across various dimensions. Then, we investigate the behavior and preference of RALMs from the following two perspectives: (1) Conflicts between internal memory and external sources: We find that stronger RALMs emerge with the Dunning-Kruger effect, persistently favoring their faulty internal memory even when correct evidence is provided. Besides, RALMs exhibit an availability bias towards common knowledge; (2) Conflicts between truthful, irrelevant and misleading evidence: We reveal that RALMs follow the principle of majority rule, leaning towards placing trust in evidence that appears more frequently. Moreover, we find that RALMs exhibit confirmation bias, and are more willing to choose evidence that is consistent with their internal memory. To solve the challenge of knowledge conflicts, we propose a method called Conflict-Disentangle Contrastive Decoding (CD2) to better calibrate the model’s confidence. Experimental results demonstrate that our CD2 can effectively resolve knowledge conflicts in RALMs. Zhuoran Jin, Yubo Chen 0001, Kang Liu 0001, Xiaojian Jiang, Jiexin Xu, Qiuxia Li, Jun Zhao 0001 |
LREC/COLING | 1 |
| 2024 | Unlocking the Future: Exploring Look-Ahead Planning Mechanistic Interpretability in Large Language ModelsabstractPlanning, as the core module of agents, is crucial in various fields such as embodied agents, web navigation, and tool using.With the development of large language models (LLMs), some researchers treat large language models as intelligent agents to stimulate and evaluate their planning capabilities.However, the planning mechanism is still unclear.In this work, we focus on exploring the look-ahead planning mechanism in large language models from the perspectives of information flow and internal representations.First, we study how planning is done internally by analyzing the multi-layer perception (MLP) and multi-head self-attention (MHSA) components at the last token.We find that the output of MHSA in the middle layers at the last token can directly decode the decision to some extent.Based on this discovery, we further trace the source of MHSA by information flow, and we reveal that MHSA mainly extracts information from spans of the goal states and recent steps.According to information flow, we continue to study what information is encoded within it.Specifically, we explore whether future decisions have been encoded in advance in the representation of flow.We demonstrate that the middle and upper layers encode a few shortterm future decisions to some extent when planning is successful.Overall, our research analyzes the look-ahead planning mechanisms of LLMs, facilitating future research on LLMs performing planning tasks. Tianyi Men, Zhuoran Jin, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001 |
EMNLP | 3 |
| 2024 | Whispers that Shake Foundations: Analyzing and Mitigating False Premise Hallucinations in Large Language ModelsabstractLarge Language Models (LLMs) have shown impressive capabilities but still suffer from the issue of hallucinations.A significant type of this issue is the false premise hallucination, which we define as the phenomenon when LLMs generate hallucinated text when confronted with false premise questions.In this paper, we perform a comprehensive analysis of the false premise hallucination and elucidate its internal working mechanism: a small subset of attention heads (which we designate as false premise heads) disturb the knowledge extraction process, leading to the occurrence of false premise hallucination.Based on our analysis, we propose FAITH (False premise Attention head constraIining for miTigating Hallucinations), a novel and effective method to mitigate false premise hallucinations.It constrains the false premise attention heads during the model inference process.Impressively, extensive experiments demonstrate that constraining only approximately 1% of the attention heads in the model yields a notable increase of nearly 20% of model performance. Hongbang Yuan, Zhuoran Jin, Yubo Chen 0001, Daojian Zeng, Kang Liu 0001, Jun Zhao 0001 |
EMNLP | 3 |
| 2024 | RWKU: Benchmarking Real-World Knowledge Unlearning for Large Language ModelsabstractLarge language models (LLMs) inevitably memorize sensitive, copyrighted, and harmful knowledge from the training corpus; therefore, it is crucial to erase this knowledge from the models. Machine unlearning is a promising solution for efficiently removing specific knowledge by post hoc modifying models. In this paper, we propose a Real-World Knowledge Unlearning benchmark (RWKU) for LLM unlearning. RWKU is designed based on the following three key factors: (1) For the task setting, we consider a more practical and challenging unlearning setting, where neither the forget corpus nor the retain corpus is accessible. (2) For the knowledge source, we choose 200 real-world famous people as the unlearning targets and show that such popular knowledge is widely present in various LLMs. (3) For the evaluation framework, we design the forget set and the retain set to evaluate the model’s capabilities across various real-world applications. Regarding the forget set, we provide four four membership inference attack (MIA) methods and nine kinds of adversarial attack probes to rigorously test unlearning efficacy. Regarding the retain set, we assess locality and utility in terms of neighbor perturbation, general ability, reasoning ability, truthfulness, factuality, and fluency. We conduct extensive experiments across two unlearning scenarios, two models and six baseline methods and obtain some meaningful findings. We release our benchmark and code publicly at http://rwku-bench.github.io for future work. Zhuoran Jin, Chenhao Wang 0004, Zhitao He 0001, Hongbang Yuan, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001 |
NeurIPS | 1 |
| 2023 | Zero-Shot Cross-Lingual Event Argument Extraction with Language-Oriented Prefix-TuningabstractEvent argument extraction (EAE) aims to identify the arguments of a given event, and classify the roles that those arguments play. Due to high data demands of training EAE models, zero-shot cross-lingual EAE has attracted increasing attention, as it greatly reduces human annotation effort. Some prior works indicate that generation-based methods have achieved promising performance for monolingual EAE. However, when applying existing generation-based methods to zero-shot cross-lingual EAE, we find two critical challenges, including Language Discrepancy and Template Construction. In this paper, we propose a novel method termed as Language-oriented Prefix-tuning Network (LAPIN) to address the above challenges. Specifically, we devise a Language-oriented Prefix Generator module to handle the discrepancies between source and target languages. Moreover, we leverage a Language-agnostic Template Constructor module to design templates that can be adapted to any language. Extensive experiments demonstrate that our proposed method achieves the best performance, outperforming the previous state-of-the-art model by 4.8% and 2.3% of the average F1-score on two multilingual EAE datasets. Zhuoran Jin, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001 |
AAAI | 2 |
| 2022 | A Good Neighbor, A Found Treasure: Mining Treasured Neighbors for Knowledge Graph Entity TypingabstractThe task of knowledge graph entity typing (KGET) aims to infer the missing types for entities in knowledge graphs.Some pioneering work has proved that neighbor information is essential for the task.However, existing methods only leverage the one-hop neighbor information of the central entity, ignoring the multi-hop neighbor information that can provide valuable clues for inference.Besides, we also observe that there are co-occurrence relations between types, which is very helpful in alleviating the false-negative problem.In this paper, we propose a novel method called Mining Treasured Neighbors (MiNer) to make use of these two characteristics.Firstly, we devise a Neighbor Information Aggregation module to aggregate the neighbor information.Then, we propose an Entity Type Inference module to mitigate the adverse impact of the irrelevant neighbor information.Finally, a Type Co-occurrence Regularization module is designed to prevent the model from overfitting the false-negative examples caused by missing types.Experimental results on two widely used datasets indicate that our approach significantly outperforms previous state-of-the-art methods. 1 Zhuoran Jin, Yubo Chen 0001, Kang Liu 0001, Jun Zhao 0001 |
EMNLP | 1 |