EDBT 2026 Demo / reviewers in the wild / expert
Weize Chen
dblp:245/7488
· DBLP profile ↗
21ranked-venue papers
7as first author
20since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 6 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Online Adaptation Framework for Enhancing Calibration-Free SSVEP-Based BCI PerformanceabstractAccomplishing a plug-and-play steady-state visual evoked potential (SSVEP)-based brain-computer interface (BCI) remains a critical challenge, due to the unsatisfying performance of calibration-free decoding algorithms.A current method called online adaptive canonical correlation analysis (OACCA) has proved efficient in enhancing calibration-free performance by self-adaptation merely with online data.However, OACCA only concerns the adaptation of spatial filters and excludes other useful adaptive procedures like individual template estimation, hindering fully exploitable model decoding and adaptation. This study proposes a new online adaptation framework termed online adaptive extended correlation analysis (OAECA) to augment the calibration-free online adaptation loop. OAECA first recalls and cleans the online trials for reliable data learning, then tunes individual templates and spatial filters for complete model updating, and finally adopts extended feature matching to improve target recognition. The simulation results on two public SSVEP datasets revealed that OAECA significantly outperformed OACCA for almost all 105 subjects, and both offline and online experiments further confirmed the effectiveness of OAECA. Particularly, OAECA achieved the highest average information transfer rate (ITR) of 202.17 bits/min in the online experiment, significantly exceeding the state-of-the-art OACCA of 177.02 bits/min. This study enhanced the calibration-free performance through comprehensive online adaptation, hopefully advancing SSVEP-based BCIs toward practical plug-and-play real-world applications. Weize Chen, Xiaolin Xiao, Lingling Tao, Kun Wang 0053, Minpeng Xu, Dong Ming |
IEEE J. Biomed. Health Informatics | 1 |
| 2025 | AgentRM: Enhancing Agent Generalization with Reward ModelingabstractYu Xia, Jingru Fan, Weize Chen, Siyu Yan, Xin Cong, Zhong Zhang, Yaxi Lu, Yankai Lin, Zhiyuan Liu, Maosong Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jingru Fan, Weize Chen, Xin Cong, Zhong Zhang 0004, Yaxi Lu, Yankai Lin 0001, Zhiyuan Liu 0001, Maosong Sun 0001 |
ACL (1) | 3 |
| 2025 | Internet of Agents: Weaving a Web of Heterogeneous Agents for Collaborative IntelligenceabstractThe rapid advancement of large language models (LLMs) has paved the way for the development of highly capable autonomous agents. However, existing multi-agent frameworks often struggle with integrating diverse capable third-party agents due to reliance on agents defined within their own ecosystems. They also face challenges in simulating distributed environments, as most frameworks are limited to single-device setups. Furthermore, these frameworks often rely on hard-coded communication pipelines, limiting their adaptability to dynamic task requirements. Inspired by the concept of the Internet, we propose the Internet of Agents (IoA), a novel framework that addresses these limitations by providing a flexible and scalable platform for LLM-based multi-agent collaboration. IoA introduces an agent integration protocol, an instant-messaging-like architecture design, and dynamic mechanisms for agent teaming and conversation flow control. Through extensive experiments on general assistant tasks, embodied AI tasks, and retrieval-augmented generation benchmarks, we demonstrate that IoA consistently outperforms state-of-the-art baselines, showcasing its ability to facilitate effective collaboration among heterogeneous agents. IoA represents a step towards linking diverse agents in an Internet-like environment, where agents can seamlessly collaborate to achieve greater intelligence and capabilities. We will release our code to facilitate further research. Weize Chen, Ziming You, Yitong Guan, Cheng Yang 0002, Ruobing Xie, Zhiyuan Liu 0001, Maosong Sun 0001 |
ICLR | 1 |
| 2025 | Scaling Large Language Model-based Multi-Agent CollaborationabstractRecent breakthroughs in large language model-driven autonomous agents have revealed that multi-agent collaboration often surpasses each individual through collective reasoning. Inspired by the neural scaling law—increasing neurons enhances performance, this study explores whether the continuous addition of collaborative agents can yield similar benefits. Technically, we utilize directed acyclic graphs to organize agents into a multi-agent collaboration network (MacNet), upon which their interactive reasoning is topologically orchestrated for autonomous task solving. Extensive evaluations reveal that it effectively supports collaboration among over a thousand agents, with irregular topologies outperforming regular ones. We also identify a collaborative scaling law—the overall performance follows a logistic growth pattern as agents scale, with collaborative emergence occurring earlier than traditional neural emergence. We speculate this may be because scaling agents catalyzes their multidimensional considerations during interactive reflection and refinement, thereby producing more comprehensive artifacts. The code is available at https://github.com/OpenBMB/ChatDev/tree/macnet. Zihao Xie, Wei Liu 0161, Kunlun Zhu, Hanchen Xia, Yufan Dang, Zhuoyun Du, Weize Chen, Cheng Yang 0002, Zhiyuan Liu 0001, Maosong Sun 0001 |
ICLR | 9 |
| 2025 | The Overthinker's DIET: Cutting Token Calories with DIfficulty-AwarE TrainingabstractRecent large language models (LLMs) exhibit impressive reasoning but often \textit{overthink}, generating excessively long responses that hinder efficiency. We introduce DIET (DIfficulty-AwarE Training), a framework that systematically cuts these "token calories" by integrating on-the-fly problem difficulty into the reinforcement learning (RL) process. DIET dynamically adapts token compression strategies by modulating token penalty strength and conditioning target lengths on estimated task difficulty, to optimize the performance-efficiency trade-off. We also theoretically analyze the pitfalls of naive reward weighting in group-normalized RL algorithms like GRPO, and propose \textit{Advantage Weighting} technique, which enables stable and effective implementation of these difficulty-aware objectives. Experimental results demonstrate that DIET significantly reduces token counts while simultaneously improving reasoning performance. Beyond raw token reduction, we show two crucial benefits largely overlooked by prior work: (1) DIET leads to superior \textbf{inference scaling}. By maintaining high per-sample quality with fewer tokens, it enables better scaling performance via majority voting under fixed computational budgets, an area where other methods falter. (2) DIET enhances the natural positive correlation between response length and problem difficulty, ensuring verbosity is appropriately allocated, unlike many existing compression methods that disrupt this relationship. Our analyses provide a principled and effective framework for developing more efficient, practical, and high-performing LLMs. Weize Chen, Jiarui Yuan, Tailin Jin, Ning Ding 0002, Zhiyuan Liu 0001, Maosong Sun 0001 |
NeurIPS | 1 |
| 2025 | Multi-Agent Collaboration via Evolving OrchestrationabstractLarge language models (LLMs) have achieved remarkable results across diverse downstream tasks, but their monolithic nature restricts scalability and efficiency in complex problem-solving. While recent research explores multi-agent collaboration among LLMs, most approaches rely on static organizational structures that struggle to adapt as task complexity and agent numbers grow, resulting in coordination overhead and inefficiencies. To this end, we propose a puppeteer-style paradigm for LLM-based multi-agent collaboration, where a centralized orchestrator ("puppeteer") dynamically directs agents ("puppets") in response to evolving task states. This orchestrator is trained via reinforcement learning to adaptively sequence and prioritize agents, enabling flexible and evolvable collective reasoning. Experiments on closed- and open-domain scenarios show that this method achieves superior performance with reduced computational costs. Analyses further reveal that the key improvements consistently stem from the emergence of more compact, cyclic reasoning structures under the orchestrator’s evolution. Our code is available at https://github.com/OpenBMB/ChatDev/tree/puppeteer. Yufan Dang, Xueheng Luo, Jingru Fan, Zihao Xie, Ruijie Shi, Weize Chen, Cheng Yang 0002, Xiaoyin Che, Xuantang Xiong, Zhiyuan Liu 0001, Maosong Sun 0001 |
NeurIPS | 7 |
| 2025 | A High-DOF BCI Control Strategy Mapping Discrete Commands to Continuous Motion for a DroneabstractObjective: Because of the non-stationary nature of electroencephalogram (EEG) signals, traditional non-invasive brain-computer interfaces (BCIs) usually only produce discrete commands, limiting their ability to control external devices continuously. This study proposes a novel BCI control strategy mapping multiple discrete commands to continuous motion, enabling real-time manipulation of a drone in four degrees of freedom (DOF).Methods: Our strategy used the fast steady state visual evoked potential (SSVEP) encoding and decoding method to convert user intentions into the drone’s flight status in near real-time. Simultaneously, the drone’s live video was embedded into the SSVEP stimuli, providing users with a first-person perspective control experience.Results: In drone control experiments, participants successfully maneuvered the drone through complex path-following tasks in simulated and physical scenarios. The mean flight trajectory bias ratio was measured as 0.81, with a mean flight smoothness of -3.31 (measured by spectral arc length) and mean Fitts’s throughput of 9.18 bits/min. Notably, the brain-to-hand ratio (BHR) for all metrics approached 1, indicating that our non-invasive control system achieved comparable performance to manual control systems.Conclusion: These results suggest the effectiveness of our proposed BCI control strategy that maps discrete commands to continuous motion and extends the capabilities of non-invasive BCIs in continuous control scenarios.Significance: This study significantly advances the applications of BCI and propels human-machine interaction towards a more direct realm. Weize Chen, Yongzhi Huang 0001, Xiaolin Xiao, Kun Wang 0053, Weibo Yi, Tzyy-Ping Jung, Minpeng Xu, Dong Ming |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2024 | Experiential Co-Learning of Software-Developing AgentsabstractChen Qian, Yufan Dang, Jiahao Li, Wei Liu, Zihao Xie, YiFei Wang, Weize Chen, Cheng Yang, Xin Cong, Xiaoyin Che, Zhiyuan Liu, Maosong Sun. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yufan Dang, Wei Liu 0161, Zihao Xie, Weize Chen, Cheng Yang 0002, Xin Cong, Xiaoyin Che, Zhiyuan Liu 0001, Maosong Sun 0001 |
ACL (1) | 7 |
| 2024 | ChatDev: Communicative Agents for Software DevelopmentabstractChen Qian, Wei Liu, Hongzhang Liu, Nuo Chen, Yufan Dang, Jiahao Li, Cheng Yang, Weize Chen, Yusheng Su, Xin Cong, Juyuan Xu, Dahai Li, Zhiyuan Liu, Maosong Sun. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Wei Liu 0161, Hongzhang Liu, Yufan Dang, Cheng Yang 0002, Weize Chen, Yusheng Su, Xin Cong, Juyuan Xu, Dahai Li, Zhiyuan Liu 0001, Maosong Sun 0001 |
ACL (1) | 8 |
| 2024 | ChatEval: Towards Better LLM-based Evaluators through Multi-Agent DebateabstractText evaluation has historically posed significant challenges, often demanding substantial labor and time cost. With the emergence of large language models (LLMs), researchers have explored LLMs' potential as alternatives for human evaluation. While these single-agent-based approaches show promise, experimental results suggest that further advancements are needed to bridge the gap between their current effectiveness and human-level evaluation quality.
Recognizing that best practices of human evaluation processes often involve multiple human annotators collaborating in the evaluation, we resort to a multi-agent debate framework, moving beyond single-agent prompting strategies.
In this paper, we construct a multi-agent referee team called $\textbf{ChatEval}$ to autonomously discuss and evaluate the quality of different texts.
Our experiments on two benchmarks illustrate that ChatEval delivers superior accuracy and correlation in alignment with human assessment. Furthermore, we find that the diverse role prompts (different personas) are essential in the multi-agent debate process; that is, utilizing the same role description in the prompts can lead to a degradation in performance. Our qualitative analysis also shows that ChatEval transcends mere textual scoring, offering a human-mimicking evaluation process for reliable assessments. Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue 0002, Shanghang Zhang, Jie Fu 0001, Zhiyuan Liu 0001 |
ICLR | 2 |
| 2024 | AgentVerse: Facilitating Multi-Agent Collaboration and Exploring Emergent BehaviorsabstractAutonomous agents empowered by Large Language Models (LLMs) have undergone significant improvements, enabling them to generalize across a broad spectrum of tasks. However, in real-world scenarios, cooperation among individuals is often required to enhance the efficiency and effectiveness of task accomplishment. Hence, inspired by human group dynamics, we propose a multi-agent framework AgentVerse that can effectively orchestrate a collaborative group of expert agents as a greater-than-the-sum-of-its-parts system. Our experiments demonstrate that AgentVerse can proficiently deploy multi-agent groups that outperform a single agent. Extensive experiments on text understanding, reasoning, coding, tool utilization, and embodied AI confirm the effectiveness of AgentVerse. Moreover, our analysis of agent interactions within AgentVerse reveals the emergence of specific collaborative behaviors, contributing to heightened group efficiency. We will release our codebase, AgentVerse, to further facilitate multi-agent research. Weize Chen, Yusheng Su, Jingwei Zuo, Cheng Yang 0002, Chenfei Yuan, Chi-Min Chan, Heyang Yu, Yaxi Lu, Yi-Hsin Hung, Yujia Qin, Xin Cong, Ruobing Xie, Zhiyuan Liu 0001, Maosong Sun 0001, Jie Zhou 0016 |
ICLR | 1 |
| 2024 | Can Large Language Models Analyze Graphs like Professionals? A Benchmark, Datasets and ModelsabstractThe need to analyze graphs is ubiquitous across various fields, from social networks to biological research and recommendation systems. Therefore, enabling the ability of large language models (LLMs) to process graphs is an important step toward more advanced general intelligence. However, current LLM benchmarks on graph analysis require models to directly reason over the prompts describing graphtopology, and are thus limited to small graphs with only a few dozens of nodes. In contrast, human experts typically write programs based on popular libraries for task solving, and can thus handle graphs with different scales. To this end, a question naturally arises: can LLMs analyze graphs like professionals? In this paper, we introduce ProGraph, a manually crafted benchmark containing 3 categories of graph tasks. The benchmark expects solutions based on programming instead of directly reasoning over raw inputs. Our findings reveal that the performance of current LLMs is unsatisfactory, with the best model achieving only 36% accuracy. To bridge this gap, we propose LLM4Graph datasets, which include crawled documents and auto-generated codes based on 6 widely used graph libraries. By augmenting closed-source LLMs with document retrieval and fine-tuning open-source ones on the codes, we show 11-32% absolute improvements in their accuracies. Our results underscore that the capabilities of LLMs in handling structured data are still under-explored, and show the effectiveness of LLM4Graph in enhancing LLMs’ proficiency of graph analysis. The benchmark, datasets and enhanced open-sourcemodels are available at https://github.com/BUPT-GAMMA/ProGraph. Xin Li 0005, Weize Chen, Qizhi Chu, Zhaojun Sun, Chuan Shi 0001, Zhiyuan Liu 0001, Maosong Sun 0001, Cheng Yang 0002 |
NeurIPS | 2 |
| 2024 | Autonomous Agents for Collaborative Task under Information AsymmetryabstractLarge Language Model Multi-Agent Systems (LLM-MAS) have greatly progressed in solving complex tasks. It communicates among agents within the system to collaboratively solve tasks, under the premise of shared information. However, when agents' collaborations are leveraged to perform multi-person tasks, a new challenge arises due to information asymmetry, since each agent can only access the information of its human user. Previous MAS struggle to complete tasks under this condition. To address this, we propose a new MAS paradigm termed iAgents, which denotes Informative Multi-Agent Systems. In iAgents, the human social network is mirrored in the agent network, where agents proactively exchange human information necessary for task resolution, thereby overcoming information asymmetry. iAgents employs a novel agent reasoning mechanism, InfoNav, to navigate agents' communication towards effective information exchange. Together with InfoNav, iAgents organizes human information in a mixed memory to provide agents with accurate and comprehensive information for exchange. Additionally, we introduce InformativeBench, the first benchmark tailored for evaluating LLM agents' task-solving ability under information asymmetry. Experimental results show that iAgents can collaborate within a social network of 140 individuals and 588 relationships, autonomously communicate over 30 turns, and retrieve information from nearly 70,000 messages to complete tasks within 3 minutes. Wei Liu 0161, Chenxi Wang 0001, Zihao Xie, Rennai Qiu, Yufan Dang, Zhuoyun Du, Weize Chen, Cheng Yang 0002 |
NeurIPS | 8 |
| 2024 | D-Bot: Database Diagnosis System using Large Language ModelsabstractDatabase administrators (DBAs) play an important role in managing database systems. However, it is hard and tedious for DBAs to manage vast database instances and give timely response (waiting for hours is intolerable in many online cases). In addition, existing empirical methods only support limited diagnosis scenarios, which are also labor-intensive to update the diagnosis rules for database version updates. Recently large language models (LLMs) have shown great potential in various fields. Thus, we propose D-Bot , an LLM-based database diagnosis system that can automatically acquire knowledge from diagnosis documents, and generate reasonable and well-founded diagnosis report (i.e., identifying the root causes and solutions) within acceptable time (e.g., under 10 minutes compared to hours by a DBA). The techniques in D-Bot include ( i ) offline knowledge extraction from documents, ( ii ) automatic prompt generation (e.g., knowledge matching, tool retrieval), ( iii ) root cause analysis using tree search algorithm, and ( iv ) collaborative mechanism for complex anomalies with multiple root causes. We verify D-Bot on real benchmarks (including 539 anomalies of six typical applications), and the results show D-Bot can effectively identify root causes of unseen anomalies and significantly outperforms traditional methods and vanilla models like GPT-4. Xuanhe Zhou, Guoliang Li 0001, Zhaoyan Sun, Zhiyuan Liu 0001, Weize Chen, Jiesi Liu, Ruohang Feng, Guoyang Zeng |
Proc. VLDB Endow. | 5 |
| 2024 | Hyperbolic Pre-Trained Language ModelabstractIn recent years, we have witnessed significant improvements in pre-trained language models (PLM) brought about by the scaling of parameter sizes and data amounts. However, this also brings high computational and storage costs. In this paper, we present a new direction to improve PLMs without scaling parameters and data: adopting a geometric feature space that is more suitable for encoding the intrinsic structured features of text. Although text is generally considered unstructured data, it possesses rich intrinsic structured features that signify syntactic and semantic relationships. Leveraging these structured features is vital for text understanding. Given that structured features are better encoded in hyperbolic spaces than in the Euclidean spaces used by conventional PLMs, we propose that PLMs should operate entirely within hyperbolic spaces. Our experiments demonstrate the superiority of hyperbolic PLMs over Euclidean PLMs across a wide variety of tasks, using the same parameter and data settings. This suggests that altering the geometry of model representation is a promising direction for model enhancement. The code is released athttps://github.com/thunlp/hyperbolic_llm Weize Chen, Xu Han 0007, Yankai Lin 0001, Kaichen He, Ruobing Xie, Jie Zhou 0016, Zhiyuan Liu 0001, Maosong Sun 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2024 | Exploring Universal Intrinsic Task Subspace for Few-Shot Learning via Prompt TuningabstractWhy can pre-trained language models (PLMs) learn universal representations and effectively adapt to broad NLP tasks differing a lot superficially? In this work, we empirically find evidence indicating that the adaptations of PLMs to various few-shot tasks can be reparameterized as optimizing only a few free parameters in a unified low-dimensionalintrinsic task subspace, which may help us understand why PLMs could easily adapt to various NLP tasks with small-scale data. To find such a subspace and examine its universality, we propose an analysis pipeline calledintrinsic prompt tuning(IPT). Specifically, we resort to the recent success of prompt tuning and decompose the soft prompts of multiple NLP tasks into the same low-dimensional nonlinear subspace, then we learn to adapt the PLM to unseen data or tasks by only tuning parameters in this subspace. In the experiments, we study diverse few-shot NLP tasks and surprisingly find that in a 250-dimensional subspace found with 100 tasks, by only tuning 250 free parameters, we can recover 97% and 83% of the full prompt tuning performance for 100 seen tasks (using different training data) and 20 unseen tasks, respectively, showing great generalization ability of the found intrinsic task subspace. Besides being an analysis tool, IPTcould further help us improve the prompt tuning stability. Yujia Qin, Xiaozhi Wang, Yusheng Su, Yankai Lin 0001, Ning Ding 0002, Jing Yi, Weize Chen, Zhiyuan Liu 0001, Juan-Zi Li, Lei Hou 0001, Peng Li 0030, Maosong Sun 0001, Jie Zhou 0016 |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2022 | Fully Hyperbolic Neural NetworksabstractWeize Chen, Xu Han, Yankai Lin, Hexu Zhao, Zhiyuan Liu, Peng Li, Maosong Sun, Jie Zhou. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Weize Chen, Xu Han 0007, Yankai Lin 0001, Hexu Zhao, Zhiyuan Liu 0001, Peng Li 0030, Maosong Sun 0001, Jie Zhou 0016 |
ACL (1) | 1 |
| 2022 | Cross-Lingual Contrastive Learning for Fine-Grained Entity Typing for Low-Resource LanguagesabstractXu Han, Yuqi Luo, Weize Chen, Zhiyuan Liu, Maosong Sun, Zhou Botong, Hao Fei, Suncong Zheng. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Xu Han 0007, Yuqi Luo, Weize Chen, Zhiyuan Liu 0001, Maosong Sun 0001, Botong Zhou, Suncong Zheng |
ACL (1) | 3 |
| 2022 | Exploring Mode Connectivity for Pre-trained Language ModelsabstractRecent years have witnessed the prevalent application of pre-trained language models (PLMs) in NLP.From the perspective of parameter space, PLMs provide generic initialization, starting from which high-performance minima could be found.Although plenty of works have studied how to effectively and efficiently adapt PLMs to high-performance minima, little is known about the connection of various minima reached under different adaptation configurations.In this paper, we investigate the geometric connections of different minima through the lens of mode connectivity, which measures whether two minima can be connected with a low-loss path.We conduct empirical analyses to investigate three questions: (1) how could hyperparameters, specific tuning methods, and training data affect PLM's mode connectivity?(2) How does mode connectivity change during pretraining?(3) How does the PLM's task knowledge change along the path connecting two minima?In general, exploring the mode connectivity of PLMs conduces to understanding the geometric connection of different minima, which may help us fathom the inner workings of PLM downstream adaptation.The codes are publicly available at https://github.com/ thunlp/Mode-Connectivity-PLM. Yujia Qin, Cheng Qian 0008, Jing Yi, Weize Chen, Yankai Lin 0001, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001, Jie Zhou 0016 |
EMNLP | 4 |
| 2022 | GACT: Activation Compressed Training for Generic Network ArchitecturesabstractTraining large neural network (NN) models requires extensive memory resources, and Activation Compression Training (ACT) is a promising approach to reduce training memory footprint. This paper presents GACT, an ACT framework to support a broad range of machine learning tasks for generic NN architectures with limited domain knowledge. By analyzing a linearized version of ACT’s approximate gradient, we prove the convergence of GACT without prior knowledge on operator type or model architecture. To make training stable, we propose an algorithm that decides the compression ratio for each tensor by estimating its impact on the gradient at run time. We implement GACT as a PyTorch library that readily applies to any NN architecture. GACT reduces the activation memory for convolutional NNs, transformers, and graph NNs by up to 8.1x, enabling training with a 4.2x to 24.7x larger batch size, with negligible accuracy loss. Lianmin Zheng, Dequan Wang, Yukuo Cen, Weize Chen, Xu Han 0007, Jianfei Chen 0001, Zhiyuan Liu 0001, Jie Tang 0001, Joey Gonzalez, Michael W. Mahoney, Alvin Cheung |
ICML | 5 |
| 2019 | Quantifying Similarity between Relations with Fact DistributionabstractWe introduce a conceptually simple and effective method to quantify the similarity between relations in knowledge bases.Specifically, our approach is based on the divergence between the conditional probability distributions over entity pairs.In this paper, these distributions are parameterized by a very simple neural network.Although computing the exact similarity is intractable, we provide a sampling-based method to get a good approximation.We empirically show the outputs of our approach significantly correlate with human judgments.By applying our method to various tasks, we also find that (1) our approach could effectively detect redundant relations extracted by open information extraction (Open IE) models, that (2) even the most competitive models for relational classification still make mistakes among very similar relations, and that (3) our approach could be incorporated into negative sampling and softmax classification to alleviate these mistakes.The source code and experiment details of this paper can be obtained from https://github.com/ thunlp/relation-similarity. Weize Chen, Hao Zhu 0006, Xu Han 0007, Zhiyuan Liu 0001, Maosong Sun 0001 |
ACL (1) | 1 |