Yongliang Shen 0001

dblp:221/5612-1 · DBLP profile ↗
← Back
49ranked-venue papers
7as first author
46since 2021 · last 2026
0000-0003-0975-3554ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 46 · 6 first-author · 43 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 10 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Test-Time Reinforcement Learning for GUI Grounding via Region Consistency
abstract
Graphical User Interface (GUI) grounding, the task of mapping natural language instructions to precise screen coordinates, is fundamental to autonomous GUI agents. While existing methods achieve strong performance through extensive supervised training or reinforcement learning with labeled rewards, they remain constrained by the cost and availability of pixel-level annotations. We observe that when models generate multiple predictions for the same GUI element, the spatial overlap patterns reveal implicit confidence signals that can guide more accurate localization. Leveraging this insight, GUI-RC (Region Consistency), a test-time scaling method that constructs spatial voting grids from multiple sampled predictions to identify consensus regions where models show highest agreement. Without any training, GUI-RC improves accuracy by 2-3% across various architectures on ScreenSpot benchmarks. We further introduce GUI-RCPO (Region Consistency Policy Optimization), transforming these consistency patterns into rewards for test-time reinforcement learning. By computing how well each prediction aligns with the collective consensus, GUI-RCPO enables models to iteratively refine their outputs on unlabeled data during inference. Extensive experiments demonstrate the generality of our approach: using only 1,272 unlabeled data, GUI-RCPO achieves 3-6% accuracy improvements across various architectures on ScreenSpot benchmarks. Our approach reveals the untapped potential of test-time scaling and test-time reinforcement learning for GUI grounding, offering a promising path toward more data-efficient GUI agents.
Fei Tang 0005, Zhengxi Lu, Chang Zong, Weiming Lu 0001, Shengpei Jiang, Yongliang Shen 0001
AAAI8
2026 Reality vs Counterfactual: Multi-World Contrastive Reinforcement Learning for Enhancing MLLM's Theory of Mind in Egocentric Videos
abstract
Theory of Mind (ToM) refers to the ability to infer others' mental states, which is an essential capability for embodied AI agents to effectively collaborate and interact with humans. While improving Large Language Models' ability to reason about characters' mental states in text-based stories/dialogues has been extensively studied, enhancing Multimodal Large Language Models' ToM capabilities, particularly in egocentric video from an embodied perspective, remains unexplored. In this paper, we propose a contrastive Reinforcement Learning (RL) paradigm that explicitly encourages models to leverage temporal and causal evolutionary patterns in user action sequences to infer user's mental states (goals, beliefs, and potential next actions). Evaluation results on in-domain and out-of-domain demonstrate that our method achieves performance improvements of (+30.00%, +2.00%) and (+5.83%, +5.00%) compared to the backbone model and vanilla Group Relative Policy Optimization (GRPO) model, respectively. Additionally, we compare the performance of two post-training paradigms (Supervise Fine-Tuning and RL) and systematically analyze the reasoning trajectories across the base model, vanilla GRPO model, and our proposed method.
Guiyang Hou, Yihui Fu, Wenqi Zhang 0001, Yongliang Shen 0001, Weiming Lu 0001
AAAI7
2026 GUI-G²: Gaussian Reward Modeling for GUI Grounding
abstract
Graphical User Interface (GUI) grounding maps natural language instructions to precise interface locations for autonomous interaction. Current reinforcement learning approaches use binary rewards that treat elements as hit-or-miss targets, creating sparse signals that ignore the continuous nature of spatial interactions. Motivated by human clicking behavior that naturally forms Gaussian distributions centered on target elements, we introduce GUI Gaussian Grounding Rewards (GUI-G2), a principled reward framework that models GUI elements as continuous Gaussian distributions across the interface plane. GUI-G2 incorporates two synergistic mechanisms: Gaussian point rewards model precise localization through exponentially decaying distributions centered on element centroids, while coverage rewards assess spatial alignment by measuring the overlap between predicted Gaussian distributions and target regions. To handle diverse element scales, we develop an adaptive variance mechanism that calibrates reward distributions based on element dimensions. This framework transforms GUI grounding from sparse binary classification to dense continuous optimization, where Gaussian distributions generate rich gradient signals that guide models toward optimal interaction positions. Extensive experiments across ScreenSpot, ScreenSpot-v2, and ScreenSpot-Pro benchmarks demonstrate that GUI-G2, substantially outperforms state-of-the-art method UI-TARS-72B, with the most significant improvement of 24.7% on ScreenSpot-Pro. Our analysis reveals that continuous modeling provides superior robustness to interface variations and enhanced generalization to unseen layouts, establishing a new paradigm for spatial reasoning in GUI interaction tasks.
Fei Tang 0005, Zhangxuan Gu, Zhengxi Lu, Shuheng Shen, Changhua Meng, Wen Wang 0009, Wenqi Zhang 0001, Yongliang Shen 0001, Weiming Lu 0001, Jun Xiao 0001, Yueting Zhuang
AAAI9
2026 Leveraging Outline-Optimized Generative Interactions and Critique for Self-Refining Outlines with Reinforcement Learning
abstract
Hengwei Liu, Haoyuan Ma, Qingqing Lyu, Daoxin Zhang, Yao Hu, Yongliang Shen, Yin Zhang, Weiming Lu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Hengwei Liu, Qingqing Lyu, Daoxin Zhang, Yongliang Shen 0001, Yin Zhang 0006, Weiming Lu 0001
ACL (1)6
2026 UI-Copilot: Advancing Long-Horizon GUI Automation via Tool-Integrated Policy Optimization
abstract
Zhengxi Lu, Fei Tang, Guangyi Liu, Jin Ma, Kaitao Song, Xu Tan, Wenqi Zhang, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhengxi Lu, Fei Tang 0005, Kaitao Song, Xu Tan 0003, Wenqi Zhang 0001, Weiming Lu 0001, Jun Xiao 0001, Yueting Zhuang, Yongliang Shen 0001
ACL (1)11
2026 Experience-driven Multi-turn Reinforcement Learning for GUI Agents
abstract
Zhengxi Lu, Jiabo Ye, Fei Tang, Yongliang Shen, Haiyang Xu, Ziwei Zheng, Weiming Lu, Ming Yan, Fei Huang, Jun Xiao, Yueting Zhuang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhengxi Lu, Jiabo Ye, Fei Tang 0005, Yongliang Shen 0001, Haiyang Xu 0001, Ziwei Zheng, Weiming Lu 0001, Ming Yan 0008, Fei Huang 0002, Jun Xiao 0001, Yueting Zhuang
ACL (1)4
2026 AutoTaskEval: Towards Domain-Specific and Fine-Grained Evaluation for LLMs
abstract
Qingqing Lyu, Linjuan Wu, Yongliang Shen, Hengwei Liu, Hao Li, Shengpei Jiang, Yin Zhang, Weiming Lu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Qingqing Lyu, Linjuan Wu, Yongliang Shen 0001, Hengwei Liu, Shengpei Jiang, Yin Zhang 0006, Weiming Lu 0001
ACL (1)3
2026 CoVerRL: Breaking the Consensus Trap in Label-Free Reasoning via Generator-Verifier Co-Evolution
abstract
Teng Pan, Yuchen Yan, Zixuan Wang, Ruiqing Zhang, Guiyang Hou, Wenqi Zhang, Weiming Lu, Jun Xiao, Yongliang Shen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Teng Pan, Ruiqing Zhang, Guiyang Hou, Wenqi Zhang 0001, Weiming Lu 0001, Jun Xiao 0001, Yongliang Shen 0001
ACL (1)9
2026 Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts
abstract
Haolei Xu, Haiwen Hong, Hongxing Li, Rui Zhou, Yang Zhang, Longtao Huang, Hui Xue, Yongliang Shen, Weiming Lu, Yueting Zhuang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Haolei Xu, Haiwen Hong, Longtao Huang, Hui Xue 0001, Yongliang Shen 0001, Weiming Lu 0001, Yueting Zhuang
ACL (1)8
2026 How Controllable Are Large Language Models? A Unified Evaluation across Behavioral Granularities
abstract
Ziwen Xu, Kewei Xu, Haoming Xu, Haiwen Hong, Longtao Huang, Hui Xue, Ningyu Zhang, Yongliang Shen, Guozhou Zheng, Huajun Chen, Shumin Deng. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Ziwen Xu, Kewei Xu, Haiwen Hong, Longtao Huang, Hui Xue 0001, Ningyu Zhang 0001, Yongliang Shen 0001, Guozhou Zheng, Huajun Chen, Shumin Deng
ACL (1)8
2026 Phun-Bench: Evaluating LLMs on Phonological Understanding in Chinese
abstract
Language is a vehicle for thought, intricately tied to sounds, symbols, and meaning.However, most large language model (LLM) research focuses on meaning (semantics) and symbols (spelling) while largely overlooking sounds.Existing benchmarks on LLMs' phonological abilities are either solvable through rote memorization or intertwined with other abilities, making them inadequate to measure LLMs' genuine ability in phonological understanding.Here, we present Phun-Bench, a purpose-built Chinese benchmark with diverse tasks and settings across three dimensions (Homophony, Rhyme, and Phonetic Similarity), designed to systematically evaluate LLMs' phonological understanding.Our results show that while LLMs excel at recalling correct pronunciations, they generally struggle to leverage phonological knowledge in the flexible and intuitive way that human speakers do.Moreover, through detailed analyses, we propose a hypothesis regarding the underlying mechanism of LLMs' phonological understanding and "perception", highlighting an underexplored frontier for future research.1
Xing Yue, Yongliang Shen 0001, Weiming Lu 0001
ACL (1)2
2026 Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks
abstract
Wenqi Zhang, Mengna Wang, Gangao Liu, Huixin Xu, Yiwei Jiang, Yongliang Shen, Guiyang Hou, Zhe Zheng, Hang Zhang, Xin Li, Jiajun Liu, Weiming Lu, Peng Li, Yueting Zhuang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Wenqi Zhang 0001, Mengna Wang, Gangao Liu, Huixin Xu, Yongliang Shen 0001, Guiyang Hou, Xin Li 0056, Weiming Lu 0001, Peng Li 0031, Yueting Zhuang
ACL (1)6
2026 Spatial Context-Aware Dynamic Fusion With Mixture-of-Experts for Wireless Localization
abstract
Multimodal learning emerges as a promising solution for high-precision localization, a cornerstone of 6G integrated sensing and communications (ISAC), by integrating measurements from different data sources. Yet its real-world deployment remains challenging because(i)the quality and relevance of different modalities fluctuate with frequency, noise, and antenna heterogeneity and(ii)spatial and fingerprint ambiguities under non-line-of-sight (NLOS) propagation obscure the mapping between channel measurements and positions. To overcome these challenges, we propose a spatial-context-aware dynamicfusion architecture built on the mixture-of-experts (SCADF-MoE) backbone. We first construct a million-scale comprehensive ray-tracing dataset measuring synchronized angle, distance, gain, and channel across diverse carrier frequencies, antenna geometries, and noise levels. A three-stage pre-processing pipeline then clusters neighboring points into short trajectories, enriching data samples with spatial context information. The resulting sequences are fed into SCADF-MoE: first, multimodal soft MoE blocks with learnable routing matrices dynamically fuse heterogeneous inputs according to their modality relevance in different environmental contexts; second, a modality-task MoE formulates position estimation as a multi-objective problem, simultaneously predicting coordinates of neighboring points to leverage their shared spatial correlations. Additionally, we introduce a regularization loss that enforces expert diversity and mitigates gradient conflicts during multi-task optimization. Simulations across three environments (dense-urban, suburban, canyon) and three heterogeneity dimensions (frequency, noise, antenna) demonstrate that SCADF-MoE achieves consistent sub-meter accuracy in all conditions, reducing overall MSE by 63%, and cuts unseen-NLOS error by 55% compared to state-of-the-art methods. To the best of our knowledge, this is the first work that leverages large-scale multimodal MoEs for high-precision ISAC localization.
Chenwei Wu 0006, Chongwen Huang, Yongliang Shen 0001, Zhaohui Yang 0001, Qianqian Yang 0002, Zhaoyang Zhang 0001, Sami Muhaidat, Chau Yuen
IEEE J. Sel. Areas Commun.4
2025 STaR-SQL: Self-Taught Reasoner for Text-to-SQL
abstract
Generating step-by-step "chain-of-thought" rationales has proven effective for improving the performance of large language models on complex reasoning tasks.However, applying such techniques to structured tasks, such as text-to-SQL, remains largely unexplored.In this paper, we introduce Self-Taught Reasoner for text-to-SQL (STaR-SQL), a novel approach that reframes SQL query generation as a reasoningdriven process.Our method prompts the LLM to produce detailed reasoning steps for SQL queries and fine-tunes it on rationales that lead to correct outcomes.Unlike traditional methods, STaR-SQL dedicates additional test-time computation to reasoning, thereby positioning LLMs as spontaneous reasoners rather than mere prompt-based agents.To further scale the inference process, we incorporate an outcomesupervised reward model (ORM) as a verifier, which enhances SQL query accuracy.Experimental results on the challenging Spider benchmark demonstrate that STaR-SQL significantly improves text-to-SQL performance, achieving an execution accuracy of 86.6%.This surpasses a few-shot baseline by 31.6% and a baseline fine-tuned to predict answers directly by 18.0%.Additionally, STaR-SQL outperforms agent-like prompting methods that leverage more powerful yet closed-source models such as GPT-4.These findings underscore the potential of reasoning-augmented training for structured tasks and open the door to extending self-improving reasoning models to text-to-SQL generation and beyond.
Mingqian He, Yongliang Shen 0001, Wenqi Zhang 0001, Qiuying Peng, Weiming Lu 0001
ACL (1)2
2025 AskToAct: Enhancing LLMs Tool Use via Self-Correcting Clarification
abstract
Xuan Zhang, Yongliang Shen, Zhe Zheng, Linjuan Wu, Wenqi Zhang, Yuchen Yan, Qiuying Peng, Jun Wang, Weiming Lu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Yongliang Shen 0001, Linjuan Wu, Wenqi Zhang 0001, Qiuying Peng, Weiming Lu 0001
EMNLP2
2025 2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining
abstract
Compared to image-text pair data, interleaved corpora enable Vision-Language Models (VLMs) to understand the world more naturally like humans. However, such existing datasets are crawled from webpage, facing challenges like low knowledge density, loose image-text relations, and poor logical coherence between images. On the other hand, the internet hosts vast instructional videos (e.g., online geometry courses) that are widely used by humans to learn foundational subjects, yet these valuable resources remain underexplored in VLM training. In this paper, we introduce a high-quality \textbf{multimodal textbook} corpus with richer foundational knowledge for VLM pretraining. It collects over 2.5 years of instructional videos, totaling 22,000 class hours. We first use an LLM-proposed taxonomy to systematically gather instructional videos. Then we progressively extract and refine visual (keyframes), audio (ASR), and textual knowledge (OCR) from the videos, and organize as an image-text interleaved corpus based on temporal order. Compared to its counterparts, our video-centric textbook offers more coherent context, richer knowledge, and better image-text alignment. Experiments demonstrate its superb pretraining performance, particularly in knowledge- and reasoning-intensive tasks like ScienceQA and MathVista. Moreover, VLMs pre-trained on our textbook exhibit outstanding interleaved context awareness, leveraging visual and textual cues in their few-shot context for task solving. Our code are available at https://github.com/DAMO-NLP-SG/multimodal_textbook.
Wenqi Zhang 0001, Xin Li 0056, Jiashuo Sun, Yongliang Shen 0001, Weiming Lu 0001, Deli Zhao, Yueting Zhuang, Lidong Bing
ICCV5
2025 SVGenius: Benchmarking LLMs in SVG Understanding, Editing and Generation
abstract
Large Language Models (LLMs) and Multimodal LLMs have shown promising capabilities for SVG processing, yet existing benchmarks suffer from limited real-world coverage, lack of complexity stratification, and fragmented evaluation paradigms. We introduce SVGenius, a comprehensive benchmark comprising 2,377 queries across three progressive dimensions: understanding, editing, and generation. Built on real-world data from 24 application domains with systematic complexity stratification, SVGenius evaluates models through 8 task categories and 18 metrics. We assess 22 mainstream models spanning different scales, architectures, training paradigms, and accessibility levels. Our analysis reveals that while proprietary models significantly outperform open-source counterparts, all models exhibit systematic performance degradation with increasing complexity, indicating fundamental limitations in current approaches; however, reasoning-enhanced training proves more effective than pure scaling for overcoming these limitations, though style transfer remains the most challenging capability across all model types. SVGenius establishes the first systematic evaluation framework for SVG processing, providing crucial insights for developing more capable vector graphics models and advancing automated graphic design applications. Appendix and supplementary materials (including all data and code) are available at https://zju-real.github.io/SVGenius.
Haolei Xu, Fei Tang 0005, Linjuan Wu, Wenqi Zhang 0001, Guiyang Hou, Yongliang Shen 0001, Weiming Lu 0001, Yueting Zhuang
ACM Multimedia11
2025 EASYTOOL: Enhancing LLM-based Agents with Concise Tool Instruction
abstract
Siyu Yuan, Kaitao Song, Jiangjie Chen, Xu Tan, Yongliang Shen, Kan Ren, Dongsheng Li, Deqing Yang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Kaitao Song, Jiangjie Chen, Xu Tan 0003, Yongliang Shen 0001, Kan Ren, Dongsheng Li 0002, Deqing Yang
NAACL (Long Papers)5
2025 Chain-of-Model Learning for Language Model
abstract
In this paper, we propose a novel learning paradigm, termed *Chain-of-Model* (CoM), which incorporates the causal relationship into the hidden states of each layer as a chain style. thereby introducing great scaling efficiency in model training and inference flexibility in deployment.We introduce the concept of *Chain-of-Representation* (CoR), which formulates the hidden states at each layer as a combination of multiple sub-representations (i.e., chains). In each layer, each chain from the output representations can only view all of its preceding chains in the input representations. Consequently, the model built upon CoM framework can progressively scale up the model size by increasing the chains based on the previous models (i.e., chains), and offer multiple sub-models at varying sizes for elastic inference by using different chain numbers. Based on this principle, we devise *Chain-of-Language-Model* (CoLM), which incorporates the idea of CoM into each layer of Transformer architecture. Based on CoLM, we further introduce CoLM-Air by introducing a *KV sharing* mechanism, that computes all keys and values within the first chain and then shares across all chains. This design demonstrates additional extensibility, such as enabling seamless LM switching, prefilling acceleration and so on. Experimental results demonstrate our CoLM family can achieve comparable performance to the standard Transformer, while simultaneously enabling greater flexiblity, such as progressive scaling to improve training efficiency and offer multiple varying model sizes for elastic inference, paving a a new way toward building language models.
Kaitao Song, Xu Tan 0003, Huiqiang Jiang, Chengruidong Zhang, Yongliang Shen 0001, Cen Lu, Zihao Li 0006, Zifan Song, Yansen Wang, Kan Ren, Xiaoqing Zheng, Tao Qin 0001, Yuqing Yang 0001, Dongsheng Li 0002, Lili Qiu
NeurIPS6
2025 Mind the Gap: Bridging Thought Leap for Improved Chain-of-Thought Tuning
abstract
Large language models (LLMs) have achieved remarkable progress on mathematical tasks through Chain-of-Thought (CoT) reasoning. However, existing mathematical CoT datasets often suffer from **Thought Leaps** due to experts omitting intermediate steps, which negatively impacts model learning and generalization. We propose the CoT Thought Leap Bridge Task, which aims to automatically detect leaps and generate missing intermediate reasoning steps to restore the completeness and coherence of CoT. To facilitate this, we constructed a specialized training dataset called **ScaleQM+**, based on the structured ScaleQuestMath dataset, and trained **CoT-Bridge** to bridge thought leaps. Through comprehensive experiments on mathematical reasoning benchmarks, we demonstrate that models fine-tuned on bridged datasets consistently outperform those trained on original datasets, with improvements of up to +5.87\% on NuminaMath. Our approach effectively enhances distilled data (+3.02\%) and provides better starting points for reinforcement learning (+3.1\%), functioning as a plug-and-play module compatible with existing optimization techniques. Furthermore, CoT-Bridge demonstrates improved generalization to out-of-domain logical reasoning tasks, confirming that enhancing reasoning completeness yields broadly applicable benefits.
Haolei Xu, Yongliang Shen 0001, Wenqi Zhang 0001, Guiyang Hou, Shengpei Jiang, Kaitao Song, Weiming Lu 0001, Jun Xiao 0001, Yueting Zhuang
NeurIPS3
2025 Let LRMs Break Free from Overthinking via Self-Braking Tuning
abstract
Large reasoning models (LRMs), such as OpenAI o1 and DeepSeek-R1, have significantly enhanced their reasoning capabilities by generating longer chains of thought, demonstrating outstanding performance across a variety of tasks. However, this performance gain comes at the cost of a substantial increase in redundant reasoning during the generation process, leading to high computational overhead and exacerbating the issue of overthinking. Although numerous existing approaches aim to address the problem of overthinking, they often rely on external interventions. In this paper, we propose a novel framework, **Self-Braking Tuning**(SBT), which tackles overthinking from the perspective of allowing the model to regulate its own reasoning process, thus eliminating the reliance on external control mechanisms. We construct a set of overthinking identification metrics based on standard answers and design a systematic method to detect redundant reasoning. This method accurately identifies unnecessary steps within the reasoning trajectory and generates training signals for learning self-regulation behaviors. Building on this foundation, we develop a complete strategy for constructing data with adaptive reasoning lengths and introduce an innovative braking prompt mechanism that enables the model to naturally learn when to terminate reasoning at an appropriate point. Experiments across mathematical benchmarks (AIME, AMC, MATH500, GSM8K) demonstrate that our method reduces token consumption by up to 60\% while maintaining comparable accuracy to unconstrained models.
Yongliang Shen 0001, Haolei Xu, Wenqi Zhang 0001, Kaitao Song, Jian Shao 0001, Weiming Lu 0001, Jun Xiao 0001, Yueting Zhuang
NeurIPS3
2024 Learning Global Controller in Latent Space for Parameter-Efficient Fine-Tuning
abstract
Zeqi Tan, Yongliang Shen, Xiaoxia Cheng, Chang Zong, Wenqi Zhang, Jian Shao, Weiming Lu, Yueting Zhuang. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Zeqi Tan, Yongliang Shen 0001, Xiaoxia Cheng, Chang Zong, Wenqi Zhang 0001, Jian Shao 0001, Weiming Lu 0001, Yueting Zhuang
ACL (1)2
2024 Insert or Attach: Taxonomy Completion via Box Embedding
abstract
Taxonomy completion, enriching existing taxonomies by inserting new concepts as parents or attaching them as children, has gained significant interest.Previous approaches embed concepts as vectors in Euclidean space, which makes it difficult to model asymmetric relations in taxonomy.In addition, they introduce pseudo-leaves to convert attachment cases into insertion cases, leading to an incorrect bias in network learning dominated by numerous pseudo-leaves.Addressing these, our framework, TAXBOX, leverages box containment and center closeness to design two specialized geometric scorers within the box embedding space.These scorers are tailored for insertion and attachment operations and can effectively capture intrinsic relationships between concepts by optimizing on a granular box constraint loss.We employ a dynamic ranking loss mechanism to balance the scores from these scorers, allowing adaptive adjustments of insertion and attachment scores.Experiments on four real-world datasets show that TAXBOX significantly outperforms previous methods, yielding substantial improvements over prior methods in real-world datasets, with average performance boosts of 6.7%, 34.9%, and 51.4% in MRR, Hit@1, and Prec@1, respectively.
Yongliang Shen 0001, Wenqi Ren, Jietian Guo, Shiliang Pu, Weiming Lu 0001
ACL (1)2
2024 Self-Contrast: Better Reflection Through Inconsistent Solving Perspectives
abstract
Wenqi Zhang, Yongliang Shen, Linjuan Wu, Qiuying Peng, Jun Wang, Yueting Zhuang, Weiming Lu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Wenqi Zhang 0001, Yongliang Shen 0001, Linjuan Wu, Qiuying Peng, Yueting Zhuang, Weiming Lu 0001
ACL (1)2
2024 Agent-Pro: Learning to Evolve via Policy-Level Reflection and Optimization
abstract
Wenqi Zhang, Ke Tang, Hai Wu, Mengna Wang, Yongliang Shen, Guiyang Hou, Zeqi Tan, Peng Li, Yueting Zhuang, Weiming Lu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Wenqi Zhang 0001, Mengna Wang, Yongliang Shen 0001, Guiyang Hou, Zeqi Tan, Peng Li 0031, Yueting Zhuang, Weiming Lu 0001
ACL (1)5
2024 Advancing Process Verification for Large Language Models via Tree-Based Preference Learning
abstract
Large Language Models (LLMs) have demonstrated remarkable potential in handling complex reasoning tasks by generating step-by-step rationales.Some methods have proven effective in boosting accuracy by introducing extra verifiers to assess these paths.However, existing verifiers, typically trained on binarylabeled reasoning paths, fail to fully utilize the relative merits of intermediate steps, thereby limiting the effectiveness of the feedback provided.To overcome this limitation, we propose Tree-based Preference Learning Verifier (Tree-PLV), a novel approach that constructs reasoning trees via a best-first search algorithm and collects step-level paired data for preference training.Compared to traditional binary classification, step-level preferences more finely capture the nuances between reasoning steps, allowing for a more precise evaluation of the complete reasoning path.We empirically evaluate Tree-PLV across a range of arithmetic and commonsense reasoning tasks, where it significantly outperforms existing benchmarks.For instance, Tree-PLV achieved substantial performance gains over the Mistral-7B selfconsistency baseline on GSM8K (67.55% → 82.79%), MATH (17.00% → 26.80%), CSQA (68.14% → 72.97%), and StrategyQA (82.86% → 83.25%).Additionally, our study explores the appropriate granularity for applying preference learning, revealing that step-level guidance provides feedback that better aligns with the evaluation of the reasoning process.
Mingqian He, Yongliang Shen 0001, Wenqi Zhang 0001, Zeqi Tan, Weiming Lu 0001
EMNLP2
2024 Multimodal Self-Instruct: Synthetic Abstract Image and Visual Reasoning Instruction Using Language Model
abstract
Wenqi Zhang, Zhenglin Cheng, Yuanyu He, Mengna Wang, Yongliang Shen, Zeqi Tan, Guiyang Hou, Mingqian He, Yanna Ma, Weiming Lu, Yueting Zhuang. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Wenqi Zhang 0001, Zhenglin Cheng, Yuanyu He, Mengna Wang, Yongliang Shen 0001, Zeqi Tan, Guiyang Hou, Mingqian He, Yanna Ma, Weiming Lu 0001, Yueting Zhuang
EMNLP5
2024 TaskBench: Benchmarking Large Language Models for Task Automation
abstract
In recent years, the remarkable progress of large language models (LLMs) has sparked interest in task automation, which involves decomposing complex tasks described by user instructions into sub-tasks and invoking external tools to execute them, playing a central role in autonomous agents. However, there is a lack of systematic and standardized benchmarks to promote the development of LLMs in task automation. To address this, we introduce TaskBench, a comprehensive framework to evaluate the capability of LLMs in task automation. Specifically, task automation can be divided into three critical stages: task decomposition, tool selection, and parameter prediction. To tackle the complexities inherent in these stages, we introduce the concept of Tool Graph to represent decomposed tasks and adopt a back-instruct method to generate high-quality user instructions. We propose TaskEval, a multi-faceted evaluation methodology that assesses LLM performance across these three stages. Our approach combines automated construction with rigorous human verification, ensuring high consistency with human evaluation. Experimental results demonstrate that TaskBench effectively reflects the capabilities of various LLMs in task automation. It provides insights into model performance across different task complexities and domains, pushing the boundaries of what current models can achieve. TaskBench offers a scalable, adaptable, and reliable benchmark for advancing LLM-based autonomous agents.
Yongliang Shen 0001, Kaitao Song, Xu Tan 0003, Wenqi Zhang 0001, Kan Ren, Weiming Lu 0001, Dongsheng Li 0002, Yueting Zhuang
NeurIPS1
2024 Specialized Mathematical Solving by a Step-By-Step Expression Chain Generation
abstract
Math Solving requires both semantic understanding and relation reasoning. Most current approaches treat it as a translation task from natural language to mathematical symbols, generating tokens one by one. However, token-level generation is usually vulnerable when confronted with diverse annotations and complex reasoning. We consider the equation is an ordered combination of multiple sub-expressions, and math reasoning should be performed on the sub-expression level rather than the token level. We treat sub-expression as a minimum generative unit and minimum reasoning node. At each step, candidate sub-expression nodes are generated in parallel, and the whole reasoning chain is deduced by combining multiple nodes in order. Besides, we can obtain multiple valid reasoning chains by sub-expression searching, further improving interpretability and precision. Experiments on multilingual datasets show our method significantly outperforms the baselines. Additionally, our approach is more stable and efficient when faced with the challenges of diverse annotation and complex reasoning with limited resources. Moreover, our approach can be seamlessly integrated with large language models (LLMs), enhancing LLM's mathematical reasoning capabilities at a minimal cost. Experiments show a synergistic collaboration between general-purpose LLM and our specialized model yields superior performance.
Wenqi Zhang 0001, Yongliang Shen 0001, Guiyang Hou, Kuangyi Wang, Weiming Lu 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 DiffusionNER: Boundary Diffusion for Named Entity Recognition
abstract
In this paper, we propose DIFFUSIONNER, which formulates the named entity recognition task as a boundary-denoising diffusion process and thus generates named entities from noisy spans.During training, DIFFUSIONNER gradually adds noises to the golden entity boundaries by a fixed forward diffusion process and learns a reverse diffusion process to recover the entity boundaries.In inference, DIFFU-SIONNER first randomly samples some noisy spans from a standard Gaussian distribution and then generates the named entities by denoising them with the learned reverse diffusion process.The proposed boundary-denoising diffusion process allows progressive refinement and dynamic sampling of entities, empowering DIFFUSIONNER with efficient and flexible entity generation capability.Experiments on multiple flat and nested NER datasets demonstrate that DIFFUSIONNER achieves comparable or even better performance than previous state-of-the-art models 1 .
Yongliang Shen 0001, Kaitao Song, Xu Tan 0003, Dongsheng Li 0002, Weiming Lu 0001, Yueting Zhuang
ACL (1)1
2023 PromptNER: Prompt Locating and Typing for Named Entity Recognition
abstract
Yongliang Shen, Zeqi Tan, Shuhui Wu, Wenqi Zhang, Rongsheng Zhang, Yadong Xi, Weiming Lu, Yueting Zhuang. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yongliang Shen 0001, Zeqi Tan, Shuhui Wu, Wenqi Zhang 0001, Yadong Xi, Weiming Lu 0001, Yueting Zhuang
ACL (1)1
2023 MProto: Multi-Prototype Network with Denoised Optimal Transport for Distantly Supervised Named Entity Recognition
abstract
Distantly supervised named entity recognition (DS-NER) aims to locate entity mentions and classify their types with only knowledge bases or gazetteers and unlabeled corpus.However, distant annotations are noisy and degrade the performance of NER models.In this paper, we propose a noise-robust prototype network named MProto for the DS-NER task.Different from previous prototype-based NER methods, MProto represents each entity type with multiple prototypes to characterize the intra-class variance among entity representations.To optimize the classifier, each token should be assigned an appropriate ground-truth prototype and we consider such token-prototype assignment as an optimal transport (OT) problem.Furthermore, to mitigate the noise from incomplete labeling, we propose a novel denoised optimal transport (DOT) algorithm.Specifically, we utilize the assignment result between Other class tokens and all prototypes to distinguish unlabeled entity tokens from true negatives.Experiments on several DS-NER benchmarks demonstrate that our MProto achieves state-of-the-art performance.The source code is now available on Github 1 .
Shuhui Wu, Yongliang Shen 0001, Zeqi Tan, Wenqi Ren, Jietian Guo, Shiliang Pu, Weiming Lu 0001
EMNLP2
2023 An Expression Tree Decoding Strategy for Mathematical Equation Generation
abstract
Generating mathematical equations from natural language requires an accurate understanding of the relations among math expressions.Existing approaches can be broadly categorized into token-level and expression-level generation.The former treats equations as a mathematical language, sequentially generating math tokens.Expression-level methods generate each expression one by one.However, each expression represents a solving step, and there naturally exist parallel or dependent relations between these steps, which are ignored by current sequential methods.Therefore, we integrate tree structure into the expression-level generation and advocate an expression tree decoding strategy.To generate a tree with expression as its node, we employ a layer-wise parallel decoding strategy: we decode multiple independent expressions (leaf nodes) in parallel at each layer and repeat parallel decoding layer by layer to sequentially generate these parent node expressions that depend on others.Besides, a bipartite matching algorithm is adopted to align multiple predictions with annotations for each layer.Experiments show our method outperforms other baselines, especially for these equations with complex structures.
Wenqi Zhang 0001, Yongliang Shen 0001, Qingpeng Nong, Zeqi Tan, Yanna Ma, Weiming Lu 0001
EMNLP2
2023 HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face
abstract
Solving complicated AI tasks with different domains and modalities is a key step toward artificial general intelligence. While there are numerous AI models available for various domains and modalities, they cannot handle complicated AI tasks autonomously. Considering large language models (LLMs) have exhibited exceptional abilities in language understanding, generation, interaction, and reasoning, we advocate that LLMs could act as a controller to manage existing AI models to solve complicated AI tasks, with language serving as a generic interface to empower this. Based on this philosophy, we present HuggingGPT, an LLM-powered agent that leverages LLMs (e.g., ChatGPT) to connect various AI models in machine learning communities (e.g., Hugging Face) to solve AI tasks. Specifically, we use ChatGPT to conduct task planning when receiving a user request, select models according to their function descriptions available in Hugging Face, execute each subtask with the selected AI model, and summarize the response according to the execution results. By leveraging the strong language capability of ChatGPT and abundant AI models in Hugging Face, HuggingGPT can tackle a wide range of sophisticated AI tasks spanning different modalities and domains and achieve impressive results in language, vision, speech, and other challenging tasks, which paves a new way towards the realization of artificial general intelligence.
Yongliang Shen 0001, Kaitao Song, Xu Tan 0003, Dongsheng Li 0002, Weiming Lu 0001, Yueting Zhuang
NeurIPS1
2022 Parallel Instance Query Network for Named Entity Recognition
abstract
Yongliang Shen, Xiaobin Wang, Zeqi Tan, Guangwei Xu, Pengjun Xie, Fei Huang, Weiming Lu, Yueting Zhuang. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Yongliang Shen 0001, Xiaobin Wang, Zeqi Tan, Pengjun Xie, Fei Huang 0002, Weiming Lu 0001, Yueting Zhuang
ACL (1)1
2022 De-Bias for Generative Extraction in Unified NER Task
abstract
Named entity recognition (NER) is a fundamental task to recognize specific types of entities from a given sentence.Depending on how the entities appear in the sentence, it can be divided into three subtasks, namely, Flat NER, Nested NER, and Discontinuous NER.Among the existing approaches, only the generative model can be uniformly adapted to these three subtasks.However, when the generative model is applied to NER, its optimization objective is not consistent with the task, which makes the model vulnerable to the incorrect biases.In this paper, we analyze the incorrect biases in the generation process from a causality perspective and attribute them to two confounders: pre-context confounder and entityorder confounder.Furthermore, we design Intra-and Inter-entity Deconfounding Data Augmentation methods to eliminate the above confounders according to the theory of backdoor adjustment.Experiments show that our method can improve the performance of the generative NER model in various datasets.
Yongliang Shen 0001, Zeqi Tan, Yiquan Wu 0001, Weiming Lu 0001
ACL (1)2
2022 Molecular Substructure-Aware Network for Drug-Drug Interaction Prediction
abstract
Concomitant administration of drugs can cause drug-drug interactions (DDIs). Some drug combinations are beneficial, but other ones may cause negative effects which are previously unrecorded. Previous works on DDI prediction usually rely on hand-engineered domain knowledge, which is laborious to obtain. In this work, we propose a novel model, Molecular Substructure-Aware Network (MSAN), to effectively predict potential DDIs from molecular structures of drug pairs. We adopt a Transformer-like substructure extraction module to acquire a fixed number of representative vectors that are associated with various substructure patterns of the drug molecule. Then, interaction strength between the two drugs' substructures will be captured by a similarity-based interaction module. We also perform a substructure dropping augmentation before graph encoding to alleviate overfitting. Experimental results from a real-world dataset reveal that our proposed model achieves the state-of-the-art performance. We also show that the predictions of our model are highly interpretable through a case study.
Yongliang Shen 0001, Weiming Lu 0001
CIKM2
2022 Towards Data-Efficient Detection Transformers
Wen Wang 0009, Jing Zhang 0037, Yang Cao 0010, Yongliang Shen 0001, Dacheng Tao
ECCV (9)4
2022 Query-based Instance Discrimination Network for Relational Triple Extraction
abstract
Joint entity and relation extraction has been a core task in the field of information extraction.Recent approaches usually consider the extraction of relational triples from a stereoscopic perspective, either learning a relation-specific tagger or separate classifiers for each relation type.However, they still suffer from error propagation, relation redundancy and lack of highlevel connections between triples.To address these issues, we propose a novel query-based approach to construct instance-level representations for relational triples.By metric-based comparison between query embeddings and token embeddings, we can extract all types of triples in one step, thus eliminating the error propagation problem.In addition, we learn the instance-level representation of relational triples via contrastive learning.In this way, relational triples can not only enclose rich classlevel semantics but also access to high-order global connections.Experimental results show that our proposed method achieves the state of the art on five widely used benchmarks.
Zeqi Tan, Yongliang Shen 0001, Xuming Hu, Wenqi Zhang 0001, Xiaoxia Cheng, Weiming Lu 0001, Yueting Zhuang
EMNLP2
2022 Prompting to Distill: Boosting Data-Free Knowledge Distillation via Reinforced Prompt
abstract
Data-free knowledge distillation (DFKD) conducts knowledge distillation via eliminating the dependence of original training data, and has recently achieved impressive results in accelerating pre-trained language models. At the heart of DFKD is to reconstruct a synthetic dataset by inverting the parameters of the uncompressed model. Prior DFKD approaches, however, have largely relied on hand-crafted priors of the target data distribution for the reconstruction, which can be inevitably biased and often incompetent to capture the intrinsic distributions. To address this problem, we propose a prompt-based method, termed as PromptDFD, that allows us to take advantage of learned language priors, which effectively harmonizes the synthetic sentences to be semantically and grammatically correct. Specifically, PromptDFD leverages a pre-trained generative model to provide language priors and introduces a reinforced topic prompter to control data synthesis, making the generated samples thematically relevant and semantically plausible, and thus friendly to downstream tasks. As shown in our experiments, the proposed method substantially improves the synthesis quality and achieves considerable improvements on distillation performance. In some cases, PromptDFD even gives rise to results on par with those from the data-driven knowledge distillation with access to the original training data.
Xinyin Ma, Xinchao Wang, Gongfan Fang, Yongliang Shen 0001, Weiming Lu 0001
IJCAI4
2022 Propose-and-Refine: A Two-Stage Set Prediction Network for Nested Named Entity Recognition
abstract
Nested named entity recognition (nested NER) is a fundamental task in natural language processing. Various span-based methods have been proposed to detect nested entities with span representations. However, span-based methods do not consider the relationship between a span and other entities or phrases, which is helpful in the NER task. Besides, span-based methods have trouble predicting long entities due to limited span enumeration length. To mitigate these issues, we present the Propose-and-Refine Network (PnRNet), a two-stage set prediction network for nested NER. In the propose stage, we use a span-based predictor to generate some coarse entity predictions as entity proposals. In the refine stage, proposals interact with each other, and richer contextual information is incorporated into the proposal representations. The refined proposal representations are used to re-predict entity boundaries and classes. In this way, errors in coarse proposals can be eliminated, and the boundary prediction is no longer constrained by the span enumeration length limitation. Additionally, we build multi-scale sentence representations, which better model the hierarchical structure of sentences and provide richer contextual information than token-level representations. Experiments show that PnRNet achieves state-of-the-art performance on four nested NER datasets and one flat NER dataset.
Shuhui Wu, Yongliang Shen 0001, Zeqi Tan, Weiming Lu 0001
IJCAI2
2022 A Closed-Loop Perception, Decision-Making and Reasoning Mechanism for Human-Like Navigation
abstract
Reliable navigation systems have a wide range of applications in robotics and autonomous driving. Current approaches employ an open-loop process that converts sensor inputs directly into actions. However, these open-loop schemes are challenging to handle complex and dynamic real-world scenarios due to their poor generalization. Imitating human navigation, we add a reasoning process to convert actions back to internal latent states, forming a two-stage closed loop of perception, decision-making, and reasoning. Firstly, VAE-Enhanced Demonstration Learning endows the model with the understanding of basic navigation rules. Then, two dual processes in RL-Enhanced Interaction Learning generate reward feedback for each other and collectively enhance obstacle avoidance capability. The reasoning model can substantially promote generalization and robustness, and facilitate the deployment of the algorithm to real-world robots without elaborate transfers. Experiments show our method is more adaptable to novel scenarios compared with state-of-the-art approaches.
Wenqi Zhang 0001, Peng Li 0031, Yongliang Shen 0001, Yanna Ma, Weiming Lu 0001
IJCAI5
2021 Locate and Label: A Two-stage Identifier for Nested Named Entity Recognition
abstract
Yongliang Shen, Xinyin Ma, Zeqi Tan, Shuai Zhang, Wen Wang, Weiming Lu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Yongliang Shen 0001, Xinyin Ma, Zeqi Tan, Wen Wang 0009, Weiming Lu 0001
ACL/IJCNLP (1)1
2021 A Sequence-to-Set Network for Nested Named Entity Recognition
abstract
Named entity recognition (NER) is a widely studied task in natural language processing. Recently, a growing number of studies have focused on the nested NER. The span-based methods, considering the entity recognition as a span classification task, can deal with nested entities naturally. But they suffer from the huge search space and the lack of interactions between entities. To address these issues, we propose a novel sequence-to-set neural network for nested NER. Instead of specifying candidate spans in advance, we provide a fixed set of learnable vectors to learn the patterns of the valuable spans. We utilize a non-autoregressive decoder to predict the final set of entities in one pass, in which we are able to capture dependencies between entities. Compared with the sequence-to-sequence method, our model is more suitable for such unordered recognition task as it is insensitive to the label order. In addition, we utilize the loss function based on bipartite matching to compute the overall training loss. Experimental results show that our proposed model achieves state-of-the-art on three nested NER corpora: ACE 2004, ACE 2005 and KBP 2017. The code is available at https://github.com/zqtan1024/sequence-to-set.
Zeqi Tan, Yongliang Shen 0001, Weiming Lu 0001, Yueting Zhuang
IJCAI2
2021 Heterogeneous Graph Neural Networks for Concept Prerequisite Relation Learning in Educational Data
abstract
Chenghao Jia, Yongliang Shen, Yechun Tang, Lu Sun, Weiming Lu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Chenghao Jia, Yongliang Shen 0001, Yechun Tang, Weiming Lu 0001
NAACL-HLT2
2021 A Trigger-Sense Memory Flow Framework for Joint Entity and Relation Extraction
abstract
Joint entity and relation extraction framework constructs a unified model to perform entity recognition and relation extraction simultaneously, which can exploit the dependency between the two tasks to mitigate the error propagation problem suffered by the pipeline model. Current efforts on joint entity and relation extraction focus on enhancing the interaction between entity recognition and relation extraction through parameter sharing, joint decoding, or other ad-hoc tricks (e.g., modeled as a semi-Markov decision process, cast as a multi-round reading comprehension task). However, there are still two issues on the table. First, the interaction utilized by most methods is still weak and uni-directional, which is unable to model the mutual dependency between the two tasks. Second, relation triggers are ignored by most methods, which can help explain why humans would extract a relation in the sentence. They’re essential for relation extraction but overlooked. To this end, we present a Trigger-Sense Memory Flow Framework (TriMF) for joint entity and relation extraction. We build a memory module to remember category representations learned in entity recognition and relation extraction tasks. And based on it, we design a multi-level memory flow attention mechanism to enhance the bi-directional interaction between entity recognition and relation extraction. Moreover, without any human annotations, our model can enhance relation trigger information in a sentence through a trigger sensor module, which improves the model performance and makes model predictions with better interpretation. Experiment results show that our proposed framework achieves state-of-the-art results by improves the relation F1 to 52.44% (+3.2%) on SciERC, 66.49% (+4.9%) on ACE05, 72.35% (+0.6%) on CoNLL04 and 80.66% (+2.3%) on ADE.
Yongliang Shen 0001, Xinyin Ma, Yechun Tang, Weiming Lu 0001
WWW1
2020 Adversarial Self-Supervised Data-Free Distillation for Text Classification
abstract
Large pre-trained transformer-based language models have achieved impressive results on a wide range of NLP tasks.In the past few years, Knowledge Distillation(KD) has become a popular paradigm to compress a computationally expensive model to a resource-efficient lightweight model.However, most KD algorithms, especially in NLP, rely on the accessibility of the original training dataset, which may be unavailable due to privacy issues.To tackle this problem, we propose a novel twostage data-free distillation method, named Adversarial self-Supervised Data-Free Distillation (AS-DFD), which is designed for compressing large-scale transformer-based models (e.g., BERT).To avoid text generation in discrete space, we introduce a Plug & Play Embedding Guessing method to craft pseudo embeddings from the teacher's hidden knowledge.Meanwhile, with a self-supervised module to quantify the student's ability, we adapt the difficulty of pseudo embeddings in an adversarial training manner.To the best of our knowledge, our framework is the first data-free distillation framework designed for NLP tasks.We verify the effectiveness of our method on several text classification datasets.
Xinyin Ma, Yongliang Shen 0001, Gongfan Fang, Chenghao Jia, Weiming Lu 0001
EMNLP (1)2
2020 Multi-hop Reading Comprehension across Documents with Path-based Graph Convolutional Network
abstract
Multi-hop reading comprehension across multiple documents attracts much attentions recently. In this paper, we propose a novel approach to tackle this multi-hop reading comprehension problem. Inspired by the human reasoning processing, we introduce a path-based graph with reasoning paths which extracted from supporting documents. The path-based graph can combine both the idea of the graph-based and path-based approaches, so it is better for multi-hop reasoning. Meanwhile, we propose Gated-GCN to accumulate evidences on the path-based graph, which contains a new question-aware gating mechanism to regulate the usefulness of information propagating across documents and add question information during reasoning. We evaluate our approach on WikiHop dataset, and our approach achieves the the-state-of-art accuracy against previous published approaches. Especially, our ensemble model surpasses the human performance by 4.2%.
Zeyun Tang, Yongliang Shen 0001, Xinyin Ma, Jiale Yu, Weiming Lu 0001
IJCAI2
2020 Boosting Cross-lingual Entity Alignment with Textual Embedding
Chenghao Jia, Yongliang Shen 0001, Xinyin Ma, Weiming Lu 0001
NLPCC (2)4