EDBT 2026 Demo / reviewers in the wild / expert
Chenhao Xie 0002
dblp:175/5479-2
· DBLP profile ↗
13ranked-venue papers
3as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Gen-SQL: Efficient Text-to-SQL By Bridging Natural Language Question And Database Schema With Pseudo-SchemaabstractWith the prevalence of Large Language Models (LLMs), recent studies have shifted paradigms and leveraged LLMs to tackle the challenging task of Text-to-SQL. Because of the complexity of real world databases, previous works adopt the retrieve-then-generate framework to retrieve relevant database schema and then to generate the SQL query. However, efficient embedding-based retriever suffers from lower retrieval accuracy, and more accurate LLM-based retriever is far more expensive to use, which hinders their applicability for broader applications. To overcome this issue, this paper proposes Gen-SQL, a novel generate-ground-regenerate framework, where we exploit prior knowledge from the LLM to enhance embedding-based retriever and reduce cost. Experiments on several datasets are conducted to demonstrate the effectiveness and scalability of our proposed method. We release our code and data at https://github.com/jieshi10/gensql. Jie Shi 0010, Bo Xu 0023, Jiaqing Liang, Yanghua Xiao, Jia Chen 0037, Chenhao Xie 0002, Peng Wang 0027, Wei Wang 0009 |
COLING | 6 |
| 2024 | Enhancing Chinese abbreviation prediction with LLM generation and contrastive evaluation
Xianyang Tian, Hanwen Tong, Chenhao Xie 0002, Tong Ruan, Baohua Wu, Haofen Wang |
Inf. Process. Manag. | 4 |
| 2022 | A Context-Enhanced Generate-then-Evaluate Framework for Chinese Abbreviation PredictionabstractAs a popular form of lexicalization, abbreviation is widely used in both oral and written language and plays an important role in various Natural Language Processing applications. However, current approaches cannot ensure that the predicted abbreviation preserves the meaning of its full form and maintains fluency. In this paper, we introduce a fresh perspective to evaluate the quality of abbreviations within their textual contexts with pre-trained language model. To this end, we propose a novel two-stage generate-then-evaluate framework enhanced by context, which consists of a generation model to generate multiple candidate abbreviations and an evaluation model to evaluate their quality within their contexts. Experimental results show that our framework consistently outperforms all the existing approaches, achieving 53.2% [email protected] performance with a 5.6 points improvement compared to its previous best result. Our code and data are publicly available at https://github.com/HavenTong/CEGE. Hanwen Tong, Chenhao Xie 0002, Jiaqing Liang, Qianyu He, Zhiang Yue, Yanghua Xiao |
CIKM | 2 |
| 2022 | Parsing Natural Language into Propositional and First-Order Logic with Dual Reinforcement LearningabstractSemantic parsing converts natural language utterances into structured logical expressions. We consider two such formal representations: Propositional Logic (PL) and First-order Logic (FOL). The paucity of labeled data is a major challenge in this field. In previous works, dual reinforcement learning has been proposed as an approach to reduce dependence on labeled data. However, this method has the following limitations: 1) The reward needs to be set manually and is not applicable to all kinds of logical expressions. 2) The training process easily collapses when models are trained with only the reward from dual reinforcement learning. In this paper, we propose a scoring model to automatically learn a model-based reward, and an effective training strategy based on curriculum learning is further proposed to stabilize the training process. In addition to the technical contribution, a Chinese-PL/FOL dataset is constructed to compensate for the paucity of labeled data in this field. Experimental results show that the proposed method outperforms competitors on several datasets. Furthermore, by introducing PL/FOL generated by our model, the performance of existing Natural Language Inference (NLI) models is further enhanced. Xuantao Lu, Zhouhong Gu, Hanwen Tong, Chenhao Xie 0002, Junyang Huang, Yanghua Xiao |
COLING | 5 |
| 2022 | Rule mining over knowledge graphs via reinforcement learning
Sihang Jiang 0001, Chao Wang 0095, Sheng Zhang 0027, Chenhao Xie 0002, Jiaqing Liang, Yanghua Xiao, Rui Song 0006 |
Knowl. Based Syst. | 6 |
| 2021 | Revisiting the Negative Data of Distantly Supervised Relation ExtractionabstractChenhao Xie, Jiaqing Liang, Jingping Liu, Chengsong Huang, Wenhao Huang, Yanghua Xiao. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Chenhao Xie 0002, Jiaqing Liang, Chengsong Huang, Yanghua Xiao |
ACL/IJCNLP (1) | 1 |
| 2021 | WebKE: Knowledge Extraction from Semi-structured Web with Pre-trained Markup Language ModelabstractThe World Wide Web contains rich up-to-date information for knowledge graph construction. However, most current relation extraction techniques are designed for free text and thus do not handle well semi-structured web content. In this paper, we propose a novel multi-phase machine reading framework, called WebKE. It processes the web content on different granularity by first detecting areas of interest at DOM tree node level and then extracting relational triples for each area. We also propose HTMLBERT as an encoder the web content. It is a pre-trained markup language model that fully leverages the visual layout information and DOM-tree structure, without the need of hand engineered features. Experimental results show that the proposed approach outperforms state-of- the-art methods by a considerable gain. The source code is available at https://github.com/redreamality/webke. Chenhao Xie 0002, Jiaqing Liang, Chengsong Huang, Yanghua Xiao |
CIKM | 1 |
| 2021 | Bootstrapping Information Extraction via ConceptualizationabstractBootstrapping enables us to use existing knowledge to find patterns and extract new knowledge from free texts, from which more patterns can be found. Due to its minimally supervised, domain-independent, and language-independent nature, it has been widely adopted in real-world applications. However, as iterations go on, semantic drift may happen. The extraction may shift from the target class to other classes and result in errors, which propagate in the succeeding iterations and hurt the performance significantly. Existing solutions simply throw away bad patterns, sacrificing recall to ensure high precision. However, we argue that most of these patterns and instances can be kept as long as being applied selectively, guided by prior knowledge. In this paper, we propose a pattern-based extraction framework with three distinguished features: (1) it uses conceptual taxonomies to guide the extraction to reduce semantic drift; (2) it uses the knowledge of existing triples to improve the precision; (3) it integrates all patterns to form a generalized pattern set with quantified confidence measurement. The proposed solution is applied on enriching two real-world knowledge bases and achieves higher precision and recall compared to existing solutions. Jiaqing Liang, Suo Feng, Chenhao Xie 0002, Yanghua Xiao, Jindong Chen, Seung-won Hwang |
ICDE | 3 |
| 2018 | Short Text Entity Linking with Fine-grained TopicsabstractA wide range of web corpora are in the form of short text, such as QA queries, search queries and news titles. Entity linking for these short texts is quite important. Most of supervised approaches are not effective for short text entity linking. The training data for supervised approaches are not suitable for short text and insufficient for low-resourced languages. Previous unsupervised methods are incapable of handling the sparsity and noisy problem of short text. We try to solve the problem by mapping the sparse short text to a topic space. We notice that the concepts of entities have rich topic information and characterize entities in a very fine-grained granularity. Hence, we use the concepts of entities as topics to explicitly represent the context, which helps improve the performance of entity linking for short text. We leverage our linking approach to segment the short text semantically, and build a system for short entity text recognition and linking. Our entity linking approach exhibits the state-of-the-art performance on several datasets for the realistic short text entity linking problem. Jiaqing Liang, Chenhao Xie 0002, Yanghua Xiao |
CIKM | 3 |
| 2018 | Timeline: A Chinese Event Extraction and Exploration SystemabstractEvent extraction plays a significant role in information extraction (IE). Compared with previous works on event extraction in English, relatively little effort has been made to extract Chinese events. Existing Chinese event extraction systems have two main drawbacks. First, they can only extract a limited number of events. Second, they don't organize their extracted events by entities or date or demonstrate them in a user-friendly way. In this paper, we propose Timeline, a Chinese event extraction system that extracts massive events from Chinese online encyclopedias . Our proposed system extracts event triples (entity, date and event description) from huge numbers of articles meanwhile generating the largest Chinese structured event base. Our system also automatizes event validation and normalization procedures to harvest high-quality events, and then organize events by corresponding entities and dates. Furthermore, we have also developed an interactive web portal that encodes events along a visual timeline, which satisfies the process of exploring historical events. We also designed comprehensive experiments to show the effectiveness of our extraction and validation work. Both extracted event triples and timeline web portal are published. Yanghua Xiao, Chenhao Xie 0002, Haiyun Jiang, Suo Feng |
SoMeT | 4 |
| 2017 | Automatic Navbox Generation by Interpretable Clustering over Linked EntitiesabstractRare efforts have been devoted to generating the structured Navigation Box (Navbox) for Wikipedia articles. A Navbox is a table in Wikipedia article page that provides a consistent navigation system for related entities. Navbox is critical for the readership and editing efficiency of Wikipedia. In this paper, we target on the automatic generation of Navbox for Wikipedia articles. Instead of performing information extraction over unstructured natural language text directly, an alternative avenue is explored by focusing on a rich set of semi-structured data in Wikipedia articles: linked entities. The core idea of this paper is as follows: If we cluster the linked entities and interpret them appropriately, we can construct a high-quality Navbox for the article entity. We propose a clustering-then-labeling algorithm to realize the idea. Experiments show that the proposed solutions are effective. Ultimately, our approach enriches Wikipedia with 1.95 million new Navboxes of high quality. Chenhao Xie 0002, Jiaqing Liang, Kezun Zhang, Yanghua Xiao, Hanghang Tong, Haixun Wang, Wei Wang 0009 |
CIKM | 1 |
| 2017 | CN-DBpedia: A Never-Ending Chinese Knowledge Extraction System
Bo Xu 0023, Jiaqing Liang, Chenhao Xie 0002, Wanyun Cui, Yanghua Xiao |
IEA/AIE (2) | 4 |
| 2016 | Learning Defining Features for Categories
Bo Xu 0023, Chenhao Xie 0002, Yanghua Xiao, Haixun Wang, Wei Wang 0009 |
IJCAI | 2 |