Xiaxia Wang 0001

dblp:126/2338-1 · DBLP profile ↗
← Back
11ranked-venue papers
6as first author
9since 2021 · last 2026
0000-0003-4184-0754ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 7 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Caddie: A prototype of content-based ad hoc RDF dataset retrieval
abstract
The rapid growth of open and structured RDF data on the Web has promoted the development of dataset search as an important research topic. The core function of existing systems is ad hoc dataset retrieval (AHDR) based on the metadata of datasets, which contains limited information and often suffers from quality issues. To overcome the limitations, in this article, we systematically investigate content-based AHDR to exploit the actual RDF data in datasets. We address three main tasks of content-based AHDR with novel methods for handling the large size and complex structure of RDF data to facilitate dataset retrieval, deduplication, and snippet extraction. These methods are integrated into an online and open-source prototype called Caddie . The effectiveness and practicability of its components are evaluated on a public test collection and by a user study.
Xiaxia Wang 0001, Qiaosheng Chen, Weiqing Luo, Jeff Z. Pan, Baifan Zhou, Evgeny Kharlamov, Gong Cheng 0001
J. Web Semant.1
2025 TARGA: Targeted Synthetic Data Generation for Practical Reasoning over Structured Data
abstract
Semantic parsing, which converts natural language questions into logic forms, plays a crucial role in reasoning within structured environments.However, existing methods encounter two significant challenges: reliance on extensive manually annotated datasets and limited generalization capability to unseen examples.To tackle these issues, we propose Targeted Synthetic Data Generation (TARGA), a practical framework that dynamically generates high-relevance synthetic data without manual annotation.Starting from the pertinent entities and relations of a given question, we probe for the potential relevant queries through layer-wise expansion and cross-layer combination.Then we generate corresponding natural language questions for these constructed queries to jointly serve as the synthetic demonstrations for in-context learning.Experiments on multiple knowledge base question answering (KBQA) datasets demonstrate that TARGA, using only a 7B-parameter model, substantially outperforms existing non-finetuned methods that utilize close-sourced model, achieving notable improvements in F1 scores on GrailQA (+7.7) and KBQA-Agent (+12.2).Furthermore, TARGA also exhibits superior sample efficiency, robustness, and generalization capabilities under non-I.I.D. settings.
Xiang Huang 0007, Jiayu Shen, Sitao Cheng, Xiaxia Wang 0001, Yuzhong Qu
ACL (1)5
2025 Transforming decoder-only models into encoder-only models with improved understanding capabilities
Zixian Huang, Xinwei Huang, Ao Wu, Xiaxia Wang 0001, Gong Cheng 0001
Knowl. Based Syst.4
2024 Faithful Rule Extraction for Differentiable Rule Learning Models
abstract
There is increasing interest in methods for extracting interpretable rules from ML models trained to solve a wide range of tasks over knowledge graphs (KGs), such as KG completion, node classification, question answering and recommendation. Many such approaches, however, lack formal guarantees establishing the precise relationship between the model and the extracted rules, and this lack of assurance becomes especially problematic when the extracted rules are applied in safety-critical contexts or to ensure compliance with legal requirements. Recent research has examined whether the rules derived from the influential Neural-LP model exhibit soundness (or completeness), which means that the results obtained by applying the model to any dataset always contain (or are contained in) the results obtained by applying the rules to the same dataset. In this paper, we extend this analysis to the context of DRUM, an approach that has demonstrated superior practical performance. After observing that the rules currently extracted from a DRUM model can be unsound and/or incomplete, we propose a novel algorithm where the output rules, expressed in an extension of Datalog, ensure both soundness and completeness. This algorithm, however, can be inefficient in practice and hence we propose additional constraints to DRUM models facilitating rule extraction, albeit at the expense of reduced expressive power.
Xiaxia Wang 0001, David Tena Cucala, Bernardo Cuenca Grau, Ian Horrocks 0001
ICLR1
2024 A Survey on Extractive Knowledge Graph Summarization: Applications, Approaches, Evaluation, and Future Directions
Xiaxia Wang 0001, Gong Cheng 0001
IJCAI1
2024 ACORDAR 2.0: A Test Collection for Ad Hoc Dataset Retrieval with Densely Pooled Datasets and Question-Style Queries
abstract
Dataset search, or more specifically, ad hoc dataset retrieval which is a trending specialized IR task, has received increasing attention in both academia and industry. While methods and systems continue evolving, existing test collections for this task exhibit shortcomings, particularly suffering from lexical bias in pooling and limited to keyword-style queries for evaluation. To address these limitations, in this paper, we construct ACORDAR 2.0, a new test collection for this task which is also the largest to date. To reduce lexical bias in pooling, we adapt dense retrieval models to large structured data, using them to find an extended set of semantically relevant datasets to be annotated. To diversify query forms, we employ a large language model to rewrite keyword queries into high-quality question-style queries. We use the test collection to evaluate popular sparse and dense retrieval models to establish a baseline for future studies. The test collection and source code are publicly available.
Qiaosheng Chen, Weiqing Luo, Zixian Huang, Tengteng Lin, Xiaxia Wang 0001, Ahmet Soylu, Basil Ell, Baifan Zhou, Evgeny Kharlamov, Gong Cheng 0001
SIGIR5
2023 BANDAR: Benchmarking Snippet Generation Algorithms for (RDF) Dataset Search
abstract
The large volume of open data on the Web is expected to be reused and create value. Finding the right data to reuse is a non-trivial task addressed by the recent dataset search systems, which retrieve datasets relevant to a keyword query. An important component of such systems is snippet generation, extracting data from a retrieved dataset to exemplify its content and explain its relevance to the query. Snippet generation algorithms have emerged but were mainly evaluated by user studies. More efficient and reproducible evaluation methods are needed. To meet this challenge, in this article, we present a set of quality metrics for assessing the usefulness of a snippet from different perspectives, and we select and aggregate them into quality profiles for different stages of a dataset search process. Furthermore, we create a benchmark from thousands of collected real-world data needs and datasets, on which we apply the presented quality metrics and profiles to evaluate snippets generated by two existing algorithms and three adapted algorithms. The results, which are reproducible as they are automatically computed without human interaction, show the pros and cons of the tested algorithms and highlight directions for future research. The benchmark data is publicly available.
Xiaxia Wang 0001, Gong Cheng 0001, Jeff Z. Pan, Evgeny Kharlamov, Yuzhong Qu
IEEE Trans. Knowl. Data Eng.1
2022 ACORDAR: A Test Collection for Ad Hoc Content-Based (RDF) Dataset Retrieval
abstract
Ad hoc dataset retrieval is a trending topic in IR research. Methods and systems are evolving from metadata-based to content-based ones which exploit the data itself for improving retrieval accuracy but thus far lack a specialized test collection. In this paper, we build and release the first test collection for ad hoc content-based dataset retrieval, where content-oriented dataset queries and content-based relevance judgments are annotated by human experts who are assisted with a dashboard designed specifically for comprehensively and conveniently browsing both the metadata and data of a dataset. We conduct extensive experiments on the test collection to analyze its difficulty and provide insights into the underlying task.
Tengteng Lin, Qiaosheng Chen, Gong Cheng 0001, Ahmet Soylu, Basil Ell, Ruoqi Zhao, Xiaxia Wang 0001, Yu Gu 0016, Evgeny Kharlamov
SIGIR8
2021 PCSG: Pattern-Coverage Snippet Generation for RDF Datasets
Xiaxia Wang 0001, Gong Cheng 0001, Tengteng Lin, Jeff Z. Pan, Evgeny Kharlamov, Yuzhong Qu
ISWC1
2019 Towards More Usable Dataset Search: From Query Characterization to Snippet Generation
abstract
Reusing published datasets on the Web is of great interest to researchers and developers. Their data needs may be met by submitting queries to a dataset search engine to retrieve relevant datasets. In this ongoing work towards developing a more usable dataset search engine, we characterize real data needs by annotating the semantics of 1,947 queries using a novel fine-grained scheme, to provide implications for enhancing dataset search. Based on the findings, we present a query-centered framework for dataset search, and explore the implementation of snippet generation and evaluate it with a preliminary user study.
Jinchi Chen, Xiaxia Wang 0001, Gong Cheng 0001, Evgeny Kharlamov, Yuzhong Qu
CIKM2
2019 A Framework for Evaluating Snippet Generation for Dataset Search
Xiaxia Wang 0001, Jinchi Chen, Gong Cheng 0001, Jeff Z. Pan, Evgeny Kharlamov, Yuzhong Qu
ISWC (1)1