Xiaxia Wang 0001

dblp:126/2338-1 · DBLP profile ↗
← Back
7ranked-venue papers in the field
4as first author
5since 2021 · last 2026
0000-0003-4184-0754ORCID · verified

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 3Knowledge Engineering, Semantic Web & Information Systems · 3 (3 first)Database Systems & Data Management · 1 (1 first)
YearPublicationVenuePosition
2026 Caddie: A prototype of content-based ad hoc RDF dataset retrieval
abstract
The rapid growth of open and structured RDF data on the Web has promoted the development of dataset search as an important research topic. The core function of existing systems is ad hoc dataset retrieval (AHDR) based on the metadata of datasets, which contains limited information and often suffers from quality issues. To overcome the limitations, in this article, we systematically investigate content-based AHDR to exploit the actual RDF data in datasets. We address three main tasks of content-based AHDR with novel methods for handling the large size and complex structure of RDF data to facilitate dataset retrieval, deduplication, and snippet extraction. These methods are integrated into an online and open-source prototype called Caddie . The effectiveness and practicability of its components are evaluated on a public test collection and by a user study.
Xiaxia Wang 0001, Qiaosheng Chen, Weiqing Luo, Jeff Z. Pan, Baifan Zhou, Evgeny Kharlamov, Gong Cheng 0001
J. Web Semant.1
2024 ACORDAR 2.0: A Test Collection for Ad Hoc Dataset Retrieval with Densely Pooled Datasets and Question-Style Queries
abstract
Dataset search, or more specifically, ad hoc dataset retrieval which is a trending specialized IR task, has received increasing attention in both academia and industry. While methods and systems continue evolving, existing test collections for this task exhibit shortcomings, particularly suffering from lexical bias in pooling and limited to keyword-style queries for evaluation. To address these limitations, in this paper, we construct ACORDAR 2.0, a new test collection for this task which is also the largest to date. To reduce lexical bias in pooling, we adapt dense retrieval models to large structured data, using them to find an extended set of semantically relevant datasets to be annotated. To diversify query forms, we employ a large language model to rewrite keyword queries into high-quality question-style queries. We use the test collection to evaluate popular sparse and dense retrieval models to establish a baseline for future studies. The test collection and source code are publicly available.
Qiaosheng Chen, Weiqing Luo, Zixian Huang, Tengteng Lin, Xiaxia Wang 0001, Ahmet Soylu, Basil Ell, Baifan Zhou, Evgeny Kharlamov, Gong Cheng 0001
SIGIR5
2023 BANDAR: Benchmarking Snippet Generation Algorithms for (RDF) Dataset Search
abstract
The large volume of open data on the Web is expected to be reused and create value. Finding the right data to reuse is a non-trivial task addressed by the recent dataset search systems, which retrieve datasets relevant to a keyword query. An important component of such systems is snippet generation, extracting data from a retrieved dataset to exemplify its content and explain its relevance to the query. Snippet generation algorithms have emerged but were mainly evaluated by user studies. More efficient and reproducible evaluation methods are needed. To meet this challenge, in this article, we present a set of quality metrics for assessing the usefulness of a snippet from different perspectives, and we select and aggregate them into quality profiles for different stages of a dataset search process. Furthermore, we create a benchmark from thousands of collected real-world data needs and datasets, on which we apply the presented quality metrics and profiles to evaluate snippets generated by two existing algorithms and three adapted algorithms. The results, which are reproducible as they are automatically computed without human interaction, show the pros and cons of the tested algorithms and highlight directions for future research. The benchmark data is publicly available.
Xiaxia Wang 0001, Gong Cheng 0001, Jeff Z. Pan, Evgeny Kharlamov, Yuzhong Qu
IEEE Trans. Knowl. Data Eng.1
2022 ACORDAR: A Test Collection for Ad Hoc Content-Based (RDF) Dataset Retrieval
abstract
Ad hoc dataset retrieval is a trending topic in IR research. Methods and systems are evolving from metadata-based to content-based ones which exploit the data itself for improving retrieval accuracy but thus far lack a specialized test collection. In this paper, we build and release the first test collection for ad hoc content-based dataset retrieval, where content-oriented dataset queries and content-based relevance judgments are annotated by human experts who are assisted with a dashboard designed specifically for comprehensively and conveniently browsing both the metadata and data of a dataset. We conduct extensive experiments on the test collection to analyze its difficulty and provide insights into the underlying task.
Tengteng Lin, Qiaosheng Chen, Gong Cheng 0001, Ahmet Soylu, Basil Ell, Ruoqi Zhao, Xiaxia Wang 0001, Yu Gu 0016, Evgeny Kharlamov
SIGIR8
2021 PCSG: Pattern-Coverage Snippet Generation for RDF Datasets
Xiaxia Wang 0001, Gong Cheng 0001, Tengteng Lin, Jeff Z. Pan, Evgeny Kharlamov, Yuzhong Qu
ISWC1
2019 Towards More Usable Dataset Search: From Query Characterization to Snippet Generation
abstract
Reusing published datasets on the Web is of great interest to researchers and developers. Their data needs may be met by submitting queries to a dataset search engine to retrieve relevant datasets. In this ongoing work towards developing a more usable dataset search engine, we characterize real data needs by annotating the semantics of 1,947 queries using a novel fine-grained scheme, to provide implications for enhancing dataset search. Based on the findings, we present a query-centered framework for dataset search, and explore the implementation of snippet generation and evaluate it with a preliminary user study.
Jinchi Chen, Xiaxia Wang 0001, Gong Cheng 0001, Evgeny Kharlamov, Yuzhong Qu
CIKM2
2019 A Framework for Evaluating Snippet Generation for Dataset Search
Xiaxia Wang 0001, Jinchi Chen, Gong Cheng 0001, Jeff Z. Pan, Evgeny Kharlamov, Yuzhong Qu
ISWC (1)1