Xiao Zhou 0009

dblp:267/2864-9 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
0009-0008-1132-6408ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2025 Benchmarking Recommendation, Classification, and Tracing Based on Hugging Face Knowledge Graph
abstract
The rapid growth of open source machine learning (ML) resources, such as models and datasets, has accelerated IR research. However, existing platforms like Hugging Face do not explicitly utilize structured representations, limiting advanced queries and analyses such as tracing model evolution and recommending relevant datasets. To fill the gap, we construct HuggingKG, the first large-scale knowledge graph built from the Hugging Face community for ML resource management. With 2.6 million nodes and 6.2 million edges, HuggingKG captures domain-specific relations and rich textual attributes. It enables us to further present HuggingBench, a multi-task benchmark with three novel test collections for IR tasks including resource recommendation, classification, and tracing. Our experiments reveal unique characteristics of HuggingKG and the derived tasks. Both resources are publicly available, expected to advance research in open source resource sharing and management.
Qiaosheng Chen, Kaijia Huang, Xiao Zhou 0009, Weiqing Luo, Yuanning Cui, Gong Cheng 0001
SIGIR3
2025 μDS: Multi-Objective Data Snippet Extraction for Dataset Search
abstract
With the continuous growth of open data on the Web, dataset search has become a prominent specialized retrieval problem to find datasets relevant to a query. Recent solutions rank datasets based on not only their metadata, but also data snippets extracted from their actual data. While the goodness of a data snippet has been studied from various aspects, in this paper we propose to, for the first time, jointly optimize compactness, relevance, representativeness, and cohesiveness in snippet extraction. To extract such multi-objective data snippets, we formulate a new combinatorial optimization problem and design an efficient algorithm with a proved worst-case approximation ratio. We evaluate the data snippets extracted by our algorithm intrinsically through a set of quality metrics and extrinsically by applying them to dataset search.
Xiao Zhou 0009, Qiaosheng Chen, Jiageng Chen, Gong Cheng 0001
SIGIR1
2024 DUNKS: Chunking and Summarizing Large and Heterogeneous Data for Dataset Search
Qiaosheng Chen, Xiao Zhou 0009, Gong Cheng 0001
ISWC (2)2
2024 Enhancing Dataset Search with Compact Data Snippets
abstract
In light of the growing availability and significance of open data, the problem of dataset search has attracted great attention in the field of information retrieval. Nevertheless, current metadata-based approaches have revealed shortcomings due to the low quality and availability of dataset metadata, while the magnitude and heterogeneity of actual data hindered the development of content-based solutions. To address these challenges, we propose to convert different formats of structured data into a unified form, from which we extract a compact data snippet that indicates the relevance of the whole data. Thanks to its compactness, we feed it into a dense reranker to improve search accuracy. We also convert it back to the original format to be presented for assisting users in relevance judgment. The effectiveness of our approach has been demonstrated by extensive experiments on two test collections for dataset search.
Qiaosheng Chen, Jiageng Chen, Xiao Zhou 0009, Gong Cheng 0001
SIGIR3