Guorui Xiao

dblp:77/10844 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0003-1440-0446ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 KathDB: Explainable Multimodal Database Management System with Human-AI Collaboration
Guorui Xiao, Enhao Zhang 0001, Nicole Sullivan, Will Hansen, Magdalena Balazinska
CIDR1
2025 CENTS: A Flexible and Cost-Effective Framework for LLM-Based Table Understanding
abstract
Large Language Models (LLMs) have recently shown impressive capabilities in a variety of applications including table understanding tasks such as column type annotation. Existing LLM-based solutions for table understanding, however, focus on developing specific framework for each individual task, or do not consider the cost-effectiveness tradeoff. In this paper, we present Cents, a unified and cost-effective framework for LLM-based solutions for table understanding tasks. Cents's key capability is an efficient and effective approach to compress the tabular LLM input in a way that reduces input token cost while improving performance compared with state-of-the-art methods. Experiment results show that Cents outperforms other LLM-based baselines on a variety of table understanding tasks at the same or lower cost.
Guorui Xiao, Dong He 0002, Jin Wang 0007, Magdalena Balazinska
Proc. VLDB Endow.1
2024 Real-Time Precise Zenith Tropospheric Delay Estimation With BDS PPP-B2b, Galileo HAS, and QZSS MADOCA-PPP Services
abstract
Real-time precise point positioning (PPP) is an innovative and useful atmosphere remote sensing tool to estimate zenith tropospheric delays (ZTDs) to support various emerging meteorological applications such as weather and climate research. However, additional costs on internet accesses or service subscription licenses are required to obtain real-time precise satellite orbit and clock products from external PPP augmentation service operators. We present the real-time precise estimation of ZTD with freely and openly accessible GNSS satellite-based PPP augmentation services, i.e., BeiDou navigation satellite system (BDS) PPP-B2b service, Galileo high accuracy service (HAS), and quasi-zenith satellite system (QZSS) Multi-GNSS ADvanced Orbit and Clock Augmentation-PPP (MADOCA-PPP) service. GNSS observations of 49 days from 15 stations in the Asia-Pacific region are used for performance assessment and comparison. It shows that the real-time PPP estimated ZTDs agree well with International GNSS Service (IGS) final ZTDs, with mean standard deviation (STD) values of 20.0, 17.5, and 9.5 mm for BDS PPP-B2b, Galileo HAS, and QZSS MADOCA-PPP, respectively. The derived real-time zenith wet delays (ZWDs) are further compared with those derived from Integrated Global Radiosonde Archive (IGRA) and the European Centre for Medium-Range Weather Forecasts (ECMWF) Reanalysis v5 (ERA5) datasets, which confirms that the real-time PPP estimated ZWDs are accurate with mean STDs of smaller than 30 mm. It concludes that real-time PPP with GNSS satellite-based real-time PPP augmentation correction services will be a flexible and cost-effective tool to derive real-time precise ZTDs/ZWDs with high spatiotemporal resolution in all-weather conditions.
Peiyuan Zhou, Zejun Liu, Daqian Lyu, Guorui Xiao
IEEE Trans. Geosci. Remote. Sens.5
2022 Demonstration of LogicLib: An Expressive Multi-Language Interface over Scalable Datalog System
abstract
With the ever-increasing volume of data, there is an urgent need to provide expressive and efficient tools to support Big Data analytics. The declarative logical language Datalog has proven very effective at expressing concisely graph, machine learning, and knowledge discovery applications via recursive queries. In this demonstration, we develop Logic Library (LLib), a library of recursive algorithms written in Datalog that can be executed in BigDatalog, a Datalog engine on top of Apache Spark developed by us. LLib encapsulates complex logic-based algorithms into high-level APIs, which simplify the development and provide a unified interface akin to the one of Spark MLlib. As LLib is fully compatible with DataFrame, it enables the integrated utilization of its built-in applications and new Datalog queries with existing Spark functions, such as those provided by \mlib and Spark SQL. With a variety of examples, we will (i) show how to write programs with LLib to express a variety of applications; (ii) illustrate its user experience in Apache Spark ecosystem; and (iii) present a user-friendly interface to interact with the LLib framework and monitor the query results.
Jin Wang 0007, Guorui Xiao, Youfu Li 0003, Carlo Zaniolo
CIKM3
2022 Highly Efficient String Similarity Search and Join over Compressed Indexes
abstract
String similarity search and join are essential op-erations in many fields. Existing solutions adopt a filter-and-verification framework and build inverted indexes based on generated signatures to prune dissimilar candidates. While existing solutions mainly focus on improving the query processing performance, little attention is paid to reducing the inverted indexes' memory consumption. In cases where the index size is larger than the memory, users have to employ more expensive disk-based algorithms rather than in-memory ones. In this paper, we propose a flexible framework CSS to reduce the index size and keep high query performance for string search and join applications. It can be easily incorporated into a broad scope of existing frameworks. We first give improved solutions for offline inverted lists construction to better support string similarity search. Nevertheless, they cannot be applied in the problem of string similarity join where indexes are constructed online. To address this issue, we further propose the first approach for online construction of compressed inverted lists. We theoretically study a benefit model to help find the best trade-off between memory consumption and execution time, and then propose an adaptive compression approach based on it. Experimental results on large-scale datasets demonstrate that CSS can reduce the memory consumption by 3 to 5 times while having similar or even better query processing performance for a variety of string similarity search and join frameworks.
Guorui Xiao, Jin Wang 0007, Chunbin Lin, Carlo Zaniolo
ICDE1
2020 RASQL: A Powerful Language and its System for Big Data Applications
abstract
There is a growing interest in supporting advanced Big Data applications on distributed data processing platforms. Most of these systems support SQL or its dialect as the query interface due to its portability and declarative nature. However, current SQL standard cannot effectively express advanced analytical queries due to its limitation in supporting recursive queries. In this demonstration, we show that this problem can be resolved via a simple SQL extension that delivers greater expressive power by allowing aggregates in recursion. To this end, we propose the Recursive-aggregate-SQL (RASQL) language and its system on top of Apache Spark to express and execute complex queries and declarative algorithms in many applications, such as graph search and machine learning. With a variety of examples, we will (i) show how complicated analytic queries can be expressed with RASQL; (ii) illustrate formal semantics of the powerful new constructs; and (iii) present a user-friendly interface to interact with the RASQL system and monitor the query results.
Jin Wang 0007, Guorui Xiao, Jiaqi Gu 0001, Jiacheng Wu 0001, Carlo Zaniolo
SIGMOD Conference2