EDBT 2026 Demo / reviewers in the wild / expert
Xiaosong Ma
dblp:m/XiaosongMa
· DBLP profile ↗
13ranked-venue papers in the field
1as first author
4since 2021 · last 2026
—ORCID · conflict
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 7 (1 first)Big Data, Cloud & Distributed Data Systems · 6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HCT-QA: A Benchmark for Question Answering on Human-Centric TablesabstractTabular data embedded in PDF files, web pages, and other types of documents is prevalent in various domains. These tables, which we call human-centric tables (HCTs for short), are dense in information but often exhibit complex structural and semantic layouts. To query these HCTs, some existing solutions focus on transforming them into relational formats. However, they fail to handle the diverse and complex layouts of HCTs, making them not amenable to easy querying with SQL-based approaches. Another emerging option is to use Large Language Models (LLMs) and Vision Language Models (VLMs). However, there is a lack of standard evaluation benchmarks to measure and compare the performance of models to query HCTs using natural language. To address this gap, we propose the HumanCentric Tables Question-Answering extensive benchmark (HCTQA) consisting of thousands of HCTs with several thousands of natural language questions with their respective answers. More specifically, HCT-QA includes 1,880 real-world HCTs with 9,835 QA pairs in addition to 4,679 synthetic HCTs with 67.7K QA pairs. Also, we show through extensive experiments the performance of 25 and 9 different LLMS and VLMs, respectively, in an answering HCT-QA's questions. In addition, we show how finetuning an LLM on HCT-QA improves F1 scores by up to 25 percentage points compared to the off-the-shelf model. Compared to existing benchmarks, HCT-QA stands out for its broad complexity and diversity of covered HCTs and generated questions, its comprehensive metadata enabling deeper insight and analysis, and its novel synthetic data and QA generator. Mohammad Shahmeer Ahmad, Zan Ahmad Naeem, Michaël Aupetit 0001, Ahmed K. Elmagarmid, Mohamed Y. Eltabakh, Xiaosong Ma, Mourad Ouzzani, Chaoyi Ruan, Hani Al-Sayeh |
ICDE | 6 |
| 2024 | PolyBase: Adapting to Data Affinity Changes in Geo-Replicated Database via Row-Level Paxos-Group Affiliation Re-AssignmentabstractTransaction performance in geo-replicated databases heavily relies on the request location: when not issued by the primary region, transactions are forced to involve costly wide-area communication. While existing systems distribute primary roles across regions, such assignment typically occurs at the shard level, making it difficult to align with geographically dispersed access to individual records. This paper introduces PolyBase, a pioneering architecture to address such misalignment, leveraging the widely adopted Paxos-based log replication mechanisms. It enables flexible row-level consensus group affiliation , which runs on an unchanged Paxos protocol , but dynamically re-assigns database rows between Paxos log replication groups, whose leaders become the primary region, enjoying faster writes and up-to-date versions for reads. With carefully designed data structures and protocols, PolyBase significantly reduces wide-area RTTs without compromising transaction or log replication consistency or reliability guarantees. We implemented PolyBase with optimized re-assignment policies and integrated it into two popular databases (RocksDB and MySQL). Our evaluation on AWS, using a production e-commerce workload and microbench-marks confirms that PolyBase offers significantly higher transaction throughput and lower average/tail latency compared to baselines. Chaoyi Ruan, Yingqiang Zhang, Juncheng Zhang, Cheng Li 0001, Xiaosong Ma, Hao Chen 0080, Feifei Li 0001, Xinjun Yang |
Proc. VLDB Endow. | 5 |
| 2021 | SpanDB: A Fast, Cost-Effective LSM-tree Based KV Store on Hybrid Storage
Hao Chen 0080, Chaoyi Ruan, Cheng Li 0001, Xiaosong Ma, Yinlong Xu 0001 |
FAST | 4 |
| 2021 | FusionRAID: Achieving Consistent Low Latency for Commodity SSD Arrays
Tianyang Jiang, Guangyan Zhang, Zican Huang, Xiaosong Ma, Junyu Wei, Zhiyue Li |
FAST | 4 |
| 2020 | QarSUMO: A Parallel, Congestion-optimized Traffic SimulatorabstractTraffic simulators are important tools for tasks such as urban planning and transportation management. Microscopic simulators allow per-vehicle movement simulation, but require longer simulation time. The simulation overhead is exacerbated when there is traffic congestion and most vehicles move slowly. This in particular hurts the productivity of emerging urban computing studies based on reinforcement learning, where traffic simulations are heavily and repeatedly used for designing policies to optimize traffic related tasks. Hao Chen 0080, Stefano Giovanni Rizzo, Giovanna Vantini, Phillip Taylor, Xiaosong Ma, Sanjay Chawla |
SIGSPATIAL/GIS | 6 |
| 2020 | LiveGraph: A Transactional Graph Storage System with Purely Sequential Adjacency List ScansabstractThe specific characteristics of graph workloads make it hard to design a one-size-fits-all graph storage system. Systems that support transactional updates use data structures with poor data locality, which limits the efficiency of analytical workloads or even simple edge scans. Other systems run graph analytics workloads efficiently, but cannot properly support transactions. This paper presents LiveGraph, a graph storage system that outperforms both the best graph transactional systems and the best solutions for real-time graph analytics on fresh data. LiveGraph achieves this by ensuring that adjacency list scans, a key operation in graph workloads, are purely sequential: they never require random accesses even in presence of concurrent transactions. Such pure-sequential operations are enabled by combining a novel graph-aware data structure, the Transactional Edge Log (TEL), with a concurrency control mechanism that leverages TEL's data layout. Our evaluation shows that LiveGraph significantly outperforms state-of-the-art (graph) database solutions on both transactional and real-time analytical workloads. Xiaowei Zhu 0001, Marco Serafini, Xiaosong Ma, Ashraf Aboulnaga, Guanyu Feng |
Proc. VLDB Endow. | 3 |
| 2019 | Automatic, Application-Aware I/O Forwarding Resource Allocation
Bin Yang 0043, Xiaosong Ma, Xiupeng Zhu, Xiyang Wang 0003, Nosayba El-Sayed, Jidong Zhai, Wei Xue 0003 |
FAST | 4 |
| 2018 | RAID+: Deterministic and Balanced Data Distribution for Large Disk Enclosures
Guangyan Zhang, Zican Huang, Xiaosong Ma, Zhufan Wang |
FAST | 3 |
| 2014 | Automatic identification of application I/O signatures from noisy server-side traces
Yang Liu 0129, Raghul Gunasekaran, Xiaosong Ma, Sudharshan S. Vazhkudai |
FAST | 3 |
| 2013 | Active flash: towards energy-efficient, in-situ data analytics on extreme-scale machines
Devesh Tiwari, Simona Boboila, Sudharshan S. Vazhkudai, Youngjae Kim 0001, Xiaosong Ma, Peter Desnoyers, Yan Solihin |
FAST | 5 |
| 2008 | Adaptive Request Scheduling for Parallel Scientific Web Services
Heshan Lin, Xiaosong Ma, Jiangtian Li, Ting Yu 0001, Nagiza F. Samatova |
SSDBM | 2 |
| 2004 | GODIVA: Lightweight Data Management for Scientific Visualization ApplicationsabstractScientific visualization applications are very data-intensive, with high demands for I/O and data management. Developers of many visualization tools hesitate to use traditional DBMSs, due to the lack of support for these DBMSs on parallel platforms and the risk of reducing the portability of their tools and the user data. We propose the GODIVA framework, which provides simple database-like interfaces to help visualization tool developers manage their in-memory data, and I/O optimizations such as prefetching and caching to improve input performance at run time. We implemented the GODIVA interfaces in a stand-alone, portable user library, which can be used by all types of visualization codes: interactive and batch-mode, sequential and parallel. Performance results from running a visualization tool using the GODIVA library on multiple platforms show that the GODIVA framework is easy to use, alleviates developers' data management burden, and can bring substantial I/O performance improvement. Xiaosong Ma, Marianne Winslett, John Norris, Xiangmin Jiao, Robert Fiedler |
ICDE | 1 |
| 2003 | Declustering Large Multidimensional Data Sets for Range Queries over Heterogeneous DisksabstractDeclustering is a technique to distribute data sets over multiple disks so that future retrievals can be well balanced over the disks and be performed in parallel. Although clusters often have heterogeneous disks, most declustering work has focused only on homogeneous environments. In this work, we investigate the declustering problem for a heterogeneous disk environment using virtual servers, and propose approaches for deciding the number of virtual servers and the mapping between virtual servers and physical disks. Our experimental results show that by combining our algorithm for choosing the number of virtual servers with a greedy algorithm for mapping virtual servers to disks, users can expect range query retrieval performance within 4% of the optimum achievable in practice on average, in all configurations studied. Compared to an intuitively natural approach to the problem, this represents an improvement of 8-31% in average fetch ratio, as well a 26-38% reduction in the standard deviation of performance for small queries. Jonghyun Lee 0001, Marianne Winslett, Xiaosong Ma, Shengke Yu |
SSDBM | 3 |