VLDB 2026 Research / reviewers in the wild / expert
Xun Song
dblp:158/6656
· DBLP profile ↗
4ranked-venue papers
2as first author
4since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Magnus: A Holistic Approach to Data Management for Large-Scale Machine Learning WorkloadsabstractMachine learning (ML) has become a cornerstone of key applications at ByteDance. As model complexity and data volumes surge, data management for large-scale ML workloads faces substantial challenges, particularly with recent advances in large recommendation models (LRMs) and large multimodal models (LMMs). Traditional approaches exhibit limitations in storage efficiency, metadata scalability, update mechanisms, and integration with ML frameworks. To address these challenges, we propose Magnus, a holistic data management system built upon Apache Iceberg. Magnus integrates innovative optimizations across resource-efficient storage formats optimized for large wide tables and multimodal data, built-in support for vector and inverted indexes to accelerate data retrieval, scalable metadata planning with Git-like branching and tagging capabilities, and high-performance update/upsert based on lightweight merge-on-read (MOR) strategies. Additionally, Magnus provides native support and specialized enhancement for LRM and LMM training workloads. Experimental results demonstrate significant performance gains in real-world ML scenarios. Magnus has been deployed at ByteDance for over five years, enabling robust and efficient data infrastructure for large-scale ML workloads. Jingyi Ding, Irshad Kandy, Yanghao Lin, Zhongjia Wei, Zhiwei Peng, Jixi Shan, Hongyue Mao, Xiuqi Huang, Xun Song, Yanjia Li, Tianhao Yang, Xiaohong Dong, Kang Lei, Pengwei Zhao, Wei Chen 0001 |
Proc. VLDB Endow. | 11 |
| 2025 | In Search of a Memory-Efficient Framework for Online Cardinality EstimationabstractEstimating per-flow cardinality from high-speed data streams has many applications such as anomaly detection and resource allocation. Yet despite tracking single flow cardinality with approximation algorithms offered, there remain algorithmical challenges for monitoring multi-flows especially under unbalanced cardinality distribution: existing methods adopt a uniform sketch layout and incur a large memory footprint to achieve high accuracy. Furthermore, they are hard to implement in the compact hardware used for line-rate processing. In this paper, we propose Couper, a memory-efficient measurement framework that can estimate cardinality for multi-flows under unbalanced cardinality distribution. We propose a two-layer structure based on a classic coupon collector's principle, where numerous mice flows are confined to the first layer and only the potential elephant flows are allowed to enter the second layer. Our two-layer structure can better fit the unbalanced cardinality distribution in practice and achieve much higher memory efficiency. We implement Couper in both software and hardware. Extensive evaluation under real-world and synthetic data traces show more than 20× improvements in terms of memory-efficiency compared to state-of-the-art. Xun Song, Jiaqi Zheng 0001, Shiju Zhao, Hongxuan Zhang, Xuntao Pan, Guihai Chen |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2024 | A Debiased Domain Adaptation Framework with Minimum Class Confusion for Motor Imagery DecodingabstractRecently, motor imagery decoding technology based on electroencephalogram (EEG) signals has made significant progress. However, there are still challenges in adapting to new sessions, mainly due to changes in data distribution between different sessions and confusion problems caused by similar oscillation patterns in various categories of EEG signals. To address these problems, this paper proposes a debiased domain adaptation framework with minimum class confusion to learn unbiased representations in motor imagery tasks. Specially, unlike the feature alignment and adversarial training methods, we explore the class predictions for domain adaptation, applying the minimum class confusion loss criterion in the target domain to reduce inter-class confusion. This approach aims to alleviate the bias issues inherent in classifiers trained on the source domain for making predictions in the target domain, achieving class-level alignment. Consequently, it enhances the model’s ability to adapt to the data distribution of the target domain. Experimental results on two public EEG datasets (BCI Competition IV datasets IIa and IIb) show that the method for cross-session decoding is significantly improved compared to the baseline, with average classification accuracy reaching 81.01% and 82.92%, respectively. Cunhang Fan, Zhen Chen 0022, Xun Song, Jun Xue 0001, Ping Li 0020, Zhao Lv |
IJCNN | 3 |
| 2023 | Couper: Memory-Efficient Cardinality Estimation under Unbalanced DistributionabstractEstimating per-flow cardinality from high-speed data streams has many applications such as anomaly detection and resource allocation. Yet despite tracking single flow cardinality with approximation algorithms offered, there remain algorithmical challenges for monitoring multi-flows especially under unbalanced cardinality distribution: existing methods adopt a uniform sketch layout and incur a large memory footprint to achieve high accuracy. Furthermore, they are hard to implement in the compact hardware used for line-rate processing.In this paper, we propose Couper, a memory-efficient measurement framework that can estimate cardinality for multi-flows under unbalanced cardinality distribution. We propose a two-layer structure based on a classic coupon collector’s principle, where numerous mice flows are confined to the first layer and only the potential elephant flows are allowed to enter the second layer. Our two-layer structure can better fit the unbalanced cardinality distribution in practice and achieve much higher memory efficiency. We implement Couper in both software and hardware. Extensive evaluation under real-world and synthetic data traces show more than 20× improvements in terms of memory-efficiency compared to state-of-the-art. Xun Song, Jiaqi Zheng 0001, Shiju Zhao, Hongxuan Zhang, Xuntao Pan, Guihai Chen |
ICDE | 1 |