EDBT 2026 Demo / reviewers in the wild / expert
Hongyue Mao
dblp:234/5239
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
2 papers |
Machine learning and data management · 34% Database system architecture and tuning · 34% Recommender systems · 10% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% |
Topics — the 5 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning › distributed training › distributed training systems
large-scale distributed training |
0.9 | 1 | 2025 | Primus: Unified Training System for Large-Scale Deep Learning Recommendation Models · USENIX ATC 2025 |
Machine learning and data management
data management for machine learning |
0.9 | 1 | 2025 | Magnus: A Holistic Approach to Data Management for Large-Scale Machine Learning Workloads · Proc. VLDB Endow. 2025 |
Information retrieval › indexing
inverted index |
0.3 | 1 | 2025 | Magnus: A Holistic Approach to Data Management for Large-Scale Machine Learning Workloads · Proc. VLDB Endow. 2025 |
Recommender systems
neural recommendation |
0.3 | 1 | 2025 | Primus: Unified Training System for Large-Scale Deep Learning Recommendation Models · USENIX ATC 2025 |
Indexing and storage engines
vector index |
0.3 | 1 | 2025 | Magnus: A Holistic Approach to Data Management for Large-Scale Machine Learning Workloads · Proc. VLDB Endow. 2025 |
Methods — techniques the papers use, named apart from their topics
metadata planning · 0.9merge-on-read · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Primus: Unified Training System for Large-Scale Deep Learning Recommendation Models
Jixi Shan, Xiuqi Huang, Hongyue Mao, Ho-Pang Hsu, Hang Cheng, Xiaofeng Gao 0001, Shiru Ren, Jiaxiao Zheng, Lele Yu, Guihai Chen |
USENIX ATC | 4 |
| 2025 | Magnus: A Holistic Approach to Data Management for Large-Scale Machine Learning WorkloadsabstractMachine learning (ML) has become a cornerstone of key applications at ByteDance. As model complexity and data volumes surge, data management for large-scale ML workloads faces substantial challenges, particularly with recent advances in large recommendation models (LRMs) and large multimodal models (LMMs). Traditional approaches exhibit limitations in storage efficiency, metadata scalability, update mechanisms, and integration with ML frameworks. To address these challenges, we propose Magnus, a holistic data management system built upon Apache Iceberg. Magnus integrates innovative optimizations across resource-efficient storage formats optimized for large wide tables and multimodal data, built-in support for vector and inverted indexes to accelerate data retrieval, scalable metadata planning with Git-like branching and tagging capabilities, and high-performance update/upsert based on lightweight merge-on-read (MOR) strategies. Additionally, Magnus provides native support and specialized enhancement for LRM and LMM training workloads. Experimental results demonstrate significant performance gains in real-world ML scenarios. Magnus has been deployed at ByteDance for over five years, enabling robust and efficient data infrastructure for large-scale ML workloads. Jingyi Ding, Irshad Kandy, Yanghao Lin, Zhongjia Wei, Zhiwei Peng, Jixi Shan, Hongyue Mao, Xiuqi Huang, Xun Song, Yanjia Li, Tianhao Yang, Xiaohong Dong, Kang Lei, Pengwei Zhao, Wei Chen 0001 |
Proc. VLDB Endow. | 9 |