EDBT 2026 Demo / reviewers in the wild / expert
Xianyan Jia
dblp:183/9294
· DBLP profile ↗
5ranked-venue papers
2as first author
3since 2021 · last 2023
0009-0006-8663-0653ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Efficient and distributed learning · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Cloud and datacenter computing · 35% GPUs and heterogeneous computing · 35% Distributed systems · 30% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning
distributed training |
1.2 | 2 | 2023 | EasyScale: Elastic Training with Consistent Accuracy and Improved Utilization on GPUs · SC 2023 Whale: Efficient Giant Model Training over Heterogeneous GPUs · USENIX ATC 2022 |
Machine learning › Efficient and distributed learning › distributed training › distributed training systems
elastic training |
0.7 | 1 | 2023 | EasyScale: Elastic Training with Consistent Accuracy and Improved Utilization on GPUs · SC 2023 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.7 | 1 | 2023 | EasyScale: Elastic Training with Consistent Accuracy and Improved Utilization on GPUs · SC 2023 |
GPUs and heterogeneous computing
GPU scheduling |
0.7 | 1 | 2023 | EasyScale: Elastic Training with Consistent Accuracy and Improved Utilization on GPUs · SC 2023 |
Distributed systems › distributed machine learning
distributed training |
0.6 | 1 | 2022 | Whale: Efficient Giant Model Training over Heterogeneous GPUs · USENIX ATC 2022 |
Methods — techniques the papers use, named apart from their topics
intra-/inter-job scheduling · 1.3data-parallel training · 1.3context switching · 1.3giant model training · 1.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | EasyScale: Elastic Training with Consistent Accuracy and Improved Utilization on GPUsabstractDistributed synchronized GPU training is commonly used for deep learning. The resource constraint of using a fixed number of GPUs makes large-scale training jobs suffer from long queuing time for resource allocation, and lowers the cluster utilization. Adapting to resource elasticity can alleviate this but often introduces inconsistent model accuracy, due to lacking of capability to decouple model training procedure from resource allocation. We propose EasyScale, an elastic training system that achieves consistent model accuracy under resource elasticity for both homogeneous and heterogeneous GPUs. EasyScale preserves the data-parallel training behaviors strictly, traces the consistency-relevant factors carefully, utilizes the deep learning characteristics for EasyScaleThread abstraction and fast context-switching. To utilize heterogeneous cluster, EasyScale dynamically assigns workers based on the intra-/inter-job schedulers, minimizing load imbalance and maximizing aggregated job throughput. Deployed in an online serving cluster, EasyScale powers the training jobs to utilize idle GPUs opportunistically, improving overall cluster utilization by 62.1%. Mingzhen Li 0001, Wencong Xiao, Hailong Yang 0002, Biao Sun 0002, Shiru Ren, Zhongzhi Luan, Xianyan Jia, Yi Liu 0013, Yong Li 0045, Wei Lin 0016, Depei Qian 0001 |
SC | 8 |
| 2022 | Whale: Efficient Giant Model Training over Heterogeneous GPUs
Xianyan Jia, Ang Wang, Wencong Xiao, Ziji Shi, Jie Zhang 0135, Langshi Chen, Yong Li 0045, Zhen Zheng, Wei Lin 0016 |
USENIX ATC | 1 |
| 2021 | EasyTransfer: A Simple and Scalable Deep Transfer Learning Platform for NLP ApplicationsabstractThe literature has witnessed the success of leveraging Pre-trained Language Models (PLMs) and Transfer Learning (TL) algorithms to a wide range of Natural Language Processing (NLP) applications, yet it is not easy to build an easy-to-use and scalable TL toolkit for this purpose. To bridge this gap, the EasyTransfer platform is designed to develop deep TL algorithms for NLP applications. EasyTransfer is backended with a high-performance and scalable engine for efficient training and inference, and also integrates comprehensive deep TL algorithms, to make the development of industrial-scale TL applications easier. In EasyTransfer, the built-in data and model parallelism strategies, combined with AI compiler optimization, show to be 4.0x faster than the community version of distributed training. EasyTransfer supports various NLP models in the ModelZoo, including mainstream PLMs and multi-modality models. It also features various in-house developed TL algorithms, together with the AppZoo for NLP applications. The toolkit is convenient for users to quickly start model training, evaluation, and online deployment. EasyTransfer is currently deployed at Alibaba to support a variety of business scenarios, including item recommendation, personalized search, conversational question answering, etc. Extensive experiments on real-world datasets and online applications show that EasyTransfer is suitable for online production with cutting-edge performance for various applications. The source code of EasyTransfer is released at Github1. Minghui Qiu, Peng Li 0056, Chengyu Wang 0001, Haojie Pan, Ang Wang, Cen Chen 0001, Xianyan Jia, Yaliang Li, Jun Huang 0007, Deng Cai 0001, Wei Lin 0016 |
CIKM | 7 |
| 2019 | BigDL: A Distributed Deep Learning Framework for Big DataabstractThispaperpresentsBigDL (adistributeddeeplearning framework for Apache Spark), which has been used by a variety of users in the industry for building deep learning applications on production big data platforms. It allows deep learning applications to run on the Apache Hadoop/Spark cluster so as to directly process the production data, and as a part of the end-to-end data analysis pipeline for deployment and management. Unlike existing deep learning frameworks, BigDL implements distributed, data parallel training directly on top of the functional compute model (with copy-on-write and coarse-grained operations) of Spark. We also share real-world experience and "war stories" of users that havead-optedBigDLtoaddresstheirchallenges(i.e., howtoeasilybuildend-to-enddataanalysisanddeep learning pipelines for their production data). Jason Jinquan Dai, Xin Qiu 0006, Yanzhang Wang, Xianyan Jia, Cherry Li Zhang, Shengsheng Huang, Zhongyuan Wu, Yang Wang 0009, Bowen She, Dongjie Shi, Guoqiong Song |
SoCC | 7 |
| 2016 | Target-Oriented Keyword Search over Temporal Databases
Xianyan Jia, Wynne Hsu, Mong-Li Lee |
DEXA (1) | 1 |