EDBT 2026 Demo / reviewers in the wild / expert
Yifan Xiong 0001
dblp:188/1149-1
· DBLP profile ↗
9ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0001-9056-1386ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorComputer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SuperBench: A Proactive Validation System for Improving Reliability of Cloud AI InfrastructureabstractReliability in cloud AI infrastructure is crucial for cloud service providers, prompting the widespread use of hardware redundancies. However, these redundancies can inadvertently lead to hidden degradation, known as “gray failure”, for AI workloads, significantly affecting end-to-end performance and concealing performance issues, which complicates root cause analysis for failures and regressions. We introduce SuperBench, a proactive validation system for AI infrastructure that mitigates hidden degradation caused by hardware redundancies and enhances overall reliability. SuperBench features a comprehensive benchmark suite, capable of evaluating individual hardware components and representing most real AI workloads. It comprises a Validator that learns benchmark criteria to pinpoint defective components clearly. Additionally, SuperBench incorporates a Selector to balance validation time and issue-related penalties, enabling optimal timing for validation execution with a tailored subset of benchmarks. Through testbed evaluation and simulation, we demonstrate that SuperBench can increase the mean time between incidents by up to 22.61×. SuperBench has been successfully deployed in Azure production, validating hundreds of thousands of GPUs every year. Yifan Xiong 0001, Ziyue Yang 0002, Guoshuai Zhao 0001, Dong Zhong, Boris Pinzur, Jie Zhang 0048, Yang Wang 0053, Hossein Pourreza, Jeff Baxter, Kushal Datta, Prabhat Ram, Luke Melton, Joe Chau, Peng Cheng 0005, Yongqiang Xiong, Lidong Zhou |
ACM Trans. Comput. Syst. | 1 |
| 2024 | SuperBench: Improving Cloud AI Infrastructure Reliability with Proactive Validation
Yifan Xiong 0001, Ziyue Yang 0002, Guoshuai Zhao 0001, Dong Zhong, Boris Pinzur, Jie Zhang 0048, Yang Wang 0053, Hossein Pourreza, Jeff Baxter, Kushal Datta, Prabhat Ram, Luke Melton, Joe Chau, Peng Cheng 0005, Yongqiang Xiong, Lidong Zhou |
USENIX ATC | 1 |
| 2023 | MSCCLang: Microsoft Collective Communication LanguageabstractMachine learning models with millions or billions of parameters are increasingly trained and served on large multi-GPU systems. As models grow in size and execute on more GPUs, collective communication becomes a bottleneck. Custom collective algorithms optimized for both particular network topologies and application-specific communication patterns can alleviate this bottleneck and help these applications scale. However, implementing correct and efficient custom algorithms is challenging. Meghan Cowan, Saeed Maleki, Madan Musuvathi, Olli Saarikivi, Yifan Xiong 0001 |
ASPLOS (2) | 5 |
| 2023 | ElasticFlow: An Elastic Serverless Training Platform for Distributed Deep LearningabstractThis paper proposes ElasticFlow, an elastic serverless training platform for distributed deep learning. ElasticFlow provides a serverless interface with two distinct features: (i) users specify only the deep neural network (DNN) model and hyperparameters for a job, but not the number of GPUs; (ii) users specify the deadline for a job, but not the amount of time to occupy GPUs. In contrast to existing server-centric platforms, ElasticFlow provides performance guarantees in terms of meeting deadlines while alleviating tedious, low-level, and manual resource management for deep learning developers. The characteristics of distributed training introduce two challenges. First, the training throughput scales non-linearly with the number of GPUs. Second, the scaling efficiency is affected by worker placement. To address these challenges, we propose Minimum Satisfactory Share to capture the resource usage of training jobs to meet deadlines, and ElasticFlow performs admission control based on it. We develop a greedy algorithm that dynamically allocates resources to admitted jobs based on diminishing returns. We apply buddy allocation to worker placement to eliminate the effect of topology. Evaluation results on a cluster of 128 GPUs show that ElasticFlow increases the number of jobs that can meet their deadlines by 1.46–7.65× compared to existing solutions. Diandian Gu, Yinmin Zhong, Yifan Xiong 0001, Zhenhua Han, Peng Cheng 0005, Fan Yang 0024, Gang Huang 0001, Xin Jin 0008, Xuanzhe Liu |
ASPLOS (2) | 4 |
| 2022 | Moneo: Non-intrusive Fine-grained Monitor for AI InfrastructureabstractCloud-based AI infrastructure is increasingly important, especially on large-scale distributed training. To improve its efficiency and serviceability, real-time monitoring of the infrastructure and profiling the workload are proved to be the effective approach empirically. However, cloud environment poses great challenges as service providers cannot interfere with their tenants' workloads or touch user data, thus previous instrumentation-based monitoring approach cannot be applied, nor does the workload trace collection.We propose Moneo, a non-intrusive cloud-friendly monitoring system for AI infrastructure. Moneo is capable of intelligently collecting the key architecture-level metrics at finer granularity in real-time without instrumenting or tracing the workloads, which has been deployed in real production cloud, Azure. We analyze the results reported by Moneo for typical large-scale distributed AI workloads from real deployment. Results demonstrate that Moneo can effectively help service providers understand the real resource usage patterns of various AI workloads and real networking requirements, so as to get valuable findings help improve the efficiency of cloud infrastructure and optimize the software stack with the consideration of the characteristic resource usage requirements for different AI workloads. Yifan Xiong 0001, Chen Tian 0001, Peng Cheng 0005, Yongqiang Xiong |
ICC | 2 |
| 2020 | HiveD: Sharing a GPU Cluster for Deep Learning with Guarantees
Zhenhua Han, Zhi Yang 0001, Quanlu Zhang, Fan Yang 0024, Lidong Zhou, Mao Yang 0004, Francis C. M. Lau 0001, Yifan Xiong 0001 |
OSDI | 10 |
| 2016 | Semantic Image Profiling for Historic Events: Linking Images to PhrasesabstractAutomatically generating image profiles for historic events is desired for history knowledge preservation and curation. However, a simple profile with groups of related images lacks explicit semantic information, such as which images correspond to which aspects of the event. In this paper, we propose to add explicit semantic information to image profiling by linking images in the profile with related phrases in the event description. We measure the relevance of an image-phrase pair via a real-valued matching score. We exploit instance-wise ranking loss function to learn the matching score and we deal with two challenges: 1) how to automatically generate labeled positive data: we leverage out-of-domain labeled datasets to generate pseudo positive in-domain labels and propose a new algorithm (WIL4PPL) to robustly learn the model from the noisy pseudo positive labels; 2) how to automatically generate negative data: we propose a negative set generation algorithm to guide the model in learning which phrases and images to distinguish. We compare our model to three baselines and conduct detailed analysis and case studies to verify the quality of learnt semantic information. The extensive experiment results show the effectiveness of our proposed algorithms which significantly outperform the baselines. Jia Chen 0001, Qin Jin, Yifan Xiong 0001 |
ACM Multimedia | 3 |
| 2016 | Describing Videos using Multi-modal FusionabstractDescribing videos with natural language is one of the ultimate goals of video understanding. Video records multi-modal information including image, motion, aural, speech and so on. MSR Video to Language Challenge provides a good chance to study multi-modality fusion in caption task. In this paper, we propose the multi-modal fusion encoder and integrate it with text sequence decoder into an end-to-end video caption framework. Features from visual, aural, speech and meta modalities are fused together to represent the video contents. Long Short-Term Memory Recurrent Neural Networks (LSTM-RNNs) are then used as the decoder to generate natural language sentences. Experimental results show the effectiveness of multi-modal fusion encoder trained in the end-to-end framework, which achieved top performance in both common metrics evaluation and human evaluation. Qin Jin, Jia Chen 0001, Shizhe Chen, Yifan Xiong 0001, Alex Hauptmann 0001 |
ACM Multimedia | 4 |
| 2016 | History Rhyme: Searching Historic Events by Multimedia KnowledgeabstractThis demo shows a novel system "History Rhyme" which searches historic events by multimedia knowledge. Different from existing historic events related works, we focus on the retrieval of historic events by semantic related multimedia knowledge. Our system can not only search historic events based on keyword queries, but also retrieve similar historic events to a given event based on chosen facets. In both cases, the system returns top retrieval results and shows the image profiles of each historic event. To build such functions, we automatically mine knowledge from multimedia data and index each historic event. Our online demo is available at http://222.29.193.164/HistoryRhyme. Yifan Xiong 0001, Jia Chen 0001, Qin Jin |
ACM Multimedia | 1 |