EDBT 2026 Demo / reviewers in the wild / expert
Liang You
dblp:217/3880
· DBLP profile ↗
7ranked-venue papers
0as first author
5since 2021 · last 2024
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Distributed systems · 72% Memory systems · 24% Cloud and datacenter computing · 4% | |
| Software engineering, system software, and programming languages
2 papers |
Software testing · 50% Program analysis · 50% | |
| Artificial intelligence
1 paper |
Deep learning architectures and training · 100% |
Topics — the 7 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Program analysis
static analysis |
0.9 | 2 | 2021 | Detecting TensorFlow Program Bugs in Real-World Industrial Environment · ASE 2021 CrashTuner: detecting crash-recovery bugs in cloud systems via meta-info analysis · SOSP 2019 |
Distributed systems › fault tolerance
checkpointing |
0.7 | 1 | 2023 | OpenEmbedding: A Distributed Parameter Server for Deep Learning Recommendation Models using Persistent Memory · ICDE 2023 |
Distributed systems › distributed machine learning
parameter server |
0.7 | 1 | 2023 | OpenEmbedding: A Distributed Parameter Server for Deep Learning Recommendation Models using Persistent Memory · ICDE 2023 |
Memory systems › non-volatile memory
persistent memory |
0.7 | 1 | 2023 | OpenEmbedding: A Distributed Parameter Server for Deep Learning Recommendation Models using Persistent Memory · ICDE 2023 |
Distributed systems › fault tolerance › checkpointing
synchronous checkpointing |
0.7 | 1 | 2023 | OpenEmbedding: A Distributed Parameter Server for Deep Learning Recommendation Models using Persistent Memory · ICDE 2023 |
Software testing › mutation testing › fault injection
fault injection testing |
0.4 | 1 | 2019 | CrashTuner: detecting crash-recovery bugs in cloud systems via meta-info analysis · SOSP 2019 |
Cloud and datacenter computing
cloud service reliability |
0.1 | 1 | 2019 | CrashTuner: detecting crash-recovery bugs in cloud systems via meta-info analysis · SOSP 2019 |
Methods — techniques the papers use, named apart from their topics
static analysis · 1.0constraint-based analysis · 1.0meta-info analysis · 0.8log-based static program analysis · 0.8fault injection · 0.8pipeline processing · 0.7checkpointing · 0.7DRAM cache · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | An efficient cloud-based elastic RDMA protocol for HPC applications
Muhui Lin, Jinhu Li, Xiangzheng Sun, Ronghui He, Liang You |
CCF Trans. High Perform. Comput. | 10 |
| 2023 | OpenEmbedding: A Distributed Parameter Server for Deep Learning Recommendation Models using Persistent MemoryabstractIn this paper, we present OpenEmbedding, a distributed parameter server system for deep learning recommendation models (DLRM) workloads. In order to support rapid growth in the number of features and the model size (Terabytes are common) of DLRM workloads, OpenEmbedding takes advantage of emerging persistent memory (PMem) to address scalability and reliability issues in training DLRMs. Compared to DRAM, PMem can have much lower per-GB cost, higher density, and non-volatility, while with slightly low access performance to DRAM. OpenEmbedding uses DRAM as cache and PMem as storage for the sparse features and develops a simple but effective pipeline processing approach to optimize the access latency of the sparse features in PMem. For reliability, we develop a lightweight synchronous checkpointing scheme that is specially co-designed with the pipelined cache to reduce the run-time overhead of checkpointing. Our evaluations on a real-world industry workload consisting of billions of parameters demonstrate 1) the effectiveness of our PMem-aware optimizations, 2) checkpointing mechanism with near-zero run-time overhead to the training performance and 3) fast recovery with up to 3.97× speedup compared to the state-of-the-art. OpenEmbedding has been deployed in hundreds of scenarios in industry within 4Paradigm, and is open-sourced1. Cheng Chen 0008, Jun Yang 0022, Mian Lu, Zhao Zheng, Bingsheng He, Weng-Fai Wong, Liang You, Penghao Sun, Yuping Zhao, Fenghua Hu, Andy Rudoff |
ICDE | 9 |
| 2022 | AIACC-Training: Optimizing Distributed Deep Learning Training through Multi-streamed and Concurrent Gradient CommunicationsabstractThere is a growing interest in training deep neural networks (DNNs) in a GPU cloud environment. This is typically achieved by running parallel training workers on multiple GPUs across computing nodes. Under such a setup, the communication overhead is often responsible for long training time and poor scalability. This paper presents AIACC-Training, a unified communication framework designed for the distributed training of DNNs in a GPU cloud environment. AIACC-Training permits a training worker to participate in multiple gradient communication operations simultaneously to improve network bandwidth utilization and reduce communication latency. It employs auto-tuning techniques to dynamically determine the right communication parameters based on the input DNN workloads and the underlying network infrastructure. AIACC-Training has been deployed to production at Alibaba GPU Cloud with 3000+ GPUs executing AIACC-Training optimized code at any time. Experiments performed on representative DNN workloads show that AIACC-Training outperforms existing solutions, improving the training throughput and scalability by a large margin. Lixiang Lin, Shenghao Qiu, Liang You, Long Xin, Jie Xu 0007, Zheng Wang 0001 |
ICDCS | 4 |
| 2021 | Detecting TensorFlow Program Bugs in Real-World Industrial EnvironmentabstractDeep learning has been widely adopted in industry and has achieved great success in a wide range of application areas. Bugs in deep learning programs can cause catastrophic failures, in addition to a serious waste of resources and time.This paper aims at detecting industrial TensorFlow program bugs. We report an extensive empirical study on 12,289 failed TensorFlow jobs, showing that existing static tools can effectively detect 72.55% of the top three types of Python bugs in industrial TensorFlow programs. In addition, we propose (for the first time) a constraint-based approach for detecting TensorFlow shape-related errors (one of the most common TensorFlow-specific bugs), together with an associated tool, ShapeTracer. Our evaluation on a set of 60 industrial TensorFlow programs shows that ShapeTracer is efficient and effective: it analyzes each program in at most 3 seconds and detects effectively 40 out of 60 industrial TensorFlow program bugs, with no false positives. ShapeTracer has been deployed in the platform-X platform and will be released soon. Jie Lu 0009, Guangwei Li, Lian Li 0002, Liang You, Jingling Xue |
ASE | 8 |
| 2021 | Group multi-scale attention pyramid network for traffic sign detection
Lili Shen, Liang You, Chuhe Zhang |
Neurocomputing | 2 |
| 2019 | CrashTuner: detecting crash-recovery bugs in cloud systems via meta-info analysisabstractCrash-recovery bugs (bugs in crash-recovery-related mechanisms) are among the most severe bugs in cloud systems and can easily cause system failures. It is notoriously difficult to detect crash-recovery bugs since these bugs can only be exposed when nodes crash under special timing conditions. This paper presents CrashTuner, a novel fault-injection testing approach to combat crash-recovery bugs. The novelty of CrashTuner lies in how we identify fault-injection points (crash points) that are likely to expose errors. We observe that if a node crashes while accessing meta-info variables, i.e., variables referencing high-level system state information (e.g., an instance of node or task), it often triggers crash-recovery bugs. Hence, we identify crash points by automatically inferring meta-info variables via a log-based static program analysis. Our approach is automatic and no manual specification is required. Jie Lu 0009, Lian Li 0002, Xiaobing Feng 0002, Liang You |
SOSP | 7 |
| 2018 | HaaS: Cloud-Based Real-Time Data Analytics with Heterogeneity-Aware SchedulingabstractReal-time data analytics has become increasingly important in modern times as many organizations and companies are generating and analyzing high volume of data constantly. Despite of the impressive technical development, it remains a challenging job to analyze the stream data effectively and efficiently because traditional hardware and software lack specific designs and optimizations for those emerging requirements. In this paper, we discuss our experience on real-time data analytics, with our in-house processing framework HaaS. HaaS is designed to exploit existing data analytics tools and libraries as well as distributed computing technologies to embrace heterogeneous computation resources in the cloud. HaaS utilizes hierarchical clustering to partition physical topology of clusters weighted with task topology information into densely connected sub-graphs. HaaS is also equipped with a heterogeneity-aware scheduling algorithm to facilitate holistic optimization over multiple running tasks with various service level agreements. To the best of our knowledge, HaaS is the first ever streaming analytical framework providing users with flexible and optimized usage with CPUs, GPUs and FPGAs in the cloud. Users with stream processing tasks can easily enjoy remarkable advantages of CPUs, GPUs and FPGAs in throughput, power consumption and monetary cost over others. In our empirical evaluations with highly diversified workloads, HaaS saves over 18% on power consumption and 24% on monetary cost over existing system design architecture, while the overall throughput of HaaS remains no lower than 90% of the theoretical limit. Jiong He, Yao Chen 0008, Tom Z. J. Fu, Marianne Winslett, Liang You |
ICDCS | 6 |