Quanlu Zhang

dblp:165/8284 · DBLP profile ↗
← Back
26ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0003-0557-1104ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 4 first-author · 3 since 2021Software engineering, systems software and programming languages · 7 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Computer networks · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 STAlloc: Enhancing Memory Efficiency in Large-Scale Model Training with Spatio-Temporal Planning
abstract
The rapid scaling of large language models (LLMs) has significantly increased GPU memory pressure, which is further aggravated by training optimization techniques such as virtual pipeline and recomputation that disrupt tensor lifespans and introduce considerable memory fragmentation. Such fragmentation stems from the use of online GPU memory allocators in popular deep learning frameworks like PyTorch, which disregard tensor lifespans. As a result, this inefficiency can waste as much as 43% of memory and trigger out-of-memory errors, undermining the effectiveness of optimization methods.
Zixiao Huang 0001, Hao Lin 0005, Chunyang Zhu, Yueran Tang, Quanlu Zhang, Zhenhua Li 0001, Shengen Yan, Zhenhua Zhu 0002, Guohao Dai 0001, Yu Wang 0002
EuroSys6
2024 You Only Cache Once: Decoder-Decoder Architectures for Language Models
abstract
We introduce a decoder-decoder architecture, YOCO, for large language models, which only caches key-value pairs once. It consists of two components, i.e., a cross-decoder stacked upon a self-decoder. The self-decoder efficiently encodes global key-value (KV) caches that are reused by the cross-decoder via cross-attention. The overall model behaves like a decoder-only Transformer, although YOCO only caches once. The design substantially reduces GPU memory demands, yet retains global attention capability. Additionally, the computation flow enables prefilling to early exit without changing the final output, thereby significantly speeding up the prefill stage. Experimental results demonstrate that YOCO achieves favorable performance compared to Transformer in various settings of scaling up model size and number of training tokens. We also extend YOCO to 1M context length with near-perfect needle retrieval accuracy. The profiling results show that YOCO improves inference memory, prefill latency, and throughput by orders of magnitude across context lengths and model sizes.
Yutao Sun, Li Dong 0004, Shaohan Huang, Wenhui Wang 0003, Shuming Ma, Quanlu Zhang, Jianyong Wang 0001, Furu Wei
NeurIPS7
2024 nnScaler: Constraint-Guided Parallelization Plan Generation for Deep Learning Training
Youshan Miao, Quanlu Zhang, Fan Yang 0024, Cheng Li 0001, Saeed Maleki, Yilei Yang, Weijiang Xu, Mao Yang 0004, Lidong Zhou
OSDI3
2024 Ladder: Enabling Efficient Low-Precision Deep Learning Computing through Hardware-aware Tensor Transformation
Lei Wang 0222, Lingxiao Ma, Shijie Cao, Quanlu Zhang, Jilong Xue, Yining Shi 0001, Ningxin Zheng, Ziming Miao, Fan Yang 0024, Ting Cao 0003, Yuqing Yang 0001, Mao Yang 0004
OSDI4
2024 Automating Cloud Deployment for Real-Time Online Foundation Model Inference
abstract
Deep neural network (DNN) foundation models are currently exhibiting high prediction accuracy and strong adaptability to broad tasks with remarkably large model scales. They are increasingly becoming the backend support of DNN-driven real-time online services, e.g., Siri and Instagram. Such services require low-latency and cost-efficiency for quality-of-service and commercial competitiveness. When deployed in a cloud environment, these services call for an appropriate selection of cloud configurations (i.e., specific types of VM instances), as well as a considerate device placement plan that places the operations of the model to multiple GPUs via model parallelism for cost-efficiency. Currently, the deployment mainly relies on service providers’ manual efforts, which is not only onerous but also far from satisfactory oftentimes due to the huge joint search space of cloud configurations and device placement plans (for a same service, a poor deployment can incur significantly more costs by tens of times). In this paper, we attempt to efficiently automate the cloud deployment for real-time foundation model inference with minimum costs under the constraint of acceptably low latency. This attempt is enabled by 1) jointly leveraging the Bayesian Optimization and Deep Reinforcement Learning to adaptively unearth the (nearly) optimal cloud configuration and device placement with limited search time, and 2) enhancing the cost-efficiency of the deployment based on the probing-informed block multiplexing mechanism and Tensor Algebra SuperOptimizer. We implement a prototype system based on TensorFlow, conduct extensive experiments on top of Microsoft Azure, and demonstrate the generality and scalability of our solution. Results show that for lightweight DNN models and foundation models, our solution essentially saves inference costs by up to 15% and 47% with 57% and 38% lower search overheads respectively, compared with non-trivial baselines.
Yang Li 0092, Zhenhua Li 0001, Zhenhua Han, Quanlu Zhang, Xiaobo Ma 0001
IEEE/ACM Trans. Netw.4
2023 SiloD: A Co-design of Caching and Scheduling for Deep Learning Clusters
abstract
Deep learning training on cloud platforms usually follows the tradition of the separation of storage and computing. The training executes on a compute cluster equipped with GPUs/TPUs while reading data from a separate cluster hosting the storage service. To alleviate the potential bottleneck, a training cluster usually leverages its local storage as a cache to reduce the remote IO from the storage cluster. However, existing deep learning schedulers do not manage storage resources thus fail to consider the diverse caching effects across different training jobs. This could degrade scheduling quality significantly.
Zhenhua Han, Zhi Yang 0001, Quanlu Zhang, Mingxia Li, Fan Yang 0024, Qianxi Zhang, Binyang Li, Yuqing Yang 0001, Lili Qiu, Lidong Zhou
EuroSys4
2023 ElasticViT: Conflict-aware Supernet Training for Deploying Fast Vision Transformer on Diverse Mobile Devices
abstract
Neural Architecture Search (NAS) has shown promising performance in the automatic design of vision transformers (ViT) exceeding 1G FLOPs. However, designing lightweight and low-latency ViT models for diverse mobile devices remains a big challenge. In this work, we propose ElasticViT, a two-stage NAS approach that trains a high-quality ViT supernet over a very large search space for covering a wide range of mobile devices, and then searches an optimal sub-network (subnet) for direct deployment. However, current supernet training methods that rely on uniform sampling suffer from the gradient conflict issue: the sampled subnets can have vastly different model sizes (e.g., 50M vs. 2G FLOPs), leading to different optimization directions and inferior performance. To address this challenge, we propose two novel sampling techniques: complexity-aware sampling and performance-aware sampling. Complexity-aware sampling limits the FLOPs difference among the subnets sampled across adjacent training steps, while covering different-sized subnets in the search space. Performance-aware sampling further selects subnets that have good accuracy, which can reduce gradient conflicts and improve supernet quality. Our discovered models, ElasticViT models, achieve top-1 accuracy from 67.2% to 80.0% on ImageNet from 60M to 800M FLOPs without extra retraining, outperforming all prior CNNs and ViTs in terms of accuracy and latency. Our tiny and small models are also the first ViT models that surpass state-of-the-art CNNs with significantly lower latency on mobile devices. For instance, ElasticViT-S1 runs 2.62× faster than EfficientNet-B0 with 0.1% higher accuracy.
Li Lyna Zhang, Huiqiang Jiang, Jiahang Xu, Ting Cao 0003, Quanlu Zhang, Yuqing Yang 0001, Mao Yang 0004
ICCV6
2023 SpaceEvo: Hardware-Friendly Search Space Design for Efficient INT8 Inference
abstract
The combination of Neural Architecture Search (NAS) and quantization has proven successful in automatically designing low-FLOPs INT8 quantized neural networks (QNN). However, directly applying NAS to design accurate QNN models that achieve low latency on real-world devices leads to inferior performance. In this work, we identify that the poor INT8 latency is due to the quantization-unfriendly issue: the operator and configuration (e.g., channel width) choices in prior art search spaces lead to diverse quantization efficiency and can slow down the INT8 inference speed. To address this challenge, we propose SpaceEvo, an automatic method for designing a dedicated, quantization-friendly search space for each target hardware. The key idea of SpaceEvo is to automatically search hardware-preferred operators and configurations to construct the search space, guided by a metric called Q-T score to quantify how quantization-friendly a candidate search space is. We further train a quantized-for-all supernet over our discovered search space, enabling the searched models to be directly deployed without extra retraining or quantization. Our discovered models, SEQnet, establish new SOTA INT8 quantized accuracy under various latency constraints, achieving up to 10.1% accuracy improvement on ImageNet than prior art CNNs under the same latency. Extensive experiments on real devices show that SpaceEvo consistently outperforms manually-designed search spaces with up to 2.5× faster speed while achieving the same accuracy.
Li Lyna Zhang, Jiahang Xu, Quanlu Zhang, Yujing Wang 0002, Yuqing Yang 0001, Ningxin Zheng, Ting Cao 0003, Mao Yang 0004
ICCV4
2023 PIT: Optimization of Dynamic Sparse Deep Learning Models via Permutation Invariant Transformation
abstract
Dynamic sparsity, where the sparsity patterns are unknown until runtime, poses a significant challenge to deep learning. The state-of-the-art sparsity-aware deep learning solutions are restricted to pre-defined, static sparsity patterns due to significant overheads associated with preprocessing. Efficient execution of dynamic sparse computation often faces the misalignment between the GPU-friendly tile configuration for efficient execution and the sparsity-aware tile shape that minimizes coverage wastes (non-zero values in tensor).
Ningxin Zheng, Huiqiang Jiang, Quanlu Zhang, Zhenhua Han, Lingxiao Ma, Yuqing Yang 0001, Fan Yang 0024, Chengruidong Zhang, Lili Qiu, Mao Yang 0004, Lidong Zhou
SOSP3
2022 Privacy-preserving Online AutoML for Domain-Specific Face Detection
abstract
Despite the impressive progress of general face detection, the tuning of hyper-parameters and architectures is still critical for the performance of a domain-specific face detector. Though existing AutoML works can speedup such process, they either require tuning from scratch for a new scenario or do not consider data privacy. To scale up, we derive a new AutoML setting from a platform perspective. In such setting, new datasets sequentially arrive at the platform, where an architecture and hyper-parameter configuration is recommended to train the optimal face detector for each dataset. This, however, brings two major challenges: (1) how to predict the best configuration for any given dataset without touching their raw images due to the privacy concern? and (2) how to continuously improve the AutoML algorithm from previous tasks and offer a better warm-up for future ones? We introduce “HyperFD”, a new privacy-preserving online AutoML framework for face detection. At its core part, a novel meta-feature representation of a dataset as well as its learning paradigm is proposed. Thanks to HyperFD, each local task (client) is able to effectively leverage the learning “experience” of previous tasks without uploading raw images to the platform; meanwhile, the meta-feature extractor is continuously learned to better trade off the bias and variance. Extensive experiments demonstrate the effectiveness and efficiency of our design.
Chenqian Yan, Yuge Zhang, Quanlu Zhang, Yaming Yang 0001, Xinyang Jiang, Yuqing Yang 0001, Baoyuan Wang
CVPR3
2022 Nesting Forward Automatic Differentiation for Memory-Efficient Deep Neural Network Training
abstract
An activation function is an element-wise mathematical function and plays a crucial role in deep neural networks (DNN). Many novel and sophisticated activation functions have been proposed to improve the DNN accuracy but also consume massive memory in the training process with back-propagation. In this study, we propose the nested forward automatic differentiation (Forward-AD), specifically for the element-wise activation function for memory-efficient DNN training. We deploy nested Forward-AD in two widely-used deep learning frameworks, TensorFlow and PyTorch, which support the static and dynamic computation graph, respectively. Our evaluation shows that nested Forward-AD reduces the memory footprint by up to 1.97× than the baseline model and outperforms the recomputation by 20% under the same memory reduction ratio.
Cong Guo 0003, Yuxian Qiu, Jingwen Leng, Chen Zhang 0001, Quanlu Zhang, Yunxin Liu 0001, Fan Yang 0024, Minyi Guo
ICCD6
2022 SparTA: Deep-Learning Model Sparsity via Tensor-with-Sparsity-Attribute
Ningxin Zheng, Quanlu Zhang, Lingxiao Ma, Yuqing Yang 0001, Fan Yang 0024, Yang Wang 0053, Mao Yang 0004, Lidong Zhou
OSDI3
2020 LadaBERT: Lightweight Adaptation of BERT through Hybrid Model Compression
abstract
BERT is a cutting-edge language representation model pre-trained by a large corpus, which achieves superior performances on various natural language understanding tasks. However, a major blocking issue of applying BERT to online services is that it is memory-intensive and leads to unsatisfactory latency of user requests, raising the necessity of model compression. Existing solutions leverage the knowledge distillation framework to learn a smaller model that imitates the behaviors of BERT. However, the training procedure of knowledge distillation is expensive itself as it requires sufficient training data to imitate the teacher model. In this paper, we address this issue by proposing a tailored solution named LadaBERT (Lightweight adaptation of BERT through hybrid model compression), which combines the advantages of different model compression methods, including weight pruning, matrix factorization and knowledge distillation. LadaBERT achieves state-of-the-art accuracy on various public datasets while the training overheads can be reduced by an order of magnitude.
Yihuan Mao, Yujing Wang 0002, Chufan Wu, Chen Zhang 0001, Yang Wang 0053, Quanlu Zhang, Yaming Yang 0001, Yunhai Tong, Jing Bai 0010
COLING6
2020 Automating Cloud Deployment for Deep Learning Inference of Real-time Online Services
abstract
Real-time online services using pre-trained deep neural network (DNN) models, e.g., Siri and Instagram, require low-latency and cost-efficiency for quality-of-service and commercial competitiveness. When deployed in a cloud environment, such services call for an appropriate selection of cloud configurations (i.e., specific types of VM instances), as well as a considerate device placement plan that places the operations of a DNN model to multiple computation devices like GPUs and CPUs. Currently, the deployment mainly relies on service providers' manual efforts, which is not only onerous but also far from satisfactory oftentimes (for a same service, a poor deployment can incur significantly more costs by tens of times). In this paper, we attempt to automate the cloud deployment for real-time online DNN inference with minimum costs under the constraint of acceptably low latency. This attempt is enabled by jointly leveraging the Bayesian Optimization and Deep Reinforcement Learning to adaptively unearth the (nearly) optimal cloud configuration and device placement with limited search time. We implement a prototype system of our solution based on TensorFlow and conduct extensive experiments on top of Microsoft Azure. The results show that our solution essentially outperforms the nontrivial baselines in terms of inference speed and cost-efficiency.
Yang Li 0092, Zhenhua Han, Quanlu Zhang, Zhenhua Li 0001, Haisheng Tan
INFOCOM3
2020 Retiarii: A Deep Learning Exploratory-Training Framework
Quanlu Zhang, Zhenhua Han, Fan Yang 0024, Yuge Zhang, Mao Yang 0004, Lidong Zhou
OSDI1
2020 HiveD: Sharing a GPU Cluster for Deep Learning with Guarantees
Zhenhua Han, Zhi Yang 0001, Quanlu Zhang, Fan Yang 0024, Lidong Zhou, Mao Yang 0004, Francis C. M. Lau 0001, Yifan Xiong 0001
OSDI4
2020 AutoSys: The Design and Operation of Learning-Augmented Systems
Chieh-Jan Mike Liang, Hui Xue 0004, Mao Yang 0004, Lidong Zhou, Lifei Zhu, Zhao Lucis Li, Qi Chen 0009, Quanlu Zhang, Chuanjie Liu, Wenjun Dai
USENIX ATC9
2018 Scheduling CPU for GPU-based Deep Learning Jobs
abstract
Deep learning (DL) is popular in data-center as an important workload for artificial intelligence. With the recent breakthrough of using graphics accelerators and the popularity of DL framework, GPU server cluster dominates DL training in current practice. Cluster scheduler simply treats DL jobs as black-boxes and allocates GPUs as per job request specified by a user. However, other resources, e.g. CPU, are often allocated with workload-agnostic approaches. Kubeflow[1] performs heuristic static CPU resource assignment based on task types (e.g., worker, parameter-server), while [2] evenly divides CPUs of a server to each GPU. Despite the traditional impression that GPU is critical in DL, our observation suggests that the importance of CPU is undervalued. Identifying an appropriate CPU core number in a heterogeneous cluster is challenging yet performance critical to DL jobs. The diverse CPU usage characteristic is not well recognized in the following three aspects.
Wencong Xiao, Zhenhua Han, Quanlu Zhang, Fan Yang 0024, Lidong Zhou
SoCC5
2018 SDPaxos: Building Efficient Semi-Decentralized Geo-replicated State Machines
abstract
Existing state machine replication protocols are confronting two major challenges in geo-replication: (1) limited performance caused by load imbalance, and (2) severe performance degradation in heterogeneous environments or under high-contention workloads. This paper presents a new semi-decentralized approach to addressing both the challenges at the same time. Our protocol, SDPaxos, divides the task of a replication protocol into two parts: durably replicating each command across replicas without global order, and ordering all commands to enforce the consistency guarantee. We decentralize the process of replicating commands, which accounts for the largest proportion of load, to provide high performance. In contrast, we centralize the process of ordering commands, which is lightweight but needs a global view, for better performance stability against heterogeneity or contention. The key novelty lies in that SDPaxos achieves the optimal one-round-trip latency under realistic configurations, despite the two separated steps, replicating and ordering, which are both based on Paxos. We also design a recovery protocol to do rapid failover under failures, and a series of optimizations to boost performance. We show via a prototype implementation the significant advantage of SDPaxos on both throughput and latency, facing different environments and workloads.
Quanlu Zhang, Zhi Yang 0001, Ming Wu 0007, Yafei Dai
SoCC2
2018 Towards Web-based Delta Synchronization for Cloud Storage Services
Zhenhua Li 0001, Ennan Zhai, Tianyin Xu, Yang Li 0092, Yunhao Liu 0001, Quanlu Zhang, Yao Liu 0001
FAST7
2018 Gandiva: Introspective Cluster Scheduling for Deep Learning
Wencong Xiao, Romil Bhardwaj, Ramachandran Ramjee, Muthian Sivathanu, Nipun Kwatra, Zhenhua Han, Pratyush Patel, Quanlu Zhang, Fan Yang 0024, Lidong Zhou
OSDI10
2017 DeltaCFS: Boosting Delta Sync for Cloud Storage Services by Learning from NFS
abstract
Cloud storage services, such as Dropbox, iCloud Drive, Google Drive, and Microsoft OneDrive, have greatly facilitated users' synchronizing files across heterogeneous devices. Among them, Dropbox-like services are particularly beneficial owing to the delta sync functionality that strives towards greater network-level efficiency. However, when delta sync trades computation overhead for network-traffic saving, the tradeoff could be highly unfavorable under some typical workloads. We refer to this problem as the abuse of delta sync. To address this problem, we propose DeltaCFS, a novel file sync framework for cloud storage services by learning from the design of conventional NFS (Network File System). Specifically, we combine delta sync with NFS-like file RPC in an adaptive manner, thus significantly cutting computation overhead on both the client and server sides while preserving the network-level efficiency. DeltaCFS also enables a neat design for guaranteeing causal consistency and fine-grained version control of files. In our FUSE-based prototype system (which is open-source), DeltaCFS outperforms Dropbox by generating up to 11x less data transfer and up to 100x less computation overhead under concerned workloads.
Quanlu Zhang, Zhenhua Li 0001, Zhi Yang 0001, Shenglong Li, Shouyang Li, Yangze Guo, Yafei Dai
ICDCS1
2015 DSwitch: a dual mode direct and network attached disk
abstract
Putting computers into low power mode (e.g., suspend-to-RAM) could potentially save significant amount of power when the computers are not in use. Unfortunately, this is often infeasible in practice because data stored on the computers (i.e., directly attached disks, DAS) might need to be accessed by others. Separating storage from computation by attaching storage on the network (e.g., NAS and SAN) could potentially solve this problem, at the cost of lower performance, more network congestion, increased peak power consumption, and higher equipment cost. Though DAS does not suffer these problems, it is not flexible for power saving. In this paper, we present DSwitch, an architecture that, depending on the workload, allows a disk to be attached either directly or through network. We design flexible workload migration based on DSwitch, and show that a wide variety of applications in both data center and home/office settings can be well supported. The experiments demonstrate that our prototype DSwitch achieves a power savings of 91.9% to 97.5% when a disk is in low power network attached mode, while incurring no performance degradation and minimal power overhead when it is in high performance directly attached mode.
Quanlu Zhang, Yafei Dai
SoCC1
2015 Understanding and Surpassing Dropbox: Efficient Incremental Synchronization in Cloud Storage Services
abstract
Cloud storage services allow files to be synchronized among multiple users or devices easily. To minimize the amount of network traffic, these services utilize incremental data synchronization techniques. However, little is known about their particular mechanisms and corresponding efficiency. In this paper, we focus on Dropbox, the most popular cloud storage service, as a case study. We examine the bandwidth consumption on the Dropbox client under typical synchronization scenarios, and find that Dropbox wastes a lot of traffic due to matching unrelated chunks to compute incremental changes. More seriously, when file conflicts among clients, Dropbox directly downloads the whole file instead of the incremental changes. To solve these problems, we design and implement an efficient incremental synchronization system named Minbox. MinBox employs an efficient locality-sensitive hash for better similar chunk matching. Moreover, Minbox could forward the incremental changes during confliction by maintaining the relation of conflicting files. In comparison with Dropbox, Minbox significantly reduces network traffic and resolves file conflict with little overhead.
Shenglong Li, Quanlu Zhang, Zhi Yang 0001, Yafei Dai
GLOBECOM2
2015 UStore: A Low Cost Cold and Archival Data Storage System for Data Centers
abstract
Recent trend in cloud computing demands vast and ever increasing storage capacity for data centers. For many cloud service providers, much of the storage capacity demand is driven by cold and archival data, such as user uploaded contents, system logs, and backups. In this paper, we describe UStore, a hard disk based storage system designed for such workloads. We make the assumption that most data centers are already populated with computer servers and networking gears, and propose a solution to attach additional disks to these servers reliably at extremely low cost. The main component of UStore is a novel fat tree interconnect fabric to connect hard disks to existing servers and network infrastructure. To reduce cost, UStore leverages the mature commodity USB 3.0 technology to build the fabric, which has extremely low amortized cost per disk while still providing sufficient throughput to satisfy cold and archival workload. The software of the UStore system abstracts the system's physical topology and provides a consistent view of the storage capacity to the upper layer services such as distributed file systems or backup services. In a sense, UStore can be regarded as external USB hard disks designed for data centers.
Quanlu Zhang, Yafei Dai, Fengqian Li
ICDCS1
2015 CHARM: A Cost-Efficient Multi-Cloud Data Hosting Scheme with High Availability
abstract
Nowadays, more and more enterprises and organizations are hosting their data into the cloud, in order to reduce the IT maintenance cost and enhance the data reliability. However, facing the numerous cloud vendors as well as their heterogenous pricing policies, customers maywell be perplexed with which cloud(s) are suitable for storing their data and what hosting strategy is cheaper.The general status quo is that customers usually put their data into a single cloud (which is subject to the vendor lock-in risk) and then simply trust to luck. Based on comprehensive analysis of various state-of-the-art cloud vendors, this paper proposes a novel data hosting scheme (named CHARM) which integrates two key functions desired. The first is selecting several suitable clouds and an appropriate redundancy strategy to store data with minimized monetary cost and guaranteed availability. The second is triggering a transition process to re-distribute data according to the variations of data access pattern and pricing of clouds. We evaluate the performance of CHARM using both trace-driven simulations and prototype experiments. The results show that compared with the major existing schemes, CHARM not only saves around 20 percent of monetary cost but also exhibits sound adaptability to data and price adjustments.
Quanlu Zhang, Shenglong Li, Zhenhua Li 0001, Yuanjian Xing, Zhi Yang 0001, Yafei Dai
IEEE Trans. Cloud Comput.1