Jiachen Mao

dblp:178/5589 · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
4since 2021 · last 2026
0000-0001-8986-0696ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 5 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 first-authorArtificial intelligence and machine learning · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 FIITED: Fine-Grained Embedding Dimension Optimization During Training for Recommender Systems
abstract
Huge embedding tables in modern deep learning recommender models (DLRM) require prohibitively large memory during training and inference. This paper proposes FIITED, a system to automatically reduce the memory footprint via FIne-grained In-Training Embedding Dimension pruning. By leveraging the key insight that embedding vectors are not equally important, FIITED adaptively adjusts the dimension of each individual embedding vector during model training, assigning larger dimensions to more important embeddings while adapting to dynamic changes in data. We prioritize embedding dimensions with higher frequencies and gradients as more important. To enable efficient pruning of embeddings and their dimensions during model training, we propose an embedding storage system based on virtually-hashed physically-indexed hash tables. Experiments on two industry models and months of realistic datasets show that FIITED can reduce DLRM embedding size by more than 65’ while preserving model quality, outperforming state-of-the-art in-training embedding pruning methods and increasing the reduction ratio by 1.3× to 1.67×. On public datasets, FIITED can reduce the size of embedding tables by 2.2× to 800× with negligible accuracy drop, achieving 1.05× to 12.5× improvement in reduction ratio compared to baselines while improving model throughput.
Qinyi Luo, Penghan Wang, Wei Zhang 0044, Fan Lai 0001, Jiachen Mao, Xiaohan Wei, Wei-Yu Tsai, Yuxi Hu 0001, Xuehai Qian
IEEE Trans. Computers5
2022 Toward Efficient and Adaptive Design of Video Detection System with Deep Neural Networks
abstract
In the past decade, Deep Neural Networks (DNNs), e.g., Convolutional Neural Networks, achieved human-level performance in vision tasks such as object classification and detection. However, DNNs are known to be computationally expensive and thus hard to be deployed in real-time and edge applications. Many previous works have focused on DNN model compression to obtain smaller parameter sizes and consequently, less computational cost. Such methods, however, often introduce noticeable accuracy degradation. In this work, we optimize a state-of-the-art DNN-based video detection framework—Deep Feature Flow (DFF) from the cloud end using three proposed ideas. First, we propose Asynchronous DFF (ADFF) to asynchronously execute the neural networks. Second, we propose a Video-based Dynamic Scheduling (VDS) method that decides the detection frequency based on the magnitude of movement between video frames. Last, we propose Spatial Sparsity Inference, which only performs the inference on part of the video frame and thus reduces the computation cost. According to our experimental results, ADFF can reduce the bottleneck latency from 89 to 19 ms. VDS increases the detection accuracy by 0.6% mAP without increasing computation cost. And SSI further saves 0.2 ms with a 0.6% mAP degradation of detection accuracy.
Jiachen Mao, Qing Yang 0011, Ang Li 0005, Kent W. Nixon, Hai Li 0001, Yiran Chen 0001
ACM Trans. Embed. Comput. Syst.1
2021 Dynamic Regularization on Activation Sparsity for Neural Network Efficiency Improvement
abstract
When deploying deep neural networks in embedded systems, it is crucial to decrease the model size and computational complexity for improving the execution speed and efficiency. In addition to conventional compression techniques, e.g., weight pruning and quantization, removing unimportant activations can also dramatically reduce the amount of data communication and the computation cost. Unlike weight parameters, the pattern of activations is directly related to input data and thereby changes dynamically. To regulate the dynamic activation sparsity (DAS), in this work, we propose a generic low-cost approach based on winners-take-all (WTA) dropout technique. The network enhanced by the proposed WTA dropout, namely DASNet , features structured activation sparsity with an improved sparsity level. Compared to the static feature map pruning methods, DASNets provide better computation cost reduction. The WTA dropout technique can be easily applied in deep neural networks without incurring additional training variables. More importantly, DASNet can be seamlessly integrated with other compression techniques, such as weight pruning and quantization, without compromising accuracy. Our experiments on various networks and datasets present significant runtime speedups with negligible accuracy losses.
Qing Yang 0011, Jiachen Mao, Zuoguan Wang, Hai Li 0001
ACM J. Emerg. Technol. Comput. Syst.2
2021 TPrune: Efficient Transformer Pruning for Mobile Devices
abstract
The invention of Transformer model structure boosts the performance of Neural Machine Translation (NMT) tasks to an unprecedented level. Many previous works have been done to make the Transformer model more execution-friendly on resource-constrained platforms. These researches can be categorized into three key fields: Model Pruning, Transfer Learning, and Efficient Transformer Variants. The family of model pruning methods are popular for their simplicity in practice and promising compression rate and have achieved great success in the field of convolution neural networks (CNNs) for many vision tasks. Nonetheless, previous Transformer pruning works did not perform a thorough model analysis and evaluation on each Transformer component on off-the-shelf mobile devices. In this work, we analyze and prune transformer models at the line-wise granularity and also implement our pruning method on real mobile platforms. We explore the properties of all Transformer components as well as their sparsity features, which are leveraged to guide Transformer model pruning. We name our whole Transformer analysis and pruning pipeline as TPrune. In TPrune, we first propose Block-wise Structured Sparsity Learning (BSSL) to analyze Transformer model property. Then, based on the characters derived from BSSL, we apply Structured Hoyer Square (SHS) to derive the final pruned models. Comparing with the state-of-the-art Transformer pruning methods, TPrune is able to achieve a higher model compression rate with less performance degradation. Experimental results show that our pruned models achieve 1.16×–1.92× speedup on mobile devices with 0%–8% BLEU score degradation compared with the original Transformer model.
Jiachen Mao, Huanrui Yang, Ang Li 0005, Hai Li 0001, Yiran Chen 0001
ACM Trans. Cyber Phys. Syst.1
2019 NeuralHMC: an efficient HMC-based accelerator for deep neural networks
abstract
In Deep Neural Network (DNN) applications, energy consumption and performance cost of moving data between memory hierarchy and computational units are significantly higher than that of the computation itself. Process-in-memory (PIM) architecture such as Hybrid Memory Cube (HMC), becomes an excellent candidate to improve the data locality for efficient DNN execution. However, it's still hard to efficiently deploy large-scale matrix computation in DNN on HMC because of its coarse grained packet protocol. In this work, we propose NeuralHMC, the first HMC-based accelerator tailored for efficient DNN execution. Experimental results show that NeuralHMC reduces the data movement by 1.4x to 2.5x (depending on the DNN data reuse strategy) compared to Von Neumann architecture. Furthermore, compared to state-of-the-art PIM-based DNN accelerator, NeuralHMC can promisingly improve the system performance by 4.1x and reduces energy by 1.5x, on average.
Chuhan Min, Jiachen Mao, Hai Li 0001, Yiran Chen 0001
ASP-DAC2
2019 MobiEye: An Efficient Cloud-based Video Detection System for Real-time Mobile Applications
abstract
In recent years, machine learning research has largely shifted focus from the cloud to the edge. While the resulting algorithm- and hardware-level optimizations have enabled local execution for the majority of deep neural networks (DNNs) on edge devices, the sheer magnitude of DNNs associated with real-time video detection workloads has forced them to remain relegated to remote execution in the cloud. This problematic when combined with the strict latency requirements that are coupled with these workloads, and imposes a unique set of challenges not directly addressed in prior works. In this work, we design MobiEye, a cloud-based video detection system optimized for deployment in real-time mobile applications. MobiEye is able to achieve up to a 32% reduction in latency when compared to a conventional implementation of video detection system with only a marginal reduction in accuracy.
Jiachen Mao, Qing Yang 0011, Ang Li 0005, Hai Li 0001, Yiran Chen 0001
DAC1
2019 HyPar: Towards Hybrid Parallelism for Deep Learning Accelerator Array
abstract
With the rise of artificial intelligence in recent years, Deep Neural Networks (DNNs) have been widely used in many domains. To achieve high performance and energy efficiency, hardware acceleration (especially inference) of DNNs is intensively studied both in academia and industry. However, we still face two challenges: large DNN models and datasets, which incur frequent off-chip memory accesses; and the training of DNNs, which is not well-explored in recent accelerator designs. To truly provide high throughput and energy efficient acceleration for the training of deep and large models, we inevitably need to use multiple accelerators to explore the coarse-grain parallelism, compared to the fine-grain parallelism inside a layer considered in most of the existing architectures. It poses the key research question to seek the best organization of computation and dataflow among accelerators. In this paper, we propose a solution HYPAR to determine layer-wise parallelism for deep neural network training with an array of DNN accelerators. HYPAR partitions the feature map tensors (input and output), the kernel tensors, the gradient tensors, and the error tensors for the DNN accelerators. A partition constitutes the choice of parallelism for weighted layers. The optimization target is to search a partition that minimizes the total communication during training a complete DNN. To solve this problem, we propose a communication model to explain the source and amount of communications. Then, we use a hierarchical layer-wise dynamic programming method to search for the partition for each layer. HYPAR is practical: the time complexity for the partition search in HYPAR is linear. We apply this method in an HMC-based DNN training architecture to minimize data movement. We evaluate HYPAR with ten DNN models from classic Lenet to large-size model VGGs, and the number of weighted layers of these models range from four to nineteen. Our evaluation finds that: the default Model Parallelism is indeed the worst; the default Data Parallelism is not the best; but hybrid parallelism can be better than either the default Data Parallelism or Model Parallelism in DNN training with an array of accelerators. Our evaluation shows that HYPAR achieves a performance gain of 3.39× and an energy efficiency gain of 1.51× compared to Data Parallelism on average, and HYPAR performs up to 2.40× better than “one weird trick”.
Linghao Song, Jiachen Mao, Youwei Zhuo, Xuehai Qian, Hai Li 0001, Yiran Chen 0001
HPCA2
2019 DASNet: Dynamic Activation Sparsity for Neural Network Efficiency Improvement
abstract
To improve the execution speed and efficiency of neural networks in embedded systems, it is crucial to decrease the model size and computational complexity. In addition to conventional compression techniques, e.g., weight pruning and quantization, removing unimportant activations can reduce the amount of data communication and the computation cost. Unlike weight parameters, the pattern of activations is directly related to input data and thereby changes dynamically. To regulate the dynamic activation sparsity (DAS), in this work, we propose a generic low-cost approach based on winners-take-all (WTA) dropout technique. The network enhanced by the proposed WTA dropout, namely DASNet, features structured activation sparsity with an improved sparsity level. Compared to the static feature map pruning methods, DASNets provide better computation cost reduction. The WTA technique can be easily applied in deep neural networks without incurring additional training variables. Our experiments on various networks and datasets present significant run-time speedups with negligible accuracy loss.
Qing Yang 0011, Jiachen Mao, Zuoguan Wang, Hai Li 0001
ICTAI2
2018 Running sparse and low-precision neural network: When algorithm meets hardware
abstract
Deep Neural Networks (DNNs) are pervasively applied in many artificial intelligence (AI) applications. The high performance of DNNs comes at the cost of larger size and higher compute complexity. Recent studies show that DNNs have much redundancy, such as the zero-value parameters and excessive numerical precision. To reduce computing complexity, many redundancy reduction techniques have been proposed, including pruning and data quantization. In this paper, we demonstrate our co-optimization of the DNN algorithm and hardware which exploits the model redundancy to accelerate DNNs.
Bing Li 0017, Wei Wen 0003, Jiachen Mao, Sicheng Li 0001, Yiran Chen 0001, Hai Li 0001
ASP-DAC3
2018 SPN dash: fast detection of adversarial attacks on mobile via sensor pattern noise fingerprinting
abstract
A concerning weakness of deep neural networks is their susceptibility to adversarial attacks. While methods exist to detect these attacks, they incur significant drawbacks, ignoring external features which could aid in the task of attack detection. In this work, we propose SPN Dash, a method for detection of adversarial attacks based on integrity of sensor pattern noise embedded in submitted images. Through experiment, we show that our SPN Dash method is capable of detecting the addition of adversarial noise with up to 94% accuracy for images of size $256\times256$. Analysis shows that SPN Dash is robust to image scaling techniques, as well as a small amount of image compression. This performance is on par with state of the art neural network-based detectors, while incurring an order of magnitude less computational and memory overhead.
Kent W. Nixon, Jiachen Mao, Juncheng Shen, Huanrui Yang, Hai Li 0001, Yiran Chen 0001
ICCAD2
2017 MoDNN: Local distributed mobile computing system for Deep Neural Network
abstract
Although Deep Neural Networks (DNN) are ubiquitously utilized in many applications, it is generally difficult to deploy DNNs on resource-constrained devices, e.g., mobile platforms. Some existing attempts mainly focus on client-server computing paradigm or DNN model compression, which require either infrastructure supports or special training phases, respectively. In this work, we propose MoDNN - a local distributed mobile computing system for DNN applications. MoDNN can partition already trained DNN models onto several mobile devices to accelerate DNN computations by alleviating device-level computing cost and memory usage. TWo model partition schemes are also designed to minimize non-parallel data delivery time, including both wakeup time and transmission time. Experimental results show that when the number of worker nodes increases from 2 to 4, MoDNN can accelerate the DNN computation by 2.17-4.28 X. Besides the parallel execution, the performance speedup also partially comes from the reduction of the data delivery time, e.g., 30.02% w.r.t. conventional 2D-grids partition.
Jiachen Mao, Xiang Chen 0010, Kent W. Nixon, Christopher D. Krieger, Yiran Chen 0001
DATE1
2017 AdaLearner: An adaptive distributed mobile learning system for neural networks
abstract
Neural networks hold a critical domain in machine learning algorithms because of their self-adaptiveness and state-of-the-art performance. Before the testing (inference) phases in practical use, sophisticated training (learning) phases are required, calling for efficient training methods with higher accuracy and shorter converging time. Many existing studies focus on the training optimization on high-performance servers or computing clusters, e.g. GPU clusters. However, training neural networks on resource-constrained devices, e.g. mobile platforms, is an important research topic barely touched. In this paper, we implement AdaLearner-an adaptive distributed mobile learning system for neural networks that trains a single network with heterogenous mobile resources under the same local network in parallel. To exploit the potential of our system, we adapt neural networks training phase to mobile device-wise resources and fiercely decrease the transmission overhead for better system scalability. On three representative neural network structures trained from two image classification datasets, AdaLearner boosts the training phase significantly. For example, on LeNet, 1.75-3.37× speedup is achieved when increasing the worker nodes from 2 to 8, thanks to the achieved high execution parallelism and excellent scalability.
Jiachen Mao, Zhuwei Qin, Kent W. Nixon, Xiang Chen 0010, Hai Li 0001, Yiran Chen 0001
ICCAD1
2017 MeDNN: A distributed mobile system with enhanced partition and deployment for large-scale DNNs
abstract
Deep Neural Networks (DNNs) are pervasively used in a significant number of applications and platforms. To enhance the execution efficiency of large-scale DNNs, previous attempts focus mainly on client-server paradigms, relying on powerful external infrastructure, or model compression, with complicated pre-processing phases. Though effective, these methods overlook the optimization of DNNs on distributed mobile devices. In this work, we design and implement MeDNN, a local distributed mobile computing system with enhanced partitioning and deployment tailored for large-scale DNNs. In MeDNN, we first propose Greedy Two Dimensional Partition (GTDP), which can adaptively partition DNN models onto several mobile devices w.r.t. individual resource constraints. We also propose Structured Model Compact Deployment (SMCD), a mobile-friendly compression scheme which utilizes a structured sparsity pruning technique to further accelerate DNN execution. Experimental results show that, GTDP can accelerate the original DNN execution time by 1.86-2.44x with 2-4 worker nodes. By utilizing SMCD, 26.5% of additional computing time and 14.2% of extra communication time are saved, on average, with negligible effect on the model accuracy.
Jiachen Mao, Zhongda Yang, Wei Wen 0003, Chunpeng Wu, Linghao Song, Kent W. Nixon, Xiang Chen 0010, Hai Li 0001, Yiran Chen 0001
ICCAD1
2016 MORPh: mobile OLED-friendly recording and playback system for low power video streaming
abstract
Even with the adoption of the latest OLED technology, the display panel remains one of the most power-hungry components in smartphones. Existing attempts for OLED power optimization have mainly focused on modifying the content that is shown on the display during the playback phase, requiring significant overhead in terms of image analysis and modification. While such methods are effective, they overlook opportunities present during the camera recording phase, where utilization of already determined camera parameters could reduce or eliminate the image processing overhead. Hence, we proposed MORPh, a cross-layer optimization system for OLED. We first analyze three fundamental parameters extracted from the smartphone camera and their impact on OLED power consumption and visual quality. We then define corresponding metrics to assess power optimization potentials, and propose a set of algorithms that optimize camera recording to be more OLED friendly. Finally, the proposed schemes are realized as part of a video recording and playback application on an existing Android smartphone. The experiments results indicate power saving of 7.3%~39.7%, and 20.3% on average while maintaining perceived visual quality.
Xiang Chen 0010, Jiachen Mao, Jiafei Gao, Kent W. Nixon, Yiran Chen 0001
DAC2
2016 MORPh: mobile OLED power friendly camera system
abstract
With superior advantages of better display quality and power efficiency, the latest OLED technology has achieved unprecedented popularity in the display screen market. However, the OLED remains one of the most power-hungry components in mobile devices. Various optimization schemes have been proposed based on the color-dependent power consumption feature of OLED pixels. These schemes mainly focus on color modification during the playback phase and require significant overhead in terms of frame analysis and real-time modification. While such schemes are effective, the power saving opportunities during the camera recording phase are overlooked. To further enhance the power optimization, the camera parameters during the recording phase could be effectively utilized to reduce or eliminate the optimization overhead. Hence, we proposed MORPh, a cross-layer optimization system for OLED display in the smartphones. We analyze three fundamental parameters of smartphone camera system and their impact on the OLED screen power consumption. We then define corresponding metrics to quantitatively assess each parameter's potential of power saving guidance. Finally, we develop a set of schemes and integrated them into a video recording and playback application on an existing Android smartphone. The experiments results indicate power saving of 7.3%∼39.7%, and 20.3% on average while maintaining perceived visual quality.
Xiang Chen 0010, Jiachen Mao, Kent W. Nixon, Yiran Chen 0001
RSP2