Jing Pu

dblp:131/5127 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
1since 2021 · last 2026
0009-0009-2032-4751ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-authorSoftware engineering, systems software and programming languages · 5Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Hardware accelerators and domain-specific architectures · 57% Parallel and multicore computing · 25% Memory systems · 11%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%
Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 100%

Topics — the 15 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator
0.722019
TANGRAM: Optimized Coarse-Grained Dataflow for Scalable NN Accelerators · ASPLOS 2019
TETRIS: Scalable and Efficient Neural Network Acceleration with 3D Memory · ASPLOS 2017
Parallel and multicore computing
dataflow computing
0.412020
Interstellar: Using Halide's Scheduling Language to Analyze DNN Accelerators · ASPLOS 2020
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN accelerator
0.412020
Interstellar: Using Halide's Scheduling Language to Analyze DNN Accelerators · ASPLOS 2020
Parallel and multicore computing › parallel scheduling
loop scheduling
0.412020
Interstellar: Using Halide's Scheduling Language to Analyze DNN Accelerators · ASPLOS 2020
Hardware accelerators and domain-specific architectures › neural network mapping
inter-layer pipelining
0.412019
TANGRAM: Optimized Coarse-Grained Dataflow for Scalable NN Accelerators · ASPLOS 2019
Memory systems
DRAM
0.312017
TETRIS: Scalable and Efficient Neural Network Acceleration with 3D Memory · ASPLOS 2017
Reconfigurable computing and FPGAs
FPGA accelerator
0.312017
Programming Heterogeneous Systems from an Image Processing DSL · ACM Trans. Archit. Code Optim. 2017
Hardware accelerators and domain-specific architectures
image processing accelerator
0.312017
Programming Heterogeneous Systems from an Image Processing DSL · ACM Trans. Archit. Code Optim. 2017
Machine learning › Efficient and distributed learning
model compression
0.212016
EIE: Efficient Inference Engine on Compressed Deep Neural Network · ISCA 2016
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN inference accelerator
0.212016
EIE: Efficient Inference Engine on Compressed Deep Neural Network · ISCA 2016
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.212016
EIE: Efficient Inference Engine on Compressed Deep Neural Network · ISCA 2016
Compilers and program optimization
scheduling language
0.112020
Interstellar: Using Halide's Scheduling Language to Analyze DNN Accelerators · ASPLOS 2020
Parallel and multicore computing › parallelization strategies
coarse-grained parallelism
0.112019
TANGRAM: Optimized Coarse-Grained Dataflow for Scalable NN Accelerators · ASPLOS 2019
Compilers and program optimization
domain-specific compilation
0.112017
Programming Heterogeneous Systems from an Image Processing DSL · ACM Trans. Archit. Code Optim. 2017
Memory systems › on-chip memory
on-chip SRAM
0.112016
EIE: Efficient Inference Engine on Compressed Deep Neural Network · ISCA 2016

Methods — techniques the papers use, named apart from their topics

halide scheduling language · 0.9hardware-software co-design · 0.6Halide DSL · 0.6zero-activation skipping · 0.5weight sharing · 0.5weight pruning · 0.5sparse matrix-vector multiplication · 0.5buffer sharing dataflow · 0.4alternate layer loop ordering · 0.43d memory stacking · 0.3
YearPublicationVenuePosition
2026 LexiAlign: a diffusion model for text alignment and font restoration based on local regeneration
Weijia Zhu, Xinjin Li, Jing Pu, Minglu Wang
Vis. Comput.3
2020 Interstellar: Using Halide's Scheduling Language to Analyze DNN Accelerators
abstract
We show that DNN accelerator micro-architectures and their program mappings represent specific choices of loop order and hardware parallelism for computing the seven nested loops of DNNs, which enables us to create a formal taxonomy of all existing dense DNN accelerators. Surprisingly, the loop transformations needed to create these hardware variants can be precisely and concisely represented by Halide's scheduling language. By modifying the Halide compiler to generate hardware, we create a system that can fairly compare these prior accelerators. As long as proper loop blocking schemes are used, and the hardware can support mapping replicated loops, many different hardware dataflows yield similar energy efficiency with good performance. This is because the loop blocking can ensure that most data references stay on-chip with good locality and the processing units have high resource utilization. How resources are allocated, especially in the memory system, has a large impact on energy and performance. By optimizing hardware resource allocation while keeping throughput constant, we achieve up to 4.2X energy improvement for Convolutional Neural Networks (CNNs), 1.6X and 1.8X improvement for Long Short-Term Memories (LSTMs) and multi-layer perceptrons (MLPs), respectively.
Mingyu Gao 0001, Qiaoyi Liu, Jeff Setter, Jing Pu, Ankita Nayak, Steven Bell, Kaidi Cao, Heonjae Ha, Priyanka Raina, Christoforos E. Kozyrakis, Mark Horowitz
ASPLOS5
2019 TANGRAM: Optimized Coarse-Grained Dataflow for Scalable NN Accelerators
abstract
The use of increasingly larger and more complex neural networks (NNs) makes it critical to scale the capabilities and efficiency of NN accelerators. Tiled architectures provide an intuitive scaling solution that supports both coarse-grained parallelism in NNs: intra-layer parallelism, where all tiles process a single layer, and inter-layer pipelining, where multiple layers execute across tiles in a pipelined manner. This work proposes dataflow optimizations to address the shortcomings of existing parallel dataflow techniques for tiled NN accelerators. For intra-layer parallelism, we develop buffer sharing dataflow that turns the distributed buffers into an idealized shared buffer, eliminating excessive data duplication and the memory access overheads. For inter-layer pipelining, we develop alternate layer loop ordering that forwards the intermediate data in a more fine-grained and timely manner, reducing the buffer requirements and pipeline delays. We also make inter-layer pipelining applicable to NNs with complex DAG structures. These optimizations improve the performance of tiled NN accelerators by 2x and reduce their energy consumption by 45% across a wide range of NNs. The effectiveness of our optimizations also increases with the NN size and complexity.
Mingyu Gao 0001, Jing Pu, Mark Horowitz, Christoforos E. Kozyrakis
ASPLOS3
2018 Performance and programming effort trade-offs of android persistence frameworks
Zheng Song 0001, Jing Pu, Junjie Cheng, Eli Tilevich
J. Syst. Softw.2
2017 TETRIS: Scalable and Efficient Neural Network Acceleration with 3D Memory
abstract
The high accuracy of deep neural networks (NNs) has led to the development of NN accelerators that improve performance by two orders of magnitude. However, scaling these accelerators for higher performance with increasingly larger NNs exacerbates the cost and energy overheads of their memory systems, including the on-chip SRAM buffers and the off-chip DRAM channels.
Mingyu Gao 0001, Jing Pu, Mark Horowitz, Christoforos E. Kozyrakis
ASPLOS2
2017 Programming Heterogeneous Systems from an Image Processing DSL
abstract
Specialized image processing accelerators are necessary to deliver the performance and energy efficiency required by important applications in computer vision, computational photography, and augmented reality. But creating, “programming,” and integrating this hardware into a hardware/software system is difficult. We address this problem by extending the image processing language Halide so users can specify which portions of their applications should become hardware accelerators, and then we provide a compiler that uses this code to automatically create the accelerator along with the “glue” code needed for the user’s application to access this hardware. Starting with Halide not only provides a very high-level functional description of the hardware but also allows our compiler to generate a complete software application, which accesses the hardware for acceleration when appropriate. Our system also provides high-level semantics to explore different mappings of applications to a heterogeneous system, including the flexibility of being able to change the throughput rate of the generated hardware. We demonstrate our approach by mapping applications to a commercial Xilinx Zynq system. Using its FPGA with two low-power ARM cores, our design achieves up to 6× higher performance and 38× lower energy compared to the quad-core ARM CPU on an NVIDIA Tegra K1, and 3.5× higher performance with 12× lower energy compared to the K1’s 192-core GPU.
Jing Pu, Steven Bell, Jeff Setter, Stephen Richardson, Jonathan Ragan-Kelley, Mark Horowitz
ACM Trans. Archit. Code Optim.1
2016 Deep compression and EIE: Efficient inference engine on compressed deep neural network
Song Han 0003, Xingyu Liu 0001, Huizi Mao, Jing Pu, Ardavan Pedram, Mark Horowitz, William J. Dally
Hot Chips Symposium4
2016 EIE: Efficient Inference Engine on Compressed Deep Neural Network
abstract
State-of-the-art deep neural networks (DNNs) have hundreds of millions of connections and are both computationally and memory intensive, making them difficult to deploy on embedded systems with limited hardware resources and power budgets. While custom hardware helps the computation, fetching weights from DRAM is two orders of magnitude more expensive than ALU operations, and dominates the required power. Previously proposed 'Deep Compression' makes it possible to fit large DNNs (AlexNet and VGGNet) fully in on-chip SRAM. This compression is achieved by pruning the redundant connections and having multiple connections share the same weight. We propose an energy efficient inference engine (EIE) that performs inference on this compressed network model and accelerates the resulting sparse matrix-vector multiplication with weight sharing. Going from DRAM to SRAM gives EIE 120x energy saving, Exploiting sparsity saves 10x, Weight sharing gives 8x, Skipping zero activations from ReLU saves another 3x. Evaluated on nine DNN benchmarks, EIE is 189x and 13x faster when compared to CPU and GPU implementations of the same DNN without compression. EIE has a processing power of 102 GOPS working directly on a compressed network, corresponding to 3 TOPS on an uncompressed network, and processes FC layers of AlexNet at 1.88x104 frames/sec with a power dissipation of only 600mW. It is 24,000x and 3,400x more energy efficient than a CPU and GPU respectively. Compared with DaDianNao, EIE has 2.9x, 19x and 3x better throughput, energy efficiency and area efficiency.
Song Han 0003, Xingyu Liu 0001, Huizi Mao, Jing Pu, Ardavan Pedram, Mark Horowitz, William J. Dally
ISCA4
2016 Understanding the Energy, Performance, and Programming Effort Trade-Offs of Android Persistence Frameworks
abstract
One of the fundamental building blocks of a mobile application is the ability to persist program data between different invocations. Referred to as persistence, this functionality is commonly implemented by means of persistence frameworks. When choosing a particular framework, Android—the most popular mobile platform—offers a wide variety of options to developers. Unfortunately, the energy, performance, and programming effort trade-offs of these frameworks are poorly understood, leaving the Android developer in the dark trying to select the most appropriate option for their applications. To address this problem, this paper reports on the results of the first systematic study of six Android persistence frameworks (i.e., ActiveAndroid, greenDAO, OrmLite, Sugar ORM, Android SQLite, and Java Realm) in their application to and performance with popular benchmarks, such as DaCapo. Having measured and analyzed the energy, performance, and programming effort trade-offs for each framework, we present a set of practical guidelines for the developer to choose between Android persistence frameworks. Our findings can also help the framework developers to optimize their products to meet the desired design objectives.
Jing Pu, Zheng Song 0001, Eli Tilevich
MASCOTS1
2013 FPU Generator for Design Space Exploration
abstract
FPUs have been a topic of research for almost a century, leading to thousands of papers and books. Each advance focuses on the virtues of some specific new technique. This paper compares the energy efficiency of both throughput-optimized and latency-sensitive designs, each employing an array of optimization techniques, through a fair "apples to apples" methodology. This comparison required us to build many optimized FP units. We accomplished this by creating a highly parameterized FPgenerator, hierarchically encompassing lower-level generators for summation trees, Booth encoders, adders, etc. Having constructed this generator we quickly relearned a number of low-level issues that are critical and are often the most neglected by papers. By exploring cascade and fused multiply-add architectures across a variety of bit widths, summation trees, booth encoders, pipelining techniques, and pipe depths, we found that for most throughput based designs, a Booth-3 fused multiply-add architecture with a Wallace combining tree is optimal. For latency designs, we found that Booth-2 cascade multiply-add architectures are better. As we describe in the paper, Wallace is not always the optimal combining network due to wire delay and track count, and the precise way the CSA's are connected in the tree can make a larger difference than the type of tree used.
Sameh Galal, Ofer Shacham, John S. Brunhaver, Jing Pu, Artem Vassiliev, Mark Horowitz
IEEE Symposium on Computer Arithmetic4