Shunning Jiang

dblp:168/6430 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
2since 2021 · last 2023
0000-0003-3439-5760ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Electronic design automation · 36% Performance modeling and evaluation · 20% Processor architecture and microarchitecture · 10%
Databases, data mining, and information retrieval
1 paper
Query processing and optimization · 100%

Topics — the 15 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Performance modeling and evaluation › simulation › architectural simulation
cycle-accurate simulation
0.512021
UMOC: Unified Modular Ordering Constraints to Unify Cycle- and Register-Transfer-Level Modeling · DAC 2021
Electronic design automation › design representation
hardware modeling
0.512021
UMOC: Unified Modular Ordering Constraints to Unify Cycle- and Register-Transfer-Level Modeling · DAC 2021
Electronic design automation › hardware description language
register-transfer-level modeling
0.512021
UMOC: Unified Modular Ordering Constraints to Unify Cycle- and Register-Transfer-Level Modeling · DAC 2021
Electronic design automation
hardware simulation
0.312018
Mamba: closing the performance gap in productive hardware development frameworks · DAC 2018
Processor architecture and microarchitecture
chip multiprocessor
0.312017
Using intra-core loop-task accelerators to improve the productivity and performance of task-based parallel programs · MICRO 2017
Parallel and multicore computing › parallel programming models
task-based programming
0.312017
Using intra-core loop-task accelerators to improve the productivity and performance of task-based parallel programs · MICRO 2017
Query processing and optimization
sorting
0.212016
A Study of Sorting Algorithms on Approximate Memory · SIGMOD Conference 2016
Performance modeling and evaluation
analytical modeling
0.212016
A performance analysis framework for optimizing OpenCL applications on FPGAs · HPCA 2016
Storage systems › energy-efficient storage
approximate storage
0.212016
A Study of Sorting Algorithms on Approximate Memory · SIGMOD Conference 2016
Memory systems
non-volatile memory
0.212016
A Study of Sorting Algorithms on Approximate Memory · SIGMOD Conference 2016
Reconfigurable computing and FPGAs › FPGA high-level synthesis
OpenCL-to-FPGA
0.212016
A performance analysis framework for optimizing OpenCL applications on FPGAs · HPCA 2016
Runtime systems and virtual machines › dynamic compilation
just-in-time compilation
0.112018
Mamba: closing the performance gap in productive hardware development frameworks · DAC 2018
Emerging computing paradigms
approximate computing
0.112016
A Study of Sorting Algorithms on Approximate Memory · SIGMOD Conference 2016
High-performance computing
code optimization
0.112016
A performance analysis framework for optimizing OpenCL applications on FPGAs · HPCA 2016
High-performance computing › lossy compression
precision-resolution trade-off
0.112016
A Study of Sorting Algorithms on Approximate Memory · SIGMOD Conference 2016

Methods — techniques the papers use, named apart from their topics

python-based framework · 0.7just-in-time compilation · 0.7unified modular ordering constraints · 0.5implicit local ordering constraints · 0.5hybrid storage · 0.5approx-refine execution · 0.5microarchitectural template · 0.3instruction set hint · 0.3static analysis · 0.2dynamic analysis · 0.2
YearPublicationVenuePosition
2023 Symbolic Elaboration: Checking Generator Properties in Dynamic Hardware Description Languages
Peitian Pan, Shunning Jiang, Yanghui Ou, Christopher Batten
MEMOCODE2
2021 UMOC: Unified Modular Ordering Constraints to Unify Cycle- and Register-Transfer-Level Modeling
abstract
We propose unified modular ordering constraints (UMOC), a novel approach that seamlessly unifies method-based cycle-level (CL) modeling and signal-based register-transfer-level (RTL) modeling. Motivated by the challenges in state-of-the-art CL modeling methodologies and existing CL/RTL composition attempts, UMOC successfully breaks the trade-off between model fidelity and scheduling modularity for CL modeling and provides seamless composition of CL and RTL models. Instead of requiring the designer to specify the global intra-cycle ordering of hardware processes, UMOC eliminates this burden using implicit local ordering constraints of RTL signals and explicit local ordering constraints of CL methods. We implement and evaluate UMOC in PyMTL3, a state-of-the-art open-source Python-based hardware modeling framework.
Shunning Jiang, Yanghui Ou, Peitian Pan, Christopher Batten
DAC1
2019 PyOCN: A Unified Framework for Modeling, Testing, and Evaluating On-Chip Networks
abstract
There is a growing interest in the open-source hardware movement to amortize non-recurring engineering costs by using plug-and-play system-on-chip (SoC) designs, where the communication among different components is provided by an on-chip interconnection network. Unfortunately, building an on-chip network (OCN) that is suitable for a specific SoC design requires the exploration of a large number of design options and involves diverse research methodologies to evaluate performance, area, energy, and timing. In this paper, we propose PyOCN, a unified framework that vertically integrates multiple research methodologies to enable productively exploring the OCN design space. PyOCN is the first comprehensive framework for modeling (e.g., functional-level, cycle-level, and register-transfer-level), testing (e.g., unit testing, integration testing, and property-based random testing), and evaluating (e.g., simulating, generating, and characterizing) on-chip interconnection networks. We use a case study based on a 64-terminal butterfly network to illustrate the key features of PyOCN and to demonstrate the framework's potential in productively modeling, testing, and evaluating OCNs.
Cheng Tan 0002, Yanghui Ou, Shunning Jiang, Peitian Pan, Christopher Torng, Shady O. Agwa, Christopher Batten
ICCD3
2018 Mamba: closing the performance gap in productive hardware development frameworks
abstract
Modern high-level languages bring compelling productivity benefits to hardware design and verification. For example, hardware generation and simulation frameworks (HGSFs) use a single "host" language for parameterization, static elaboration, test bench generation, behavioral modeling, and simulation. Unfortunately, HGSFs often suffer from slow simulator performance which undermines their potential productivity benefits. In this paper, we introduce Mamba, a new Python-based HGSF that co-optimizes both the framework and a general-purpose just-in-time compiler. We conduct a quantitative comparison of Mamba vs. traditional and emerging hardware development frameworks across both simple and complex designs. Our results suggest Mamba is able to match the performance of commercial Verilog simulators and is 10× faster than existing HGSFs while still maintaining the productivity of using a high-level language in hardware design.
Shunning Jiang, Berkin Ilbeyi, Christopher Batten
DAC1
2017 Using intra-core loop-task accelerators to improve the productivity and performance of task-based parallel programs
abstract
Task-based parallel programming frameworks offer compelling productivity and performance benefits for modern chip multi-processors (CMPs). At the same time, CMPs also provide packed-SIMD units to exploit fine-grain data parallelism. Two fundamental challenges make using packed-SIMD units with task-parallel programs particularly difficult: (1) the intra-core parallel abstraction gap; and (2) inefficient execution of irregular tasks. To address these challenges, we propose augmenting CMPs with intra-core loop-task accelerators (LTAs). We introduce a lightweight hint in the instruction set to elegantly encode loop-task execution and an LTA microarchitectural template that can be configured at design time for different amounts of spatial/temporal decoupling to efficiently execute both regular and irregular loop tasks. Compared to an in-order CMP baseline, CMP+LTA results in an average speedup of 4.2X (1.8X area normalized) and similar energy efficiency. Compared to an out-of-order CMP baseline, CMP+LTA results in an average speedup of 2.3X (1.5X area normalized) and also improves energy efficiency by 3.2X. Our work suggests augmenting CMPs with lightweight LTAs can improve performance and efficiency on both regular and irregular loop-task parallel programs with minimal software changes.
Ji Kim, Shunning Jiang, Christopher Torng, Moyang Wang, Shreesha Srinath, Berkin Ilbeyi, Khalid Al-Hawaj, Christopher Batten
MICRO2
2017 Bank Stealing for a Compact and Efficient Register File Architecture in GPGPU
abstract
Modern general-purpose graphic processing units (GPGPUs) have emerged as pervasive alternatives for parallel high-performance computing. The extreme multithreading in modern GPGPUs demands a large register file (RF), which is typically organized into multiple banks to support the massive parallelism. Although a heavily banked structure benefits RF throughput, its associated area and energy costs with diminishing performance gains greatly limit the future RF scaling. In this paper, we propose an improved RF design with bank stealing techniques, which enable a high RF throughput with compact area. By deeply investigating the GPGPU microarchitecture, we find that the state-of-the-art RF designs' is far from optimal due to the deficiency in bank utilization, which is the intrinsic limitation to a high RF throughput and a compact RF area. We investigate the causes for bank conflicts and identify that most conflicts can be eliminated by leveraging the fact that the highly banked RF oftentimes experiences underutilization. This is especially true in GPGPUs, where multiple ready warps are available at the scheduling stage with their operands to be wisely coordinated. In this paper, we propose two lightweight bank stealing techniques that can opportunistically fill the idle banks and register entries for better operand service. Using the proposed architecture, the average GPGPU performance can be improved under a smaller energy budget with significant area saving, which makes it promising for sustainable RF scaling.
Naifeng Jing, Shunning Jiang, Shuang Chen 0002, Jingjie Zhang, Li Jiang 0002, Chao Li 0009, Xiaoyao Liang
IEEE Trans. Very Large Scale Integr. Syst.2
2016 A performance analysis framework for optimizing OpenCL applications on FPGAs
abstract
Recently, FPGA vendors such as Altera and Xilinx have released OpenCL SDK for programming FPGAs. However, the architecture of FPGA is significantly different from that of CPU/GPU, for which OpenCL is originally designed. Tuning the OpenCL code for good performance on FPGAs is still an open problem, since the existing OpenCL tools and models designed for CPUs/GPUs are not directly applicable to FPGAs. In the paper, we present an FPGA-based performance analysis framework that can shed light on the performance bottleneck and thus guide the code tuning for OpenCL applications on FPGAs. Particularly, we leverage static and dynamic analysis to develop an analytical performance model, which has captured the key architectural features of FPGA abstractions under OpenCL. Then, we provide four programmer-interpretable metrics to quantify the performance potentials of the OpenCL program with input optimization combination for the next optimization step. We evaluate our framework with a number of user cases, and demonstrate that 1) our analytical performance model can accurately predict the performance of OpenCL programs with different optimization combinations on FPGAs, and 2) our tool can be used to effectively guide the code tuning on alleviating the performance bottleneck.
Zeke Wang, Bingsheng He, Wei Zhang 0012, Shunning Jiang
HPCA4
2016 A Study of Sorting Algorithms on Approximate Memory
abstract
Hardware evolution has been one of the driving factors for the redesign of database systems. Recently, approximate storage emerges in the area of computer architecture. It trades off precision for better performance and/or energy consumption. Previous studies have demonstrated the benefits of approximate storage for applications that are tolerant to imprecision such as image processing. However, it is still an open question whether and how approximate storage can be used for applications that do not expose such intrinsic tolerance. In this paper, we study one of the most basic operations in database--sorting on a hybrid storage system with both precise storage and approximate storage. Particularly, we start with a study of three common sorting algorithms on approximate storage. Experimental results show that a 95% sorted sequence can be obtained with up to 40% reduction in total write latencies. Thus, we propose an approx-refine execution mechanism to improve the performance of sorting algorithms on the hybrid storage system to produce precise results. Our optimization gains the performance benefits by offloading the sorting operation to approximate storage, followed by an efficient refinement to resolve the unsortedness on the output of the approximate storage. Our experiments show that our approx-refine can reduce the total memory access time by up to 11%. These studies shed light on the potential of approximate hardware for improving the performance of applications that require precise results.
Shuang Chen 0002, Shunning Jiang, Bingsheng He, Xueyan Tang
SIGMOD Conference2
2015 Bank stealing for conflict mitigation in GPGPU Register File
abstract
Modern General Purpose Graphic Processing Unit (GPGPU) demands a large Register File (RF), which is typically organized into multiple banks to support the massive parallelism. Although heavy banking benefits RF throughput, its associated area and energy costs with diminishing performance gains greatly limit future RF s-caling. In this paper, we propose an improved RF design with a bank stealing technique, which enables a high RF throughput with compact area. By deeply investigating the GPGPU microarchitecture, we identify the deficiency in the state-of-the-art RF designs as the bank conflict problem, while the majority of conflicts can be eliminated leveraging the fact that the highly-banked RF oftentimes experiences under-utilization. This is especially true in GPGPU where multiple ready warps are available at the scheduling stage with their operands to be wisely coordinated. Our lightweight bank stealing technique can opportunistically fill the idle banks for better operand service, and the average GPGPU performance can be improved under smaller energy budget with significant area saving, which makes it promising for sustainable RF scaling.
Naifeng Jing, Shuang Chen 0002, Shunning Jiang, Li Jiang 0002, Chao Li 0009, Xiaoyao Liang
ISLPED3