VLDB 2026 Research / reviewers in the wild / expert
Peitian Pan
dblp:209/9032
· DBLP profile ↗
10ranked-venue papers
4as first author
6since 2021 · last 2024
0000-0001-6147-9092ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Theory of computation · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Scalable, Programmable and Dense: The HammerBlade Open-Source RISC-V ManycoreabstractExisting tiled manycore architectures propose to convert abundant silicon resources into general-purpose parallel processors with unmatched computational density and programmability. However, as we approach 100 K cores in one chip, conventional manycore architectures struggle to navigate three key axes: scalability, programmability, and density. Many manycores sacrifice programmability for density; or scalability for programmability. In this paper, we explore HammerBlade, which simultaneously achieves scalability, programmability and density. HammerBlade is a fully open-source RISC-V manycore architecture, which has been silicon-validated with a 2048-core ASIC implementation using a 14/16nm process. We evaluate the system using a suite of parallel benchmarks that captures a broad spectrum of computation and communication patterns. Dai Cheol Jung, Max Ruttenberg, Paul Gao 0001, Scott Davidson 0004, Daniel Ruelas-Petrisko, Kangli Li, Aditya K. Kamath, Shaolin Xie, Peitian Pan, Zhongyuan Zhao 0004, Zichao Yue, Bandhav Veluri, Sripathi Muralitharan, Adrian Sampson, Andrew Lumsdaine, Zhiru Zhang, Christopher Batten, Mark Oskin, Dustin Richmond, Michael B. Taylor |
ISCA | 10 |
| 2023 | Formal Verification of the Stall Invariant Property for Latency-Insensitive RTL Modules
Peitian Pan, Christopher Batten |
MEMOCODE | 1 |
| 2023 | Symbolic Elaboration: Checking Generator Properties in Dynamic Hardware Description Languages
Peitian Pan, Shunning Jiang, Yanghui Ou, Christopher Batten |
MEMOCODE | 1 |
| 2022 | A Tensor Processing Framework for CPU-Manycore Heterogeneous SystemsabstractFuture CPU-manycore heterogeneous systems can provide high peak throughput by integrating thousands of simple, independent, energy-efficient cores in a single die. However, there are two key challenges to translating this high peak throughput into improved end-to-end workload performance: 1) manycore co-processors rely on simple hardware putting significant demands on the software programmer and 2) manycore co-processors use in-order cores that struggle to tolerate long memory latencies. To address the manycore programmability challenge, this article presents a dense and sparse tensor processing framework based on PyTorch that enables domain experts to easily accelerate off-the-shelf workloads on CPU-manycore heterogeneous systems. To address the manycore memory latency challenge, we use our extended PyTorch framework to explore the potential for decoupled access/execute (DAE) software and hardware mechanisms. More specifically, we propose two software-only techniques, naïve-software DAE and systolic-software DAE, along with a lightweight hardware access accelerator to further improve area-normalized throughput. We evaluate our techniques using a combination of PyTorch operator microbenchmarking and real-world PyTorch workloads running on a detailed register-transfer-level model of a 128-core manycore architecture. Our evaluation on three real-world dense and sparse tensor workloads suggests these workloads can achieve approximately 2–$6\times $performance improvement when scaled to a future 2000-core CPU-manycore heterogeneous system compared to an 18-core out-of-order CPU baseline, while potentially achieving higher area-normalized throughput and improved energy efficiency compared to general-purpose graphics processing units. Peitian Pan, Zhongyuan Zhao 0004, Krithik Ranjan, Jack Weber, Bandhav Veluri, Seyed Borna Ehsani, Max Ruttenberg, Dai Cheol Jung, Preslav Ivanov, Dustin Richmond, Michael B. Taylor, Zhiru Zhang, Christopher Batten |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | UMOC: Unified Modular Ordering Constraints to Unify Cycle- and Register-Transfer-Level ModelingabstractWe propose unified modular ordering constraints (UMOC), a novel approach that seamlessly unifies method-based cycle-level (CL) modeling and signal-based register-transfer-level (RTL) modeling. Motivated by the challenges in state-of-the-art CL modeling methodologies and existing CL/RTL composition attempts, UMOC successfully breaks the trade-off between model fidelity and scheduling modularity for CL modeling and provides seamless composition of CL and RTL models. Instead of requiring the designer to specify the global intra-cycle ordering of hardware processes, UMOC eliminates this burden using implicit local ordering constraints of RTL signals and explicit local ordering constraints of CL methods. We implement and evaluate UMOC in PyMTL3, a state-of-the-art open-source Python-based hardware modeling framework. Shunning Jiang, Yanghui Ou, Peitian Pan, Christopher Batten |
DAC | 3 |
| 2021 | Ultra-Elastic CGRAs for Irregular Loop SpecializationabstractReconfigurable accelerator fabrics, including coarse-grain reconfigurable arrays (CGRAs), have experienced a resurgence in interest because they allow fast-paced software algorithm development to continue evolving post-fabrication. CGRAs traditionally target regular workloads with data-level parallelism (e.g., neural networks, image processing), but once integrated into an SoC they remain idle and unused for irregular workloads. An emerging trend towards repurposing these idle resources raises important questions for how to efficiently map and execute general-purpose loops which may have irregular memory accesses, irregular control flow, and inter-iteration loop dependencies. Recent work has increasingly leveraged elasticity in CGRAs to mitigate the first two challenges, but elasticity alone does not address inter-iteration loop dependencies which can easily bottleneck overall performance. In this paper, we address all three challenges for irregular loop specialization and propose ultra-elastic CGRAs (UE-CGRAs), a novel elastic CGRA that accelerates true-dependency bottlenecks and saves energy in irregular loops by overcoming traditional VLSI challenges. UE-CGRAs allow configurable fine-grain dynamic voltage and frequency scaling (DVFS) for each of potentially hundreds of tiny processing elements (PEs) in the CGRA, enabling chains of connected PEs to “rest” at lower voltages and frequencies to save energy, while other chains of connected PEs can “sprint” at higher voltages and frequencies to accelerate through true-dependency bottlenecks. UE-CGRAs rely on a novel ratiochronous clocking scheme carefully overlaid on the inter-PE elastic interconnect to enable low-latency crossings while remaining fully verifiable with commercial static timing analysis tools. We present the UE-CGRA analytical model, compiler, architectural template, and VLSI circuitry, and we demonstrate how UE-CGRAs can specialize for irregular loops and improve performance ($ 1.42-1.50\times$) or energy efficiency $(1.24-2.32\times)$ with reasonable area overhead compared to traditional inelastic and elastic CGRAs, while also improving performance ($ 1.35-3.38\times$) or energy efficiency (up to $1.53\times$) compared to a RISC-V core. Christopher Torng, Peitian Pan, Yanghui Ou, Cheng Tan 0002, Christopher Batten |
HPCA | 2 |
| 2019 | PyOCN: A Unified Framework for Modeling, Testing, and Evaluating On-Chip NetworksabstractThere is a growing interest in the open-source hardware movement to amortize non-recurring engineering costs by using plug-and-play system-on-chip (SoC) designs, where the communication among different components is provided by an on-chip interconnection network. Unfortunately, building an on-chip network (OCN) that is suitable for a specific SoC design requires the exploration of a large number of design options and involves diverse research methodologies to evaluate performance, area, energy, and timing. In this paper, we propose PyOCN, a unified framework that vertically integrates multiple research methodologies to enable productively exploring the OCN design space. PyOCN is the first comprehensive framework for modeling (e.g., functional-level, cycle-level, and register-transfer-level), testing (e.g., unit testing, integration testing, and property-based random testing), and evaluating (e.g., simulating, generating, and characterizing) on-chip interconnection networks. We use a case study based on a 64-terminal butterfly network to illustrate the key features of PyOCN and to demonstrate the framework's potential in productively modeling, testing, and evaluating OCNs. Cheng Tan 0002, Yanghui Ou, Shunning Jiang, Peitian Pan, Christopher Torng, Shady O. Agwa, Christopher Batten |
ICCD | 4 |
| 2019 | CongraPlus: Towards Efficient Processing of Concurrent Graph Queries on NUMA MachinesabstractGraph analytics has been routinely used to solve problems in a wide range of real-life applications. Efficiently processing concurrent graph analytics queries in a multiuser environment is highly desirable as we enter a world of edge device oriented services. Existing research, however, primarily focuses on analyzing a single, large graph dataset and leaves the efficient processing of multiple mid-sized graph analytics queries an intriguing yet challenging open problem. In this work, we investigate the scheduling of concurrent graph analytics queries on NUMA machines. We analyze the performance of several graph analytics algorithms and observe that they have diminishing performance returns as the number of processor cores increases. With concurrent graph analytics, such diminishing returns translate to no or even negative performance gains because of increasing contention on shared hardware resources. We also demonstrate the unpredictability of memory bandwidth usage for numerous graph analytics algorithms, which can lead to sub-optimal performance due to its potential to cause severe memory bandwidth contention. Motivated by the above observations, we propose CongraPlus, a NUMA-aware scheduler that intelligently manages concurrent graph analytics queries for better system throughput and memory bandwidth efficiency. CongraPlus collects the memory bandwidth consumption characteristics of graph analytics queries via offline profiling and eliminates memory bandwidth contention by computing the optimal sequence to launch queries. It also avoids computation resource contention by assigning a certain number of processor cores to the individual queries. We implement CongraPlus in C++ on top of the Ligra graph processing framework and test it with judiciously selected graph processing query combinations. Our results reveal that CongraPlus-based schemes improve query throughput by 30 percent compared to the conventional approach. It also exhibits a much better quality of service and scalability. Peitian Pan, Chao Li 0009, Minyi Guo |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2018 | A Network-Centric Hardware/Algorithm Co-Design to Accelerate Distributed Training of Deep Neural NetworksabstractTraining real-world Deep Neural Networks (DNNs) can take an eon (i.e., weeks or months) without leveraging distributed systems. Even distributed training takes inordinate time, of which a large fraction is spent in communicating weights and gradients over the network. State-of-the-art distributed training algorithms use a hierarchy of worker-aggregator nodes. The aggregators repeatedly receive gradient updates from their allocated group of the workers, and send back the updated weights. This paper sets out to reduce this significant communication cost by embedding data compression accelerators in the Network Interface Cards (NICs). To maximize the benefits of in-network acceleration, the proposed solution, named INCEPTIONN (In-Network Computing to Exchange and Process Training Information Of Neural Networks), uniquely combines hardware and algorithmic innovations by exploiting the following three observations. (1) Gradients are significantly more tolerant to precision loss than weights and as such lend themselves better to aggressive compression without the need for the complex mechanisms to avert any loss. (2) The existing training algorithms only communicate gradients in one leg of the communication, which reduces the opportunities for in-network acceleration of compression. (3) The aggregators can become a bottleneck with compression as they need to compress/decompress multiple streams from their allocated worker group. To this end, we first propose a lightweight and hardware-friendly lossy-compression algorithm for floating-point gradients, which exploits their unique value characteristics. This compression not only enables significantly reducing the gradient communication with practically no loss of accuracy, but also comes with low complexity for direct implementation as a hardware block in the NIC. To maximize the opportunities for compression and avoid the bottleneck at aggregators, we also propose an aggregator-free training algorithm that exchanges gradients in both legs of communication in the group, while the workers collectively perform the aggregation in a distributed manner. Without changing the mathematics of training, this algorithm leverages the associative property of the aggregation operator and enables our in-network accelerators to (1) apply compression for all communications, and (2) prevent the aggregator nodes from becoming bottlenecks. Our experiments demonstrate that INCEPTIONN reduces the communication time by 70.9~80.7% and offers 2.2~3.1x speedup over the conventional training system, while achieving the same level of accuracy. Youjie Li, Jongse Park, Mohammad Alian, Zheng Qu 0002, Peitian Pan, Ren Wang 0001, Alexander G. Schwing, Hadi Esmaeilzadeh, Nam Sung Kim |
MICRO | 6 |
| 2017 | Congra: Towards Efficient Processing of Concurrent Graph Queries on Shared-Memory MachinesabstractParallel graph processing has been routinely used to solve problems in a wide range of real life applications today. Efficiently handling many concurrent graph processing queries in a multi-user environment is highly desirable as we enter a world full of edge device oriented services. Existing research, however, primarily focuses on processing a single, large graph dataset and leaves the efficient handling of multiple mid-sized graph queries an interesting yet challenging open problem. In this work, we investigate the management of multiple graph processing queries on shared-memory machines. We carefully analyze the performance of several graph algorithms and find some of them do not scale well with more CPU cores (more threads) while others have diminishing returns as the number of CPU cores grow. Further, we identify the trade-off relationship between the increasing cost of atomic operations and the growing available memory bandwidth with more CPU cores. Motivated by the above observations, we propose Congra, a dynamic graph scheduler that intelligently manages multiple concurrent graph queries for better system throughput and resource efficiency. Congra collects the memory bandwidth consumption and atomic operations characteristics of graph queries via offline profiling and decides which subset of CPU cores for the query to run, as long as the query is not blocked due to heavy resources contention. We implement Congra in C++ on top of the Ligra graph processing framework and test it with judiciously selected graph processing query combinations. Our results show the Congra improves query throughput by 60% compared to existing designs. It exhibits much better quality of service and readily supports a scale-out computing environment. Peitian Pan, Chao Li 0009 |
ICCD | 1 |