VLDB 2026 Research / reviewers in the wild / expert
Jeff Zhang 0001
dblp:29/4363-1 · also Jeff Jun Zhang
· DBLP profile ↗
46ranked-venue papers
7as first author
37since 2021 · last 2026
0000-0001-7411-8923ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 43 · 7 first-author · 35 since 2021Software engineering, systems software and programming languages · 7 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Chiplet-NAS: Chiplet-aware Neural Architecture Search for Efficient AI Inference on 2.5D IntegrationabstractThe co-design of neural network architectures and their target chiplet-based hardware systems presents a significant challenge due to the vast and combinatorial design space. Identifying solutions that are Pareto-optimal across competing objectives of task accuracy, system latency, and power consumption requires solutions beyond manual design and brute-force methods. This paper proposes a closed-loop chiplet-aware neural architecture search (Chiplet-NAS) framework to automate the exploration and discover hardware-optimized models for efficient AI inference on 2.5 D chiplet-based systems. The framework integrates a Tree-structured Parzen Estimator (TPE) for sampleefficient search with CLAIRE, a chiplet-based library and fast performance benchmarking tool, to provide direct hardware feedback on latency and energy consumption, along with accuracy optimization. The framework is evaluated by co-designing ResNet-based model architectures with chiplet based hardware systems. Compared to a baseline NAS that optimizes only on the task accuracy, our Chiplet-NAS achieves significant power and performance benefits at the iso-accuracy. Pragnya Sudershan Nalla, Nikhil K. Cherukuri, Sachin S. Sapatnekar, Chaitali Chakrabarti, Yu Cao 0001, Jeff Zhang 0001 |
ASP-DAC | 8 |
| 2026 | Neura: A Unified Framework for Hierarchical and Adaptive CGRAsabstractCoarse-Grained Reconfigurable Arrays (CGRAs) are a promising solution for energy-efficient acceleration across multiple application domains. Yet, CGRAs face significant scalability challenges that hinder their widespread adoption, stemming from three main concerns: (1) Mapping Scalability — existing mapping algorithms struggle to find feasible and optimal solutions as the design complexity grows; (2) Architectural Limitations — rigid mapping granularity and memory access restrict flexibility and performance; and (3) Dynamic Multi-Kernel Support — dynamic and simultaneous execution of multiple kernels are not thoroughly explored, limiting the applicability of CGRAs in complex multi-kernel scenarios. Cheng Tan 0002, Miaomiao Jiang, Ruihong Yin, Yanghui Ou, Lei Ju 0001, Jeff Zhang 0001 |
ASPLOS (2) | 8 |
| 2026 | Exploring Heterogeneity-Aware Optimizations for Resource Efficient Edge RecommendationabstractRecommendation systems are widely deployed on edge devices to enable personalized user experience. While recommendation inference has traditionally been performed on centralized servers, recent advances in mobile SoCs have motivated a shift toward on-device execution. However, achieving efficient recommendation inference on edge devices remains challenging due to edge-specific execution characteristics and heterogeneity. In this paper, we characterize resource inefficiencies under realistic edge constraints and propose optimization strategies. Yerin Lee, Gyudong Kim, Eunjin Lee, Jeff Zhang 0001, Young-Ho Gong, Carole-Jean Wu |
DATE | 4 |
| 2026 | vFPGA: Towards Sub-µs Reconfiguration via 3D FPGA and Packaging Co-Design
Nikhil K. Cherukuri, Sharad Nag, Pragnya Sudershan Nalla, Ashish K. Kola, Chetan S. Gadireddi, Kevin Dai, Jae-sun Seo, Zhenman Fang, Jeff Zhang 0001, Yu Cao 0001 |
FPGA | 9 |
| 2026 | CAPO: Certification-Guided Agentic Workflow for Physical Design Parameter OptimizationabstractVLSI Physical design is a long, stage-coupled flow with multiple tunable parameters. Efficient optimization is challenging because each evaluation is expensive, and decisions made in one stage can strongly affect final metric such as power, performance, and area (PPA). Traditional design space exploration (DSE) methods, such as Bayesian optimization, rely on repeated full-flow evaluations guided by mathematical surrogate models. Although effective in some settings, these methods can be costly and may waste iterations on failed or low-quality configurations. Zesong Jiang, Qihang Wu, Bing-Yue Wu, Jeff Zhang 0001 |
ACM Great Lakes Symposium on VLSI | 4 |
| 2026 | A 22nm Reconfigurable Systolic Array for FFT and AI Inference
John Stolzberg-Schray, Sharad Nag, Jacob Johnson, Nikhil K. Cherukuri, Ashish K. Kola, Gopikrishnan Raveendran Nair, Jeff Zhang 0001, Jae-sun Seo, Yu Cao 0001 |
ISCAS | 7 |
| 2026 | 2.5D/3D Chiplet-based Integration: New Dimensions in Design and Testing
Ganap A. Tewary, Partho Bhoumik, Pragnya S. Nalla, Yu Cao 0001, Krishnendu Chakrabarty, Jeff Zhang 0001 |
VTS | 6 |
| 2025 | PICACHU: Plug-In CGRA Handling Upcoming Nonlinear Operations in LLMsabstractLarge language models (LLMs) have revolutionized natural language processing (NLP) domain by achieving state-of-the-art performance across a range of benchmarks. However, nonlinear operations in LLMs significantly contribute to inference latency and present unique challenges that have not been encountered previously. Addressing these challenges requires accelerators that combine efficiency, flexibility, and support for user-defined precision. Our analysis reveals that Coarse-Grained Reconfigurable Arrays (CGRAs) provide an effective solution, offering a balance of performance and flexibility tailored to domain-specific workloads. Jiajun Qin, Tianhua Xia, Cheng Tan 0002, Jeff Zhang 0001, Sai Qian Zhang |
ASPLOS (2) | 4 |
| 2025 | Time-Period-Aware Embedding Regeneration for Session-Based RecommendationabstractSession-based recommender systems typically focus on intra-session user behavior but often overlook the macro-level temporal evolution of items themselves. To address this gap, we introduce a model that explicitly captures item dynamics by regenerating time-period-aware embeddings. Our approach partitions the data timeline into several distinct periods and employs a simple yet effective module, TEG, which uses a GRU and a causal attention layer to recurrently learn how item representations evolve from one period to the next. This allows our model to capture global, long-term trends while its gating mechanism naturally accommodates both dynamic and static items. Extensive experiments on three real-world datasets show that our method achieves highly competitive performance against complex state-of-the-art models. More importantly, it demonstrates the significant and complementary value of modeling global item evolution, providing a new dimension for improving session-based recommendation. Jeff Zhang 0001 |
CIKM | 3 |
| 2025 | Invited: EDA for Heterogeneous IntegrationabstractThe advent of heterogeneous integration (HI) places new demands on EDA tooling. Building large systems requires (1) methods for chiplet disaggregation that map the system to smaller chiplets, working in conjunction with system-technology co-optimization to determine the right design decisions that optimize computation and communication, together with the choice of substrate and chiplet technologies; (2) multiphysics and multiscale analyses that incorporate thermomechanical aspects into performance analysis, ranging from fast machine-learningdriven analyses in early stages to signoff-quality multiphysics-based analysis; (3) physical design techniques for placing and routing chiplets and embedded active/passive elements on and within the substrate, including the design of thermal and power delivery solutions; and (4) underlying infrastructure required to facilitate HI-based design, including the design and characterization of chiplet libraries and the establishment of data formats and standards. This paper overviews these issues and lays out a set of EDA needs for HI designs. Emad Haque, Pragnya Sudershan Nalla, Chetal Choppali Sudarshan, Divya Yogi, Chaitali Chakrabarti, Vidya A. Chhabria, Ramesh Harjani, Jeff Zhang 0001, Sachin S. Sapatnekar |
DAC | 9 |
| 2025 | CHORD: Composable Hybrid Optical Reconfigurable Diffractive Framework For Optical Neural NetworkabstractDiffractive optical neural networks (DONNs), leveraging freespace light wave propagation for ultra-parallel, high-efficiency computing, have emerged as promising artificial intelligence (AI) accelerators. However, their inherent lack of reconfigurability due to fixed optical structures postfabrication hinders practical deployment in the face of dynamic AI workloads and evolving applications. To overcome this challenge, we introduce, for the first time, a composable hybrid optical reconfigurable diffractive framework (CHORD), a physically composable architecture that unlocks a new degree of freedom and unprecedented versatility in DONNs. By leveraging full-system learnability, CHORD repurposes fixed fabricated optical hardware, achieving exponentially expanded functionality and superior task adaptability through the differentiable learning of system variables. Furthermore, CHORD adopts a hybrid optical/photonic design, combining the reconfigurability of integrated photonics with the ultra-parallelism of free-space diffractive systems. Extensive evaluations demonstrate that CHORD has digital-comparable accuracy on various task adaptations with $74 \times$ faster speed and $194 \times$ lower energy. Compared to prior DONNs, CHORD shows exponentially larger functional space with $5 \times$ faster training speed, paving the way for a new paradigm of versatile, composable, hybrid optical/photonic AI computing. Our code is open-sourced at link1.1github.com/ScopeX-ASU/CHORD Ziang Yin, Jeff Zhang 0001, Jiaqi Gu 0002 |
DAC | 3 |
| 2025 | SimPhony: A Device-Circuit-Architecture Cross-Layer Modeling and Simulation Framework for Heterogeneous Electronic-Photonic AI SystemabstractElectronic-photonic integrated circuits (EPICs) offer transformative potential for next-generation high-performance AI, but they require interdisciplinary advances across devices, circuits, architecture, and design automation. The complexity of these hybrid systems makes it challenging even for domain experts to understand distinct behaviors and interactions across the design stack. The lack of a flexible, accurate, fast, and easy-to-use EPIC AI system simulation framework significantly limits the exploration of hardware innovations and system evaluations on common benchmarks. To address this gap, we propose SimPhony, a cross-layer modeling and simulation framework for heterogeneous electronic-photonic AI systems. SimPhony offers a platform that enables (1) generic, extensible hardware topology representation that supports heterogeneous multi-core architectures with diverse photonic tensor core designs; (2) optics-specific dataflow modeling with unique multi-dimensional parallelism and reuse beyond spatial/temporal dimensions; (3) data-aware energy modeling with realistic device responses, layout-aware area estimation, link budget analysis, and bandwidth-adaptive memory modeling; and (4) seamless integration with model training framework for hardware/software co-simulation. By providing a unified, versatile, and high-fidelity simulation platform, SimPhony enables researchers to innovate and evaluate EPIC AI hardware across multiple domains, facilitating the next leap in emerging AI hardware. Our code is open-sourced at link1.1https://github.com/ScopeX-ASU/SimPhony Ziang Yin, Meng Zhang 0023, Nicholas Gangi, Z. Rena Huang, Jeff Zhang 0001, Jiaqi Gu 0002 |
DAC | 5 |
| 2025 | CLAIRE: Composable Chiplet Libraries for AI InferenceabstractArtificial intelligence has made a significant impact on fields like computer vision, Natural Language Processing (NLP), healthcare, and robotics. However, recent AI models, such as GPT-4 and LLaMAv3, demand significant number of computational resources, pushing monolithic chips to their technological and practical limits. 2.5D chiplet-based heterogeneous architectures have been proposed to address these technological and practical limits. While chiplet optimization for models like Convolutional Neural Networks (CNNs) is well-established, scaling this approach to accommodate diverse AI inference models with different computing primitives, data volumes, and different chiplet sizes is very challenging. A set of hardened IPs and chiplet libraries optimized for a broad range of AI applications is proposed in this work. We derive the set of chiplet configurations that are composable, scalable and reusable by employing an analytical framework trained on a diverse set of AI algorithms. Testing these set of library synthesized configurations on a different set of algorithms, we achieve a$1.99\times-3.99\times$improvement in non-recurring engineering (NRE) chiplet design costs, with minimal performance overhead compared to custom chiplet-based ASIC designs. Similar to soft IPs for SoC development, the library of chiplets improves flexibility, reusability, and efficiency for AI hardware designs. Pragnya Sudershan Nalla, Emad Haque, Yaotian Liu, Sachin S. Sapatnekar, Jeff Zhang 0001, Chaitali Chakrabarti, Yu Cao 0001 |
DATE | 5 |
| 2025 | FP-SMR: A Fully Digital Floating-Point Processing-in-SAS-MRAM for Session-based Recommender System
Asmer Hamid Ali, Amitesh Sridharan, William Hwang, Wilman Tsai, Jeff Zhang 0001, Yiran Chen 0001, Shan X. Wang, Deliang Fan |
ACM Great Lakes Symposium on VLSI | 6 |
| 2025 | ML4SODA: A Decision Tree Guided Design Space Exploration for Fast and High Quality MLIR-based HLS
Darshith Manjunath, Nicolas Bohm Agostini, Antonino Tumeo, Jeff Zhang 0001, Chaitali Chakrabarti |
ACM Great Lakes Symposium on VLSI | 4 |
| 2025 | HSMU-SpGEMM: Achieving High Shared Memory Utilization for Parallel Sparse General Matrix-Matrix Multiplication on Modern GPUsabstractSparse general matrix-matrix multiplication (SpGEMM) is a core primitive for numerous scientific applications. Traditional hash-based approaches fail to strike a balance between reducing hash collisions and efficiently utilizing fast shared memory, which significantly undermines the performance of executing SpGEMM on GPUs. To address this issue, this paper introduces a novel accumulator design that achieves high shared memory utilization on modern GPUs. For the proposed high shared memory utilization algorithm, i.e., HSMU-SpGEMM1, we further optimize different symbolic stages. Our evaluations with four state-of-the-art hash-based SpGEMM libraries (Nsparse, spECK, OpSparse, and NVIDIA’s cuSPARSE) on three NVIDIA GPUs (Ampere, Ada Lovelace, Turing) demonstrate significant performance benefits from HSMU-SpGEMM.1HSMU-SpGEMM is available at https://github.com/wuminqaq/HSMUSpGEMM Huizhang Luo, Fenfang Li, Zhuo Tang, Kenli Li 0001, Jeff Zhang 0001, Chubo Liu |
HPCA | 7 |
| 2025 | Optimizing both performance and tail latency for B+tree on persistent memory
Xianyu He, Chaoshu Yang, Runyu Zhang 0002, Huizhang Luo, Zhichao Cao 0002, Jeff Zhang 0001 |
J. Syst. Archit. | 6 |
| 2025 | HISIM: Analytical Performance Modeling and Design Space Exploration of 2.5D/3D Integration for AI ComputingabstractMonolithic designs face significant fabrication cost and data movement challenges, especially when executing complex and diverse AI models. Advanced 2.5D/3D packaging promises high bandwidth and connection density to overcome these challenges, yet it also introduces new electro-thermal constraints. This article develops a suite of analytical performance models to enable efficient benchmarking of a 2.5D/3D heterogeneous system for energy-efficient AI computing. These models encompass various performance metrics related to computing units, network-on-chip (NoC), and network-on-package (NoP). The results are summarized into a new tool, HISIM, which is$10^{4} \times $–$10^{6} \times $faster than state-of-the-art AI benchmark tools. Furthermore, HISIM integrates rapid thermal simulation for the 2.5D/3D system, helping shed light on both the potential and limitations of 2.5D/3D heterogeneous integration (HI) on representative AI algorithms. The code of HISIM is available athttps://github.com/mec-UMN/HISIM. Zhenyu Wang 0016, Pragnya Sudershan Nalla, Jingbo Sun 0003, A. Alper Goksoy, Sumit K. Mandal, Jae-sun Seo, Vidya A. Chhabria, Jeff Zhang 0001, Chaitali Chakrabarti, Ümit Y. Ogras, Yu Cao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2024 | Exploiting 2.5D/3D Heterogeneous Integration for AI ComputingabstractThe evolution of AI algorithms has not only revolutionized many application domains, but also posed tremendous challenges on the hardware platform. Advanced packaging technology today, such as 2.5D and 3D interconnection, provides a promising solution to meet the ever-increasing demands of bandwidth, data movement, and system scale in AI computing. This work presents HISIM, a modeling and benchmarking tool for chiplet-based heterogeneous integration. HISIM emphasizes the hierarchical interconnection that connects various chiplets through network-on-package. It further integrates technology roadmap, power/latency prediction, and thermal analysis together to support electro-thermal co-design. Leveraging HISIM with in-memory computing chiplets, we explore the advantages and limitations of 2.5D and 3D heterogenous integration on representative AI algorithms, such as DNNs, transformers, and graph neural networks. Zhenyu Wang 0016, Jingbo Sun 0003, A. Alper Goksoy, Sumit K. Mandal, Yaotian Liu, Jae-sun Seo, Chaitali Chakrabarti, Ümit Y. Ogras, Vidya A. Chhabria, Jeff Zhang 0001, Yu Cao 0001 |
ASPDAC | 10 |
| 2024 | zeroTT: A Two-Step State Transition Avoidance Scheme for MLC STT-RAMabstractCompared with conventional SRAM, Spin-Transfer Torque Random Access Memory(STT-RAM) is expected to play a crucial role in future memory technologies with the increasing demands for higher storage density and lower power consumption for modern embedded systems. Moreover, Multi-Level Cell (MLC) STT-RAM outperforms Single-Level Cell (SLC) STT-RAM since it has higher bit density. However, MLC STT-RAM suffers from write performance due to the two-step state transitions (TTs) in memory cells' soft domain. State-of-the-art approaches mitigate this issue by reducing TTs with efficient data coding. Unfortunately, none of the existing works can fully eliminate the TTs. In this work, zeroTT, an optimal (3, 4)-based expansion coding method that eliminates TTs for MLC STT-RAM. The design of ZeroTT considers space overhead and coding complexity, and our experimental results demonstrate that zeroTT can completely avoid TTs, leading to a more efficient MLC STT-RAM memory in terms of access latency, energy consumption, and device lifetime. Huizhang Luo, Jeff Zhang 0001, Mingxing Duan, Wangdong Yang, Zhuo Tang, Kenli Li 0001 |
DAC | 3 |
| 2024 | SCATTER: Algorithm-Circuit Co-Sparse Photonic Accelerator with Thermal-Tolerant, Power-Efficient In-situ Light Redistribution
Ziang Yin, Nicholas Gangi, Meng Zhang 0023, Jeff Zhang 0001, Z. Rena Huang, Jiaqi Gu 0002 |
ICCAD | 4 |
| 2024 | A 16nm Heterogeneous Accelerator for Energy-Efficient Sparse and Dense AI ComputingabstractArtificial intelligence (AI) has evolved from dense Deep Neural Networks (DNNs) toward a diverse set of models, such as sparse graph convolutional neural networks (GCNs). These new models differ in model size, processing flow, memory access patterns, and data/model sparsity. Hardware platforms optimized for dense DNNs with a regular data structure are inefficient to manage new unstructured, sparse workloads, such as GCNs. For instance, in-memory computing (IMC) units that is suitable for dense matrix/vector computation, but significantly underutilized for sparse data. Gopikrishnan Raveendran Nair, Fengyang Jiang, Jeff Zhang 0001, Yu Cao 0001 |
ISLPED | 3 |
| 2024 | ICED: An Integrated CGRA Framework Enabling DVFS-Aware AccelerationabstractCoarse-grained reconfigurable arrays (CGRAs) are a promising solution to enable energy-efficient acceleration of applications from different domains. By leveraging reconfiguration at the functional level, they can adapt to significantly different computational patterns. However, the relationships of voltage and frequency with the utilization of CGRA resources and the dynamic management of them are not well explored, leading to inefficient designs. CGRAs have also been successful in accelerating data-dependent streaming applications. However, in these applications, the execution time of each kernel in the pipeline might dynamically vary depending on the characteristics of the input. This also leads to under-utilization of resources for the dynamically changing kernels that do not limit the application throughput. DVFS can also improve energy efficiency for these applications by dynamically changing the voltage and frequency levels of tiles that host non-performance-constraining kernels. This paper proposes ICED - an integrated DVFS-aware framework to map applications on CGRAs that support power islands. ICED proposes a CGRA architecture supporting DVFS islands at varying granularity (from a single tile to a group of tiles) and the related DVFS-aware compilation and mapping toolchain. ICED is the first work that introduces DVFS support for spatio-temporal CGRAs at power-island levels. The experimental evaluation shows that ICED improves average utilization by$\mathbf{2}.\mathbf{3}\times$and energy-efficiency by$\mathbf{1}.\mathbf{32}\times$over a conventional CGRA. With streaming applications, ICED can achieve up to$\mathbf{1}.\mathbf{26}\times$energy-efficiency compared with a state-of-the-art CGRA that introduces partial dynamic reconfiguration to adapt to variations in kernels' throughput. Cheng Tan 0002, Miaomiao Jiang, Deepak Patil, Yanghui Ou, Zhaoying Li 0004, Lei Ju 0001, Tulika Mitra, Antonino Tumeo, Jeff Zhang 0001 |
MICRO | 10 |
| 2024 | Intelligent Networking for Energy Harvesting Powered IoT SystemsabstractAs the next-generation battery substitute for IoT system, energy harvesting (EH) technology revolutionizes the IoT industry with environmental friendliness, ubiquitous accessibility, and sustainability, which enables various self-sustaining IoT applications. However, due to the weak and intermittent nature of EH power, the performance of EH-powered IoT systems as well as its collaborative routing mechanism can severely deteriorate, rendering unpleasant data package loss during each power failure. Such a phenomenon makes conventional routing policies and energy allocation strategies impractical. Given the complexity of the problem, reinforcement learning (RL) appears to be one of the most promising and applicable methods to address this challenge. Nevertheless, although the energy allocation and routing policy are jointly optimized by the RL method, due to the energy restriction of EH devices, the inappropriate configuration of multi-hop network topology severely degrades the data collection performance. Therefore, this article first conducts a thorough mathematical discussion and develops the topology design and validation algorithm under energy harvesting scenarios. Then, this article develops DeepIoTRouting , a distributed and scalable deep reinforcement learning (DRL)-based approach, to address the routing and energy allocation jointly for the energy harvesting powered distributed IoT system. The experimental results show that with topology optimization, DeepIoTRouting achieves at least 38.71% improvement on the amount of data delivery to sink in a 20-device IoT network, which significantly outperforms state-of-the-art methods. Tao Liu 0023, Jeff Zhang 0001, Mehdi Sookhak, Mimi Xie |
ACM Trans. Sens. Networks | 4 |
| 2023 | VecPAC: A Vectorizable and Precision-Aware CGRAabstractCoarse-grained reconfigurable arrays (CGRAs) are a promising solution to accelerate applications from several domains, thanks to their balance between the performance achieved through specialization and the adaptability to obtain different computational patterns through dynamic reconfiguration. Several state-of-the-art CGRA designs try to further exploit domain specialization by integrating additional specialized functional units to support custom numeric formats and/or vector functional units. While this approach can improve performance and efficiency for kernels coming from a single application domain, it lowers the overall utilization of the hardware resources and reduces the adaptability of the accelerator. This paper proposes VecPAC - a vectorizable and precision-aware coarse-grained reconfigurable array (CGRA) design. Vec-PAC integrates CGRA tiles with scalar functional units and specialized tiles with vector functional units that can trade off the number of vector lanes for the accuracy of the computation. We discuss the architecture design and present the related compilation framework. The experimental evaluation on a set of applications from three different domains (embedded, machine learning, and high-performance computing) shows that the hybrid design of VecPAC outperforms CGRAs with only scalar functional units by 1.48 ×, while providing higher scalability (evaluated on 2×2, 4×4, and 8×8). Moreover, VecPAC achieves better area-efficiency (1.74×) over a CGRA with only vector functional units. Cheng Tan 0002, Deepak Patil, Antonino Tumeo, Gabriel Weisz, Steven K. Reinhardt, Jeff Zhang 0001 |
ICCAD | 6 |
| 2023 | Path Planning Under Uncertainty to Localize mmWave SourcesabstractIn this paper, we study a navigation problem where a mobile robot needs to locate a mmWave wireless signal. Using the directionality properties of the signal, we propose an estimation and path planning algorithm that can efficiently navigate in cluttered indoor environments. We formulate Extended Kalman filters for emitter location estimation in cases where the signal is received in line-of-sight or after reflections. We then propose to plan motion trajectories based on belief-space dynamics in order to minimize the uncertainty of the position estimates. The associated non-linear optimization problem is solved by a state-of-the-art constrained iLQR solver. In particular, we propose a method that can handle a large number of obstacles (∼ 300) with reasonable computation times. We validate the approach in an extensive set of simulations. We show that our estimators can help increase navigation success rate and that planning to reduce estimation uncertainty can improve the overall task completion speed. Kai Pfeiffer, Yuze Jia, Mingsheng Yin, Akshaj Kumar Veldanda, Yaqi Hu, Amee Trivedi, Jeff Zhang 0001, Siddharth Garg, Elza Erkip, Sundeep Rangan, Ludovic Righetti |
ICRA | 7 |
| 2023 | COFFEE: Cross-Layer Optimization for Fast and Efficient Executions of Sinkhorn-Knopp Algorithm on HPC SystemsabstractIn this paper, we present COFFEE, cross-layer optimization for fast and efficient executions of the Sinkhorn-Knopp (SK) algorithm on HPC systems with clusters of compute nodes by exploring some architectural features of the system. By analyzing the performance of a typical implementation of the SK algorithm on such a system, a huge performance gap is observed between the row rescaling and column rescaling of the algorithm, where the latter requires much more time than the former. We also found that the costly MPI communication of the column rescaling seriously hinders the exploitation of parallelism. By observing and leveraging unique architectural characteristics across different system optimizations, such as column rescaling redesign, data blocking, micro-kernel design, enhanced intra-node and inter-node communication in MPI, etc., COFFEE is able to explore cross-layer optimization opportunities that enable fast and efficient execution of the SK algorithm. Our experimental results show that COFFEE provides up to 7.5X with an average of 2.0X performance improvement over the typical implementation on a single node, and up to 2.9X with an average of 1.6X performance improvement over the state-of-the-art MPI Allreduce algorithms on Tianhe-1 supercomputer. Chengyu Sun 0001, Huizhang Luo, Hong Jiang 0001, Jeff Zhang 0001, Kenli Li 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2022 | Energy-Efficient Brain-Inspired Hyperdimensional Computing Using Voltage ScalingabstractRecently, brain-inspired hyperdimensional computing (HDC) has demonstrated promising capability in a wide range of applications such as medical diagnosis, human activity recognition, and voice classification, etc. Despite the growing popularity of HDC, its memory-centric computing characteristics make the associative memory implementation under significant energy consumption due to the massive data storage and processing. In this paper, we present a systematic case study to leverage the application-level error resilience of HDC to reduce the energy consumption of HDC associative memory by using voltage scaling. Evaluation results on various applications show that our proposed approach can achieve 47.6% energy saving on associative memory with a 1% accuracy loss. We further explore two low-cost error masking methods: word masking and bit masking, to mitigate the impact of voltage scaling-induced errors. Experimental results show that the proposed word masking (bit masking) method can further enhance energy saving up to 62.3% (72.5%) with accuracy loss ≤1%. Sizhe Zhang, Dongning Ma, Jeff Zhang 0001, Xunzhao Yin, Xun Jiao 0002 |
DATE | 4 |
| 2022 | M2M-Routing: Environmental Adaptive Multi-agent Reinforcement Learning based Multi-hop Routing Policy for Self-Powered IoT SystemsabstractEnergy harvesting (EH) technologies facilitate the trending proliferation of IoT devices with sustainable power supplies. However, the intrinsic weak and unstable nature of EH results in frequent and unpredictable power interruptions in EH IoT devices, which further causes unpleasant packet loss or reconnection failures in IoT network. Therefore, conventional routing and energy allocation methods are inefficient in the EH environments. The complexity of the EH environment caused a stumbling block to an intelligent routing policy and energy allocation. To address the problems, this work proposes an environment adaptive Deep Reinforcement Learning (DRL)-based multi-hop routing policy, M2M-Routing, to jointly optimize energy allocation and routing policy and mitigate these challenges through leveraging the offline computation resources. We prepare multi-models for the complex energy harvesting environment offline. By searching a historically similar power trace to identify the model ID, the prepared DRL model is selected to manage energy allocation and routing policy on the query power traces. Simulation results indicate that M2M-Routing improves the amount of data delivery by ~ 3 × to ~ 4 × compared with baselines. Jeff Zhang 0001, Mimi Xie, Tao Liu 0023, Wenlu Wang |
DATE | 2 |
| 2022 | From High-Level Frameworks to custom Silicon with SODAabstractPresents a powerpoint on the topic of high level frameworks to custom silicon with SODA. Serena Curzel, Nicolas Bohm Agostini, Reece Neff, Ankur Limaye, Jeff Zhang 0001, Vinay Amatya, Marco Minutoli, Vito Giovanni Castellana, Joseph B. Manzano, David Brooks 0001, Gu-Yeon Wei, Fabrizio Ferrandi, Antonino Tumeo |
HCS | 5 |
| 2022 | A Scalable Methodology for Agile Chip Development with Open-Source Hardware ComponentsabstractWe present a scalable methodology for the agile physical design of tile-based heterogeneous system-on-chip (SoC) architectures that simplifies the reuse and integration of open-source hardware components. The methodology leverages the regularity of the on-chip communication infrastructure, which is based on a multi-plane network-on-chip (NoC), and the modularity of socket interfaces, which connect the tiles to the NoC. Each socket also provides its tile with a set of platform services, including independent clocking and voltage control. As a result, the physical design of each tile can be decoupled from its location in the top-level floorplan of the SoC and the overall SoC design can benefit from a hierarchical timing-closure flow, design reuse and, if necessary, fast respin. With the proposed methodology we completed two SoC tapeouts of increasing complexity, which illustrate its capabilities and the resulting gains in terms of design productivity. Maico Cassel, Martin Cochet, Karthik Swaminathan, Joseph Zuckerman, Paolo Mantovani, Davide Giri, Jeff Zhang 0001, Erik Jens Loscalzo, Gabriele Tombesi, Kevin Tien, Nandhini Chandramoorthy, John-David Wellman, David Brooks 0001, Gu-Yeon Wei, Kenneth L. Shepard, Luca P. Carloni, Pradip Bose |
ICCAD | 8 |
| 2022 | ASAP: automatic synthesis of area-efficient and precision-aware CGRAsabstractCoarse-grained reconfigurable accelerators (CGRAs) are a promising accelerator design choice that strikes a balance between performance and adaptability to different computing patterns across various applications domains. Designing a CGRA for a specific application domain involves enormous software/hardware engineering effort. Recent research works explore loop transformations, functional unit types, network topology, and memory size to identify optimal CGRA designs given a set of kernels from a specific application domain. Unfortunately, the impact of functional units with different precision support has rarely been investigated. To address this gap, we propose ASAP - a hardware/software co-design framework that automatically identifies and synthesizes optimal precision-aware CGRA for a set of applications of interest. Our evaluation shows that ASAP generates specialized designs 3.2X, 4.21X, and 5.8X more efficient (in terms of performance per unit of energy or area) than non-specialized homogeneous CGRAs, for the scientific computing, embedded, and edge machine learning domains, respectively, with limited accuracy loss. Moreover, ASAP provides more efficient designs than other state-of-the-art synthesis frameworks for specialized CGRAs. Cheng Tan 0002, Thierry Tambe, Jeff Zhang 0001, Bo Fang 0002, Tong Geng, Gu-Yeon Wei, David Brooks 0001, Antonino Tumeo, Ganesh Gopalakrishnan, Ang Li 0006 |
ICS | 3 |
| 2022 | End-to-End Synthesis of Dynamically Controlled Machine Learning AcceleratorsabstractEdge systems are required to autonomously make real-time decisions based on large quantities of input data under strict power, performance, area, and other constraints. Meeting these constraints is only possible by specializing systems through hardware accelerators purposefully built for machine learning and data analysis algorithms. However, data science evolves at a quick pace, and manual design of custom accelerators has high non-recurrent engineering costs: general solutions are needed to automatically and rapidly transition from the formulation of a new algorithm to the deployment of a dedicated hardware implementation. Our solution is the SOftware Defined Architectures (SODA) Synthesizer, an end-to-end, multi-level, modular, extensible compiler toolchain providing a direct path from machine learning tools to hardware. The SODA Synthesizer frontend is based on the multilevel intermediate representation (MLIR) framework; it ingests pre-trained machine learning models, identifies kernels suited for acceleration, performs high-level optimizations, and prepares them for hardware synthesis. In the backend, SODA leverages state-of-the-art high-level synthesis techniques to generate highly efficient accelerators, targeting both field programmable devices (FPGAs) and application-specific circuits (ASICs). In this paper, we describe how the SODA Synthesizer can also assemble the generated accelerators (based on the finite state machine with datapath model) in a custom system driven by a distributed controller, building a coarse-grained dataflow architecture that does not require a host processor to orchestrate parallel execution of multiple accelerators. We show the effectiveness of our approach by automatically generating ASIC accelerators for layers of popular deep neural networks (DNNs). Our high-level optimizations result in up to 74x speedup on isolated accelerators for individual DNN layers, and our dynamically scheduled architecture yields an additional 3x performance improvement when combining accelerators to handle streaming inputs. Serena Curzel, Nicolas Bohm Agostini, Vito Giovanni Castellana, Marco Minutoli, Ankur Limaye, Joseph B. Manzano, Jeff Zhang 0001, David Brooks 0001, Gu-Yeon Wei, Fabrizio Ferrandi, Antonino Tumeo |
IEEE Trans. Computers | 7 |
| 2021 | OpenCGRA: Democratizing Coarse-Grained Reconfigurable ArraysabstractReconfigurable architectures are today experiencing a renewed interest for their ability to provide specialization without sacrificing the capability to adapt to disparate workloads. Coarse-grained reconfigurable arrays (CGRAs) provide higher flexibility than application-specific integrated circuits (ASICs) while offering increased hardware efficiency with respect to field-programmable gate arrays (FPGAs). This makes CGRAs a promising alternative to enable power-/area-efficient acceleration across different application domains. Unfortunately, specializing and implementing a CGRA for a specific application domain requires the exploration in a large design space (e.g., applying appropriate loop transformation on each application, specializing the reconfigurable processing elements of the CGRA, refining the network topology, deciding the size of the data memory, etc.) and involves enormous software/hardware engineering effort (e.g., modeling, testing, and evaluating the CGRA, map operations onto the CGRA, etc). In this paper, we discuss a hardware/software co-design framework*to automatically specialize and implement optimal CGRA designs given a set of applications of interest. Cheng Tan 0002, Nicolas Bohm Agostini, Jeff Zhang 0001, Marco Minutoli, Vito Giovanni Castellana, Chenhao Xie 0001, Tong Geng, Ang Li 0006, Kevin J. Barker, Antonino Tumeo |
ASAP | 3 |
| 2021 | Towards Automatic and Agile AI/ML Accelerator Design with End-to-End SynthesisabstractDomain-specific designs offer greater energy efficiency and performance gain than general-purpose processors. For this reason, modern system-on-chips have a significant portion of their silicon area with custom accelerators. However, designing hardware by hand is laborious and time-consuming, given the large design space and the performance, power, and area constraints that are not realized in the software. Moreover, domain-specific algorithms (e.g., machine learning models) are evolving quickly, challenging the accelerator design further. To address these issues, this paper presents SODA Synthesizer, an automated open-source high-level ML framework to Verilog modular compiler targeting AI/ML Application-Specific Integrated Circuits (ASICs) accelerators. SODA tightly couples the Multi-Level Intermediate Representation (MLIR) compiler infrastructure [24] and open-source HLS approaches. Thus, SODA can support various ML frameworks and algorithms and can perform optimizations that combine specialized architecture templates and conventional HLS to generate the hardware modules. In addition, SODA’s closed-loop design space exploration (DSE) engine allows developers to perform end-to-end design space explorations on different metrics and technology nodes. Jeff Zhang 0001, Nicolas Bohm Agostini, Shihao Song, Cheng Tan 0002, Ankur Limaye, Vinay Amatya, Joseph B. Manzano, Marco Minutoli, Vito Giovanni Castellana, Antonino Tumeo, Gu-Yeon Wei, David Brooks 0001 |
ASAP | 1 |
| 2021 | Assessing Robustness of Hyperdimensional Computing Against Errors in Associative Memory : (Invited Paper)abstractBrain-inspired hyperdimensional computing (HDC) is an emerging computational paradigm that has achieved success in various domains. HDC mimics brain cognition and lever-ages hyperdimensional vectors with fully distributed holographic representation and (pseudo)randomness. Compared to the traditional machine learning methods, HDC offers several critical advantages, including smaller model size, less computation cost, and one-shot learning capability, making it a promising candidate in low-power platforms. Despite the growing popularity of HDC, the robustness of HDC models has not been systematically explored. This paper presents a study on the robustness of HDC to errors in associative memory—the key component storing the class representations in HDC. We perform extensive error injection experiments to the associative memory in a number of HDC models (and datasets), sweeping the error rates and varying HDC configurations (i.e., dimension and data width). Empirically, we observe that HDC is considerably robust to errors in the associative memory, opening up opportunities for further optimizations. Further, results show that HDC robustness varies significantly with different HDC configurations such as data width. Moreover, we explore a low-cost error masking mechanism in the associative memory to enhance its robustness. Sizhe Zhang, Jeff Zhang 0001, Abbas Rahimi, Xun Jiao 0002 |
ASAP | 3 |
| 2021 | RecPipe: Co-designing Models and Hardware to Jointly Optimize Recommendation Quality and PerformanceabstractDeep learning recommendation systems must provide high quality, personalized content under strict tail-latency targets and high system loads. This paper presents RecPipe, a system to jointly optimize recommendation quality and inference performance. Central to RecPipe is decomposing recommendation models into multi-stage pipelines to maintain quality while reducing compute complexity and exposing distinct parallelism opportunities. RecPipe implements an inference scheduler to map multi-stage recommendation engines onto commodity, heterogeneous platforms (e.g., CPUs, GPUs). While the hardware-aware scheduling improves ranking efficiency, the commodity platforms suffer from many limitations requiring specialized hardware. Thus, we design RecPipeAccel (RPAccel), a custom accelerator that jointly optimizes quality, tail-latency, and system throughput. RPAccel is designed specifically to exploit the distinct design space opened via RecPipe. In particular, RPAccel processes queries in sub-batches to pipeline recommendation stages, implements dual static and dynamic embedding caches, a set of top-k filtering units, and a reconfigurable systolic array. Compared to previously proposed specialized recommendation accelerators and at iso-quality, we demonstrate that RPAccel improves latency and throughput by 3 × and 6 ×. Udit Gupta 0001, Samuel Hsia, Jeff Zhang 0001, Mark Wilkening, Javin Pombra, Hsien-Hsin S. Lee, Gu-Yeon Wei, Carole-Jean Wu, David Brooks 0001 |
MICRO | 3 |
| 2019 | Building Robust Machine Learning Systems: Current Progress, Research Challenges, and OpportunitiesabstractMachine learning, in particular deep learning, is being used in almost all the aspects of life to facilitate humans, specifically in mobile and Internet of Things (IoT)-based applications. Due to its state-of-the-art performance, deep learning is also being employed in safety-critical applications, for instance, autonomous vehicles. Reliability and security are two of the key required characteristics for these applications because of the impact they can have on human's life. Towards this, in this paper, we highlight the current progress, challenges and research opportunities in the domain of robust systems for machine learning-based applications. Jeff Zhang 0001, Kang Liu 0017, Faiq Khalid, Muhammad Abdullah Hanif, Semeen Rehman, Theocharis Theocharides, Alessandro Artussi, Muhammad Shafique 0001, Siddharth Garg |
DAC | 1 |
| 2019 | Split Manufacturing-Based Register Transfer-Level ObfuscationabstractFabrication-less integrated circuit (IC) design houses outsource fabrication to third-party foundries to reduce cost of manufacturing. The outsourcing of IC fabrication, beyond our expectation, raises concerns regarding intellectual property (IP) piracy and theft by rogue elements in the third-party foundries. Obfuscation techniques have been proposed to increase resistance to reverse engineering, IP recovery, IP theft, and piracy. However, prior work on obfuscation for IP protection has primarily applied to the gate level or the layout level. As a result, it can significantly impact the performance of the original design in addition to requiring redesign of standard cells. In this article, we propose a high-level synthesis and analysis (HLSA)-based obfuscation approach for IP protection. The proposed method is based on split manufacturing. Additional dummy units and MUXes can be added to further obfuscate the design. The proposed technique aligns with the standard-cell-based design methodologies and does not significantly impact the performance of the original design. Our experimental results confirm that the proposed approach can provide high levels of IC obfuscation with moderate area cost. Xiaotong Cui, Jeff Zhang 0001, Kaijie Wu 0001, Siddharth Garg, Ramesh Karri |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2019 | CompAct: On-chip <underline>Com</underline>pression of <underline>Act</underline>ivations for Low Power Systolic Array Based CNN AccelerationabstractThis paper addresses the design of systolic array (SA) based convolutional neural network (CNN) accelerators for mobile and embedded domains. On- and off-chip memory accesses to the large activation inputs (sometimes called feature maps) of CNN layers contribute significantly to total energy consumption for such accelerators; while prior has proposed off-chip compression, activations are still stored on-chip in uncompressed form, requiring either large on-chip activation buffers or slow and energy-hungry off-chip accesses. In this paper, we propose CompAct, a new architecture that enables on-chip compression of activations for SA based CNN accelerators. CompAct is built around several key ideas. First, CompAct identifies an SA schedule that has nearly regular access patterns, enabling the use of a modified run-length coding scheme (RLC). Second, CompAct improves compression ratio of the RLC scheme using Sparse-RLC in later CNN layers and Lossy-RLC in earlier layers. Finally, CompAct proposes look-ahead snoozing that operates synergistically with RLC to reduce the leakage energy of activation buffers. Based on detailed synthesis results, we show that CompAct enables up to 62% reduction in activation buffer energy, and 34% reduction in total chip energy. Jeff Zhang 0001, Parul Raj, Shuayb Zarar, Amol Ambardekar, Siddharth Garg |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2018 | Thundervolt: enabling aggressive voltage underscaling and timing error resilience for energy efficient deep learning acceleratorsabstractHardware accelerators are being increasingly deployed to boost the performance and energy efficiency of deep neural network (DNN) inference. In this paper we propose Thundervolt, a new framework that enables aggressive voltage underscaling of high-performance DNN accelerators without compromising classification accuracy even in the presence of high timing error rates. Using post-synthesis timing simulations of a DNN accelerator modeled on the Google TPU, we show that Thundervolt enables between 34%-57% energy savings on state-of-the-art speech and image recognition benchmarks with less than 1% loss in classification accuracy and no performance loss. Further, we show that Thundervolt is synergistic with and can further increase the energy efficiency of commonly used run-time DNN pruning techniques like Zero-Skip. Jeff Zhang 0001, Kartheek Rangineni, Zahra Ghodsi, Siddharth Garg |
DAC | 1 |
| 2018 | FATE: fast and accurate timing error prediction framework for low power DNN accelerator designabstractDeep neural networks (DNN) are increasingly being accelerated on application-specific hardware such as the Google TPU designed especially for deep learning. Timing speculation is a promising approach to further increase the energy efficiency of DNN accelerators. Architectural exploration for timing speculation requires detailed gate-level timing simulations that can be time-consuming for large DNNs which execute millions of multiply-and-accumulate (MAC) operations. In this paper we propose FATE, a new methodology for fast and accurate timing simulations of DNN accelerators like the Google TPU. FATE proposes two novel ideas: (i) DelayNet, a DNN based timing model for MAC units; and (ii) a statistical sampling methodology that reduces the number of MAC operations for which timing simulations are performed. We show that FATE results in between 8 × −58× speed-up in timing simulations, while introducing less than 2% error in classification accuracy estimates. We demonstrate the use of FATE by comparing a conventional DNN accelerator that uses 2's complement (2C) arithmetic with one that uses signed magnitude representation (SMR). We show that that the SMR implementation provides 18% more energy savings for the same classification accuracy than 2C, a result that might be of independent interest. Jeff Zhang 0001, Siddharth Garg |
ICCAD | 1 |
| 2018 | Analyzing and mitigating the impact of permanent faults on a systolic array based neural network acceleratorabstractDue to their growing popularity and computational cost, deep neural networks (DNNs) are being targeted for hardware acceleration. A popular architecture for DNN acceleration, adopted by the Google Tensor Processing Unit (TPU), utilizes a systolic array based matrix multiplication unit at its core. This paper deals with the design of fault-tolerant, systolic array based DNN accelerators for high defect rate technologies. To this end, we empirically show that the classification accuracy of a baseline TPU drops significantly even at extremely low fault rates (as low as 0.006%). We then propose two novel strategies, fault-aware pruning (FAP) and fault-aware pruning+retraining (FAP+T), that enable the TPU to operate at fault rates of up to 50%, with negligible drop in classification accuracy (as low as 0.1%) and no run-time performance overhead. The FAP+T does introduce a one-time retraining penalty per TPU chip before it is deployed, but we propose optimizations that reduce this one-time penalty to under 12 minutes. The penalty is then amortized over the entire lifetime of the TPU's operation. Jeff Zhang 0001, Tianyu Gu, Kanad Basu, Siddharth Garg |
VTS | 1 |
| 2017 | BandiTS: Dynamic timing speculation using multi-armed bandit based optimizationabstractTiming speculation has recently been proposed as a method for increasing performance beyond that achievable by conventional worst-case design techniques. Starting with the observation of fast temporal variations in timing error probabilities, we propose a run-time technique to dynamically determine the optimal degree of timing speculation (i.e., how aggressively the processor is over-clocked) based on a novel formulation of the dynamic timing speculation problem as a multi-armed bandit problem. By conducting detailed post-synthesis timing simulations on a 5-stage MIPS processor running a variety of workloads, the proposed adaptive mechanism improves processor's performance significantly comparing with a competing approach (about 8.3% improvement); on the other hand, it shows only about 2.8% performance loss on average, compared with the oracle results. Jeff Zhang 0001, Siddharth Garg |
DATE | 1 |
| 2016 | Synergistic timing speculation for multi-threaded programsabstractIn this paper, we address the problem of timing speculation for multi-threaded workloads executing on a multi-core processor. Our approach is based on a new observation --- heterogeneity in path sensitization delays across different threads in multi-threaded programs. Leveraging this heterogeneity, we propose Synergistic Timing Speculation (SynTS) to jointly optimize the energy and execution time of multithreaded applications. In particular, SynTS uses a sampling based online error probability estimation technique, coupled with a polynomial time algorithm, to optimally determine the voltage, frequency and the amount of timing speculation for each thread. Our experimental evaluations, based on detailed cross-layer simulations, demonstrate that SynTS reduces energy delay product by up to 21%, compared to existing timing speculation schemes. Atif Yasin, Jeff Zhang 0001, Siddharth Garg, Sanghamitra Roy, Koushik Chakraborty |
DAC | 2 |
| 2015 | An Adaptive Invasive Weed Optimization AlgorithmabstractWith regards to the low search accuracy of the basic invasive weed optimization algorithm which is easy to get into local extremum, this paper proposes an adaptive invasive weed optimization (AIWO) algorithm. The algorithm sets the initial step size and the final step size as the adaptive step size to guide the global search of the algorithm, and it is applied to 20 famous benchmark functions for a test, the results of which show that the AIWO algorithm owns better global optimization search capacity, faster convergence speed and higher computation accuracy compared with other advanced algorithms. Shuo Peng, Aijia Ouyang, Jeff Zhang 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 3 |