EDBT 2026 Demo / reviewers in the wild / expert
Jingchen Zhu
dblp:45/7775
· DBLP profile ↗
6ranked-venue papers
3as first author
5since 2021 · last 2025
0000-0002-4321-7694ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Defect-Aware Task Scheduling and Mapping for Redundancy-Enhanced Spatial Accelerators
Jingchen Zhu, Guangyu Sun 0003 |
APPT | 1 |
| 2025 | H2-LLM: Hardware-Dataflow Co-Exploration for Heterogeneous Hybrid-Bonding-based Low-Batch LLM InferenceabstractLow-batch large language model (LLM) inference has been extensively applied to edge-side generative tasks, such as personal chat helper, virtual assistant, reception bot, private edge server, etc.To efficiently handle both prefill and decoding stages in LLM inference, near-memory processing (NMP) enabled heterogeneous computation paradigm has been proposed.However, existing NMP designs typically embed processing engines into DRAM dies, resulting in limited computation capacity, which in turn restricts their ability to accelerate edge-side low-batch LLM inference.To tackle this problem, we propose H 2 -LLM, a Hybrid-bondingbased Heterogeneous accelerator for edge-side low-batch LLM inference.To balance the trade-off between computation capacity and bandwidth intrinsic to hybrid-bonding technology, we propose * Co-corresponding authors. Cong Li 0008, Yihan Yin, Xintong Wu, Jingchen Zhu, Zhutianya Gao, Dimin Niu, Qiang Wu 0012, Xin Si, Yuan Xie 0001, Chen Zhang 0001, Guangyu Sun 0003 |
ISCA | 4 |
| 2025 | AIM: Software and Hardware Co-design for Architecture-level IR-drop Mitigation in High-performance PIMabstractSRAM Processing-in-Memory (PIM) has emerged as the most promising implementation for high-performance PIM, delivering superior computing density, energy efficiency, and computational precision.However, the pursuit of higher performance necessitates more complex circuit designs and increased operating frequencies, which exacerbate IR-drop issues.Severe IR-drop can significantly degrade chip performance and even threaten reliability.Conventional circuit-level IR-drop mitigation methods, such as back-end optimizations, are resource-intensive and often compromise power, performance, and area (PPA).To address these challenges, we propose AIM, comprehensive software and hardware co-design for architecture-level IR-drop mitigation in high-performance PIM.Initially, leveraging the bit-serial and in-situ dataflow processing properties of PIM, we introduce R tog and HR, which establish a direct correlation between PIM workloads and IR-drop.Building on this foundation, we propose LHR and WDS, enabling extensive exploration of architecture-level IR-drop mitigation while maintaining computational accuracy through software optimization.Subsequently, we develop IR-Booster, a dynamic adjustment mechanism that integrates software-level HR information with hardwarebased IR-drop monitoring to adapt the V-f pairs of the PIM macro, achieving enhanced energy efficiency and performance.Finally, we propose the HR-aware task mapping method, bridging software and hardware designs to achieve optimal improvement.Post-layout simulation results on a 7nm 256-TOPS PIM chip demonstrate that AIM achieves up to 69.2% IR-drop mitigation, resulting in 2.29× energy efficiency improvement and 1.152× speedup. Yuanpeng Zhang 0002, Xing Hu 0010, Xi Chen 0107, Zhihang Yuan, Cong Li 0008, Jingchen Zhu, Xin Si, Wei Gao 0058, Qiang Wu 0012, Runsheng Wang, Guangyu Sun 0003 |
ISCA | 6 |
| 2025 | Theseus: Exploring Efficient Wafer-Scale Chip Design for Large Language ModelsabstractThe emergence of the large language model (LLM) poses an exponential growth of demand for computation throughput, memory capacity, and communication bandwidth. Such a demand growth has significantly surpassed the improvement of corresponding chip designs. With the advancement of fabrication and integration technologies, designers have been developing Wafer-Scale Chips (WSCs) to scale up and exploit the limits of computation density, memory capacity, and communication bandwidth at the level of a single chip. Existing solutions have demonstrated the significant advantages of WSCs over traditional designs, showing potential to effectively support LLM workloads. Despite the benefits, exploring the early-stage design space of WSCs for LLMs is a crucial yet challenging task due to the enormous and complicated design space, time-consuming evaluation methods, and inefficient exploration strategies. To address these challenges, we propose Theseus, an efficient WSC design space exploration framework for LLMs. We construct the design space of WSCs with various constraints considering the unique characteristics of WSCs. We propose efficient evaluation methodologies for large-scale NoC-based WSCs and introduce multi-fidelity Bayesian optimization to efficiently explore the design space. Evaluation results demonstrate the efficiency of Theseus that the searched Pareto optimal results outperform GPU cluster and existing WSC designs by up to 62.8%/73.7% in performance (with the same or lower power) and 38.6%/42.4% in power consumption (with the same or higher performance) for LLM training, while improving up to 23.2× and 15.7× for the performance and power of inference tasks. Furthermore, we conduct case studies to address the design tradeoffs in WSCs and provide insights to facilitate WSC designs for LLMs. Jingchen Zhu, Chenhao Xue, Chen Zhang 0001, Yu Shen 0003, Zekang Cheng, Yibo Lin, Wei Hu 0003, Bin Cui 0001, Runsheng Wang, Yun Liang 0001, Guangyu Sun 0003 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | FD-CNN: A Frequency-Domain FPGA Acceleration Scheme for CNN-Based Image-Processing ApplicationsabstractIn the emerging edge-computing scenarios, FPGAs have been widely adopted to accelerate convolutional neural network (CNN)–based image-processing applications, such as image classification, object detection, and image segmentation, and so on. A standard image-processing pipeline first decodes the collected compressed images from Internet of Things (IoTs) to RGB data, then feeds them into CNN engines to compute the results. Previous works mainly focus on optimizing the CNN inference parts. However, we notice that on the popular ZYNQ FPGA platforms, image decoding can also become the bottleneck due to the poor performance of embedded ARM CPUs. Even with a hardware accelerator, the decoding operations still incur considerable latency. Moreover, conventional RGB-based CNNs have too few input channels at the first layer, which can hardly utilize the high parallelism of CNN engines and greatly slows down the network inference. To overcome these problems, in this article, we propose FD-CNN, a novel CNN accelerator leveraging the partial-decoding technique to accelerate CNNs directly in the frequency domain. Specifically, we omit the most time-consuming IDCT (Inverse Discrete Cosine Transform) operations of image decoding and directly feed the DCT coefficients (i.e., the frequency data) into CNNs. By this means, the image decoder can be greatly simplified. Moreover, compared to the RGB data, frequency data has a narrower input resolution but has 64× more channels. Such an input shape is more hardware friendly than RGB data and can substantially reduce the CNN inference time. We then systematically discuss the algorithm, architecture, and command set design of FD-CNN. To deal with the irregularity of different CNN applications, we propose an image-decoding-aware design-space exploration (DSE) workflow to optimize the pipeline. We further propose an early stopping strategy to tackle the time-consuming progressive JPEG decoding. Comprehensive experiments demonstrate that FD-CNN achieves, on average, 3.24×, 4.29× throughput improvement, 2.55×, 2.54× energy reduction and 2.38×, 2.58× lower latency on ZC-706 and ZCU-102 platforms, respectively, compared to the baseline image-processing pipelines. Xiaoyang Wang 0006, Zhe Zhou 0002, Zhihang Yuan, Jingchen Zhu, Kangrui Sun, Guangyu Sun 0003 |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2020 | Fork Path: Batching ORAM Requests to Remove Redundant Memory AccessesabstractOutsourcing data to a third-party cloud provider has become quite common with the increasing use of cloud computing. This brings convenience, as well as the concern for data security and privacy. It is believed that data encryption alone is often not enough to protect users' privacy from the cloud provider. According to previous work, the sequence of storage locations accessed by the client can leak up to 90% of the sensitive information, even with data encrypted. In this context, Oblivious RAM (ORAM) is proposed. ORAM algorithms allow the client to hide its access pattern from the service provider while introducing a lot of extra operations. Among all the prototypes, Path ORAM is one of the most promising designs. However, there are still redundant memory accesses that can be removed without harming the security of traditional ORAM as we observed. We came up with three optimization techniques, including path merging, ORAM request scheduling, and merging aware caching. We also propose a prefetching technique to further decreasing the access overhead. Moreover, we also illustrate the compatibility of Fork Path and some state-of-the-art Path ORAM optimizations. Compared to traditional Path ORAM approaches, our Fork Path ORAM can reduce overall performance overhead and power consumption of memory system by 65% and 44%, while the design overhead is trivial. Jingchen Zhu, Guangyu Sun 0003, Xian Zhang 0001, Chao Zhang 0007, Yun Liang 0001, Tao Wang 0004, Yiran Chen 0001, Jia Di |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |