EDBT 2026 Demo / reviewers in the wild / expert
Jianfeng Zhu 0001
dblp:69/6211-1
· DBLP profile ↗
35ranked-venue papers
5as first author
24since 2021 · last 2026
0000-0002-0485-8034ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 32 · 5 first-author · 21 since 2021Software engineering, systems software and programming languages · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EDWAC: A Deadlock-Free Scheme for Compiling Whole Programs Onto Dynamically Reconfigurable Dataflow ArchitecturesabstractCoarse-grained Reconfigurable Arrays (CGRAs) have become prevailing to accelerate regular kernels coupled with a host processor. As the end-to-end applications are increasingly complex, it is necessary to consider mapping whole programs onto a monolithic CGRA, to avoid the bottleneck of host communication implied by Amdahl’s law.State-of-the-art studies have developed compiling methods that spatially pipeline the whole program over hardware with many cores. However, these methods mainly focus on static reconfigurable dataflow architectures and fail to exploit the dynamic reconfiguration potential of dataflow architectures, resulting in suboptimal performance and underutilization of hardware resources. Nevertheless, it is nontrivial to generate a performant mapping on dynamic reconfigurable dataflow architectures since instruction-level deadlocks are introduced. To address this challenge, this paper proposes EDWAC, a whole-program compiler that generates high-quality configurations for dynamic reconfigurable dataflow architectures. EDWAC resolves the deadlock problem by a two-stage deadlock-prevention mechanism, which comprises a shared-resource-constrained Place and Route (PnR) stage, and a Finite State Machine(FSM)-based resource reallocation stage. Together with a gated control flow Intermediate Representation (IR) design and throughput-oriented optimization methods, EDWAC achieves exceptional resource utilization and PnR feasibility. Jianfeng Zhu 0001, Xingchen Man, Guihuan Song, Zijiao Ma, Shanxin Chen, Chunyang Feng, Yang Liu 0326, Shaojun Wei, Leibo Liu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2025 | Chameleon-SAT: An Adaptive Boolean Satisfiability Accelerator Using Mixed-Signal In-Memory Computing for Versatile SAT ProblemsabstractBoolean satisfiability (SAT), the first proven nondeterministic polynominal-complete problem, is crucial in dataintensive applications. Different applications have a wide spectrum of SAT problem sets (scale, complexity) and also various solution requirements (algorithm completeness, speed). Current SAT solvers are insufficient for providing ideal solutions under different scenarios. This work presents the Chameleon-SAT, the first ASIC-based SAT accelerator that can support local search, Davis-Putnam- Logemann-Lovel, Conflict-Driven Clause Learning algorithms, while leveraging the efficient mixed-signal inmemory computing architecture to achieve orders-of-magnitude improvements in speed compared to the prior SAT solvers. By judiciously selecting the reconfiguration mode, Chameleon-SAT is able to solve a wide range of the SAT problems to achieve smallscale, high-complexity cases ($\geq 90 \times$ for 20 variables/ 86 clauses, satisfiable problems), medium-scale, structured cases ($\geq 19 \times$ for 50 variables/ 215 clauses, unsatisfiable problems), and largescale, high-complexity cases ($\geq 7 \times$ for 100 variables/ 430 clauses, satisfiable problems). Iris Ying Chou, Hao Kong 0003, Yi Huang 0036, Jianfeng Zhu 0001, Wenping Zhu, Shaojun Wei, Aoyang Zhang, Leibo Liu |
DAC | 4 |
| 2025 | EFFACT: A Highly Efficient Full-Stack FHE Acceleration PlatformabstractFully Homomorphic Encryption (FHE) is a set of powerful cryptographic schemes that allows computation to be performed directly on encrypted data with an unlimited depth. Despite FHE’s promising in privacy-preserving computing, yet in most FHE schemes, ciphertext generally blows up thousands of times compared to the original message, and the massive amount of data load from off-chip memory for bootstrapping and privacy-preserving machine learning applications (such as HELR, ResNet-20), both degrade the performance of FHE-based computation. Several hardware designs have been proposed to address this issue, however, most of them require enormous resources and power. An acceleration platform with easy programmability, high efficiency, and low overhead is a prerequisite for practical application. This paper proposes EFFACT, a highly efficient full-stack FHE acceleration platform with a compiler that provides comprehensive optimizations and vector-friendly hardware. We start by examining the computational overhead across different real-world benchmarks to highlight the potential benefits of reallocating computing resources for efficiency enhancement. Then we make a design space exploration to find an optimal SRAM size with high utilization and low cost. On the other hand, EFFACT features a novel optimization named streaming memory access which is proposed to enable high throughput with limited SRAMs. Regarding the software-side optimization, we also propose a circuit-level function unit reuse scheme, to substantially reduce the computing resources without performance degradation. Moreover, we design novel NTT and automorphism units that are suitable for a cost-sensitive and highly efficient architecture, leading to low area. For generality, EFFACT is also equipped with an ISA and a compiler backend that can support several FHE schemes like CKKS, BGV, and BFV. We provide both FPGA and ASIC versions of EFFACT. On account of our full stack design, FPGA-EFFACT outperforms the SOTA FPGA accelerators in gmean by $1.22 \times$. Meanwhile, ASIC-EFFACT shows increased improvements in terms of the performance per chip area and the performance per Watt compared with the SOTA ASIC works. Yi Huang 0036, Xinsheng Gong, Dibei Chen, Jianfeng Zhu 0001, Wenping Zhu, Liangwei Li, Mingyu Gao 0001, Shaojun Wei, Aoyang Zhang, Leibo Liu |
HPCA | 5 |
| 2025 | Lincoln: Real-Time 50~100B LLM Inference on Consumer Devices with LPDDR-Interfaced, Compute-Enabled Flash MemoryabstractWith the widespread use of large language models (LLMs), and with the privacy and cost concerns on cloud-based services, vendors are now pushing LLM inference to consumer devices. However, current attempts only enable real-time inference of low-quality small-sized LLMs. Large-sized LLMs have to load most of their weights from Flash storage for every execution iteration, which dominates the execution time of both the prefill and the generation phase. This performance bottleneck is attributed to both the low internal Flash memory bandwidth and the low transmission bandwidth between Flash and the Neural Processing Unit (NPU). To tackle these two challenges, we present Lincoln, a device-architecture co-design solution with LPDDR-interfaced, Compute-Enabled Flash Memory. On the device level, we boost the Flash internal bandwidth by improving upon existing array shrinking methods, to enable lower read latency and more parallel Flash planes within each Flash die. We specifically leverage 3D hybrid bonding, which is already adopted in consumer Flash products, to maintain high area efficiency and low density loss. On the architecture level, to leverage such increased internal bandwidth for resolving the transmission bottleneck, we propose two solutions for the two distinct phases of LLMs. For the compute-intensive prefill phase, we let Flash devices use the existing high-speed LPDDR interface (originally for DRAM), which offers much higher transmission bandwidth to the NPU than the conventional Flash interface, while maintaining good cost and area efficiency. For the memory-intensive generation phase, we rely on hybrid-bonding-based near-Flash computing to fully utilize the internal Flash bandwidth, and further equip with speculative decoding to eventually reach the real-time latency goal. Our evaluation shows that Lincoln enables real-time inference, with up to $13.23 \times$ and $254.1 \times$ speedups for LLM prefill and generation phases over conventional SSD-based systems. Weiyi Sun, Mingyu Gao 0001, Zhaoshi Li, Aoyang Zhang, Iris Ying Chou, Jianfeng Zhu 0001, Shaojun Wei, Leibo Liu |
HPCA | 6 |
| 2025 | PointISA: ISA-Extensions for Efficient Point Cloud Analytics via Architecture and Algorithm Co-Design
Liang Wang 0020, Limin Xiao 0001, Hao Zhang 0205, Xilong Xie, Jianfeng Zhu 0001, Shaojun Wei, Leibo Liu |
MICRO | 7 |
| 2025 | Software-defined process-near-memory architecture using 3D hybrid bonding integration
Anlin Xu, Chenchen Deng, Jianfeng Zhu 0001, Shaojun Wei, Leibo Liu |
Sci. China Inf. Sci. | 3 |
| 2025 | Exploiting Fine-Grained Task-Level Parallelism for Variant Calling AccelerationabstractVariant calling, which identifies genomic differences relative to a reference genome, is critical for understanding disease mechanisms, identifying therapeutic targets, and advancing precision medicine. However, as two critical stages in this process, serial processing in local assembly and the computational dependencies in Pair-HMM make variant calling highly time-consuming. Moreover, optimizing only one of these stages often shifts the performance bottleneck to the other. This paper observes that the similarity between reads allows parallel processing in the local assembly and that alignment information from the local assembly can significantly diminish the burdensome computations in Pair-HMM. Accordingly, this paper co-optimizes the software and hardware for both steps to achieve the best performance. First, we collect$k$-mer locations in each read during the local assembly process and utilize the similarity between reads to make it parallel. Second, we propose the mPair-HMM algorithm, leveraging location information to split a Pair-HMM computation task into multiple independent sub-tasks, improving the computation's parallelism. To fully exploit the parallelism stemming from the novel algorithms, we propose an end-to-end accelerator VCAx for variant calling that accelerates both stages in collaboration. Evaluation results demonstrate that our implementation achieves up to a 7× speedup over the GPU baseline for local assembly and a 3.16× performance improvement compared to the state-of-the-art ASIC implementation for Pair-HMM. Longlong Chen, Hongyi Guan, Shaojun Wei, Jianfeng Zhu 0001, Leibo Liu |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2025 | Raccoon: Lightweight Support for Comprehensive Control Flows in Reconfigurable Spatial ArchitecturesabstractCoarse-grained reconfigurable arrays (CGRAs) have emerged as promising candidates for digital signal processing, biomedical, and automotive applications, where energy efficiency and flexibility are paramount. Yet existing CGRAs suffer from the Amdahl bottleneck caused by constrained control handling via either off-device communication or expensive tag-matching mechanisms. More importantly, mapping control flow onto CGRAs is extremely arduous and time-consuming due to intricate instruction structures and hardware mechanisms. To counteract these limitations, we propose Raccoon, a portable and lightweight framework for CGRAs targeting vast control flows. Raccoon comprises a comprehensive approach that spans microarchitecture, HW/SW interface, and compiler aspects. Regarding microarchitecture, Raccoon incorporates specialized infrastructure for branch- and loop-level control patterns with concise execution mechanisms. The HW/SW interface of Raccoon includes well-characterized abstractions and instruction sets tailored for easy compilation, featuring custom operators and architectural models for control-oriented units. On the compiler front, Raccoon integrates advanced control handling techniques and employs a portable mapper leveraging reinforcement learning and Monte Carlo tree search. This enables agile mapping and optimization of the entire program, ensuring efficient execution and high-quality results. Through the cohesive co-design, Raccoon can empower various CGRAs with robust control-flow handling capabilities, surpassing conventional tagged mechanisms in terms of hardware efficiency and compiler adaptability. Evaluation results show that Raccoon achieves up to a 5.78× improvement in energy efficiency and a 2.24× reduction in cycle count over state-of-the-art CGRAs. Raccoon stands out for its versatility in managing intricate control flows and showcases remarkable portability across diverse CGRA architectures. Yi Huang 0036, Longlong Chen, Jianfeng Zhu 0001, Liangwei Li, Xingchen Man, Mingyu Gao 0001, Shaojun Wei, Leibo Liu |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2025 | SSS-DIMM: Removing Redundant Data Movement in Trusted DIMM-Based Near-Memory-Processing Kernel Offloading via Secure Space SharingabstractDIMM-based Near-Memory-Processing (NMP) kernel offloading enables a program to execute in computation-enabled DIMM buffer chips, bypassing the bandwidth-constrained CPU main memory bus for high performance. Yet, it also enables programs to access memory without restrictions and protection from CPU, resulting in potential security hazards. To protect general NMP kernel offloading even with malicious privileged software, a heterogeneous TEE is required. However, the conventional heterogeneous TEE design results in severe data movement bottleneck for DIMM-based NMP. Concretely, it isolates host CPU process from NMP kernel's memory and vice versa, such that CPU TEE and trusted NMP driver can protect CPU processes and NMP kernels in complete separation, simplifying the architectural design. Such isolation results in redundant input/output data movement between the two isolated memory spaces, with half of the movement performed by host CPU. Worsened by limited CPU memory bandwidth, we identify that such redundancy severely bottlenecks the performance of many potential NMP applications. To overcome this bottleneck, we propose to abandon isolation and share the NMP kernel memory with its host CPU process. Considering security, however, two challenges exist that fundamentally contradict the conventional separation-oriented TEE design. First, for protection against software attacks on the shared memory, consistent security guarantees have to be offered by the CPU TEE and the NMP driver respectively on CPU processes and NMP kernels, in terms of both memory ownership (allocation) and views (mapping). Second, to enable shared memory access while offering protection against physical attacks, cryptography metadata like keys and Merkle tree root have to be securely shared and synchronized between CPU and NMP unit. To overcome these challenges, we designSSS-DIMM, an efficient TEE for DIMM-based NMP kernel offloading that removes the redundant data movement viaSecureSpaceSharing. At its core, we devise secure, general and complexity-minimized instruction interfaces, which empower the trusted NMP driver with restricted authority to access the memory allocation/mapping recordings of CPU TEE, and to set and access cryptography metadata of the shared memory in both NMP unit and CPU. Along with carefully designed software workflows, these interfaces enable full resolve of the challenges. Compared with conventional heterogeneous TEE and the unprotected baseline, our evaluation shows that SSS-DIMM maintains both security and performance, achieving a geomean speedup of 9.1× for NMP kernel offloading over conventional TEE design. Weiyi Sun, Jianfeng Zhu 0001, Mingyu Gao 0001, Zhaoshi Li, Shaojun Wei, Leibo Liu |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2024 | Harp: Leveraging Quasi-Sequential Characteristics to Accelerate Sequence-to-Graph Mapping of Long ReadsabstractRead mapping is a crucial task in computational genomics. Recently, there has been a significant paradigm shift from sequence-to-sequence mapping (S2S) to sequence-to-graph mapping (S2G). The S2G mapping incurs high graph processing overheads and leads to an unnoticed shift of performance hotspots. This presents a substantial challenge to current software implementations and hardware accelerators. Dibei Chen, Jianfeng Zhu 0001, Zhaoshi Li, Longlong Chen, Shaojun Wei, Leibo Liu |
ASPLOS (3) | 4 |
| 2024 | CATCAM: a 28 nm constant-time alteration TCAM enabling less than 50 ns update latency
Chenchen Deng, Tianzhu Xiong, Zhaoshi Li, Jianfeng Zhu 0001, Jun Yang 0006, Shaojun Wei, Leibo Liu |
Sci. China Inf. Sci. | 6 |
| 2024 | A High-Performance Genomic Accelerator for Accurate Sequence-to-Graph Alignment Using Dynamic Programming AlgorithmabstractThe rapid mutation of viruses, such as SARS-CoV-2, highlights the urgent need for fast and precise genomic sequencing. The traditional sequencing technique maps the DNA fragments collected from an individual to a known linear reference genome sequence. The linear reference cannot express the genetic diversity of the population, which leads to mapping bias. Therefore, researchers proposed to use a graph reference together with long reads for sequence mapping so that the mapping bias can be avoided to the greatest extent. However, the graph reference introduces irregular edges making memory access of alignment a bottleneck and meanwhile the long read quadratically increases the storage pressure in the alignment process. Therefore, there is a pressing need for a high-performance hardware accelerator for accurate sequence-to-graph alignment. To our best knowledge, this paper presents ASGDP, the first hardware accelerator designed for aligning sequences of arbitrary length reads to a graph. It is based on the traditional dynamic programming algorithm and supports flexible penalty scoring strategies. ASGDP has proposed an efficient memory access pattern in hardware and a hierarchical prediction pruning strategy in algorithm. This combined software-hardware strategy effectively alleviates the storage bottleneck of multi-edge access and improves the accuracy of pruning strategies. We demonstrate that ASGDP provides significant improvements for long reads of the sequence-to-graph alignment. For a typical 10 K long read, a single ASGDP accelerator outperforms state-of-the-art S2G mapping tools by 70.8×, 168.1×. Jianfeng Zhu 0001, Ganhui Chen, Zhenhai Yuan, Shaojun Wei, Leibo Liu |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2023 | Orinoco: Ordered Issue and Unordered Commit with Non-Collapsible QueuesabstractModern out-of-order processors call for more aggressive scheduling techniques such as priority scheduling and out-of-order commit to make use of increasing core resources. Since these approaches prioritize the issue or commit of certain instructions, they face the conundrum of providing the capacity efficiency of scheduling structures while preserving the ideal ordering of instructions. Traditional collapsible queues are too expensive for today's processors, while state-of-the-art queue designs compromise with the pseudo-ordering of instructions, leading to performance degradation as well as other limitations. Dibei Chen, Tairan Zhang, Yi Huang 0036, Jianfeng Zhu 0001, Yang Liu 0326, Pengfei Gou, Chunyang Feng, Shaojun Wei, Leibo Liu |
ISCA | 4 |
| 2023 | MapZero: Mapping for Coarse-grained Reconfigurable Architectures with Reinforcement Learning and Monte-Carlo Tree SearchabstractCoarse-grained reconfigurable architecture (CGRA) has become a promising candidate for data-intensive computing due to its flexibility and high energy efficiency. CGRA compilers map data flow graphs (DFGs) extracted from applications onto CGRAs, playing a fundamental role in fully exploiting hardware resources for acceleration. Yet the existing compilers are time-demanding and cannot guarantee optimal results due to the traversal search of enormous search spaces brought about by the spatio-temporal flexibility of CGRA structures and the complexity of DFGs. Inspired by the amazing progress in reinforcement learning (RL) and Monte-Carlo tree search (MCTS) for real-world problems, we consider constructing a compiler that can learn from past experiences and comprehensively understand the target DFG and CGRA. Yi Huang 0036, Jianfeng Zhu 0001, Xingchen Man, Yang Liu 0326, Chunyang Feng, Pengfei Gou, Minggui Tang, Shaojun Wei, Leibo Liu |
ISCA | 3 |
| 2023 | Shogun: A Task Scheduling Framework for Graph Mining AcceleratorsabstractGraph mining is an emerging application of great importance to big data analytic. Graph mining algorithms are bottle-necked by both computation complexity and memory access, hence necessitating specialized hardware accelerators to improve the processing efficiency. Current accelerators have extensively exploited task-level and fine-grained parallelism in these algorithms. However, their task scheduling still has room for optimization. They use either breadth-first search, depth-first search or a combination of both, leading to either poor intermediate data locality, low parallelism or inter-depth barriers. Jianfeng Zhu 0001, Wenrui Wei, Longlong Chen, Liang Wang 0020, Shaojun Wei, Leibo Liu |
ISCA | 2 |
| 2023 | CASA: An Energy-Efficient and High-Speed CAM-based SMEM Seeding Accelerator for Genome AlignmentabstractGenome analysis is a critical tool in medical and bioscience research, clinical diagnostics and treatment, and disease control and prevention. Seed and extension-based alignment is the main approach in the genome analysis pipeline, and BWA-MEM2, a widely acknowledged tool for genome alignment, performs seeding by searching for super maximal exact match (SMEM). The computation of SMEM searching requires high memory bandwidth and energy consumption, which becomes the main performance bottleneck in BWA-MEM2. State-of-the-Art designs like ERT and GenAx have achieved impressive speed-ups of SMEM-based genome alignment. However, they are constrained by frequent DRAM fetches or computationally intensive intersection calculations for all possible k-mers at every read position. Yi Huang 0036, Lingkun Kong, Dibei Chen, Zhiyu Chen 0003, Jianfeng Zhu 0001, Konstantinos Mamouras, Shaojun Wei, Kaiyuan Yang 0001, Leibo Liu |
MICRO | 6 |
| 2023 | QuickFPS: Architecture and Algorithm Co-Design for Farthest Point Sampling in Large-Scale Point CloudsabstractPoint clouds have been employed extensively in machine perception applications. Farthest point sampling (FPS) is a critical kernel for point cloud processing. With the rapid growth of point cloud scale, FPS introduces a large number of memory accesses, which become the bottleneck of the large-scale point cloud processing. In this article, we present QuickFPS, an architecture and algorithm co-design of FPS in large-scale point clouds. First, we systemically analyze the characteristics of FPS and put forward a bucket-based FPS algorithm. The algorithm introduces a two-level tree data structure to organize the large-scale point cloud into multiple buckets. By using two mechanisms named merged computation and implicit computation for the buckets, the external memory accesses and compute cost are significantly reduced. Then, we design an efficient domain-specific accelerator for FPS in large-scale point clouds. The accelerator takes advantage of different forms of parallelism and further improves the accelerator’s efficiency. Finally, we evaluate QuickFPS with several widely used point cloud datasets, which include small-scale and large-scale point clouds (up to 120 000 points). Overall, QuickFPS achieves performance speedups of$43.4\times$and$12.2\times$compared to GTX 1080Ti GPU and state-of-the-art point cloud accelerator PointAcc, respectively. Liang Wang 0020, Limin Xiao 0001, Hao Zhang 0205, Xiangrong Xu 0002, Jianfeng Zhu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2023 | M2STaR: A Multimode Spatio-Temporal Redundancy Design for Fault-Tolerant Coarse-Grained Reconfigurable ArchitecturesabstractCoarse-grained reconfigurable architectures (CGRAs) can provide both energy efficiency and performance for embedded systems, and thus they are increasingly deployed in the areas of aerospace, automotive engineering, and security where reliability is also a main criterion. However, the state-of-the-art fault-tolerant strategies for CGRAs apply either temporal or spatial scheme, including redundancy, periodic detection, workload balancing, and reconfiguration, failing to exploit the feature of dynamic and partial reconfiguration of CGRAs. Also, vulnerable judging circuits and inflexible mode shifting bottleneck the reliability design of fault-tolerant CGRAs. This article proposes a novel multimode fault-tolerant framework for CGRAs, which combines spatial-redundant data paths with temporal-redundant voters and thus reduces the vulnerable judging circuits while balancing the performance and reliability. This framework can also enable a changing reliability level at runtime via an online configuration transformation method based on precompiled patterns. Within the proposed framework, we systematically searched the design space spanning various combinations of the mainstream schemes with a Markov process model to compare the effectiveness and accordingly selected five points as available modes in our design after comprehensive consideration of fault tolerance and time overhead on CGRA. The framework is comprehensively evaluated on a cycle-accurate CGRA simulator, considering both permanent and transient faults. The experimental results show that the fault coverage rate of single transient faults or permanent faults has increased from 71.74% to 93.84%, which means the fault tolerance of the system has been increased by 31.03% compared with the state-of-the-art methods. There is also a great improvement in mean-time-to-failure (MTTF) and reconfiguration latency over baseline designs. Jianfeng Zhu 0001, Xingchen Man, Guihuan Song, Yi Huang 0036, Chenchen Deng, Pengfei Gou, Shouyi Yin, Shaojun Wei, Leibo Liu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | GEM: Ultra-Efficient Near-Memory Reconfigurable Acceleration for Read Mapping by Dividing and Predictive ScatteringabstractRead mapping, which maps billions of reads to a reference DNA, poses a significant performance bottleneck in genomic analysis. Current accelerators for read mapping are primarily bounded by the intensive and random memory access to huge datasets. Near-data processing (NDP) infrastructures are promising to provide extremely high bandwidth. However, existing frameworks failed to reach this potential due to poor locality and high redundancy. Our idea is to introduce prediction under the insight that candidate mapping positions become predictable when the reference is organized in coarse-grain slices. We present GEM (GenomicMemory), an ultra-efficient near-memory accelerator for read mapping. GEM adopts a novel data-centric framework, named dividing-and-predictive-scattering (DPS), which synthesizes information of seed existence to predict the target mapping locations to reduce memory access redundancy. During preparation, DPS divides the reference into coarse-grained slices and creates predictive filters to assess the likelihood of reads belonging to each slice. During mapping, DPS predicts and scatters reads to considerably fewer slices compared than without prediction. By employing small on-chip SRAM-based predictors with high accuracy, DPS minimizes unnecessary DRAM access and data movement from remote memory. In essence, DPS trades pre-seeding predictors for localized access patterns and low redundancy, hence achieving high throughput for data-intensive applications. We implement GEM by integrating coarse-grain reconfigurable architectures (CGRAs) in the logic layer of a 3D-stacked DRAM infrastructure, utilizing the massive banks as slices. GEM leverages CGRAs for their flexibility in supporting various algorithms tailored to different datasets. Bloom filters are leveraged for slice prediction, providing an error rate below 1%. Evaluation results demonstrate that GEM reduces memory requests by 95% and alignments by 87%, achieving a throughput improvement of 15.3× and 11.0× compared to compute-centric and broadcast-based baselines on the same NDP platform. Overall, GEM achieves a$3.5\times$throughput improvement and$2.1\times$energy efficiency compared to state-of-the-art ASIC accelerators. Longlong Chen, Jianfeng Zhu 0001, Guiqiang Peng, Mingxu Liu, Shaojun Wei, Leibo Liu |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2022 | Upward Packet Popup for Deadlock Freedom in Modular Chiplet-Based SystemsabstractMonolithic SoCs can be decomposed into disparate chiplets that support integration with advanced pack-aging technologies. This concept is promising in reducing the manufacturing cost of large scale SoCs due to the higher yield rate and reusability of chiplets. The chiplets should be designed in a modular manner without holistic system knowledge so that they can be reused in different SoCs. However, the design modularity is a major challenge to the networks-on-chip (NoCs) of chiplets.New deadlocks may occur across both the chiplets and the interposer due to the integration, even if the NoC of each individually designed chiplet is deadlock free. However, conventional deadlock freedom approaches are unsuitable to handle such deadlocks because they require holistic knowledge and violate the modularity. Although there are several modular approaches that specifically target at integration-induced deadlocks, their routing is overly restricted and the injection control incurs additional latency. They also lack flexibility in dynamically changing topologies due to their complex software algorithm and the hard-wired components.In this paper, a key insight on the chiplet integration-induced deadlocks is gained, inspired by which a deadlock recovery framework (named UPP) is proposed. Specifically, it is verified that an integration-induced deadlock always involves a stalled upward packet moving from the interposer to the connected chiplet via the vertical link. Thus, UPP detects a deadlock by discovering the upward packet and recovers the system from deadlock by transmitting the upward packet to its destination. Hybrid flow control mechanisms are proposed to enable the upward packet to bypass the buffers and be transmitted via the normal router datapath. To guarantee the ejection of the upward packet after transmission, a lightweight protocol is proposed to reserve ejection queue entries of the network interface. Experimental results show that while adhering to design modularity, UPP provides an average runtime speedup of 3.1%∼10.3% with an area overhead of less than 4%. Liang Wang 0020, Xiaohang Wang 0001, Jie Han 0001, Jianfeng Zhu 0001, Honglan Jiang, Shouyi Yin, Shaojun Wei, Leibo Liu |
HPCA | 5 |
| 2022 | CaSMap: agile mapper for reconfigurable spatial architectures by automatically clustering intermediate representations and scattering mapping processabstractToday, reconfigurable spatial architectures (RSAs) have sprung up as accelerators for compute- and data-intensive domains because they deliver energy and area efficiency close to ASICs and still retain sufficient programmability to keep the development cost low. The mapper, which is responsible for mapping algorithms onto RSAs, favors a systematic backtracking methodology because of high portability for evolving RSA designs. However, exponentially scaling compilation time has become the major obstacle. The key observation of this paper is that the key limiting factor to the systematic backtracking mappers is the waterfall mapping model which resolves all mapping variables and constraints at the same time using single-level intermediate representations (IRs). Xingchen Man, Jianfeng Zhu 0001, Guihuan Song, Shouyi Yin, Shaojun Wei, Leibo Liu |
ISCA | 2 |
| 2022 | An energy-efficient dynamically reconfigurable cryptographic engine with improved power/EM-side-channel-attack resistance
Chenchen Deng, Min Zhu 0001, Jinjiang Yang, Youyu Wu, Jiaji He 0001, Bohan Yang 0001, Jianfeng Zhu 0001, Shouyi Yin, Shaojun Wei, Leibo Liu |
Sci. China Inf. Sci. | 7 |
| 2022 | Dynamic-II Pipeline: Compiling Loops With Irregular Branches on Static-Scheduling CGRAabstractCoarse-grained reconfigurable architecture (CGRA) is a promising programmable hardware with high power-efficiency and high performance. However, compiling and optimizing loops with irregular branches on CGRAs is a challenge to fulfill the performance potential. Existing predication techniques, such as partial predication (PP) and full predication (FP), conservatively implement software pipeline with a static initiation interval (II) obtained from the maximum graph, and thus only parts of the graph in each loop iteration will be actually executed, resulting in underexploited performance. To exploit more loop-level parallelism for irregular branches, this article proposes a novel dynamic-II pipeline (DIP) scheme, which realizes a pipeline with variable II by accommodating multiple iterations of short path in one static configuration. Since the DIP scheme is effective to only certain types of branches, this article designs a hybrid compilation framework integrating other complementary methods, which selects the appropriate method for source programs according to a proposed performance evaluation model. Experimental results show that: 1) the hybrid compilation framework can effectively extract branch features, correctly choose and implement corresponding branch processing methods within acceptable compile time and 2) as compared to PP and FP, DIP brings a significant total execution time (TET) reduction by 27.21% and 22.04% on average when the execution probability of a short branch is 50%. Baofen Yuan, Jianfeng Zhu 0001, Xingchen Man, Zijiao Ma, Shouyi Yin, Shaojun Wei, Leibo Liu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | An Elastic Task Scheduling Scheme on Coarse-Grained Reconfigurable ArchitecturesabstractCoarse-grained reconfigurable architectures (CGRAs) are increasingly employed as domain-specific accelerators due to their efficiency and flexibility. A CGRA typically relies on compilers to perform task scheduling. The longstanding problem of static scheduling is that it suffers from insufficient parallelism in handling irregularities due to over-serialization and workload imbalance, which leads to severe resource underutilization and performance loss. To counteract the limitations of static scheduling in CGRAs, it is essential to exploit dynamic parallelism automatically and manage hardware resources adaptively. However, existing dynamic scheduling mechanisms, e.g., work stealing, often reschedule aggressively for instant performance but sacrifice efficiency, which is unfavorable to CGRAs that emphasize efficiency and fewer reconfigurations. This article proposes an elastic task scheduling scheme that enables lightweight dynamic scheduling in CGRAs. Tasks are rescheduled at runtime according to the classic tagged-token dataflow paradigm to enable dynamic task-level parallelism. Meanwhile, tasks are dynamically resized according to run-time throughputs via duplication, combination, and substitution operators for balanced multitask execution. We implement the elastic task scheduling scheme on a well-known reconfigurable architecture - triggered instruction architecture (TIA). Evaluation on the MachSuite benchmarks shows that the proposed scheme is effective in improving performance and energy efficiency. The average speedup is 2× over the baseline. Also, our design attains a 57 percent improvement in the area-normalized performance and a 49 percent better energy efficiency. Compared with a state-of-the-art dynamic scheduling method, our scheme achieves 1.6× speedup and 1.6× energy efficiency than work-stealing mechanism on the same substrate. Longlong Chen, Jianfeng Zhu 0001, Yangdong Deng, Zhaoshi Li, Xiaowei Jiang, Shouyi Yin, Shaojun Wei, Leibo Liu |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2020 | Pattern-Based Dynamic Compilation System for CGRAs With Online Configuration TransformationabstractPrevailing data-intensive applications, such as artificial intelligence and internet of things, demand considerable compute capability. Coarse-grained reconfigurable architectures (CGRAs) can meet this demand via providing abundant compute resources. However, compilation has become an essential problem because the increasing resources need to be orchestrated efficiently. Static compilation is insufficient due to conservative resource allocation and exponentially increasing time cost while state-of-the-art dynamic compilation still performs poorly in both generality and efficiency. This article proposes a dynamic compilation system for CGRAs through online pattern-based configuration transformation, which enables virtualization to improve resource utilization and flexibility. It utilizes statically-generated patterns to straightforwardly determine dynamic placement of registers and operations so that the transformation algorithm has a low complexity. Domain-specific features are extracted by a k-means clustering algorithm to help improve the quality of patterns. The experimental results show that statically compiled applications can be transformed onto arbitrary resources at runtime, reserving 73.5 (22.8-163.3 percent) of the original performance/resource on average, 9.1 (0-52.9 percent) better than the state-of-theart non-general methods. Leibo Liu, Xingchen Man, Jianfeng Zhu 0001, Shouyi Yin, Shaojun Wei |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2019 | A General Pattern-Based Dynamic Compilation Framework for Coarse-Grained Reconfigurable ArchitecturesabstractCompilation has become a major challenge to the usability of coarse-grained reconfigurable architectures as increasing programmable resources must be orchestrated. Static compilation is insufficient for prohibitive time cost while dynamic compilation still performs poorly in both generality and efficiency. This paper proposes a general pattern-based dynamic compilation framework, which utilizes statically-generated patterns to straightforwardly determine runtime re-placement and routing so that runtime configuration creation algorithm has low complexity. Domain-specific communication characteristics are harnessed to help improve the efficiency of patterns. The experimental results show that compiled general applications can be transformed onto arbitrary resources at runtime, reserving 97% (39%~163%) of the original performance/resource on average, 7% (0~17%) better than the state-of-the-art non-general methods. Xingchen Man, Leibo Liu, Jianfeng Zhu 0001, Shaojun Wei |
DAC | 3 |
| 2019 | Jintide®: A Hardware Security Enhanced Server CPU with Xeon® Cores under Runtime Surveillance by an In-Package Dynamically Reconfigurable ProcessorabstractThis article consists of a collection of slides from the author's conference presentation. Leibo Liu, Ao Luo, Guanhua Li, Jianfeng Zhu 0001, Gang Shan, Jianfeng Pan, Shouyi Yin, Shaojun Wei |
Hot Chips Symposium | 4 |
| 2016 | TLIA: Efficient Reconfigurable Architecture for Control-Intensive Kernels with Triggered-Long-InstructionsabstractCoarse-Grained Reconfigurable Architectures (CGRAs), which provide high performance, low power and flexibility, is viewed as a promising trend for computing. CGRAs are mostly employed to process compute-intensive kernels because of their inefficiency for control flows. Various methods have been proposed to alleviate this problem, and triggered instruction is one of the state-of-the-art techniques. In this paper, a reconfigurable architecture called Triggered-Long-Instruction Architecture (TLIA) is proposed to enhance the triggered instructions with parallel condition method. In the proposed architecture, triggered instruction set is employed on processing elements (PEs). In this way, over-serialized execution and branch instructions are both eliminated. In the meanwhile, each PE has an improved data-path with three ALUs which is inspired by the parallel condition method. In this way, the amount of parallelism inside each control flow is increased by paralleling predicate computations and predicated operations. Moreover, multiple triggered instructions, which may have internal control dependence, can be executed on PEs in parallel. The strategy of issuing instructions is implemented in hardware, and verified by FPGA. Experimental results show that the performance is improved by 20.9 to 140.0 percent, the area is reduced by 24.5 percent, and the power is reduced by 32.5 percent over the equivalent Triggered Instruction Architecture (TIA). Leibo Liu, Jianfeng Zhu 0001, Chenchen Deng, Shouyi Yin, Shaojun Wei |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2015 | Acceleration of control flows on reconfigurable architecture with a composite methodabstractControl-intensive kernels are becoming the bottleneck that limits the performance of Coarse-Grained Reconfigurable Architecture. Some methods, such as predicated execution, speculative execution, and dual-issue-single-execution, have been proposed to alleviate this problem. But they cannot be always efficient for various control flows. This paper proposes a new architecture, which combines the techniques of triggered instruction and parallel condition, in order to solve the problem completely. The architecture utilizes the basic framework of the triggered instruction to avoid over-serialized execution and branch instruction. Meanwhile, it takes the mechanism of the parallel condition to explore the parallelism between predicate and compute instructions without reconciliation operations. The mechanism of executing multiple instructions that have internal control dependence in parallel is discussed as well. The experiment result shows that the proposed architecture can achieve 20.9% to 140.0% higher performance than that of triggered instruction architecture in terms of cycle count. Leibo Liu, Jianfeng Zhu 0001, Shouyi Yin, Shaojun Wei |
DAC | 3 |
| 2015 | A Novel Composite Method to Accelerate Control Flow on Reconfigurable Architecture (Abstract Only)abstractReconfigurable Architecture provides a promising solution for embedded systems for high performance, low power and flexibility. Control dependence and control divergence are critical problems that impact the performance. Many methods were proposed to handle control flows efficiently, such as predicated execution and speculative execution. However, they exhibit different performances for different types of control flows, so composite methods are required to provide overall optimal performance. In this paper, a novel architecture is proposed which combines Triggered Instruction and parallel condition. It is designed on the basis of triggered instruction architecture (TIA) while each PE incorporates multiple arithmetic logic units with fast mutual control as in the technique of parallel condition. It can remove branch instructions as well as parallelize control and compute instructions without reconciliation operation, so it explores parallelism in branch level while avoids over-serialization execution in program-counter-based PE. The experiment was conducted on a model in C language and the result shows that the proposed architecture can achieve 80.0% higher performance on average than TIA. Leibo Liu, Jianfeng Zhu 0001, Shouyi Yin, Shaojun Wei |
FPGA | 3 |
| 2015 | A Hybrid Reconfigurable Architecture and Design Methods Aiming at Control-Intensive KernelsabstractWith the development of parallel computing, the compute-intensive part of an application could be accelerated so dramatically that the control intensive part, usually processed by a sequential processor, is becoming more and more critical in terms of performance and power consumption. To address this problem, this paper proposes a novel reconfigurable architecture to execute control-intensive kernels efficiently. The architecture applies three key design methods. The first one, parallel condition, exploits the instruction level parallelism of conditional branches with hardware design. The second one, configuration branch, enables the architecture to independently execute an entire application that has loops and other control flows. The third one, compound configuration, combines multiple configurations of low hardware utilization, which are common in sequential codes particularly, and thus reduces the reconfiguring times. Therefore, to offload control-intensive kernels onto the proposed architecture will speed up these workloads and boost the overall performance. The experiments were conducted on a benchmark that contains various branches, loops, and sequential codes. The results showed that the proposed architecture alone could implement the benchmark correctly. In addition, the proposed methods can improve performance by over 40% compared with the conventional techniques. The power efficiency is two orders larger than general purpose processors. Jianfeng Zhu 0001, Leibo Liu, Shouyi Yin, Shaojun Wei |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2014 | A Fast Application-Based Supply Voltage Optimization Method for Dual Voltage FPGAabstractDual supply voltage was a mature method to reduce the dynamic power of specific and programmable circuits, and the unsettled low voltage level (VL) was proved to have impact on its effect. In this paper, a circuit-level power model is developed to estimate the optimal VLfast for field-programmable gate array (FPGA). The model is mainly based on the path delay distribution of applications and the delay function of the integrated circuit technology. It can also count minor factors, such as path overlap, transition density, and capacitance. Experiment was conducted on a 90-nm FPGA model using MCNC benchmark. The results showed that the proposed method could generate near optimum VLfor most benchmarks. The best power reduction ratio is only 5.6% less than the gate-level heuristic method, which is relatively precise, but our method is ~100-10000 times faster. It implies that the dual voltage design with variable VL is a possible and promising low power method for field-programmable devices. Jianfeng Zhu 0001, Liyang Pan, Yaru Yan, Hu He 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2011 | A chip-level path-delay-distribution based Dual-VDD method for low power FPGA (abstract only)abstractDual-VDD FPGA architecture has been proposed to reduce the FPGA's power consumption, where a low VDD (VDDL) is assigned to non-critical resources and unused resources are power-gated. In this paper, a path-delay-distribution (PDD) based design method of supply voltage in dual-VDD FPGA is developed, which gives an estimated optimal VDD solution for the required applications. Meanwhile, an improved tree-based VDD assignment algorithm is accordingly designed. Thus chip-level optimization of dual-VDD FPGA is achieved on the chosen granularity with the power consumption minimized. Based on MCNC benchmark circuits at 90nm technology node, our experimental result shows that: the power reduction rate depends on VDDL level; the design method proposed in this work gives the optimal one automatically. This design method could be utilized to guide the FPGA automatic design, saving the time to search for the system's optimal supply voltage, and the proposed assignment algorithm is more efficient in dynamic power reduction. Jianfeng Zhu 0001, Yaru Yan, Hu He 0001, Liyang Pan |
FPGA | 1 |
| 2011 | A cost-efficient self-configurable BIST technique for testing multiplexer-based FPGA interconnect
Jianfeng Zhu 0001, Hu He 0001, Liyang Pan |
J. Electron. Test. | 1 |
| 2011 | Erratum to: A Cost-Efficient Self-Configurable BIST Technique for Testing Multiplexer-Based FPGA Interconnect
Jianfeng Zhu 0001, Hu He 0001, Liyang Pan |
J. Electron. Test. | 1 |