VLDB 2026 Research / reviewers in the wild / expert
Zijun Jiang
dblp:354/8083
· DBLP profile ↗
8ranked-venue papers
4as first author
8since 2021 · last 2026
0009-0000-8832-9035ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 4 first-author · 8 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AutoVeriFix: Automatically Correcting Errors and Enhancing Functional Correctness in LLM-Generated Verilog CodeabstractLarge language models (LLMs) have demonstrated impressive capabilities in generating software code for high-level programming languages such as Python and C++. However, their application to hardware description languages, such as Verilog, is challenging due to the scarcity of high-quality training data. Current approaches to Verilog code generation using LLMs often focus on syntactic correctness, resulting in code with functional errors. To address these challenges, we present AutoVeriFix, a novel Python-assisted two-stage framework designed to enhance the functional correctness of LLM-generated Verilog code. In the first stage, LLMs are employed to generate high-level Python reference models that define the intended circuit behavior. In the second stage, these Python models facilitate the creation of automated tests that guide the generation of Verilog RTL implementations. Simulation discrepancies between the reference model and the Verilog code are iteratively used to identify and correct errors, thereby improving the functional accuracy and reliability of the LLM-generated Verilog code. Experimental results demonstrate that our approach significantly outperforms existing state-of-the-art methods in improving the functional correctness of generated Verilog code. Xiangchen Meng, Zijun Jiang, Yangdi Lyu |
ASP-DAC | 3 |
| 2025 | MACO: A HW-Mapping Co-optimization Framework for DNN AcceleratorsabstractDeep neural network (DNN) accelerators have been developed to enhance the effectiveness of DNN models, particularly in resource-constrained devices. Achieving high-throughput and energy-efficient inference within area constraints requires careful consideration of design choices in both hardware (HW) and mapping spaces. To get better performance and energy efficiency, it is important to optimize hardware and mapping together. However, this co-optimization process presents a considerable challenge due to the expansive combined HW-Mapping design space. To find the optimal configuration in this large design space, we formulate the exploration of the hardware configuration as an optimization problem, and embed the exploration of the mapping design space into the evaluation stage of optimization. We implement a HW-Mapping co-optimization framework called MACO to find optimal configurations for both hardware and mapping, and provide a generic interface to integrate different optimization algorithms, including multi-objective Bayesian optimization (MOBO), non-dominated sorting genetic algorithms (NSGA), and random search. We evaluate our framework with four popular DNN models of different properties. Our evaluation shows that the MOBO-based approach can achieve a 30% energy reduction and a 37% latency reduction with the same area as the state-of-the-art HW-Mapping optimization framework. Wujie Zhong, Zijun Jiang, Yangdi Lyu |
ASP-DAC | 2 |
| 2025 | An Enhanced Data Packing Method for General Matrix Multiplication in Brakerski/Fan-Vercauteren SchemeabstractGeneral Matrix-Matrix Multiplication (GEMM) stands as the most ubiquitous operation in machine learning applications. However, performing GEMM within Fully Homomorphic Encryption (FHE) is inefficient due to high computational demands and significant data migration constrained by limited bandwidth. Additionally, the inherent limitations of FHE schemes restrict the widespread application of machine learning, as standard activation functions are incompatible. This incompatibility necessitates alternative nonlinear functions, which lead to notable accuracy reductions. To address these challenges, we introduce a polynomial encoding methodology for GEMM under the Brakerski/Fan-Vercauteren (BFV) scheme and extend the method to inference with packing inputs and weights for different sizes. Furthermore, we design specialized hardware to accelerate the inference process through optimized scheduling between the hardware and the host system. In experiments, we implemented our hardware on an FPGA U250 platform. Compared to existing solutions, our method achieves superior performance, achieving the highest $4.22 \times$ and $3.99 \times$ speedups on MNIST and CIFAR-10. Xiangchen Meng, Zijun Jiang, Yangdi Lyu |
DAC | 3 |
| 2025 | CPP-SGS: Cycle-Accurate Power Prediction Framework via SNN and Genetic Signal SelectionabstractEffective power management is crucial for optimizing the performance and longevity of integrated circuits. Cycle-accurate power prediction can help power management during runtime. This paper introduces a Cycle-accurate Power Prediction framework via Spiking neural networks (SNNs) and Genetic signal Selection (CPP-SGS), which integrates SNNs and Genetic Algorithms (GAs) to predict real-time power consumption of chips. We apply GAs to select the most relevant signals as the input to SNNs to reduce the model size and inference time, making it well-suited for dynamic power estimation in real-time scenarios. The experimental results show that CCP-SGS outperforms the state-of-the-art approaches, with a normalized root mean squared error (NRMSE) of less than 1.6%. Zijun Jiang, Yangdi Lyu |
DATE | 2 |
| 2025 | MiCo: End-to-End Mixed Precision Neural Network Co-Exploration Framework for Edge AIabstractQuantized Neural Networks (QNN) with extremely low-bitwidth data have proven promising in efficient storage and computation on edge devices. To further reduce the accuracy drop while increasing speedup, layer-wise mixed-precision quantization (MPQ) becomes a popular solution. However, existing algorithms for exploring MPQ schemes are limited in flexibility and efficiency. Comprehending the complex impacts of different MPQ schemes on post-training quantization and quantization-aware training results is a challenge for conventional methods. Furthermore, an end-to-end framework for the optimization and deployment of MPQ models is missing in existing work.In this paper, we propose the MiCo framework, a holistic MPQ exploration and deployment framework for edge AI applications. The framework adopts a novel optimization algorithm to search for optimal quantization schemes with the highest accuracies while meeting latency constraints. Hardware-aware latency models are built for different hardware targets to enable fast explorations. After the exploration, the framework enables direct deployment from PyTorch MPQ models to bare-metal C codes, leading to end-to-end speedup with minimal accuracy drops. Zijun Jiang, Yangdi Lyu |
ICCAD | 1 |
| 2025 | BNRV: A Lightweight SIMD Extension for Efficient BitNet Inference on RISC-V CPUsabstractAI models utilizing extremely low-bitwidth weights have shown promise in efficient storage and computation while maintaining satisfactory results through proper training processes. By converting floating-point multiplication into simpler addition and shifting operations, these models are well-suited for deployment on resource-constrained devices. However, traditional CPU architectures often fail to fully exploit the advantages of multiplication-free operations due to a lack of hardware support. In this paper, we propose BNRV, a lightweight SIMD extension designed for the RISC-V Instruction Set Architecture (ISA) that specifically targets multiplication-free operations involving 8-bit data and low-bitwidth weights (ranging from 1 to 2 bits). We have also developed an accompanying library with optimized kernels for deploying various AI models using BNRV. Our proposed extension significantly accelerates low-bitwidth quantized multiplication (up to$10.95 \times$times faster) and lowbitwidth transformer model inference (up to$3.11 \times$faster), with minimal power overhead (less than 2%) and area overhead (less than 4%) compared to processors without BNRV support. Zijun Jiang, Yangdi Lyu |
ICCD | 1 |
| 2024 | Microprocessor Design Space Exploration via Space Partitioning and Bayesian OptimizationabstractDesign space exploration (DSE) has long been a very important topic in electronic design automation (EDA), but the growing diversity of applications and the complexity of integrated circuits make conventional DSE frameworks less effective and efficient. Therefore, an exploration algorithm that can find the optimal designs with fewer samples is demanded. This paper proposes a DSE framework for microprocessors that integrates a novel optimization algorithm with EDA flows. The proposed optimization algorithm utilizes space partitioning and Bayesian optimization to explore diverse and high-dimensional design spaces in microprocessors efficiently. Using the framework, we explore the design space of VexRiscv CPUs for TinyML workloads, where our proposed optimization algorithm obtains more Pareto-optimal designs and higher hypervolume with fewer samples. Zijun Jiang, Yangdi Lyu |
DATE | 1 |
| 2024 | Efficient Microprocessor Design Space Exploration via Space PartitioningabstractDesign space exploration (DSE) has long been a very important topic in electronic design automation (EDA). As the diversity of applications and the complexity of integrated circuits have grown rapidly in recent years, conventional DSE frameworks become less effective and efficient. This is due to the time-consuming nature of design point evaluation and the challenge of exploring high-dimensional design spaces. To address these issues, this paper proposes a DSE framework for microprocessors with a novel multi-objective optimization algorithm to find optimal designs with fewer samples. The proposed algorithm utilizes space partitioning and Bayesian optimization to efficiently explore high-dimensional design spaces in microprocessors. Zijun Jiang, Yangdi Lyu |
ICCD | 1 |