VLDB 2026 Research / reviewers in the wild / expert
Ziqing Zeng
dblp:258/6785
· DBLP profile ↗
8ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0002-6981-2299ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 1 first-author · 7 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Security and privacy · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Hybrid Ising FPGA-COBI Architecture with Hardware-Based Problem DecompositionabstractMany combinatorial optimization problems map naturally to Ising Hamiltonians, $H({\text{s}}) = - \sum\nolimits_{i,j} {{J_{ij}}} {s_i}{s_j} - \sum\nolimits_i {{h_i}} {s_i}$ , and CMOS ring-oscillator Ising machines solve them in microseconds at milliwatts [1] , [2] . Their key limitation is capacity : the number of spins one solver core can process in a single solve. Because each hardware spin represents one binary Ising variable, capacity directly sets the largest problem solvable in one shot. Our 28 nm five-core COBI chip solves a 45-spin all-to-all subproblem per core in 77.5 µ s, so larger instances require iterative decomposition. This shifts the bottleneck from analog solving to digital orchestration: a CPU-based decomposer needs ∼321 µ s/iter over PCIe, 4× the core solve time, leaving the solver idle 84.9% of the time. We instead co-locate an FPGA decomposer with the chip and derive sizing laws for the required parallelism, achieving 1.93× geomean speedup and > 40× energy reduction vs. an optimized C++ baseline. Ruihong Yin, Chaohui Li, Ahmet Efe, Abhimanyu Kumar, Ziqing Zeng, Ulya R. Karpuzcu, Sachin S. Sapatnekar, Chris H. Kim |
FCCM | 6 |
| 2026 | SATIC: An Optimizing Ising Compiler for SAT(isfiability)
Ahmet Efe, M. Hüsrev Cilasun, Abhimanyu Kumar, Nafisa Sadaf Prova, Ziqing Zeng, Tahmida Islam, Ruihong Yin, Chaohui Li, Peter Kreye, Chris H. Kim, Sachin S. Sapatnekar, Ulya R. Karpuzcu |
ISCA | 5 |
| 2024 | An Open-Source ML-Based Full-Stack Optimization Framework for Machine Learning AcceleratorsabstractParameterizable machine learning (ML) accelerators are the product of recent breakthroughs in ML. To fully enable their design space exploration (DSE), we propose a physical-design-driven, learning-based prediction framework for hardware-accelerated deep neural network (DNN) and non-DNN ML algorithms. It adopts a unified approach that combines power, performance, and area (PPA) analysis with frontend performance simulation, thereby achieving a realistic estimation of both backend PPA and system metrics such as runtime and energy. In addition, our framework includes a fully automated DSE technique, which optimizes backend and system metrics through an automated search of architectural and backend parameters. Experimental studies show that our approach consistently predicts backend PPA and system metrics with an average 7% or less prediction error for the ASIC implementation of two deep learning accelerator platforms, VTA and VeriGOOD-ML, in both a commercial 12 nm process and a research-oriented 45 nm process. Hadi Esmaeilzadeh, Soroush Ghodrati, Andrew B. Kahng, Joon Kyung Kim, Sean Kinzer, Sayak Kundu, Rohan Mahapatra, Susmita Dey Manasi, Sachin S. Sapatnekar, Zhiang Wang, Ziqing Zeng |
ACM Trans. Design Autom. Electr. Syst. | 11 |
| 2023 | Energy-efficient Hardware Acceleration of Shallow Machine Learning ApplicationsabstractML accelerators have largely focused on building general platforms for deep neural networks (DNNs), but less so on shallow machine learning (SML) algorithms. This paper proposes Axiline, a compact, configurable, template-based generator for SML hardware acceleration. Axiline identifies computational kernels as templates that are common to these algorithms and builds a pipelined accelerator for efficient execution. The dataflow graphs of individual ML instances, with different data dimensions, are mapped to the pipeline stages and then optimized by customized algorithms. The approach generates energy-efficient hardware for training and inference of various ML algorithms, as demonstrated with post-layout FPGA and ASIC results. Ziqing Zeng, Sachin S. Sapatnekar |
DATE | 1 |
| 2023 | A Multicore GNN Training AcceleratorabstractGraph neural networks (GNN) are vital for analytics on real-world problems with graph models. This work develops a multicore GNN training accelerator and develops multicore-specific optimizations for superior performance. It uses enhanced multicore-specific dynamic caching to circumvent the costs of irregular DRAM access patterns of graph-structured data. A novel feature vector segmentation approach is used to maximize on-chip data reuse with high on-chip computation per memory access, reducing data access latency, using a machine learning model for optimal performance. The work presents a major advance over prior FPGA/ASIC GNN accelerators by handling significantly larger datasets (with up to 8.6M vertices) on a variety of GNN models. On average, training speedup of 17× and energy efficiency improvement of 322× is achieved over DGL on a GPU; a speedup of 14× with 268× lower energy is shown over GPU-based GNNAdvisor; and 11× and 24× speedups are obtained over ASIC-based Rubik and FPGA-based GraphACT. Sudipta Mondal, Ramprasath Srinivasa Gopalakrishnan, Ziqing Zeng, Kishor Kunal, Sachin S. Sapatnekar |
ISLPED | 3 |
| 2023 | A Unified Engine for Accelerating GNN Weighting/Aggregation Operations, With Efficient Load Balancing and Graph-Specific CachingabstractGraph neural networks (GNNs) analysis engines are vital for real-world problems that use large graph models. Challenges for a GNN hardware platform include the ability to 1) host a variety of GNNs; 2) handle high sparsity in input vertex feature vectors and the graph adjacency matrix and the accompanying random memory access patterns; and 3) maintain load-balanced computation in the face of uneven workloads, induced by high sparsity and power-law vertex degree distributions. This article proposes GNNIE, an accelerator designed to run a broad range of GNNs. It tackles workload imbalance by 1) splitting vertex feature operands into blocks; 2) reordering and redistributing computations; and 3) using a novel flexible MAC architecture. It adopts a graph-specific, degree-aware caching policy that is well suited to real-world graph characteristics. The policy enhances on-chip data reuse and avoids random memory access to DRAM. GNNIE achieves average speedups of$7197\times $over a CPU and$17.81\times $over a GPU over multiple datasets on graph attention networks (GATs), graph convolutional networks (GCNs), GraphSAGE, GINConv, and DiffPool. Compared to prior approaches, GNNIE achieves an average speedup of$5\times $over HyGCN (which cannot implement GATs) for GCN, GraphSAGE, and GINConv. GNNIE achieves an average speedup of$1.3\times $over AWB-GCN (which runs only GCNs), despite using$3.4\times $fewer processing units. Sudipta Mondal, Susmita Dey Manasi, Kishor Kunal, Ramprasath Srinivasa Gopalakrishnan, Ziqing Zeng, Sachin S. Sapatnekar |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2021 | VeriGOOD-ML: An Open-Source Flow for Automated ML Hardware SynthesisabstractThis paper introduces VeriGOOD-ML, an automated methodology for generating Verilog with no human in the loop, starting from a high-level description of a machine learning (ML) algorithm in a standard format such as ONNX. The Verilog RTL is then translated through a back-end design flow to GDSII, driven by a design planning approach that is well tailored to the macro-intensive nature of ML platforms. VeriGOOD-ML uses three approaches to build ML hardware: the TABLA platform uses a dataflow architecture that is well suited to non-DNN ML algorithms; the GeneSys platform, with a systolic array and a SIMD array, is optimized for implementing DNNs; and the Axiline approach synthesizes small ML algorithms by hardcoding the structure of the algorithm into hardware, thus trading off flexibility for performance and power. The overall approach explores the design space of platform configurations and Pareto-optimal-PPA back-end implementations to yield designs that represent different tradeoffs at the algorithmic level between area, power, performance, and execution time. The overall methodology, from architecture to back-end design to hardware implementation, is described in this paper, and the results of VeriGOOD-ML are demonstrated on a set of ML benchmarks. Hadi Esmaeilzadeh, Soroush Ghodrati, Jie Gu 0003, Andrew B. Kahng, Joon Kyung Kim, Sean Kinzer, Rohan Mahapatra, Susmita Dey Manasi, Edwin Mascarenhas, Sachin S. Sapatnekar, Ravi Varadarajan, Zhiang Wang, Hanyang Xu 0002, Brahmendra Reddy Yatham, Ziqing Zeng |
ICCAD | 16 |
| 2020 | Homogeneous and Heterogeneous Feed-Forward XOR Physical Unclonable FunctionsabstractPhysical unclonable functions (PUFs) are hardware security primitives that are used for device authentication and cryptographic key generation. Standard XOR PUFs typically contain multiple standard arbiter PUFs as components, and are more secure than standard arbiter PUFs or feed-forward (FF) arbiter PUFs (FF PUFs). This paper proposes design of feed-forward XOR PUFs (FFXOR PUFs) where each component PUF is a FF PUF. Various homogeneous and heterogeneous FFXOR PUFs are presented and evaluated in terms of four fundamental properties of PUFs: uniqueness, attack-resistance, reliability and randomness. Certain key issues pertaining to XOR PUFs such as their vulnerability to machine learning attacks and instability in responses are investigated. Other important challenges like the lack of uniqueness in FF PUFs and the asymmetry in FPGA arbiter PUFs are addressed and it is shown that FFXOR PUFs can naturally overcome these problems. It is shown that heterogeneous FFXOR PUFs (i.e., FFXOR PUFs with non-identical components) can be resilient to state-of-the-art machine learning attacks. We also present systematic reliability analysis of FFXOR PUFs and demonstrate that soft-response thresholding can be used as an effective countermeasure to overcome the degraded reliability bottleneck. Observations from simulations are further verified through hardware implementation of 64-bit FFXOR PUFs on Xilinx Artix-7 FPGA. S. V. Sandeep Avvaru, Ziqing Zeng, Keshab K. Parhi |
IEEE Trans. Inf. Forensics Secur. | 2 |