VLDB 2026 Research / reviewers in the wild / expert
Zheng Zhang 0005
dblp:181/2621-5
· DBLP profile ↗
44ranked-venue papers
12as first author
18since 2021 · last 2026
0000-0002-2292-0030ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 35 · 11 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Ultra Memory-Efficient On-FPGA Training of Transformers via Tensor-Compressed OptimizationabstractTransformer models have achieved state-of-the-art performance across a wide range of machine learning tasks. There is growing interest in training transformers on resource-constrained edge devices due to considerations such as privacy, domain adaptation, and on-device scientific machine learning. However, the significant computational and memory demands required for transformer training often exceed the capabilities of an edge device. Leveraging low-rank tensor compression, this paper presents the first on-FPGA accelerator for transformer training. On the algorithm side, we present a bi-directional contraction flow for tensorized transformer training, significantly reducing the computational FLOPS and intra-layer memory costs compared to existing tensor operations. On the hardware side, we store all highly compressed model parameters and gradient information on chip, creating an on-chip-memory-only framework for each stage in training. This reduces off-chip communication and minimizes latency and energy costs. Additionally, we implement custom computing kernels for each training stage and employ intra-layer parallelism and pipe-lining to further enhance run-time and memory efficiency. Through experiments on transformer models within 36.7 to 93.5 MB using FP-32 data formats on the ATIS dataset, our tensorized FPGA accelerator could conduct single-batch end-to-end training on the AMD Alevo U50 FPGA, with a memory budget of less than 6-MB BRAM and 22.5-MB URAM. Compared to uncompressed training on the NVIDIA RTX 3090 GPU, our on-FPGA training achieves a memory reduction of 30× to 51×. Our FPGA accelerator also achieves up to 4.0× less energy cost per epoch compared with tensor transformer training on an NVIDIA RTX 3090 GPU. As an initial result, this work highlights the significant potential of large-scale tensor training on edge devices. Jinming Lu, Hai Li 0008, Cong Hao, Ian A. Young, Zheng Zhang 0005 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2025 | Enhanced Operator Learning for Scalable and Ultra-fast Thermal Simulation in 3D-IC DesignabstractThermal simulation plays a critical role in the design of 3D integrated circuits (3D-ICs), where accurate and efficient temperature predictions are essential to ensure component reliability and performance. Recently, deep learning methods have shown great potential in accelerating these simulations. DeepOHeat [1] is one such approach, designed to learn solution operators that map single or multiple configurations of heat equations---such as surface power, volume power, or heat transfer coefficients---directly to the 3D temperature distribution. By utilizing a physics-informed DeepONet framework[2], DeepOHeat effectively captures complex relationships between design parameters and temperature fields, even when no training data is available and only PDE constraints are imposed. Xinling Yu, Ziyue Liu 0003, Hai Li 0008, Ian A. Young, Zheng Zhang 0005 |
ASP-DAC | 5 |
| 2025 | QuZO: Quantized Zeroth-Order Fine-Tuning for Large Language ModelsabstractJiajun Zhou, Yifan Yang, Kai Zhen, Ziyue Liu, Yequan Zhao, Ershad Banijamali, Athanasios Mouchtaris, Ngai Wong, Zheng Zhang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Jiajun Zhou 0004, Kai Zhen, Ziyue Liu 0003, Yequan Zhao, Seyed Ershad Banijamali, Athanasios Mouchtaris, Ngai Wong 0001, Zheng Zhang 0005 |
EMNLP | 9 |
| 2025 | Lagrange Coding for Tensor Network Contraction: Achieving Polynomial Recovery Thresholds
Kerong Wang, Zheng Zhang 0005 |
ISIT | 2 |
| 2025 | Poor Man's Training on MCUs: A Memory-Efficient Quantized Back-Propagation-Free ApproachabstractBack propagation (BP) is the default solution for gradient computation in neural network training. However, implementing BP-based training on various edge devices such as FPGA, microcontrollers (MCUs), and analog computing platforms faces multiple major challenges, such as the lack of hardware resources, long time-to-market, and dramatic errors in a low-precision setting. This article presents a simple BP-free training scheme on an MCU, which makes edge training hardware design as easy as inference hardware design. We adopt a quantized zeroth-order method to estimate the gradients of quantized model parameters, which can overcome the error of a straight-through estimator in a low-precision BP scheme. We further employ a few dimension reduction methods (e.g., node perturbation, sparse training) to improve the convergence of zeroth-order training. Experiment results show that our BP-free training achieves comparable performance as BP-based training on adapting a pre-trained image classifier to various corrupted data on resource-constrained edge devices (e.g., an MCU with 1024-KB SRAM for dense full-model training, or an MCU with 256-KB SRAM for sparse training). This method is most suitable for application scenarios where memory cost and time-to-market are the major concerns, but longer latency can be tolerated. Yequan Zhao, Hai Li 0008, Ian A. Young, Zheng Zhang 0005 |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2024 | AdaZeta: Adaptive Zeroth-Order Tensor-Train Adaption for Memory-Efficient Large Language Models Fine-TuningabstractFine-tuning large language models (LLMs) has achieved remarkable performance across various natural language processing tasks, yet it demands more and more memory as model sizes keep growing.To address this issue, the recently proposed Memory-efficient Zerothorder (MeZO) methods attempt to fine-tune LLMs using only forward passes, thereby avoiding the need for a backpropagation graph.However, significant performance drops and a high risk of divergence have limited their widespread adoption.In this paper, we propose the Adaptive Zeroth-order Tensor-Train Adaption (AdaZeta) framework, specifically designed to improve the performance and convergence of the ZO methods.To enhance dimension-dependent ZO estimation accuracy, we introduce a fast-forward, low-parameter tensorized adapter.To tackle the frequently observed divergence issue in large-scale ZO finetuning tasks, we propose an adaptive query number schedule that guarantees convergence.Detailed theoretical analysis and extensive experimental results on Roberta-Large and Llama-2-7B models substantiate the efficacy of our AdaZeta framework in terms of accuracy, memory efficiency, and convergence speed. 1 Kai Zhen, Seyed Ershad Banijamali, Athanasios Mouchtaris, Zheng Zhang 0005 |
EMNLP | 5 |
| 2024 | DeepZero: Scaling Up Zeroth-Order Optimization for Deep Model TrainingabstractZeroth-order (ZO) optimization has become a popular technique for solving machine learning (ML) problems when first-order (FO) information is difficult or impossible to obtain. However, the scalability of ZO optimization remains an open problem: Its use has primarily been limited to relatively small-scale ML problems, such as sample-wise adversarial attack generation. To our best knowledge, no prior work has demonstrated the effectiveness of ZO optimization in training deep neural networks (DNNs) without a significant decrease in performance. To overcome this roadblock, we develop DeepZero, a principled and practical ZO deep learning (DL) framework that can scale ZO optimization to DNN training from scratch through three primary innovations. First, we demonstrate the advantages of coordinate-wise gradient estimation (CGE) over randomized vector-wise gradient estimation in training accuracy and computational efficiency. Second, we propose a sparsity-induced ZO training protocol that extends the model pruning methodology using only finite differences to explore and exploit the sparse DL prior in CGE. Third, we develop the methods of feature reuse and forward parallelization to advance the practical implementations of ZO training. Our extensive experiments show that DeepZero achieves state-of-the-art (SOTA) accuracy on ResNet-20 trained on CIFAR-10, approaching FO training performance for the first time. Furthermore, we show the practical utility of DeepZero in applications of certified adversarial defense and DL-based partial differential equation error correction, achieving 10-20% improvement over SOTA. We believe our results will inspire future research on scalable ZO optimization and contribute to advancing deep learning. Aochuan Chen, Jinghan Jia, James Diffenderfer, Konstantinos Parasyris, Jiancheng Liu, Zheng Zhang 0005, Bhavya Kailkhura, Sijia Liu 0001 |
ICLR | 8 |
| 2024 | Coded Computing Meets Quantum Circuit Simulation: Coded Parallel Tensor Network Contraction AlgorithmabstractParallel tensor network contraction algorithms have emerged as the pivotal benchmarks for assessing the classical limits of computation, exemplified by Google's demonstration of quantum supremacy through random circuit sampling. However, the massive parallelization of the algorithm makes it vulnerable to computer node failures. In this work, we apply coded computing to a practical parallel tensor network contraction algorithm. To the best of our knowledge, this is the first attempt to code tensor network contractions. Inspired by matrix multiplication codes, we provide two coding schemes: 2-node code for practicality in quantum simulation and hyperedge code for generality. Our 2-node code successfully achieves significant gain for f-resilient number compared to naive replication, proportional to both the number of node failures and the dimension product of sliced indices. Our hyperedge code can cover tensor networks out of the scope of quantum, with degraded gain in the exchange of its generality. Zheng Zhang 0005, Sofía González-García, Haewon Jeong |
ISIT | 2 |
| 2024 | LoRETTA: Low-Rank Economic Tensor-Train Adaptation for Ultra-Low-Parameter Fine-Tuning of Large Language ModelsabstractYifan Yang, Jiajun Zhou, Ngai Wong, Zheng Zhang. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Jiajun Zhou 0004, Ngai Wong 0001, Zheng Zhang 0005 |
NAACL-HLT | 4 |
| 2023 | DeepOHeat: Operator Learning-based Ultra-fast Thermal Simulation in 3D-IC DesignabstractThermal issue is a major concern in 3D integrated circuit (IC) design. Thermal optimization of 3D IC often requires massive expensive PDE simulations. Neural network-based thermal prediction models can perform real-time prediction for many unseen new designs. However, existing works either solve 2D temperature fields only or do not generalize well to new designs with unseen design configurations (e.g., heat sources and boundary conditions). In this paper, for the first time, we propose DeepOHeat, a physics-aware operator learning framework to predict the temperature field of a family of heat equations with multiple parametric or non-parametric design configurations. This framework learns a functional map from the function space of multiple key PDE configurations (e.g., boundary conditions, power maps, heat transfer coefficients) to the function space of the corresponding solution (i.e., temperature fields), enabling fast thermal analysis and optimization by changing key design configurations (rather than just some parameters). We test DeepOHeat on some industrial design cases and compare it against Celsius 3D from Cadence Design Systems. Our results show that, for the unseen testing cases, a well-trained DeepOHeat can produce accurate results with 1000× to 300000× speedup. Ziyue Liu 0003, Yixing Li, Xinling Yu, Shinyu Shiau, Xin Ai 0007, Zhiyu Zeng, Zheng Zhang 0005 |
DAC | 8 |
| 2023 | Distributionally Robust Circuit Design Optimization under Variation ShiftsabstractDue to the significant process variations, designers have to optimize the statistical performance distribution of nano-scale IC design in most cases. This problem has been investigated for decades under the formulation of stochastic optimization, which minimizes the expected value of a performance metric while assuming that the distribution of process variation is exactly given. This paper rethinks the variation-aware circuit design optimization from a new perspective. First, we discuss the variation shift problem, which means that the actual density function of process variations almost always differs from the given model and is often unknown. Consequently, we propose to formulate the variation-aware circuit design optimization as a distributionally robust optimization problem, which does not require the exact distribution of process variations. By selecting an appropriate uncertainty set for the probability density function of process variations, we solve the shift-aware circuit optimization problem using distributionally robust Bayesian optimization. This method is validated with both a photonic IC and an electronics IC. Our optimized circuits show excellent robustness against variation shifts: the optimized circuit has excellent performance under many possible distributions of process variations that differ from the given statistical model. This work has the potential to enable a new research direction and inspire subsequent research at different levels of the EDA flow under the setting of variation shift. Yifan Pan, Zichang He, Nanlin Guo, Zheng Zhang 0005 |
ICCAD | 4 |
| 2022 | Evolutionary Tensor Train Decomposition for Hyper-Spectral Remote Sensing ImagesabstractHyper-spectral images are widely used for mapping and remote sensing of the Earth's surface. Different tensor decomposition methods have been applied for hyper-spectral image decomposition. In this study we present an evolutionary tensor train (ETT) decomposition. The ETT technique defines a combinatorial optimization model to find an optimal shape for the tensor train (TT) decomposition. The optimization model maximizes the compression ratio of the TT decomposition given an error bound. A genetic algorithm (GA) linked with the TT-SVD algorithm is applied to find the optimal shape. We adopt the ETT for the decomposition of hyper-spectral images and study the performance of the ETT with respect to the error bound. The results demonstrate the effectiveness of the proposed evolutionary tensor shape search for the the decomposition of the hyper-spectral images. Ryan Solgi, Hugo A. Loáiciga, Zheng Zhang 0005 |
IGARSS | 3 |
| 2022 | Efficient Processing of Sparse Tensor Decomposition via Unified Abstraction and PE-Interactive ArchitectureabstractWe propose a novel architecture to efficiently perform sparse tensor decomposition/completion. As the generalization of vectors and matrices, tensors are widely used to process high-dimensional data. Sparse tensor decomposition (SpTD) is not only an emerging tensor analysis technique but also an effective tool to reduce the storage and computation costs of tensors. However, conventional general-purpose processors are inefficient to perform SpTD, mainly due to: i) variable sparsity degree and flexible buffer size requirement; ii) difficulties of fusing multiple execution kernels to pursue better performance. For domain-specific accelerator designers on the other hand, the diversity of decomposition algorithms is also an important problem that must be considered. To solve these challenges, we propose a unified abstraction for SpTD algorithms and design a specialized accelerator. First, we formulate two types of core kernels (SpLrMM and LrSampling) that serve as a standard form to fit a broad range of SpTD algorithms. Second, we design a sparse tensor engine (STE) to efficiently perform SpTD. STE uses a processing element (PE)-interactive architecture where PEs can be flexibly grouped together via Network-on-Chip (NoC) to share the buffer capacity, bandwidth, and compute resources. We evaluate our accelerator with extensive experiments, and it can achieve an average speedup of 45× over CPU and 29× over GPU. Bangyan Wang, Lei Deng 0003, Zheng Qu 0002, Shuangchen Li, Zheng Zhang 0005, Yuan Xie 0001 |
IEEE Trans. Computers | 5 |
| 2022 | PoBO: A Polynomial Bounding Method for Chance-Constrained Yield-Aware Optimization of Photonic ICsabstractConventional yield optimization algorithms try to maximize the success rate of a circuit under process variations. These methods often obtain a high yield but reach a design performance that is far from the optimal value. This article investigates an alternative yield-aware optimization for photonic ICs: we will optimize the circuit design performance while ensuring a high yield requirement. This problem was recently formulated as a chance-constrained optimization, and the chance constraint was converted to a stronger constraint with statistical moments. Such a conversion reduces the feasible set and sometimes leads to an over-conservative design. To address this fundamental challenge, this article proposes a carefully designed polynomial function, called optimal polynomial kinship function, to bound the chance constraint more accurately. We modify existing kinship functions via relaxing the independence and convexity requirements, which fits our more general uncertainty modeling and tightens the bounding functions. The proposed method enables a global optimum search for the design variables via polynomial optimization. We validate this method with a synthetic function and two photonic IC design benchmarks, showing that our method can obtain better design performance while meeting a prespecified yield requirement. Many other advanced problems of yield-aware optimization and more general safety-critical design/control can be solved based on this work in the future. Zichang He, Zheng Zhang 0005 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2022 | Hardware-Enabled Efficient Data Processing With Tensor-Train DecompositionabstractIn recent years, tensor computation has become a promising tool for solving big data analysis, machine learning, medical image, and EDA problems. To ease the memory and computation intensity of tensor processing, decomposition techniques, especially tensor-train decomposition (TTD), are widely adopted to compress the extremely high-dimensional tensor data. Despite TTD’s potential to break the curse of dimensionality, researchers have not yet leveraged its full computational potential, mainly because of two reasons: 1) executing TTD itself is time- and energy-consuming due to the singular value decomposition (SVD) operation inside each of TTD’s iteration and 2) additional software/hardware optimizations are often required to process the obtained TT-format data in certain applications such as deep learning inference. In this article, we address these challenges with two approaches. First, we propose an algorithm-hardware co-design with customized architecture, namely, TTD Engine to accelerate TTD. We use MRI image compression as a demo application to illustrate the efficacy of the proposed accelerator. Second, we present a case study demonstrating the benefit of TT-format data processing and the efficacy of using TTD Engine. In the case study, we use the TT approach to realize convolution operation, which is difficult and nontrivial for TT-format data. Experimental results show that, TTD Engine achieves, on average,$14.9 \times $–$36.9 \times $speedup over CPU implementations and$4.1\times $–$9.9\times $speedup compared to the GPU baseline. The energy efficiency is also improved by at least$14.4\times $and$5.4\times $over CPU and GPU, respectively. Moreover, our hardware-enabled TT-format data processing further leads to more efficient implementations of complicated operations and applications. Zheng Qu 0002, Lei Deng 0003, Bangyan Wang, Hengnu Chen, Jilan Lin, Ling Liang 0003, Guoqi Li 0002, Zheng Zhang 0005, Yuan Xie 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2021 | 3U-EdgeAI: Ultra-Low Memory Training, Ultra-Low Bitwidth Quantization, and Ultra-Low Latency AccelerationabstractThe deep neural network (DNN) based AI applications on the edge require both low-cost computing platforms and high-quality services. However, the limited memory, computing resources, and power budget of the edge devices constrain the effectiveness of the DNN algorithms. Developing edge-oriented AI algorithms and implementations (e.g., accelerators) is challenging. In this paper, we summarize our recent efforts for efficient on-device AI development from three aspects, including both training and inference. First, we present on-device training with ultra-low memory usage. We propose a novel rank-adaptive tensor-based tensorized neural network model, which offers orders-of-magnitude memory reduction during training. Second, we introduce an ultra-low bitwidth quantization method for DNN model compression, achieving the state-of-the-art accuracy under the same compression ratio. Third, we introduce an ultra-low latency DNN accelerator design, practicing the software/hardware co-design methodology. This paper emphasizes the importance and efficacy of training, quantization and accelerator design, and calls for more research breakthroughs in the area for AI on the edge. Yao Chen 0008, Cole Hawkins, Kaiqi Zhang 0002, Zheng Zhang 0005, Cong Hao |
ACM Great Lakes Symposium on VLSI | 4 |
| 2021 | Bayesian tensorized neural networks with automatic rank selection
Cole Hawkins, Zheng Zhang 0005 |
Neurocomputing | 2 |
| 2021 | Sparse Tucker Tensor Decomposition on a Hybrid FPGA-CPU PlatformabstractRecommendation systems, social network analysis, medical imaging, and data mining often involve processing sparse high-dimensional data. Such high-dimensional data are naturally represented as tensors, and they cannot be efficiently processed by conventional matrix or vector computations. Sparse Tucker decomposition is an important algorithm for compressing and analyzing these sparse high-dimensional datasets. When energy efficiency and data privacy are major concerns, hardware accelerators on resource-constraint platforms become crucial for the deployment of tensor algorithms. In this work, we propose a hybrid computing framework containing CPU and FPGA to accelerate sparse Tucker factorization. This algorithm has three main modules: 1) tensor-times-matrix (TTM); 2) Kronecker products; and 3) QR decomposition with column pivoting (QRP). In addition, we accelerate the former two modules on a Xilinx FPGA and the latter one on a CPU. Our hybrid platform achieves$23.6 \times \sim 1091\times $speedup and over 93.519% ~ 99.514% energy savings compared with CPU on the synthetic and real-world datasets. Weiyun Jiang, Kaiqi Zhang 0002, Colin Yu Lin, Feng Xing, Zheng Zhang 0005 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2020 | High-Dimensional Uncertainty Quantification of Electronic and Photonic IC With Non-Gaussian Correlated Process VariationsabstractUncertainty quantification based on generalized polynomial chaos has been used in many applications. It has also achieved great success in variation-aware design automation. However, almost all existing techniques assume that the parameters are mutually independent or Gaussian correlated, which is rarely true in real applications. For instance, in chip manufacturing, many process variations are actually correlated. Recently, some techniques have been developed to handle non-Gaussian correlated random parameters, but they are time-consuming for high-dimensional problems. We present a new framework to solve uncertainty quantification problems with many non-Gaussian correlated uncertainties. First, we propose a set of smooth basis functions to well capture the impact of non-Gaussian correlated process variations. We develop a tensor approach to compute these basis functions in a high-dimension setting. Second, we investigate the theoretical aspect and practical implementation of a sparse solver to compute the coefficients of all basis functions. We provide some theoretical analysis for the exact recovery condition and error bound of this sparse solver in the context of uncertainty quantification. We present three adaptive sampling approaches to improve the performance of the sparse solver. Finally, we validate our methods by synthetic and practical electronic/photonic ICs with 19 to 57 non-Gaussian correlated variation parameters. Our approach outperforms Monte Carlo by thousands of times in terms of efficiency. It can also accurately predict the output density functions with multiple peaks caused by non-Gaussian correlations, which are hard to capture by existing methods. Chunfeng Cui, Zheng Zhang 0005 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2020 | Chance-Constrained and Yield-Aware Optimization of Photonic ICs With Non-Gaussian Correlated Process VariationsabstractUncertainty quantification has become an efficient tool for uncertainty-aware prediction, but its power in yield-aware optimization has not been well explored from either theoretical or application perspectives. Yield optimization is a much more challenging task. On the one side, optimizing the generally nonconvex probability measure of performance metrics is difficult. On the other side, evaluating the probability measure in each optimization iteration requires massive simulation data, especially, when the process variations are non-Gaussian correlated. This article proposes a data-efficient framework for the yield-aware optimization of photonic ICs. This framework optimizes the design performance with a yield guarantee, and it consists of two modules: 1) a modeling module that builds stochastic surrogate models for design objectives and chance constraints with a few simulation samples and 2) a novel yield optimization module that handles probabilistic objectives and chance constraints in an efficient deterministic way. This deterministic treatment avoids repeatedly evaluating probability measures at each iteration, thus it only requires a few simulations in the whole optimization flow. We validate the accuracy and efficiency of the whole framework by a synthetic example and two photonic ICs. Our optimization method can achieve more than$30\times $reduction of simulation cost and better design performance on the test cases compared with a Bayesian yield optimization approach developed recently. Chunfeng Cui, Zheng Zhang 0005 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2020 | Prediction of Multidimensional Spatial Variation Data via Bayesian Tensor CompletionabstractThis paper presents a multidimensional computational method to predict the spatial variation data inside and across multiple dies of a wafer. This technique is based on tensor computation. A tensor is a high-dimensional generalization of a matrix or a vector. By exploiting the hidden low-rank property of a high-dimensional data array, the large amount of unknown variation testing data may be predicted from a few random measurement samples. The tensor rank, which decides the complexity of a tensor representation, is decided by an available variational Bayesian approach. Our approach is validated by a practical chip testing data set, and it can be easily generalized to characterize the process variations of multiple wafers. Our approach is more efficient than the previous virtual probe techniques in terms of memory and computational cost when handling high-dimensional chip testing data. Jiali Luan, Zheng Zhang 0005 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | Tensor Methods for Generating Compact Uncertainty Quantification and Deep Learning ModelsabstractTensor methods have become a promising tool to solve high-dimensional problems in the big data era. By exploiting possible low-rank tensor factorization, many high-dimensional model-based or data-driven problems can be solved to facilitate decision making or machine learning. In this paper, we summarize the recent applications of tensor computation in obtaining compact models for uncertainty quantification and deep learning. In uncertainty analysis where obtaining data samples is expensive, we show how tensor methods can significantly reduce the simulation or measurement cost. To enable the deployment of deep learning on resource-constrained hardware platforms, tensor methods can be used to significantly compress an over-parameterized neural network model or directly train a small-size model from scratch via optimization or statistical techniques. Recent Bayesian tensorized neural networks can automatically determine their tensor ranks in the training process. Chunfeng Cui, Cole Hawkins, Zheng Zhang 0005 |
ICCAD | 3 |
| 2019 | Efficient Uncertainty Modeling for System Design via Mixed Integer ProgrammingabstractThe post-Moore era casts a shadow of uncertainty on many aspects of computer system design. Managing that uncertainty requires new algorithmic tools to make quantitative assessments. While prior uncertainty quantification methods, such as generalized polynomial chaos (gPC), show how to work precisely under the uncertainty inherent to physical devices, these approaches focus solely on variables from a continuous domain. However, as one moves up the system stack to the architecture level many parameters are constrained to a discrete (integer) domain. This paper proposes an efficient and accurate uncertainty modeling technique, named mixed generalized polynomial chaos (M-gPC), for architectural uncertainty analysis. The M-gPC technique extends the generalized polynomial chaos (gPC) theory originally developed in the uncertainty quantification community, such that it can efficiently handle the mixed-type (i.e., both continuous and discrete) uncertainties in computer architecture design. Specifically, we employ some stochastic basis functions to capture the architecture-level impact caused by uncertain parameters in a simulator. We also develop a novel mixed-integer programming method to select a small number of uncertain parameter samples for detailed simulations. With a few highly informative simulation samples, an accurate surrogate model is constructed in place of cycle-level simulators for various architectural uncertainty analysis. In the chip-multiprocessor (CMP) model, we are able to estimate the propagated uncertainties with only 95 samples whereas Monte Carlo requires$5\times 10^{4}$samples to achieve the similar accuracy. We also demonstrate the efficiency and effectiveness of our method on a detailed DRAM subsystem. Zichang He, Weilong Cui, Chunfeng Cui, Timothy Sherwood, Zheng Zhang 0005 |
ICCAD | 5 |
| 2019 | Tucker Tensor Decomposition on FPGAabstractTensor computation has emerged as a powerful mathematical tool for solving high-dimensional and/or extreme-scale problems in science and engineering. The last decade has witnessed tremendous advancement of tensor computation and its applications in machine learning and big data. However, its hardware optimization on resource-constrained devices remains an (almost) unexplored field. This paper presents an hardware accelerator for a classical tensor computation framework, Tucker decomposition. We study three modules of this architecture: tensor-times-matrix (TTM), matrix singular value decomposition (SVD), and tensor permutation, and implemented them on Xilinx FPGA for prototyping. In order to further reduce the computing time, a warm-start algorithm for the Jacobi iterations in SVD is proposed. A fixed-point simulator is used to evaluate the performance of our design. Some synthetic data sets and a real MRI data set are used to validate the design and evaluate its performance. We compare our work with state-of-the-art software toolboxes running on both CPU and GPU, and our work shows 2.16 – 30.2× speedup on the cardiac MRI data set. Kaiqi Zhang 0002, Xiyuan Zhang 0003, Zheng Zhang 0005 |
ICCAD | 3 |
| 2019 | Wafer Pattern Recognition Using Tucker DecompositionabstractIn production test data analytics, it is often that an analysis involves the recognition of a conceptual pattern on a wafer map. A wafer pattern may hint a particular issue in the production by itself or guide the analysis into a certain direction. In this work, we introduce a novel approach to recognize patterns on a wafer map of pass/fail locations. Our approach utilizes Tucker decomposition to find projection matrices that are able to project a wafer pattern represented by a small set of training samples into a nearly-diagonal matrix. Properties of such a matrix are utilized to recognize wafers with a similar pattern. Also included in our approach is a novel method to select the wafer samples that are more suitable to be used together to represent a conceptual pattern in view of the proposed approach. Ahmed Wahba, Li-C. Wang, Zheng Zhang 0005, Nik Sumikawa |
VTS | 3 |
| 2018 | Uncertainty quantification of electronic and photonic ICs with non-Gaussian correlated process variationsabstractSince the invention of generalized polynomial chaos in 2002, uncertainty quantification has impacted many engineering fields, including variation-aware design automation of integrated circuits and integrated photonics. Due to the fast convergence rate, the generalized polynomial chaos expansion has achieved orders-of-magnitude speedup than Monte Carlo in many applications. However, almost all existing generalized polynomial chaos methods have a strong assumption: the uncertain parameters are mutually independent or Gaussian correlated. This assumption rarely holds in many realistic applications, and it has been a long-standing challenge for both theorists and practitioners. Chunfeng Cui, Zheng Zhang 0005 |
ICCAD | 2 |
| 2018 | Variational Bayesian Inference for Robust Streaming Tensor Factorization and CompletionabstractStreaming tensor factorization is a powerful tool for processing high-volume and multi-way temporal data in Internet networks, recommender systems and image/video data analysis. Existing streaming tensor factorization algorithms rely on least-squares data fitting and they do not possess a mechanism for tensor rank determination. This leaves them susceptible to outliers and vulnerable to over-fitting. This paper presents a Bayesian robust streaming tensor factorization model to identify sparse outliers, automatically determine the underlying tensor rank and accurately fit low-rank structure. We implement our model in Matlab and compare it with existing algorithms on tensor datasets generated from dynamic MRI and Internet traffic. Zheng Zhang 0005, Cole Hawkins |
ICDM | 1 |
| 2018 | Variation-Aware Modeling of Integrated Capacitors Based on Floating Random Walk ExtractionabstractThis paper presents an effective approach to variability analysis of integrated capacitors due to manufacturing process uncertainty. The proposed approach combines the generalized polynomial chaos method for uncertainty quantification with the efficient floating random walk algorithm for capacitance extraction. For applications where detailed statistical descriptions are required, the method allows achieving a 1000× acceleration compared to standard Monte Carlo analysis. Application to variability analysis in digital to analog converters is illustrated. Paolo Maffezzoni, Zheng Zhang 0005, Salvatore Levantino, Luca Daniel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2017 | Tensor Computation: A New Framework for High-Dimensional Problems in EDAabstractMany critical electronic design automation (EDA) problems suffer from the curse of dimensionality, i.e., the very fast-scaling computational burden produced by large number of parameters and/or unknown variables. This phenomenon may be caused by multiple spatial or temporal factors (e.g., 3-D field solvers discretizations and multirate circuit simulation), nonlinearity of devices and circuits, large number of design or optimization parameters (e.g., full-chip routing/placement and circuit sizing), or extensive process variations (e.g., variability /reliability analysis and design for manufacturability). The computational challenges generated by such high-dimensional problems are generally hard to handle efficiently with traditional EDA core algorithms that are based on matrix and vector computation. This paper presents “tensor computation” as an alternative general framework for the development of efficient EDA algorithms and tools. A tensor is a high-dimensional generalization of a matrix and a vector, and is a natural choice for both storing and solving efficiently high-dimensional EDA problems. This paper gives a basic tutorial on tensors, demonstrates some recent examples of EDA applications (e.g., nonlinear circuit modeling and high-dimensional uncertainty quantification), and suggests further open EDA problems where the use of tensor computation could be of advantage. Zheng Zhang 0005, Kim Batselier, Luca Daniel, Ngai Wong 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2015 | Analysis and Design of Weakly Coupled LC Oscillator Arrays Based on Phase-Domain MacromodelsabstractAn array of weakly coupled oscillators can generate multiphase signals, i.e., multiple sinusoidal signals with specific phase separations. Multiphase oscillators are attractive solutions in many electronic applications such as the synchronization of multiple processing units in digital electronics and the frequency synthesis in mixed-signal radio frequency circuits. Due to the complexity of multiphase oscillators and the large number of design parameters, novel simulation techniques are highly desired to efficiently handle such large-scale problems. In this paper, an efficient phase-domain simulation technique is proposed to calculate the phase response of inductance capacitance oscillator array. By some practical examples, it is shown how the proposed method can be exploited to identify the array topologies and parameter settings that guarantee stable phase separations. It is also shown how the proposed technique can be used to evaluate phase-noise performance. Paolo Maffezzoni, Bichoy Bahr, Zheng Zhang 0005, Luca Daniel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2015 | Enabling High-Dimensional Hierarchical Uncertainty Quantification by ANOVA and Tensor-Train DecompositionabstractHierarchical uncertainty quantification can reduce the computational cost of stochastic circuit simulation by employing spectral methods at different levels. This paper presents an efficient framework to simulate hierarchically some challenging stochastic circuits/systems that include high-dimensional subsystems. Due to the high parameter dimensionality, it is challenging to both extract surrogate models at the low level of the design hierarchy and to handle them in the high-level simulation. In this paper, we develop an efficient analysis of variance-based stochastic circuit/microelectromechanical systems simulator to efficiently extract the surrogate models at the low level. In order to avoid the curse of dimensionality, we employ tensor-train decomposition at the high level to construct the basis functions and Gauss quadrature points. As a demonstration, we verify our algorithm on a stochastic oscillator with four MEMS capacitors and 184 random parameters. This challenging example is efficiently simulated by our simulator at the cost of only 10min in MATLAB on a regular personal computer. Zheng Zhang 0005, Xiu Yang, Ivan V. Oseledets, George Em Karniadakis, Luca Daniel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2014 | Calculation of Generalized Polynomial-Chaos Basis Functions and Gauss Quadrature Rules in Hierarchical Uncertainty QuantificationabstractStochastic spectral methods are efficient techniques for uncertainty quantification. Recently they have shown excellent performance in the statistical analysis of integrated circuits. In stochastic spectral methods, one needs to determine a set of orthonormal polynomials and a proper numerical quadrature rule. The former are used as the basis functions in a generalized polynomial chaos expansion. The latter is used to compute the integrals involved in stochastic spectral methods. Obtaining such information requires knowing the density function of the random input a-priori. However, individual system components are often described by surrogate models rather than density functions. In order to apply stochastic spectral methods in hierarchical uncertainty quantification, we first propose to construct physically consistent closed-form density functions by two monotone interpolation schemes. Then, by exploiting the special forms of the obtained density functions, we determine the generalized polynomial-chaos basis functions and the Gauss quadrature rules that are required by a stochastic spectral simulator. The effectiveness of our proposed algorithm is verified by both synthetic and practical circuit examples. Zheng Zhang 0005, Tarek A. El-Moselhy, Ibrahim M. Elfadel, Luca Daniel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2013 | Uncertainty quantification for integrated circuits: stochastic spectral methodsabstractDue to significant manufacturing process variations, the performance of integrated circuits (ICs) has become increasingly uncertain. Such uncertainties must be carefully quantified with efficient stochastic circuit simulators. This paper discusses the recent advances of stochastic spectral circuit simulators based on generalized polynomial chaos (gPC). Such techniques can handle both Gaussian and non-Gaussian random parameters, showing remarkable speedup over Monte Carlo for circuits with a small or medium number of parameters. We focus on the recently developed stochastic testing and the application of conventional stochastic Galerkin and stochastic collocation schemes to nonlinear circuit problems. The uncertainty quantification algorithms for static, transient and periodic steady-state simulations are presented along with some practical simulation results. Some open problems in this field are discussed. Zheng Zhang 0005, Ibrahim M. Elfadel, Luca Daniel |
ICCAD | 1 |
| 2013 | Stochastic Testing Method for Transistor-Level Uncertainty Quantification Based on Generalized Polynomial ChaosabstractUncertainties have become a major concern in integrated circuit design. In order to avoid the huge number of repeated simulations in conventional Monte Carlo flows, this paper presents an intrusive spectral simulator for statistical circuit analysis. Our simulator employs the recently developed generalized polynomial chaos expansion to perform uncertainty quantification of nonlinear transistor circuits with both Gaussian and non-Gaussian random parameters. We modify the nonintrusive stochastic collocation (SC) method and develop an intrusive variant called stochastic testing (ST) method. Compared with the popular intrusive stochastic Galerkin (SG) method, the coupled deterministic equations resulting from our proposed ST method can be solved in a decoupled manner at each time point. At the same time, ST requires fewer samples and allows more flexible time step size controls than directly using a nonintrusive SC solver. These two properties make ST more efficient than SG and than existing SC methods, and more suitable for time-domain circuit simulation. Simulation results of several digital, analog and RF circuits are reported. Since our algorithm is based on generic mathematical models, the proposed ST algorithm can be applied to many other engineering problems. Zheng Zhang 0005, Tarek A. El-Moselhy, Ibrahim M. Elfadel, Luca Daniel |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2012 | Passivity Enforcement for Descriptor Systems Via Matrix Pencil PerturbationabstractPassivity is an important property of circuits and systems to guarantee stable global simulation. Nonetheless, nonpassive models may result from passive underlying structures due to numerical or measurement error/inaccuracy. A postprocessing passivity enforcement algorithm is therefore desirable to perturb the model to be passive under a controlled error. However, previous literature only reports such passivity enforcement algorithms for pole-residue models and regular systems (RSs). In this paper, passivity enforcement algorithms for descriptor systems (DSs, a superset of RSs) with possibly singular direct term (specifically,D+DTorI-DDT) are proposed. The proposed algorithms cover all kinds of state-space models (RSs or DSs, with direct terms being singular or nonsingular, in the immittance or scattering representation) and thus have a much wider application scope than existing algorithms. The passivity enforcement is reduced to two standard optimization problems that can be solved efficiently. The objective functions in both optimization problems are the error functions, hence perturbed models with adequate accuracy can be obtained. Numerical examples then verify the efficiency and robustness of the proposed algorithms. Yuanzhe Wang, Zheng Zhang 0005, Cheng-Kok Koh, Guoyong Shi, Grantham Pang, Ngai Wong 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2011 | Balanced truncation for time-delay systems via approximate GramiansabstractIn circuit simulation, when a large RLC network is connected with delay elements, such as transmission lines, the resulting system is a time-delay system (TDS). This paper presents a new model order reduction (MOR) scheme for TDSs with state time delays. It is the first time to reduce a TDS using balanced truncation. The Lyapunov-type equations for TDSs are derived, and an analysis of their computational complexity is presented. To reduce the computational cost, we approximate the controllability and observability Gramians in the frequency domain. The reduced-order models (ROMs) are then obtained by balancing and truncating the approximate Gramians. Numerical examples are presented to verify the accuracy and efficiency of the proposed algorithm. Qing Wang 0051, Zheng Zhang 0005, Quan Chen 0007, Ngai Wong 0001 |
ASP-DAC | 3 |
| 2011 | A moment-matching scheme for the passivity-preserving model order reduction of indefinite descriptor systems with possible polynomial partsabstractPassivity-preserving model order reduction (MOR) of descriptor systems (DSs) is highly desired in the simulation of VLSI interconnects and on-chip passives. One popular method is PRIMA, a Krylov-subspace projection approach which preserves the passivity of positive semidefinite (PSD) structured DSs. However, system passivity is not guaranteed by PRIMA when the system is indefinite. Furthermore, the possible polynomial parts of singular systems are normally not captured. For indefinite DSs, positive-real balanced truncation (PRBT) can generate passive reduced-order models (ROMs), whose main bottleneck lies in solving the dual expensive generalized algebraic Riccati equations (GAREs). This paper presents a novel moment-matching MOR for indefinite DSs, which preserves both the system passivity and, if present, also the improper polynomial part. This method only requires solving one GARE, therefore it is cheaper than existing PRBT schemes. On the other hand, the proposed algorithm is capable of preserving the passivity of indefinite DSs, which is not guaranteed by traditional moment-matching MORs. Examples are finally presented showing that our method is superior to PRIMA in terms of accuracy. Zheng Zhang 0005, Qing Wang 0051, Ngai Wong 0001, Luca Daniel |
ASP-DAC | 1 |
| 2011 | A block-diagonal structured model reduction scheme for power grid networksabstractWe propose a block-diagonal structured model order reduction (BDSM) scheme for fast power grid analysis. Compared with existing power grid model order reduction (MOR) methods, BDSM has several advantages. First, unlike many power grid reductions that are based on terminal reduction and thus error-prone, BDSM utilizes an exact column-by-column moment matching to provide higher numerical accuracy. Second, with similar accuracy and macromodel size, BDSM generates very sparse block-diagonal reduced-order models (ROMs) for massive-port systems at a lower cost, whereas traditional algorithms such as PRIMA produce full dense models inefficient for the subsequent simulation. Third, different from those MOR schemes based on extended Krylov subspace (EKS) technique, BDSM is input-signal independent, so the resulting ROM is reusable under different excitations. Finally, due to its blockdiagonal structure, the obtained ROM can be simulated very fast. The accuracy and efficiency of BDSM are verified by industrial power grid benchmarks. Zheng Zhang 0005, Chung-Kuan Cheng, Ngai Wong 0001 |
DATE | 1 |
| 2011 | Model order reduction of fully parameterized systems by recursive least square optimizationabstractThis paper presents an approach for the model order reduction of fully parameterized linear dynamic systems. In a fully parameterized system, not only the state matrices, but also can the input/output matrices be parameterized. The algorithm presented in this paper is based on neither conventional moment-matching nor balanced-truncation ideas. Instead, it uses “optimal (block) vectors” to construct the projection matrix, such that the system errors in the whole parameter space are minimized. This minimization problem is formulated as a recursive least square (RLS) optimization and then solved at a low cost. Our algorithm is tested by a set of multi-port multi-parameter cases with both intermediate and large parameter variations. The numerical results show that high accuracy is guaranteed, and that very compact models can be obtained for multi-parameter models due to the fact that the ROM size is independent of the number of parameters in our approach. Zheng Zhang 0005, Ibrahim M. Elfadel, Luca Daniel |
ICCAD | 1 |
| 2010 | An extension of the generalized Hamiltonian method to S-parameter descriptor systemsabstractA generalized Hamiltonian method (GHM) was recently proposed for the passivity test of hybrid descriptor systems. This paper extends the GHM theory to its S-parameter counterpart. Based on the S-parameter GHM, a passivity test flow is proposed, which is capable of detecting nonpassive regions of descriptor-form physical models. The proposed method is applicable to S-parameter and hybrid systems either in the standard state-space or descriptor forms. Experimental results confirm the effectiveness and accuracy of the proposed method. Zheng Zhang 0005, Ngai Wong 0001 |
ASP-DAC | 1 |
| 2010 | Design space exploration for sparse matrix-matrix multiplication on FPGAsabstractThe design and implementation of a sparse matrix-matrix multiplication architecture on FPGAs is presented. Performance of the design, in terms of computational latency, as well as the associated power-delay and energy-delay tradeoff are studied. Taking advantage of the sparsity of the input matrices, the proposed design allows user-tunable power-delay and energy-delay tradeoffs by employing different number of processing elements (PEs) in the architecture design and different block size in the blocking decomposition. Such ability allows designers to employ different on-chip computational architecture for different system power-delay and energy-delay requirements. It is in contrast to conventional dense matrix-matrix multiplication architectures that always favor the maximum number of PEs and largest block size. In our implementation, the better energy consumption and power-delay product favors less PEs and smaller block size for the 90%-sparsity matrix-matrix multiplications. While in order to achieve better energy-delay product, more PEs and larger block size are preferred. Colin Yu Lin, Zheng Zhang 0005, Ngai Wong 0001, Hayden Kwok-Hay So |
FPT | 2 |
| 2010 | PEDS: Passivity enforcement for descriptor systems via Hamiltonian-symplectic matrix pencil perturbationabstractPassivity is a crucial property of macromodels to guarantee stable global (interconnected) simulation. However, weakly nonpassive models may be generated for passive circuits and systems in various contexts, such as data fitting, model order reduction (MOR) and electromagnetic (EM) macromodeling. Therefore, a post-processing passivity enforcement algorithm is desired. Most existing algorithms are designed to handle pole-residue models. The few algorithms for state space models only handle regular systems (RSs) with a nonsingular D+DTterm. To the authors' best knowledge, no algorithm has been proposed to enforce passivity for more general descriptor systems (DSs) and state space models with singular D+DTterms. In this paper, a new post-processing passivity enforcement algorithm based on perturbation of Hamiltonian-symplectic matrix pencil, PEDS, is proposed. PEDS, for the first time, can enforce passivity for DSs. It can also handle all kinds of state space models (both RSs and DSs) with singular D+DTterms. Moreover, a criterion to control the error of perturbation is devised, with which the optimal passive models with the best accuracy can be obtained. Numerical examples then verify that PEDS is efficient, robust and relatively cheap for passivity enforcement of DSs with mild passivity violations. Yuanzhe Wang, Zheng Zhang 0005, Cheng-Kok Koh, Grantham Pang, Ngai Wong 0001 |
ICCAD | 2 |
| 2010 | An Efficient Projector-Based Passivity Test for Descriptor SystemsabstractAn efficient passivity test based on canonical projector techniques is proposed for descriptor systems (DSs) widely encountered in circuit and system modeling. The test features a natural flow that first evaluates the index of a DS, followed by possible decoupling into its proper and improper subsystems. Explicit state-space formulations for respective subsystems are derived to facilitate further processing such as model order reduction and/or passivity enforcement. Efficient projector construction and a fast generalized Hamiltonian test for the proper-part passivity are also elaborated. Numerical examples then confirm the superiority of the proposed method over existing passivity tests for DSs based on linear matrix inequalities or skew-Hamiltonian/Hamiltonian matrix pencils. Zheng Zhang 0005, Ngai Wong 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2009 | GHM: A generalized Hamiltonian method for passivity test of impedance/admittance descriptor systemsabstractA generalized Hamiltonian method (GHM) is proposed for passivity test of descriptor systems (DSs) which describe impedance or admittance input-output responses. GHM can test passivity of DSs with any system index without minimal realization. This frequency-independent method can avoid the time-consuming system decomposition as required in many existing DS passivity test approaches. Furthermore, GHM can test systems with singular D + DT where traditional Hamiltonian method fails, and enjoys a more accurate passivity violation identification compared to frequency sweeping techniques. Numerical results have verified the effectiveness of GHM. The proposed method constitutes a versatile tool to speed up passivity check and enforcement of DSs and subsequently ensures globally stable simulations of electrical circuits and components. Zheng Zhang 0005, Chi-Un Lei, Ngai Wong 0001 |
ICCAD | 1 |