VLDB 2026 Research / reviewers in the wild / expert
Zhengya Zhang
dblp:17/5332
· DBLP profile ↗
44ranked-venue papers
5as first author
14since 2021 · last 2025
0000-0001-5963-9018ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 27 · 1 first-author · 12 since 2021Computer networks · 8 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Security and privacy · 2Graphics, computer vision, multimedia, augmented reality and games · 2Artificial intelligence and machine learning · 1 · 1 first-authorTheory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Practical MU-MIMO Detection and LDPC Decoding Through Digital AnnealingabstractDigital annealing has been successfully applied to solving combinatorial optimization (CO) problems. It is more flexible, robust, and easier to deploy on edge platforms compared to its counterparts including quantum annealing and analog and in-memory Ising machines. In this work, we apply digital annealing to compute-intensive communication digital signal processing problems, including multi-user detection in multiple-input and multiple-output (MU-MIMO) wireless communication systems and decoding low-density parity-check (LDPC) codes. We show that digital annealing can achieve near maximum likelihood (ML) accuracy for MIMO detection with even lower complexity than the conventional minimum mean square error (MMSE) detection. In LDPC decoding, we enhance digital annealing by introducing a new cost function that improves decoding accuracy and reduces computational complexity compared to the standard formulations. Po-Shao Chen, Zhengya Zhang |
DATE | 3 |
| 2025 | Panacea: Novel DNN Accelerator using Accuracy-Preserving Asymmetric Quantization and Energy-Saving Bit-Slice SparsityabstractLow bit-precisions and their bit-slice sparsity have recently been studied to accelerate general matrix-multiplications (GEMM) during large-scale deep neural network (DNN) inferences. While the conventional symmetric quantization facilitates low-resolution processing with bit-slice sparsity for both weight and activation, its accuracy loss caused by the activation’s asymmetric distributions cannot be acceptable, especially for largescale DNNs. In efforts to mitigate this accuracy loss, recent studies have actively utilized asymmetric quantization for activations without requiring additional operations. However, the cuttingedge asymmetric quantization produces numerous nonzero slices that cannot be compressed and skipped by recent bit-slice GEMM accelerators, naturally consuming more processing energy to handle the quantized DNN models.To simultaneously achieve high accuracy and hardware efficiency for large-scale DNN inferences, this paper proposes an Asymmetrically-Quantized bit-Slice GEMM (AQS-GEMM) for the first time. In contrast to the previous bit-slice computing, which only skips operations of zero slices, the AQS-GEMM compresses frequent nonzero slices, generated by asymmetric quantization, and skips their operations. To increase the slicelevel sparsity of activations, we also introduce two algorithm-hardware co-optimization methods: a zero-point manipulation and a distribution-based bit-slicing. To support the proposed AQS-GEMM and optimizations at the hardware-level, we newly introduce a DNN accelerator, Panacea, which efficiently handles sparse/dense workloads of the tiled AQS-GEMM to increase data reuse and utilization. Panacea supports a specialized dataflow and run-length encoding to maximize data reuse and minimize external memory accesses, significantly improving its hardware efficiency. Numerous benchmark evaluations show that Panacea outperforms existing DNN accelerators, e.g., $1.97 \times$ and $3.26 \times$ higher energy efficiency, and $1.88 \times$ and $2.41 \times$ higher throughput than the recent bit-slice accelerator Sibia and the SIMD design, respectively, on OPT-2.7B, while providing better algorithm performance with asymmetric quantization. Dongyun Kam, Myeongji Yun, Sunwoo Yoo, Seungwoo Hong, Zhengya Zhang, Youngjoo Lee 0002 |
HPCA | 5 |
| 2025 | HiPER: Hierarchically-Composed Processing for Efficient Robot Learning-Based ControlabstractLearning-Based Model Predictive Control (LMPC) is a class of algorithms that enhances Model Predictive Control (MPC) by including machine learning methods, improving robot navigation in complex environments.However, the combination of machine learning and MPC computation in LMPC creates a unique workload that cannot be efficiently handled by a simple GPU and CPU integration.We present HiPER, a hierarchically-composed processing array that provides temporal and spatial mapping capabilities, allowing efficient adaptation to workload changes at runtime.To simplify control, HiPER employs a pointer queue hierarchy to compose and orchestrate program execution.Additionally, HiPER utilizes a fractal interconnect topology that combines local systolic interconnects and their hierarchical extensions to efficiently support the workload's traffic characteristics.To evaluate the performance and efficiency of HiPER, we synthesized a 16.37 mm 2 design in 16nm CMOS.The design consists of 6 pointer queue levels and 1024 PEs.The prototype was assessed using a representative LMPC workload, demonstrating 10.75× improvement in performance compared to a GTX 1080 GPU, 12.80× improvement in energy efficiency compared to a Jetson Orin Nano embedded GPU, and 11.6×/22.2×improvement in performance compared to the RoboX/Plasticine accelerators. Justin Ting, Minsik Kim 0006, Junkang Zhu, Haotian Sheng, Zhengya Zhang |
ISCA | 5 |
| 2025 | SteROI-D: System Design and Mapping for Stereo Depth Inference on Regions of Interest
Jack Erhardt, Reid Pinkham, Andrew Berkovich, Zhengya Zhang |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2024 | Design Approach for Die-to-Die Interfaces to Enable Energy-Efficient Chiplet SystemsabstractHeterogeneous chiplet integration and advanced packaging have given a new lease to scaling of compute in a post-Moore era. A critical aspect of designing chiplet systems is die-to-die interfaces for aggregation of smaller disaggregated chiplets. In recent years, interconnects and packaging have made a huge leap forward, thereby, enabling high bandwidth and energy-efficient parallel die-to-die (D2D) interfaces. Instead of bespoke solutions, building a standardized D2D interface provides a mechanism for interoperability between heterogeneous chiplets and facilitates low power and energy-efficient design by prescribing implementation strategy. In this paper, we discuss some of the recent efforts in the standardization of die-to-die interfaces. Starting with an overview of Advanced Interface Bus (AIB) PHY, we emphasize its simplicity and high energy efficiency. Followed by a case study that demonstrates an effective chiplet integration employing AIB. A more recent open standard for chiplet integration is the Universal Chiplet Interconnect Express (UCIe). We discuss the key aspects of UCIe, including its electrical and packaging characteristics, as well as low-power features that lead to a 10x power reduction compared to typical off-package I/O. Finally, we discuss the UCIe-lite controller, an effort to democratize the chiplet infrastructure by providing a simplified open-source RTL generator of the D2D interface. The generator is highly parameterizable and lightweight, enabling chiplet systems for energy-efficient applications. Vikram Jain, Wei Tang 0010, Zuoguo Wu, Viansa Schmulbach, Sophia Shao, Zhengya Zhang, Borivoje Nikolic |
ISLPED | 6 |
| 2024 | Tiled Beamspace Processing for Scaling mmWave Massive MU-MIMOabstractRecent progress in millimeter wave (mmWave) silicon technologies has given rise to a new possibility: digital beamforming for truly massive multiuser (MU-) MIMO. However, there are two key challenges in scaling: packaging a vast array of antennas with the corresponding radio frequency integrated circuits (RFICs), and controlling the complexity of digital signal processing (DSP) for MU detection. In this paper, we show that modular tiled architectures, which simplify the task of RF packaging, also enable significant reduction of the communication and computational burden of DSP for MU-MIMO by utilizing beamspace techniques that take advantage of the sparsity of the mmWave channel. Specifically, we propose and investigate Linear Minimum Mean Squared Error (LMMSE) adaptive MU detection via novel tiled beamspace architectures in which the bulk of the DSP occurs in-place at each tile. The dimensionality reduction and parallelization enabled by such architectures not only reduce the computational burden of inference and training relative to a traditional "full array" baseline, but also significantly reduce the length of the required training period. We consider three different training strategies with differing requirements for computation and inter-tile communication: independent training for each tile, coordinated training across tiles, and hierarchical training based on independent training as a first stage. Simulation results show that these approaches can actually outperform the full array baseline when we limit the length of the training period. Jiyoon Han, Canan Cebeci, Wei Tang 0010, Zhengya Zhang, Upamanyu Madhow |
VTC Fall | 4 |
| 2024 | TetriX: Flexible Architecture and Optimal Mapping for Tensorized Neural Network ProcessingabstractThe continuous growth of deep neural network model size and complexity hinders the adoption of large models in resource-constrained platforms. Tensor decomposition has been shown effective in reducing the model size by large compression ratios, but the resulting tensorized neural networks (TNNs) require complex and versatile tensor shaping for tensor contraction, causing a low processing efficiency for existing hardware architectures. This work presents TetriX, a co-design of flexible architecture and optimal workload mapping for efficient and flexible TNN processing. TetriX adopts a unified processing architecture to support both inner and outer product. A hybrid mapping scheme is proposed to eliminate complex tensor shaping by alternating between inner and outer product in a sequence of tensor contractions. Finally, a mapping-aware contraction sequence search (MCSS) is proposed to identify the contraction sequence and workload mapping for achieving the optimal latency on TetriX. Remarkably, combining TetriX with MCSS outperforms the single-mode inner-product and outer-product baselines by up to 46.8× in performance across the collected TNN workloads. TetriX is the first work to support all existing tensor decomposition methods. Compared to a TNN accelerator designed for the hierarchical Tucker method, TetriX achieves improvements of 6.5× and 1.1× in inference throughput and efficiency, respectively. Jie-Fang Zhang, Cheng-Hsun Lu, Zhengya Zhang |
IEEE Trans. Computers | 3 |
| 2024 | TT-CIM: Tensor Train Decomposition for Neural Network in RRAM-Based Compute-in-Memory SystemsabstractCompute-in-Memory (CIM) implemented with Resistive-Random-Access-Memory (RRAM) crossbars is a promising approach for accelerating Convolutional Neural Network (CNN) computations. The growing size in the number of parameters in state-of-the-art CNN models, however, creates challenge for on-chip weight storage for CIM implementations, and CNN compression becomes a crucial topic of exploration. Tensor Train (TT) decomposition can be used to decompose a tensor into smaller ones with fewer parameters, at the cost of increased number of computations. In this work we propose a technique to minimize intermediate operations across the full convolution operation and improve hardware utilization to implement TT-CNNs in CIM systems. We first use an iterative decompose-and-fine-tune method to prepare TT-CNNs. We then propose an inter-convolutional-step reuse scheme to reduce the required operation count and post-mapping RRAM count for TT-CNN implementation in tiled-CIM architecture. We demonstrate that through proper mapping, pipelining, and reuse, effective compression ratio of 12 and 20 with 0.8% and 1.4% accuracy drop, respectively for WRN; and effective compression ratio of 6 and 11 with 0.9% and 1.2% accuracy drop for VGG8. We also show that around 30% higher hardware utilization than the original CNN format can be achieved using the proposed TT-CIM approaches. Fan-Hsuan Meng, Zhengya Zhang, Wei Lu 0003 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2023 | eNODE: Energy-Efficient and Low-Latency Edge Inference and Training of Neural ODEsabstractNeural ordinary differential equations (NODEs) provide better modeling performance with smaller amount of model parameters in many tasks by embedding neural networks (NNs) in ordinary differential equations (ODEs). They have been shown to outperform in representing continuous-time data and learning dynamic systems, and are promising for on-device inference and training. However, an edge device is limited by area and energy budget, and real-time operations have a tight latency requirement. State-of-the-art NN accelerators are not optimized for the area- and power-hungry memory storage and access for NODE inference and training, and lack the flexibility to incorporate dynamic latency reduction techniques. We present eNODE by architecture-algorithm co-design to achieve efficient and fast inference and training of NODEs. eNODE adopts compact-size depth-first integration and depth-first training for higher energy efficiency. Through function reuse, packetized processing and a unified NN core design, the efficiency of eNODE’s depth-first processing is further enhanced. We propose algorithm innovations, including slope-adaptive stepsize search and priority processing with early stop, to substantially shorten the latency. A hardware prototype is synthesized in a 28 nm CMOS technology for evaluation and benchmarking. eNODE demonstrates up to 6.59× better energy efficiency, 2.38× higher speed, and better area scalability over a SIMD ASIC baseline. Junkang Zhu, Yaoyu Tao, Zhengya Zhang |
HPCA | 3 |
| 2023 | AR-PIM: An Adaptive-Range Processing-in-Memory ArchitectureabstractThe crossbar-based processing-in-memory (PIM) architecture has garnered considerable attention for its potential in achieving high energy efficiency for deep neural networks (DNNs). The PIM hardware's accuracy depends heavily on the design and resolution of the analog-to-digital converters (ADCs). Regrettably, high-resolution ADCs tend to be costly and often dominate the overall energy and area of the PIM designs. We propose adaptive-range PIM (AR-PIM) architecture that enables the use of lower-resolution ADCs without sacrificing accuracy. This is achieved by leveraging sparsity in the weights and input activations and dynamically adjusting the number of input activations and distributing MAC operations across multiple cycles during runtime. We perform our evaluations using a commercial 7nm FinFET PDK and show that AR-PIM offers an appealing trade-off, delivering 1.7 × higher energy efficiency and 4.3 × better area benefits without losing accuracy. The latency overhead is modest, only 10% over a baseline PIM architecture. Teyuh Chou, Fernando García-Redondo, Paul N. Whatmough, Zhengya Zhang |
ISLPED | 4 |
| 2023 | ANSA: Adaptive Near-Sensor Architecture for Dynamic DNN Processing in Compact Form FactorsabstractAdvanced edge sensing/computing devices, such as AR/VR devices, have a uniquely challenging adaptive baseline workload and camera sensor structure. These devices must process images in real-time from multiple sensors, placing a large burden on a typical centralized mobile SoC processor. Augmenting the sensors with a package-integrated near-sensor processor can improve the device’s processing performance as well as reduce energy consumption. This near-sensor processor must adapt to the dynamic workloads, fit within a limited silicon footprint and energy envelope, and satisfy the real-time requirement. In this work, we present ANSA, a near-sensor processor architecture supporting flexible processing schemes and dataflows to maintain high efficiency for dynamic CNN workloads. ANSA is scalable to sub-mm2 sizes to match the footprint of advanced image sensors. ANSA supports module-level power gating to adapt the compute capacity to dynamic workloads. Finally, ANSA leverages recent advancements in high-density non-volatile memory and 3D packaging to support weight storage within the area constraints of an image sensor. Overall, ANSA achieves inference energy consumption up to$30\times $lower than a standard SIMD baseline. Additionally, our design’s scalability allows it to achieve up to$2.76\times $lower average inference energy at$4.5\times $lower silicon area compared to competing edge accelerator designs. Reid Pinkham, Jack Erhardt, Barbara De Salvo, Andrew Berkovich, Zhengya Zhang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2021 | DNC-Aided SCL-Flip Decoding of Polar CodesabstractSuccessive-cancellation list (SCL) decoding of polar codes has been adopted for 5G wireless communications. How-ever, the performance of moderate code length is not satisfactory. Heuristic or deep-learning-aided (DL-aided) flip algorithms have been developed to improve the performance by locating error bit positions after SCL decoding. In this work, we propose a new flip algorithm with the help of differentiable neural computer (DNC). New state and action encoding are developed to improve DNC training and inference efficiency. The proposed two-phase method is done by a flip DNC (F-DNC) to rank the most likely flip positions for multi-bit flipping, and if decoding still fails, a flip-validate DNC (FV-DNC) is applied to re-select error bit positions in successive flip decoding trials. Supervised training methods are designed for the two DNCs. Simulation results show that the proposed DNC-aided SCL-Flip (DNC-SCLF) decoding demonstrates up to 0.34 dB coding gain or 54.2% reduction in the average number of decoding attempts over prior work. Yaoyu Tao, Zhengya Zhang |
GLOBECOM | 2 |
| 2021 | HiMA: A Fast and Scalable History-based Memory Access Engine for Differentiable Neural ComputerabstractMemory-augmented neural networks (MANNs) provide better inference performance in many tasks with the help of an external memory. The recently developed differentiable neural computer (DNC) is a MANN that has been shown to outperform in representing complicated data structures and learning long-term dependencies. DNC’s higher performance is derived from new history-based attention mechanisms in addition to the previously used content-based attention mechanisms. History-based mechanisms require a variety of new compute primitives and state memories, which are not supported by existing neural network (NN) or MANN accelerators. We present HiMA, a tiled, history-based memory access engine with distributed memories in tiles. HiMA incorporates a multi-mode network-on-chip (NoC) to reduce the communication latency and improve scalability. An optimal submatrix-wise memory partition strategy is applied to reduce the amount of NoC traffic; and a two-stage usage sort method leverages distributed tiles to improve computation speed. To make HiMA fundamentally scalable, we create a distributed version of DNC called DNC-D to allow almost all memory operations to be applied to local memories with trainable weighted summation to produce the global memory output. Two approximation techniques, usage skimming and softmax approximation, are proposed to further enhance hardware efficiency. HiMA prototypes are created in RTL and synthesized in a 40nm technology. By simulations, HiMA running DNC and DNC-D demonstrates 6.47 × and 39.1 × higher speed, 22.8 × and 164.3 × better area efficiency, and 6.1 × and 61.2 × better energy efficiency over the state-of-the-art MANN accelerator. Compared to an Nvidia 3080Ti GPU, HiMA demonstrates speedup by up to 437 × and 2,646 × when running DNC and DNC-D, respectively. Yaoyu Tao, Zhengya Zhang |
MICRO | 2 |
| 2021 | Point-X: A Spatial-Locality-Aware Architecture for Energy-Efficient Graph-Based Point-Cloud Deep LearningabstractDeep learning on point clouds has attracted increasing attention in the fields of 3D computer vision and robotics. In particular, graph-based point-cloud deep neural networks (DNNs) have demonstrated promising performance in 3D object classification and scene segmentation tasks. However, the scattered and irregular graph-structured data in a graph-based point-cloud DNN cannot be computed efficiently by existing SIMD architectures and accelerators. We present Point-X, an energy-efficient accelerator architecture that extracts and exploits the spatial locality in point cloud data for efficient processing. Point-X uses a clustering method to extract fine-grained and coarse-grained spatial locality from the input point cloud. The clustering maps the point cloud into distributed compute tiles to maximize intra-tile computational parallelism and minimize inter-tile data movement. Point-X employs a chain network-on-chip (NoC) to further reduce the NoC traffic and achieve up to 3.2 × speedup over a traditional mesh NoC. Point-X’s multi-mode dataflow can support all common operations in a graph-based point-cloud DNN, i.e., edge convolution, shared multi-layer perceptron, and fully-connected layers. Point-X is synthesized in a 28nm technology and it demonstrates a throughput of 1307.1 inference/s and an energy efficiency of 604.5 inference/J on the DGCNN workload. Compared to the Nvidia GTX-1080Ti GPU, Point-X shows 4.5 × and 342.9 × improvement in throughput and efficiency, respectively. Jie-Fang Zhang, Zhengya Zhang |
MICRO | 2 |
| 2020 | QuickNN: Memory and Performance Optimization of k-d Tree Based Nearest Neighbor Search for 3D Point CloudsabstractThe use of Light Detection And Ranging (LiDAR) has enabled the continued improvement in accuracy and performance of autonomous navigation. The latest applications require LiDAR's of the highest spatial resolution, which generate a massive amount of 3D point clouds that need to be processed in real time. In this work, we investigate the architecture design for k-Nearest Neighbor (kNN) search, an important processing kernel for 3D point clouds. An approximate kNN search based on a k-dimensional (k-d) tree is employed to improve performance. However, even for today's moderate-sized problems, this approximate kNN search is severely hindered by memory bandwidth due to numerous random accesses and minimal data reuse opportunities. We apply several memory optimization schemes to alleviate the bandwidth bottleneck: 1) the k-d tree data structure is partitioned to two sets: tree nodes and point buckets, based on their distinct characteristics - tree nodes that have high reuse are cached for their lifetime to facilitate search, while point buckets with low reuse are organized in regular contiguous segments in external memory to facilitate efficient burst access; 2) write and read caches are added to gather random accesses to transform them to sequential accesses; and 3) tree construction and tree search are interleaved to cut redundant access streams. With optimized memory bandwidth, the kNN search can be further accelerated by two new processing schemes: 1) parallel tree traversal that utilizes multiple workers with minimal tree duplication overhead, and 2) incremental tree building that minimizes the overhead of tree construction by dynamically updating the tree instead of building it from scratch every time. We demonstrate the performance and memory-optimized QuickNN architecture on FPGA and perform exhaustive benchmarking, showing that up to a 19× and 7.3× speedup over k-d tree searches performed on a modern CPU and GPU, respectively, and a 14.5× speedup over a comparable sized architecture performing an exact search. Finally, we show that QuickNN achieves two orders of magnitude performance per watt increase over CPU and GPU methods. Reid Pinkham, Shuqing Zeng, Zhengya Zhang |
HPCA | 3 |
| 2020 | Control of Magnetically-Driven Screws in a Viscoelastic MediumabstractMagnetically-driven screws operating in soft-tissue environments could be used to deploy localized therapy or achieve minimally invasive interventions. In this work, we characterize the closed-loop behavior of magnetic screws in an agar gel tissue phantom using a permanent magnet-based robotic system with an open-configuration. Our closed-loop control strategy capitalizes on an analytical calculation of the swimming speed of the screw in viscoelastic fluids and the magnetic point-dipole approximation of magnetic fields. The analytical solution is based on the Stokes/Oldroyd-B equations and its predictions are compared to experimental results at different actuation frequencies of the screw. Our measurements matches the theoretical prediction of the analytical model before the step-out frequency of the screw owing to the linearity of the analytical model. We demonstrate open-loop control in two-dimensional space, and point-to-point closed-loop motion control of the screw (length and diameter of 6 mm and 2 mm, respectively) with maximum positioning error of 1.8 mm. Zhengya Zhang, Anke Klingner, Sarthak Misra, Islam S. M. Khalil |
IROS | 1 |
| 2019 | CASCADE: Connecting RRAMs to Extend Analog Dataflow In An End-To-End In-Memory Processing ParadigmabstractProcessing in memory (PIM) is a concept to enable massively parallel dot products while keeping one set of operands in memory. PIM is ideal for computationally demanding deep neural networks (DNNs) and recurrent neural networks (RNNs). Processing in resistive RAM (RRAM) is particularly appealing due to RRAM's high density and low energy. A key limitation of PIM is the cost of multibit analog-to-digital (A/D) conversions that can defeat the efficiency and performance benefits of PIM. In this work, we demonstrate the CASCADE architecture that connects multiply-accumulate (MAC) RRAM arrays with buffer RRAM arrays to extend the processing in analog and in memory: dot products are followed by partial-sum buffering and accumulation to implement a complete DNN or RNN layer. Design choices are made and the interface is designed to enable a variation-tolerant, robust analog dataflow. A new memory mapping scheme named R-Mapping is devised to enable the in-RRAM accumulation of partial sums; and an analog summation scheme is used to reduce the number of A/D conversions required to obtain the final sum. CASCADE is compared with recent in-RRAM computation architectures using state-of-the-art DNN and RNN benchmarks. The results demonstrate that CASCADE improves the energy efficiency by 3.5× while maintaining a competitive throughput. Teyuh Chou, Wei Tang 0010, Jacob Botimer, Zhengya Zhang |
MICRO | 4 |
| 2019 | Editorial TVLSI Positioning - Continuing and Accelerating an Upward TrajectoryabstractI. VLSI Systems: A Glance Into The Last Decades Since their inception in 1970s, VLSI systems have enabled several new technological capabilities and made them accessible to an unceasingly wider range of users, reaching a scale that has been exponentially increasing over the decades[1](seeFig. 1). Relentless integration of more complex systems has driven such remarkable evolution, as made possible by the inexorable miniaturization. As shown inFig. 1, more functionality has been crammed in a consistently smaller form factor, as exemplified by the physical volume shrinking of computers by 100 X/decade[2],[3]. At the same time, the energy per task has been decreasing at 10–100 X/decade, as shown inFig. 2, for several systems and system-on-chip subsystems[4]. This allowed packing more capabilities into the same power envelope, as generally observed in the electronic systems, even before the advent of the integrated circuit[5]. Massimo Alioto, Magdy S. Abadir, Tughrul Arslan, Chirn Chye Boon, Andreas Peter Burg, Chip-Hong Chang, Meng-Fan Chang, Yao-Wen Chang, Poki Chen, Pasquale Corsonello, Paolo Crovetti, Shiro Dosho, Rolf Drechsler, Ibrahim M. Elfadel, Ruonan Han 0001, Masanori Hashimoto, Chun-Huat Heng, Deuk Hyoun Heo, Tsung-Yi Ho, Houman Homayoun, Yuh-Shyan Hwang, Ajay Joshi, Rajiv V. Joshi, Tanay Karnik, Chulwoo Kim, Tony Tae-Hyoung Kim, Jaydeep P. Kulkarni, Volkan Kursun, Yoonmyung Lee, Hai Li 0001, Huawei Li 0001, Prabhat Mishra 0001, Baker Mohammad, Mehran Mozaffari Kermani, Makoto Nagata, Koji Nii, Partha Pratim Pande, Bipul Chandra Paul, Vasilis F. Pavlidis, José Pineda de Gyvez, Ioannis Savidis, Patrick Schaumont, Fabio Sebastiano, Anirban Sengupta 0003, Mingoo Seok, Mircea R. Stan, Mark Tehranipoor, Aida Todri, Marian Verhelst, Valerio Vignoli, Xiaoqing Wen, Jiang Xu 0001, Wei Zhang 0012, Zhengya Zhang, Jun Zhou 0017, Mark Zwolinski, Stacey Weber |
IEEE Trans. Very Large Scale Integr. Syst. | 54 |
| 2017 | Post-Processing Methods for Improving Coding Gain in Belief Propagation Decoding of Polar CodesabstractBelief propagation (BP) is a high-throughput, low- latency decoding algorithm for polar codes, but the error-correcting performance is known to be inferior than successive cancellation (SC) decoding. To improve the error-correcting performance of BP decoding, we design post- processing methods targeting false converged errors, oscillation errors, and unconverged errors that determine the performance of BP decoding. False convergence can be resolved by perturbing, or gradually freezing the information bits, followed by error cleanup using BP. Oscillations can be resolved by enhancing the stable bits and perturbing the unstable bits, followed by error cleanup using BP. Unconverged errors can be resolved by enhancing the reliably stable bits and weakening the unstable bits. Results show that the error rates of BP decoding can be improved by an order of magnitude or more, allowing it to overtake SC in error rate and coding gain. Post-processing can be implemented very efficiently, costing less than 4.3% overhead in silicon area, and it does not affect the throughput or latency of BP decoding. Shuanghong Sun, Sung-Gun Cho, Zhengya Zhang |
GLOBECOM | 3 |
| 2017 | EditorialabstractAs I start my second two-year term (2017–2018) as the Editor-in-Chief (EIC) of the IEEE Transactions on Very Large Scale Integration Systems (TVLSI), I wish the TVLSI readership a very happy new year and continued professional success. It gives me great pleasure to report on the state of the journal and our performance metrics. Over the past two years, TVLSI has seen a healthy increase in the number of submissions—from 687 in 2014 to 770 in 2015, and at the time of writing of this editorial, we are at 760 submissions for 2016. We expect the number of submissions for 2016 to cross 800 before the end of the year. TVLSI, therefore, continues to be the premier archival journal for university researchers and industry practitioners in the broad area of VLSI system design. Krishnendu Chakrabarty, Massimo Alioto, Bevan M. Baas, Chirn Chye Boon, Meng-Fan Chang, Naehyuck Chang, Yao-Wen Chang, Chip-Hong Chang, Shih-Chieh Chang 0001, Poki Chen, Masud H. Chowdhury, Pasquale Corsonello, Ibrahim M. Elfadel, Said Hamdioui, Masanori Hashimoto, Tsung-Yi Ho, Houman Homayoun, Yuh-Shyan Hwang, Rajiv V. Joshi, Tanay Karnik, Mehran Mozaffari Kermani, Chulwoo Kim, Jaydeep P. Kulkarni, Eren Kursun, Erik Larsson, Hai Li 0001, Huawei Li 0001, Patrick P. Mercier, Prabhat Mishra 0001, Makoto Nagata, Arun Natarajan 0001, Koji Nii, Partha Pratim Pande, Ioannis Savidis, Mingoo Seok, Sheldon X.-D. Tan, Mark Tehranipoor, Aida Todri, Miroslav N. Velev, Xiaoqing Wen, Jiang Xu 0001, Wei Zhang 0012, Zhengya Zhang, Stacey Weber |
IEEE Trans. Very Large Scale Integr. Syst. | 44 |
| 2016 | A 180 GHz prototype for a geostationary microwave imager/sounder-GeoSTAR-IIIabstractGeoSTAR-III, a 180 GHz prototype for the Precipitation and All-weather Temperature and Humidity Sounder (PATH), is the culmination of a decade of technology development funding. The interferometric radiometer comprises 144 receivers operating from 165-183 GHz and utilizes a 192×192 input, ASIC based mixed signal correlator. The demonstration of this instrument raises the technology readiness of the radiometer subsystem to level 6 (TRL 6) and the correlator subsystem to TRL 5. We demonstrate the full functionality of this system with observations of the Sun and Moon as well as nearby thermally emissive objects. This represents the final milestones in the development effort of pre-mission technologies for this decadal survey mission. Todd Gaier, Pekka Kangaslahti, Bjorn Lambrigtsen, Isaac Ramos-Pérez, Alan B. Tanner, Darren McKague, Christopher Ruf, Michael J. Flynn, Zhengya Zhang, Roger Backhus, David Austerberry |
IGARSS | 9 |
| 2016 | Architecture and optimization of high-throughput belief propagation decoding of polar codesabstractBelief propagation (BP) is one common decoding algorithm for polar codes. BP can be implemented using a forward-backward flooding schedule that removes the data dependency. The intrinsic parallelism of the BP algorithm permits a high throughput. In this paper, we present a parallel BP decoder design space exploration. Through gate-level implementations, we analyze the silicon area, throughput, and power consumption of each design. We present adaptive quantization and early convergence detection to improve the error-correcting performance and throughput of BP decoders. Shuanghong Sun, Zhengya Zhang |
ISCAS | 2 |
| 2016 | A Simple Generic Attack on Text Captchas
Haichang Gao, Jeff Yan, Zhengya Zhang, Mengyun Tang, Xuqin Wang |
NDSS | 4 |
| 2016 | Robustness of text-based completely automated public turing test to tell computers and humans apartabstractText‐based completely automated public turing tests to tell computers and humans apart (CAPTCHAs) have been widely deployed across the Internet to defend against undesirable or malicious bot programmes. In this study, the authors provide a systematic analysis of text‐based CAPTCHAs and innovatively improve their earlier attack on hollow CAPTCHAs to expand applicability to attack all the text CAPTCHAs. With this improved attack, they have successfully broken the CAPTCHA schemes adopted by 19 out of the top 20 web sites in Alexa including two versions of the famous ReCAPTCHA. With success rates ranging from 12 to 88.8% (note that the success rate for Yandex CAPTCHA is 0%), they demonstrate the effectiveness of their attack method. It is not only applicable to hollow CAPTCHAs, but also to non‐hollow ones. As their attack casts serious doubt on the viability of current designs, they offer lessons and guidelines for designing better text‐based CAPTCHAs. Haichang Gao, Xuqin Wang, Zhengya Zhang, Jiao Qi |
IET Inf. Secur. | 4 |
| 2014 | Memristive devices for stochastic computingabstractWe show resistive switching effects in memristive devices exhibit significant stochasticity. When the switching is dominated by a single filament, the switching time is fully random and shows a broad distribution. However, the switching distribution can be predicted and responds well to controlled changes in the programming conditions. The native stochastic characteristic can be used to generate random bit streams with predictable biases that can lead to efficient and error-tolerant computing. Siddharth Gaba, Phil Knag, Zhengya Zhang, Wei Lu 0003 |
ISCAS | 3 |
| 2014 | Design and Evaluation of Confidence-Driven Error-Resilient SystemsabstractDeeply scaled CMOS circuits are increasingly susceptible to transient faults and soft errors; emerging post-CMOS devices can be more vulnerable, sometimes exhibiting erratic errors of arbitrary duration. Applying timing and supply voltage margin is wasteful and becoming ineffective, and conventional checking and sparing techniques provide only a limited error coverage against widely varying errors. We propose a confidence-driven computing (CDC) model for an adaptive protection against nondeterministic errors. The CDC model employs fine-grained temporal redundancy and confidence checking for a faster adaptation and tunable reliability. The CDC model can be extended to deeply scaled CMOS circuits that are mainly affected by transient faults and soft errors, where an early checking (EC) technique can be used to perform independent error checking for more flexibility and better performance. To evaluate the CDC model, we apply a sample-based field-programmable gate array emulation along with real-time error injection. The CDC model is shown to adapt to fluctuating error rates and enhance the system reliability by effectively trading off performance. To evaluate the EC technique at a finer time scale, we create a new event-based simulation to capture path delay distribution, error model, and their interactions. The EC technique improves the system reliability by more than four orders of magnitude when errors are of short duration. Both the CDC model and the EC technique are synthesized in a 45-nm CMOS technology for cost estimates: 1) the area overhead is as low as 12% and 2) energy overhead can be limited to 19%. David T. Blaauw, Dennis Sylvester, Zhengya Zhang |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2013 | An FPGA-based transient error simulator for evaluating resilient system designs (abstract only)abstractError-resilient designs have become more important with the continued device scaling. One critical challenge of designing error-resilient systems is the lack of tools to quickly and accurately evaluate the effectiveness and performance of such systems. We propose an FPGA-based transient error simulator to accelerate transient error simulations incorporating accurate datapath delay models and realistic error models. Compared to conventional digital error simulators, the FPGA-based transient error simulator operates at a finer time step and captures intricate interactions between errors and datapath under different circuit-level error detection and correction techniques. The error simulator is constructed using configurable datapath delay model and error model, making it general-purpose and widely applicable. We demonstrate the capability of this simulator in the evaluation of two popular error-resilient design techniques, pre-edge and post-edge detection and correction, using a synthesized CORDIC processor and an Alpha processor that operate under soft error, coupling noise and voltage droop models. The proposed error simulator uncovers insights to guide practical designs, including the choice of checking window in pre-edge designs and the optimal operating frequency in post-edge designs. The FPGA-based transient simulation will complement circuit simulation and system emulation for resilient system designs. Shiming Song 0002, Zhengya Zhang |
FPGA | 3 |
| 2013 | A low-power VGA full-frame feature extraction processorabstractThis paper proposes an energy-efficient VGA full-frame feature extraction processor design. It is based on the SURF algorithm and makes various algorithmic modifications to improve efficiency and reduce hardware overhead while maintaining extraction performance. Low clock frequency and deep parallelism derived from a one-sample-per-cycle matched-throughput architecture provide significantly larger room for voltage scaling and enables full-frame extraction. The proposed design consumes 4.7mW at 400mV and achieves 72% higher energy efficiency than prior work. Dongsuk Jeon, Yejoong Kim, Inhee Lee 0001, Zhengya Zhang, David T. Blaauw, Dennis Sylvester |
ICASSP | 4 |
| 2013 | Efficient in situ error detection enabling diverse path coverageabstractTechnology scaling continues to improve density, but also reduces the critical charge to hold a logic state, causing devices to become more susceptible to accidental disruptions due to noise and soft errors. Increased process variation adds to the reliability challenge, resulting in over designs and extra timing margins at the cost of power consumption, silicon area and performance degradation. We present efficient in situ error detection techniques to exploit datapath characteristics for monitoring circuit errors: pre-edge checking in non-critical paths without hold time constraints; post-edge checking in critical paths without sacrificing performance; and cross-edge checking in moderate paths for the optimal trade-off. The techniques are all realized using the inherent redundancy within a conventional flip-flop design and do not require any logic or sample duplication as done by most existing methods. The detection-enabled flip-flop is implemented using only 31 transistors as a competitive and low-cost solution. Yaoyu Tao, Zhengya Zhang |
ISCAS | 3 |
| 2013 | Minimum supply voltage for sequential logic circuits in a 22nm technologyabstractThe minimum supply voltage (Vmin) is explored for sequential logic circuits by statistically simulating the impact of within-die process variations and gate-dielectric soft breakdown on data retention and hold time. As supply voltage (Vcc) scales, statistical circuit simulations demonstrate that hold time increases faster than circuit delay or cycle time, consequently the required number of min-delay buffers increases. For this reason, a new hold-time violation metric defines Vminas the Vccin which the hold time exceeds a target percentage of the cycle time. Simulation results in a 22nm tri-gate CMOS technology indicate a data-retention Vminof 0.61Vnorm and a hold-time Vminof 0.73Vnorm, where Vnormrepresents a normalized voltage for the process technology node. A key insight reveals that upsizing the first clock inverter in the sequential circuit reduces the hold-time Vminby 18% and the overall Vminby 16%. Keith A. Bowman, Charles Augustine, Zhengya Zhang, James W. Tschanz |
ISLPED | 4 |
| 2012 | Reconfigurable architecture and automated design flow for rapid FPGA-based LDPC code emulationabstractMultitude of design freedoms of LDPC codes and practical decoders require fast simulations. FPGA emulation is attractive but inaccessible due to its design complexity. We propose a library and script based approach to automate the construction of FPGA emulations. Code parameters and design parameters are programmed either during run time or by script in design time. We demonstrate the architecture and design flow using the LDPC codes for the latest wireless communication standards: each emulation model was auto-constructed within one minute and the peak emulation throughput reached 3.8 Gb/s on a BEE3 platform. Youn Sung Park, Zhengya Zhang |
FPGA | 3 |
| 2012 | High-throughput architecture and implementation of regular (2, dc) nonbinary LDPC decodersabstractNonbinary LDPC codes have shown superior performance, but decoding nonbinary codes is complex, incurring a long latency and a much degraded throughput. We propose a low-latency variable processing node by a skimming algorithm, together with a low-latency extended min-sum check processing node by prefetching and relaxing redundancy control. The processing nodes are jointly designed for an optimal pipeline schedule. This low-latency, high-throughput architecture is applied to a class of high-performance (2, dc)-regular nonbinary LDPC codes constructed based on their binary images. A conflict-free memory is proposed to resolve data hazards caused by the non-structured nature of these codes. A complete (2, 4)-regular, (960, 480) GF(64) nonbinary LDPC decoder is demonstrated on a Xilinx Virtex-5 FPGA. The decoder delivers an excellent error-correcting performance at a 9.76 Mb/s coded throughput, representing a significant improvement of state-of-the-art extended min-sum decoder implementations. Yaoyu Tao, Youn Sung Park, Zhengya Zhang |
ISCAS | 3 |
| 2011 | A confidence-driven model for error-resilient computingabstractWe propose an adaptive reliability enhancement structure for deeply-scaled CMOS and future devices that exhibit nondeterministic behavior. This structure forms the basis of a confidence-driven computing model that can be implemented in either a rollback recovery or an iterative dual modular redundancy method incorporating synchronous handshake schemes. The performance and cost of the computing model are estimated using a 45 nm CMOS technology and the functionality is verified by FPGA-based emulation. The confidence-driven computing model is demonstrated using a 16-bit, 12-stage CORDIC processor operating under random, transient errors. The confidence-driven computing model adapts to the fluctuating error rates at the device substrate level to guarantee the reliability of computation at the system level. This computing model costs 4.2 times smaller area and 2.7 times less energy overhead than triple modular redundancy to guarantee a system-level mean time to failure of two years. Yejoong Kim, Zhengya Zhang, David T. Blaauw, Dennis Sylvester, Helia Naeimi, Sumeet Sandhu |
DATE | 3 |
| 2011 | Hardware acceleration of iterative image reconstruction for X-ray computed tomographyabstractX-ray computed tomography (CT) images could be improved using iterative image reconstruction if the 3D cone-beam forward- and back-projection computations can be accelerated significantly. We investigated the feasibility of a field-programmable gate array (FPGA) implementation of the separable footprint (SF) forward projector. A 16-bit fixed-point quantization introduces negligible numerical errors without affecting the perceptual image quality. The SF-based 3D cone-beam projector can be efficiently parallelized and its memory bandwidth reduced by exploiting projection geometry and data locality. We demonstrate a fully pipelined, 75-way parallel hardware architecture of the SF forward projector on a Xilinx Virtex-5 FPGA that can complete one forward projection of a 320×320×61 object over 3,625 views in 6.3 seconds. Jung Kuk Kim, Zhengya Zhang, Jeffrey A. Fessler |
ICASSP | 2 |
| 2011 | LDPC decoder architecture for high-data rate personal-area networksabstractEmerging standards for wireless communications in the 60GHz band, such as WiGig, IEEE 802.11ad, and IEEE 802.15.3c, require throughputs between 1.5 and 6Gb/s and use rate adaptive low-density parity-check (LDPC) codes as the main form of forward error correction. State-of-the-art flexible LDPC decoders cannot simultaneously achieve the high throughput mandated by these standards and the low power needed for mobile applications. This work develops a flexible, fully pipelined architecture for the IEEE 802.11ad standard capable of achieving both goals. Based on a decoder synthesized in a low-power 65nm CMOS technology, the decoder dissipates 42mW at the 1.5Gb/s throughput and 84mW at the 3Gb/s throughput for the worst-case matrix in the standard. Matthew Weiner, Borivoje Nikolic, Zhengya Zhang |
ISCAS | 3 |
| 2011 | Absorbing set spectrum approach for practical code designabstractThis paper focuses on controlling the absorbing set spectrum for a class of regular LDPC codes known as separable, circulant-based (SCB) codes. For a specified circulant matrix, SCB codes all share a common mother matrix, examples of which are array-based LDPC codes and many common quasi-cyclic codes. SCB codes retain the standard properties of quasi-cyclic LDPC codes such as girth, code structure, and compatibility with efficient decoder implementations. In this paper, we define a cycle consistency matrix (CCM) for each absorbing set of interest in an SCB LDPC code. For an absorbing set to be present in an SCB LDPC code, the associated CCM must not be full column-rank. Our approach selects rows and columns from the SCB mother matrix to systematically eliminate dominant absorbing sets by forcing the associated CCMs to be full column-rank. We use the CCM approach to select rows from the SCB mother matrix to design SCB codes of column weight 5 that avoid all low-weight absorbing sets (4; 8), (5; 9), and (6; 8). Simulation results demonstrate that the newly designed code has a steeper error-floor slope and provides at least one order of magnitude of improvement in the low error rate region as compared to an elementary array-based code. Lara Dolecek, Zhengya Zhang, Richard D. Wesel |
ISIT | 3 |
| 2010 | Analysis of absorbing sets and fully absorbing sets of array-based LDPC codesabstractThe class of low-density parity-check (LDPC) codes is attractive, since such codes can be decoded using practical message-passing algorithms, and their performance is known to approach the Shannon limits for suitably large block lengths. For the intermediate block lengths relevant in applications, however, many LDPC codes exhibit a so-called “error floor,” corresponding to a significant flattening in the curve that relates signal-to-noise ratio (SNR) to the bit-error rate (BER) level. Previous work has linked this behavior to combinatorial substructures within the Tanner graph associated with an LDPC code, known as (fully) absorbing sets. These fully absorbing sets correspond to a particular type of near-codewords or trapping sets that are stable under bit-flipping operations, and exert the dominant effect on the low BER behavior of structured LDPC codes. This paper provides a detailed theoretical analysis of these (fully) absorbing sets for the class of$C_{p, \gamma}$array-based LDPC codes, including the characterization of all minimal (fully) absorbing sets for the array-based LDPC codes for$\gamma = 2,3,4$, and moreover, it provides the development of techniques to enumerate them exactly. Theoretical results of this type provide a foundation for predicting and extrapolating the error floor behavior of LDPC codes. Lara Dolecek, Zhengya Zhang, Venkat Anantharam, Martin J. Wainwright, Borivoje Nikolic |
IEEE Trans. Inf. Theory | 2 |
| 2009 | Predicting error floors of structured LDPC codes: deterministic bounds and estimatesabstractThe error-correcting performance of low-density parity check (LDPC) codes, when decoded using practical iterative decoding algorithms, is known to be close to Shannon limits for codes with suitably large blocklengths. A substantial limitation to the use of finite-length LDPC codes is the presence of an error floor in the low frame error rate (FER) region. This paper develops a deterministic method of predicting error floors, based on high signal-to-noise ratio (SNR) asymptotics, applied to absorbing sets within structured LDPC codes. The approach is illustrated using a class of array-based LDPC codes, taken as exemplars of high-performance structured LDPC codes. The results are in very good agreement with a stochastic method based on importance sampling which, in turn, matches the hardware-based experimental results. The importance sampling scheme uses a mean-shifted version of the original Gaussian density, appropriately centered between a codeword and a dominant absorbing set, to produce an unbiased estimator of the FER with substantial computational savings over a standard Monte Carlo estimator. Our deterministic estimates are guaranteed to be a lower bound to the error probability in the high SNR regime, and extend the prediction of the error probability to as low as 10-30. By adopting a channel-independent viewpoint, the usefulness of these results is demonstrated for both the standard Gaussian channel and a channel with mixture noise. Lara Dolecek, Pamela Lee, Zhengya Zhang, Venkat Anantharam, Borivoje Nikolic, Martin J. Wainwright |
IEEE J. Sel. Areas Commun. | 3 |
| 2009 | Design of LDPC decoders for improved low error rate performance: quantization and algorithm choicesabstractMany classes of high-performance low-density parity-check (LDPC) codes are based on parity check matrices composed of permutation submatrices. We describe the design of a parallel-serial decoder architecture that can be used to map any LDPC code with such a structure to a hardware emulation platform. High-throughput emulation allows for the exploration of the low bit-error rate (BER) region and provides statistics of the error traces, which illuminate the causes of the error floors of the (2048, 1723) Reed-Solomon based LDPC (RS-LDPC) code and the (2209, 1978) array-based LDPC code. Two classes of error events are observed: oscillatory behavior and convergence to a class of non-codewords, termed absorbing sets. The influence of absorbing sets can be exacerbated by message quantization and decoder implementation. In particular, quantization and the log-tanh function approximation in sum-product decoders Zhengya Zhang, Lara Dolecek, Borivoje Nikolic, Venkat Anantharam, Martin J. Wainwright |
IEEE Trans. Commun. | 1 |
| 2008 | Lowering LDPC Error Floors by PostprocessingabstractA class of combinatorial structures, called absorbing sets, strongly influences the performance of low-density parity-check (LDPC) decoders at low error rates. Past experiments have shown that a class of (8,8) absorbing sets determines the error floor performance of the (2048,1723) Reed-Solomon based LDPC code (RS-LDPC). A postprocessing approach is formulated to exploit the structure of the absorbing set by biasing the reliabilities of selected messages in a message-passing decoder. The approach converges quickly and can be efficiently implemented with minimal overhead. Hardware emulation of the decoder with postprocessing shows more than two orders of magnitude improvement in the very low bit error rate performance and error- floor-free operation below a BER of 10-12. Zhengya Zhang, Lara Dolecek, Borivoje Nikolic, Venkat Anantharam, Martin J. Wainwright |
GLOBECOM | 1 |
| 2008 | Error floors in LDPC codes: Fast simulation, bounds and hardware emulationabstractAbstract — The error-correcting performance of low-density parity check (LDPC) codes, when decoded using practical iterative decoding, is known to approach Shannon limits in the asymptotic limit of large blocklengths. A substantial limitation to the use of finite-length LDPC codes is the presence of an error floor in the low frame error rate (FER) region. This paper develops a method, based on importance sampling and high SNR asymptotics as applied to suitably defined absorbing structures within the LDPC code, to predict error floors. Our results are in very close agreement with hardware-based experimental results, and moreover extend the prediction of the error probability to even lower regions. We compute both importance sampling estimates of error probabilities and deterministic estimates that are guaranteed to lower bound the error probability in the high SNR regime. I. Pamela Lee, Lara Dolecek, Zhengya Zhang, Venkat Anantharam, Borivoje Nikolic, Martin J. Wainwright |
ISIT | 3 |
| 2007 | Analysis of Absorbing Sets for Array-Based LDPC CodesabstractLow density parity check codes (LDPC) are known to perform very well under iterative decoding. However, these codes also exhibit a change in the slope of the bit error rate (BER) vs. signal to noise ratio (SNR) curve in the very low BER region. In our earlier work using hardware emulation in this deep BER regime we argue that this behavior can be attributed to specific structures within the Tanner graph associated with an LDPC code, called absorbing sets. In this paper we provide a detailed theoretical analysis of absorbing sets for array-based LDPC codes Cp.gamma. Specifically, we identify and enumerate all the smallest absorbing sets for these array-based LDPC codes with gamma = 2,3,4 with standard parity check matrix. Experiments carried out on the emulation platform show excellent agreement with our theoretical results. Lara Dolecek, Zhengya Zhang, Venkat Anantharam, Martin J. Wainwright, Borivoje Nikolic |
ICC | 2 |
| 2007 | Quantization Effects in Low-Density Parity-Check DecodersabstractA. class of combinatorial structures, called absorbing sets, strongly influences the performance of low-density parity- check (LDPC) decoders. In particular, the quantization scheme strongly affects which absorbing sets dominate in the error-floor region. Absorbing sets may be characterized as weak or strong. They are a characteristic of the parity check matrix of a code. Conventional quantization schemes applied to a (2209,1978) array-based LDPC code can induce low-weight weak absorbing sets and, as a result, elevate the error floor. Adaptive quantization schemes alleviate the effects of weak absorbing sets, and, as a result, only the strong ones dominate the error floor of an optimized decoder implementation. Another benefit of an adaptive quantization scheme is that it performs well even in very few iterations. Zhengya Zhang, Lara Dolecek, Martin J. Wainwright, Venkat Anantharam, Borivoje Nikolic |
ICC | 1 |
| 2006 | Investigation of Error Floors of Structured Low-Density Parity-Check Codes by Hardware EmulationabstractSeveral high performance LDPC codes have parity-check matrices composed of permutation submatrices. We design a parallel-serial architecture to map the decoder of any structured LDPC code in this large family to a hardware emulation platform. A peak throughput of 240 Mb/s is achieved in decoding the (2048,1723) Reed-Solomon based LDPC (RS-LDPC) code. Experiments in the low bit error rate (BER) region provide statistics of the error traces, which are used to investigate the causes of the error floor. In a low precision implementation, the error floors are dominated by the fixed-point decoding effects, whereas in a higher precision implementation the errors are attributed to special configurations within the code, whose effect is exacerbated in a fixed-point decoder. This new characterization leads to an improved decoding strategy and higher performance. Zhengya Zhang, Lara Dolecek, Borivoje Nikolic, Venkat Anantharam, Martin J. Wainwright |
GLOBECOM | 1 |