EDBT 2026 Demo / reviewers in the wild / expert
Ray C. C. Cheung
dblp:96/2421 · also Chak-Chung Cheung, Ray Chak-Chung Cheung
· DBLP profile ↗
119ranked-venue papers
10as first author
46since 2021 · last 2026
0000-0002-6764-0729ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 79 · 9 first-author · 26 since 2021Artificial intelligence and machine learning · 14 · 7 since 2021Computer networks · 9 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 since 2021Security and privacy · 4 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MoToRec: Sparse-Regularized Multimodal Tokenization for Cold-Start RecommenderabstractGraph neural networks (GNNs) have revolutionized recommender systems by effectively modeling complex user-item interactions, yet data sparsity and the item cold-start problem significantly impair performance, particularly for new items with limited or no interaction history. While multimodal content offers a promising solution, existing methods result in suboptimal representations for new items due to noise and entanglement in sparse data. To address this, we transform multimodal recommendation into discrete semantic tokenization. We present Sparse-Regularized Multimodal Tokenization for Cold-Start Recommender Systems (MoToRec), a framework centered on a sparsely-regularized Residual Quantized Variational Autoencoder (RQ-VAE) that generates a compositional semantic code of discrete, interpretable tokens, promoting disentangled representations. MoToRec’s architecture is enhanced by three synergistic components: (1) a sparsely regularized RQ-VAE that promotes disentangled representations, (2) a novel adaptive rarity amplification that promotes prioritized learning for cold-start items, and (3) a hierarchical multi-source graph encoder for robust signal fusion with collaborative signals. Extensive experiments on three large-scale datasets demonstrate MoToRec’s superiority over state-of-the-art methods in both overall and cold-start scenarios. Our work validates that discrete tokenization provides an effective and scalable alternative for mitigating the long-standing cold-start challenge. Ray C. C. Cheung |
AAAI | 3 |
| 2026 | Chariot: Compiler-Aware Heterogeneous Graph Representation Learning for Automated HLS OptimizationabstractHigh-level synthesis (HLS) design space exploration (DSE) aims to find Pareto-optimal designs but is hindered by slow synthesis evaluations. Existing graph neural network (GNN) surrogates struggle with homogeneous-style graph representations (causing signal over-squashing) and imprecise source-level heuristics for pragma mapping. We propose Chariot, an automated HLS optimization framework. Chariot leverages LLVM-based static analysis for high-fidelity Use-Def chain tracking, modeling HLS designs as semantic-rich heterogeneous graphs that explicitly map directives to true hardware targets. Our framework achieves state-of-the-art QoR prediction, identifying Pareto-optimal solutions with drastically reduced ranking regret while delivering orders-of-magnitude DSE speedup. Jierui Liu, Yuhan She, Rongliang Fu, Tsung-Yi Ho, Hong Yan 0001, Ray C. C. Cheung |
FCCM | 7 |
| 2026 | ViM-Q: Scalable Algorithm-Hardware Co-Design for Vision Mamba Model Inference on FPGAabstractVision Mamba (ViM) models offer a compelling efficiency advantage over Transformers by leveraging the linear complexity of State Space Models (SSMs), yet efficiently deploying them on FPGAs remains challenging. Linear layers struggle with dynamic activation outliers that render static quantization ineffective, while uniform quantization fails to capture the weight distribution at low bit-widths. Furthermore, while associative scan accelerates SSMs on GPUs, its memory access patterns are misaligned with the streaming dataflow required by FPGAs. To address these challenges, we present ViM-Q1, a scalable algorithm-hardware co-design for end-to-end ViM inference on the edge. We introduce a hardware-aware quantization scheme combining dynamic per-token activation quantization and per-channel smoothing to mitigate outliers, alongside a custom 4-bit per-block Additive Power-of-Two (APoT) weight quantization. The models are deployed on a runtime-parameterizable FPGA accelerator featuring a linear engine employing a Lookup-Table (LUT) unit to replace multiplications with shift-add operations, and a fine-grained pipelined SSM engine that parallelizes the state dimension while preserving sequential recurrence. Crucially, the hardware supports runtime configuration, adapting to diverse dimensions and input resolutions across the ViM family. Implemented on an AMD ZCU102 FPGA, ViM-Q achieves an average 4.96× speedup and 59.8× energy efficiency gain over a quantized NVIDIA RTX 3090 GPU baseline for low-batch inference on ViM-tiny. This co-design shows a viable path for deploying ViM models on resource-constrained edge devices. Shengzhe Lyu, Yuhan She, Patrick S. Y. Hung, Ray C. C. Cheung, Weitao Xu |
FCCM | 4 |
| 2026 | ViM-Q: Energy Efficient Algorithm-Hardware Co-Design for Dynamically Quantized Vision Mamba ModelsabstractState-space models (SSMs), such as Mamba, provide an efficient alternative to Transformers for vision tasks by replacing their quadratic-cost self-attention with linear complexity state update. However, efficiently deploying Vision Mamba (ViM) models on FPGA platforms is challenging, as the latency is dominated by two key components: linear layers and the selective SSM. For the linear layers, highly dynamic activation outliers across tokens render conventional static quantization techniques ineffective. Meanwhile, while the associative scan algorithm is effective in accelerating SSM on GPUs, its data access pattern is fundamentally mismatched with FPGA architectures when mapping the model's inherently sequential recurrence, creating a critical dataflow bottleneck. Shengzhe Lyu, Yuhan She, Patrick S. Y. Hung, Ray C. C. Cheung, Weitao Xu |
FPGA | 4 |
| 2026 | SNTT: A Sparsity-aware NTT Accelerator Based on FPGA for Zero-Knowledge Proof
Zihang Guo, Qiang Liu 0011, Ray C. C. Cheung, Zhaohui Guo |
ISCAS | 3 |
| 2026 | SR-TCUR: Scalable and robust tubal CUR decomposition for large-scale multidimensional tensors
Muhammad A. A. Abdelgawad, Ray C. C. Cheung, Hong Yan 0001 |
Neurocomputing | 2 |
| 2026 | Memory-efficient neural network training via gradient compression through continuous basis tracking
Xinmin Meng, Muhammad A. A. Abdelgawad, Ray C. C. Cheung, Hong Yan 0001 |
Neurocomputing | 4 |
| 2026 | Less is More: Latent Diffusion for Efficient IoT Side-Channel AnalysisabstractThe proliferation of cryptographic primitives in resource-constrained Internet of Things (IoT) devices has made them prime targets for Side-Channel Analysis (SCA). However, designing effective defenses against these attacks has become increasingly complex, as traditional deep learning approaches rely heavily on extensive profiling datasets that are difficult to obtain in the context of widely distributed and physically restricted IoT environments. This challenge is further exacerbated by countermeasures such as clock jitter and random delays. To overcome this limitation, this paper introduces a novel and data-efficient three-stage framework for generating high-fidelity synthetic side-channel traces. First, we employ a Supervised Variational Autoencoder (S-VAE) to map noisy, high-dimensional raw traces into a compact and denoised latent space, effectively creating an information-rich manifold. Second, a conditional Denoising Diffusion Implicit Model (DDIM), powered by an advanced attention-augmented U-Net, is trained exclusively on this computationally tractable latent space to learn the complex conditional data distribution. Finally, we empirically validate our framework on the public ASCAD benchmark and ChipWhisperer CW308T UFO platform. The results are compelling: an attack model trained solely on our synthetic data successfully recovers the secret key in the most challenging ASCAD_desync100 scenario using only 4107 traces and using only 2560 traces, 97.7% accuracy can be achieved on the Chipwhisphere platform. This work provides a practical and efficient pathway for the robust security evaluation of cryptographic implementations in data-scarce IoT environments, significantly lowering the barrier for thorough side-channel vulnerability analysis. Donald Donglong Chen, Wangchen Dai, Jinfa Hong, Yu Hin Chan, Çetin Kaya Koç, Patrick S. Y. Hung, Ray C. C. Cheung |
IEEE Internet Things J. | 8 |
| 2026 | HTCNN: High-Throughput Batch CNN Inference With Homomorphic EncryptionabstractHomomorphic Encryption (HE) technology allows for processing encrypted data, breaking through data isolation barriers and providing a promising solution for privacy-preserving computation. The integration of HE technology into Convolutional Neural Network (CNN) inference shows potential in addressing privacy issues in identity verification, medical imaging diagnosis, and various other applications. The CKKS HE algorithm stands out as a popular option for homomorphic CNN inference due to its capability to handle real number computations. However, challenges such as computational delays and resource overhead present significant obstacles to the practical implementation of homomorphic CNN inference, largely due to the complex nature of HE operations. In addition, current methods for speeding up homomorphic CNN inference primarily address individual images or large batches of input images, lacking a solution for efficiently processing a moderate number of input images with fast homomorphic inference capabilities. In response to these challenges, we introduce a novel leveled homomorphic CNN inference scheme aimed at reducing latency and improving throughput using the CKKS scheme. Our proposed inference strategy involves mapping multiple inputs to a set of ciphertext by exploiting the sliding window properties of convolutions to utilize CKKS's inherent Single-Instruction-Multiple-Data (SIMD) capability. To mitigate the delay associated with homomorphic CNN inference, we introduce optimization techniques, including mask-weight merging, rotation multiplexing, stride convolution segmentation, and folding rotations. The efficacy of our homomorphic inference scheme is demonstrated through evaluations carried out on the MNIST and CIFAR-10 datasets. Specifically, results from the MNIST dataset on a single CPU thread show that inference for 163 images can be completed in 10.4 seconds with an accuracy of 98.95%, which is a 6.9× throughput improvement over state-of-the-art works. Comparative analysis with existing methodologies highlights the superior performance of our proposed inference scheme in terms of latency, throughput, communication overhead, and memory utilization. Zewen Ye, Tianyu Wang 0037, Tianshun Huang, Yonggen Li, Chengxuan Wang, Ray C. C. Cheung, Kejie Huang |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2026 | SwiftChannel: Algorithm-Hardware Co-Design for Deep Learning-Based 5G Channel EstimationabstractChannel estimation is crucial in 5G communication networks for optimizing transmission parameters and ensuring reliable, high-speed communication. However, the use of multiple-input and multiple-output (MIMO) and millimeter-wave (mmWave) in 5G networks presents challenges in achieving accurate estimation under strict latency requirements on resource-limited hardware platforms. To address these challenges, we proposeSwiftChannel, an algorithm-hardware co-design framework that integrates a hardware-friendly deep learning-based channel estimator with a dedicated accelerator. Our approach employs a convolutional neural network enhanced with a parameter-free attention mechanism, which effectively reconstructs full-resolution spatial-frequency domain channel matrices from low-resolution least squares (LS) estimates. We further develop a multi-stage model compression pipeline combining knowledge distillation, convolution re-parameterization, and quantization-aware training, resulting in substantial model size reduction with negligible accuracy loss. The hardware accelerator, implementing the compressed model and the LS estimator on FPGA platforms using High-level Synthesis (HLS), features a fine-grained pipeline architecture and optimized dataflow strategies. Tested on a Zynq UltraScale+ RFSoC, the accelerator achieves sub-millisecond latency, providing up to 24x speed-up and over 33x improvement in energy efficiency compared to GPU-based solutions. Extensive evaluations demonstrate that the proposed design generalizes not only across various noise levels and user mobilities, but also to a variety of unseen channel profiles, outperforming state-of-the-art baselines. By unifying algorithmic innovation with hardware-aware design, our work presents a future-proof channel estimation solution for 5G MIMO systems. The source codes for the dataset synthesis, deep learning algorithm, and HLS-based FPGA design are accessible via GitHub. Shengzhe Lyu, Yuhan She, Di Duan, Tao Ni 0003, Yu Hin Chan, Chengwen Luo 0001, Ray C. C. Cheung, Weitao Xu |
IEEE Trans. Mob. Comput. | 7 |
| 2025 | CAHLS: Source-to-Source Transformation to Generate Cycle Accurate Models for High-Level SynthesisabstractHigh-Level Synthesis (HLS) empowers the ability to synthesize a customized hardware description from an untimed software description. However, the quality of the generated hardware is affected by the HLS tool. Current state-of-the-art commercial HLS tools adopt static-scheduling-based algorithms, which perform well for the regular designs but suffer performance degradation for the control-dominant designs. Dynamic scheduling, on the other hand, performs well for control flows but loses certain optimizations, like resource sharing and critical path optimizations, resulting in area overhead and frequency drop. In this paper, we propose a source-to-source transformation to generate an equivalent pseudo cycle-accurate model, so that 1) the transformed code runs dynamically based on different control conditions, and 2) the transformed code still fits in the static HLS tool. As future work, this transformation can be integrated into a compiler to automatically optimize the control-dominant designs in the static-scheduling HLS flow. Yuhan She, Jierui Liu, Ray C. C. Cheung, Hong Yan 0001 |
CODES+ISSS | 4 |
| 2025 | FastViT: Real-Time Linear Attention Accelerator for Dense Predictions of Vision Transformer (ViT)abstractThe commercial success of generative artificial intelligence (GenAI) has driven an exponential surge in demand for real-time inference in Vision Transformer (ViT) applications, including latency-sensitive domains in autonomous driving, medical imaging and computational photography. This paper introduces FastViT, a high-performance and energy-efficient hardware accelerator for emerging kernel function-based linear attention mechanisms. By leveraging cost-efficient multiplication, mixed-precision quantisation and optimised data flow, FastViT improves real-time performance for high-resolution dense prediction tasks. Compared to existing approaches, experiments demonstrate that FastViT achieves higher throughput and energy efficiency while maintaining negligible accuracy degradation and balanced resource allocation. In the future, we will improve its scalability for next-generation hardware equipped with advanced DSP cores. Zhuoheng Ran, Zewen Ye, Chong Wu 0007, Ray C. C. Cheung, Hong Yan 0001 |
ISCAS | 4 |
| 2025 | Efficient CUR decomposition for interpretable low-rank approximations and imaging applications
Muhammad A. A. Abdelgawad, Ray C. C. Cheung, Hong Yan 0001 |
Neurocomputing | 2 |
| 2025 | High-Radix/Mixed-Radix NTT Multiplication Algorithm/Architecture Co-Design Over Fermat ModulusabstractPolynomial multiplication using Number Theoretic Transform (NTT) is crucial in lattice-based post-quantum cryptography (PQC) and fully homomorphic encryption (FHE), with modulusqsignificantly affecting performance. Fermat moduli of the form$2^{2^{n}} + 1$, such as 65537, offer efficiency gains due to simplified modular reduction and powers-of-2 twiddle factors in NTT. While Fermat moduli have been directly applied or explored for incorporation into existing schemes, Fermat NTT-based polynomial multiplication designs remain underexplored in fully exploiting the benefits of Fermat moduli. This work presents a high-radix/mixed-radix NTT architecture tailored for Fermat moduli, which improves the utilization of the powers-of-2 twiddle factors in large transform sizes. In most cases, our design achieves a 30%–85% reduction in DSP area-time product (ATP) and a 70%–100% reduction in BRAM ATP compared to state-of-the-art designs with smaller or equivalent modulus, while maintaining competitive LUT and FF ATP, underscoring the potential of Fermat NTT-based polynomial multipliers in lattice-based cryptography. Yile Xing, Guangyan Li, Zewen Ye, Ryan W. L. Luk, Donald Donglong Chen, Hong Yan 0001, Ray C. C. Cheung |
IEEE Trans. Computers | 7 |
| 2025 | PQNTRU: Acceleration of NTRU-Based Schemes via Customized Post-Quantum ProcessorabstractPost-quantum cryptography (PQC) has rapidly evolved in response to the emergence of quantum computers, with the US National Institute of Standards and Technology (NIST) selecting four finalist algorithms for PQC standardization in 2022, including the Falcon digital signature scheme. Hawk is currently the only lattice-based candidate in NIST Round 2 additional signatures. Falcon and Hawk are based on the NTRU lattice, offering compact signatures, fast generation, and verification suitable for deployment on resource-constrained Internet-of-Things (IoT) devices. Despite the popularity of ML-DSA and ML-KEM, research on NTRU-based schemes has been limited due to their complex algorithms and operations. Falcon and Hawk's performance remains constrained by the lack of parallel execution in crucial operations like the Number Theoretic Transform (NTT) and Fast Fourier Transform (FFT), with data dependency being a significant bottleneck. This paper enhances NTRU-based schemes Falcon and Hawk through hardware/software co-design on a customized Single-Instruction-Multiple-Data (SIMD) processor, proposing new SIMD hardware units and instructions to expedite these schemes along with software optimizations to boost performance. Our NTT optimization includes a novel layer merging technique for SIMD architecture to reduce memory accesses, and the use of modular algorithms (Signed Montgomery and Improved Plantard) targets various modulus data widths to enhance performance. We explore applying layer merging to accelerate fixed-point FFT at the SIMD instruction level and devise a dual-issue parser to streamline assembly code organization to maximize dual-issue utilization. A System-on-chip (SoC) architecture is devised to improve the practical application of the processor in real-world scenarios. Evaluation on 28$nm$technology and field programmable gate array (FPGA) platform shows that our design and optimizations can increase the performance of Hawk signature generation and verification by over 7$\times$. Zewen Ye, Junhao Huang 0001, Tianshun Huang, Yudan Bai, Guangyan Li, Donald Donglong Chen, Ray C. C. Cheung, Kejie Huang |
IEEE Trans. Computers | 9 |
| 2025 | A Novel AIoT-Based and User Behavior-Driven Dockless Bike-Sharing Management System for Chaotic Operations in a Condensed CityabstractRapid urbanization and the rising demand for sustainable mobility have accelerated the adoption of dockless bike-sharing systems. However, these systems frequently encounter challenges such as oversupply, inefficient resource allocation, and limited responsiveness to localized demand. To address these issues, this paper proposes the novel AIoT-based and user behavior-driven management system (AUMS)—a comprehensive AIoT-based framework that integrates demand prediction, dynamic clustering, and rebalancing algorithms to optimize the management of dockless bike-sharing services. Unlike existing approaches that focus narrowly on demand prediction, AUMS combines real-time IoT sensor data with user behavior analytics to support continuous, data-driven decision-making. The system employs a closed-loop control structure, enabling the dynamic reconfiguration of bike distribution in response to shifting urban mobility patterns. Its architecture includes three core modules: 1) demand prediction based on spatiotemporal behavioral data, 2) dynamic clustering to localize operational zones, and 3) intelligent rebalancing to minimize idle rates and unmet demand. The significance of this work lies in its holistic design and rigorous real-world validation. Over a 15-month period, AUMS was deployed in three escalating phases: a localized pilot, a district-level deployment in Tseung Kwan O, and a city-wide experiment across 13 districts in Hong Kong. Field results demonstrate notable performance improvements, including a 10.8% increase in bike utilization, 11.6% growth in trip frequency, and a 29.2% reduction in unmet demand. Ken Chun Ho Ching, Chaoqiang Jiang, Hassan Chun Wai Ching, Steve Kwan Po Ng, Ray C. C. Cheung, Haoliang Li, Alan H. F. Lam |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | Randomized tensor decomposition using parallel reconfigurable systemsabstractAbstract Tensor decomposition algorithms are essential for extracting meaningful latent variables and uncovering hidden structures in real-world data tensors. Unlike conventional deterministic tensor decomposition algorithms, randomized methods offer higher efficiency by reducing memory requirements and computational complexity. This paper proposes an efficient hardware architecture for a randomized tensor decomposition implemented on a field-programmable gate array (FPGA) using high-level synthesis (HLS). The proposed architecture integrates random projection, power iteration, and subspace approximation via QR decomposition to achieve low-rank approximation of multidimensional datasets. The proposed architecture utilizes the capabilities of reconfigurable systems to accelerate tensor computation. It includes three central units: (1) tensor times matrix chain (TTMc), (2) tensor unfolding unit, and (3) QR decomposition unit to implement a three-stage algorithm. Experimental results demonstrate that our FPGA design achieves up to 14.56 times speedup compared to the well-implemented tensor decomposition using software library Tensor Toolbox on an Intel i7-9700 CPU. For a large input tensor of size $$512 \times 512 \times 512$$ 512 × 512 × 512 , the proposed design achieves a 5.55 times speedup compared to an Nvidia Tesla T4 GPU. Furthermore, we utilize our hardware-based high-order singular value decomposition (HOSVD) accelerator for two real applications: background subtraction of dynamic video datasets and data compression. In both applications, our proposed design shows high efficiency regarding accuracy and computational time. Ajita Misra, Muhammad A. A. Abdelgawad, Ray C. C. Cheung, Hong Yan 0001 |
J. Supercomput. | 4 |
| 2025 | IncTSVD: Incremental Tensor Singular Value Decomposition of Multidimensional Streaming DataabstractIn this article, we develop an online method called IncTSVD to incrementally compute the tensor singular value decomposition (TSVD) of a given sequence of third-order tensors based on the tensor-tensor concept. This can be considered an extension of incremental SVD based on updating matrices to tensors. IncTSVD is suitable for streamed tensor data and where memory resources are limited. Most existing methods to compute TSVD focus on approximating it using randomized or sketching techniques in a batch setting to decrease the storage and computational costs required. The IncTSVD extends the computation of TSVD to streaming by maintaining the basis tensors of previously arrived data and incrementally updating the approximation using the tensor of incoming data. The computational cost and approximation error of the proposed method were analyzed theoretically and through extensive numerical experiments, which included using synthetic and real-world datasets under streaming scenarios. The IncTSVD method was superior to existing deterministic and randomized tensor decompositions (TDs) based on the t-product for computational and storage costs, and had comparable accuracy to the standard TSVD method. Muhammad A. A. Abdelgawad, Ray C. C. Cheung, Hong Yan 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | A Speculative Loop Pipeline Framework with Accurate Path Modeling for High-Level SynthesisabstractLoop pipelining is a key optimization in high-level synthesis (HLS), aimed at overlapping the execution of iterations. Static scheduling, dominant in commercial HLS tools, configures the pipeline based on compile-time analysis, proving conservative for designs with irregular control flow and memory access due to imbalanced recurrences. Speculative Loop pipeline (SLP) is a novel concept that addresses the problem by introducing the speculation and recovery mechanism at the source level to improve the throughput. Although proven promising, it has a significant gap from practical application: It requires accurate early-stage modeling of the pipeline configuration for each path, which is unable to obtain with classic HLS scheduling methods because the SLP process itself interferes with the path length. In this work, we made a step forward by proposing a practical SLP framework with accurate path modeling ability through iterative tuning. We further optimize the SLP technology by combining automatic dataflow extraction with speculative source-level transformation to further boost the performance in specific design patterns. Our framework works on the source level and is easy to be plugged into existing downstream HLS tools. Experiment results demonstrate significant performance improvements over commercial HLS tools and better resource trade-offs compared to the state-of-the-art dynamic-scheduling-based solutions. Yuhan She, Jierui Liu, Ray C. C. Cheung, Hong Yan 0001 |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2025 | RVSLH: Acceleration of Postquantum Standard SLH-DSA With Customized RISC-V ProcessorabstractPostquantum cryptography (PQC) has developed quickly in response to the rise of quantum computers. The US National Institute of Standards and Technology (NIST) recently released three PQC standards, one of which is the hash-based standard stateless hash-based digital signature standard (SLH-DSA), built on SPHINCS+ selected during the NIST Round 3 submissions. Despite its potential, SLH-DSA’s performance is hindered by inefficient execution in the hash function and extensive memory accesses, with the data dependency of the hash function presenting a notable bottleneck. This brief aims to enhance the efficiency of SHAKE-based SLH-DSA schemes using hardware/software co-design on a customized RISC-V processor. We incorporate tightly coupled hardware units and instructions on RISC-V to expedite SLH-DSA, coupled with memory optimizations to enhance overall performance. The contributions of this brief are twofold. First, our design introduces customized single-instruction-multiple-data (SIMD) instructions and corresponding computation hardware units to accelerate Keccak, the essential operation of SHAKE256. In addition, our design streamlines hash operation memory accesses by reorganizing memory space. The proposed processor’s performance is evaluated on Artix-7 field-programmable gate array (FPGA) and 28-nm technology application-specific integrated circuit (ASIC), demonstrating approximately 15 times acceleration compared with the baseline design. Furthermore, it surpasses recent state-of-the-art works in terms of the performance-power–area product. Zewen Ye, Xin Li 0177, Chuhui Wang, Ray C. C. Cheung, Kejie Huang |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2024 | MSCA: A Multi-Grained Sparse Convolution Accelerator for DNN TrainingabstractTraining deep neural networks (DNNs) on edge devices is appealing for its adaptability and privacy benefits, but it faces challenges due to the limited resources and energy available on edge devices. In this paper, we propose MSCA, a Multigrained Sparsity Convolution Accelerator. MSCA exploits both coarse-grained and fine-grained sparsity during the DNN training phases through two types of well-designed units. Experimental results show that MSCA implemented on FPGA achieves 218.03 GOPS throughput, 39.8 GOPS/W energy efficiency, and 4.0-6.2x speedup over dense accelerators for training VGG-8 and ResNet-10 on the CIFAR-10 and SVHN datasets. Yingchang Mao, Qiang Liu 0011, Ray C. C. Cheung |
ASAP | 3 |
| 2024 | RO-SVD: A Reconfigurable Hardware Copyright Protection Framework for AIGC ApplicationsabstractThe dramatic surge in the utilisation of generative artificial intelligence (GenAI) underscores the need for a secure and efficient mechanism to responsibly manage, use and disseminate multidimensional data generated by artificial intelligence (AI). In this paper, we propose a blockchain-based copyright traceability framework called ring oscillator-singular value decomposition (RO-SVD), which introduces decomposition computing to approximate low-rank matrices generated from hardware entropy sources and establishes an AI-generated content (AIGC) copyright traceability mechanism at the device level. By leveraging the parallelism and reconfigurability of field-programmable gate arrays (FPGAs), our framework can be easily constructed on existing AI -accelerated devices and provide a low-cost solution to emerging copyright issues of AIGC. We developed a hardware-software (HW /SW) co-design prototype based on comprehensive analysis and on-board experiments with multiple AI-applicable FPGAs. Using AI-generated images as a case study, oyr framework demonstrated effectiveness and emphasised customisation, unpredictability, efficiency, manage-ment and reconfigurability. To the best of our knowledge, this is the first practical hardware study discussing and implementing copyright traceability specifically for AI -generated conten t. Zhuoheng Ran, Muhammad A. A. Abdelgawad, Ray C. C. Cheung, Hong Yan 0001 |
ASAP | 4 |
| 2024 | Design of Light-Weight Encryption Algorithm Based on RISC-V Platform: (PhD Forum Paper)abstractIn this work, we aim to design a lightweight en-cryption algorithm, GIFT, suitable for IoT systems. The design concept of the GIFT lightweight encryption algorithm is based on the PRESENT lightweight encryption algorithm. GIFT has a higher efficiency (in terms of area and timing) than PRESENT, making it one of the most energy-efficient encryption algorithms at present. GIFT can support two kinds of block data simultaneously, 64-bit block data and 128-bit block data, and the key length is 128 bits for both. Pulpino is an IoT platform based on RISC- V lightweight core. Our proposed method can improve the security encryption core used in the original Pulpino platform, reducing its power consumption by 1/4 and resource usage by 1/2. The operating speed of the platform has been greatly improved, and at the same time, more hardware space can be reserved for other applications. Ray C. C. Cheung |
ASAP | 2 |
| 2024 | Enhanced Black-Scholes Option Pricing: Bit-Width Optimization with Automatic Differentiation and Lagrange Multipliers
Sylvia Siyi Xiang, Ray C. C. Cheung |
TENCON | 2 |
| 2024 | An AIoT LoRaWAN Control System With Compression and Image Recovery Algorithm (CIRA) for Extreme WeatherabstractPromoting smart city applications can facilitate sustainable development to achieve carbon neutrality and solve existing problems, such as the frequent occurrence of extreme weather events caused by climate change. However, monitoring a large area in detail is very challenging, and there is currently no cost-effective solution. This research aims to design an artificial intelligence (AI) of Things (AIoT) LoRaWAN-based low-power, extensive coverage monitoring and alarm system at a low cost. Using the LoRaWAN communication system, it can provide point-to-point communication distances of more than 1 km, forming a low-cost remote network. Additionally, low-power consumption cameras and temperature and humidity sensors monitor the environment for a long time, then use the image compression and data transmission methods developed in this research to achieve stable and low data rate transmission of images and data. On the other hand, AI image analysis algorithms are used for image monitoring and object detection to provide alarm functions. Through this research, this system successfully monitored the environmental data and images under different extreme weather conditions in Hong Kong and provided effective warnings. Based on this research, a low-cost remote monitoring network can be further formed to automatically and effectively provide environmental monitoring and alerts to local governments over a long period. Fred F. Z. Cai, Chaoqiang Jiang, Ray C. C. Cheung, Alan H. F. Lam |
IEEE Internet Things J. | 3 |
| 2024 | REALISE-IoT: RISC-V-Based Efficient and Lightweight Public-Key System for IoT ApplicationsabstractLoRa is a promising choice for deploying an IoT network due to its lightweight feature and the extensive support by LoRa Alliance. However, as a fundamental part of LoRa, the typical LoRaWAN protocol confronts severe security challenges because it insecurely utilizes AES-128 to support the low-cost feature. In this paper, we propose a systematic solution that is compatible with LoRaWAN for IoT applications. We extend the standard LoRaWAN protocol with public-key infrastructures. Public-key features like Key exchange and authentication are supported by lightweight hardware implementations of SHA-2, ECDH, EdDSA, and TRNG. A lightweight RISC-V processor with a security coprocessor is implemented and verified using FPGA technology. The security protocol and the prototype hardware system are validated and evaluated on practical applications from our industrial partner. The prototyped development board consumes a static power of 0.116Wand a dynamic power of 0.206 W. The proposed system can achieve a 5.6x-144.7x speed up and reduce memory usage by 2.4x-12.3x for security computations. Gaoyu Mao, Yao Liu 0006, Wangchen Dai, Guangyan Li, Alan H. F. Lam, Ray C. C. Cheung |
IEEE Internet Things J. | 7 |
| 2024 | Yet Another Improvement of Plantard Arithmetic for Faster Kyber on Low-End 32-bit IoT DevicesabstractIn 2022, the National Institute of Standards and Technology (NIST) made an announcement regarding the standardization of Post-Quantum Cryptography (PQC) candidates. Out of all the Key Encapsulation Mechanism (KEM) schemes, the CRYSTAL-Kyber emerged as the sole winner. This paper presents another improved version of Plantard arithmetic that could speed up Kyber implementations on two low-end 32-bit IoT platforms (ARM Cortex-M3 and RISC-V) without SIMD extensions. Specifically, we further enlarge the input range of the Plantard arithmetic without modifying its computation steps. After tailoring the Plantard arithmetic for Kyber’s modulus, we show that the input range of the Plantard multiplication by a constant is at least 2.14× larger than the original design in TCHES2022. Then, two optimization techniques for efficient Plantard arithmetic on Cortex-M3 and RISC-V are presented.We show that the Plantard arithmetic supersedes both Montgomery and Barrett arithmetic on low-end 32-bit platforms. With the enlarged input range and the efficient implementation of the Plantard arithmetic on these platforms, we propose various optimization strategies for NTT/INTT. We minimize or entirely eliminate the modular reduction of coefficients in NTT/INTT by taking advantage of the larger input range of the proposed Plantard arithmetic on low-end 32-bit platforms. Furthermore, we propose two memory optimization strategies that reduce 23.50%~28.31% stack usage for the speed-version Kyber implementation when compared to its counterpart on Cortex-M4. The proposed optimizations make the speed-version implementation more feasible on low-end IoT devices. Thanks to the aforementioned optimizations, our NTT/INTT implementation shows considerable speedups compared to the state-of-the-art work. Overall, we demonstrate the applicability of the speed-version Kyber implementation on memory-constrained IoT platforms and set new speed records for Kyber on these platforms. Junhao Huang 0001, Haosong Zhao, Jipeng Zhang 0001, Wangchen Dai, Lu Zhou 0002, Ray C. C. Cheung, Çetin Kaya Koç, Donald Donglong Chen |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2024 | ProgramGalois: A Programmable Generator of Radix-4 Discrete Galois Transformation Architecture for Lattice-Based CryptographyabstractLattice-based cryptography (LBC) has been established as a prominent research field, with particular attention on post-quantum cryptography (PQC) and fully homomorphic encryption (FHE). As the implementing bottleneck of PQC and FHE, number theoretic transform (NTT) has been extensively studied. However, current works struggled with scalability, hindering their adaptation to various parameters, such as bit width and polynomial length. In this article, we proposed a novel Discrete Galois Transformation (DGT) algorithm utilizing the radix-4 variant to achieve a higher level of parallelism to the existing NTT. Furthermore, to implement the efficient radix-4 DGT adapting more LBCs, we proposed a set of scalable building blocks, including a modified Barrett modular multiplier accepting arbitrary modulus with only one integer multiplier, a radix-4 DGT butterfly unit, and a stream permutation network. The proposed modules are implemented on the Xilinx Virtex-7 and U250 FPGA to evaluate resource utilization and performance. Lastly, a design space exploration framework is proposed to generate optimized radix-4 DGT hardware constrained by polynomial and platform parameters. The sensitivity analysis showcases the generated hardware’s performance and scalability. The implementation results on the Xilinx Virtex-7 and U250 FPGA show significant performance improvements over the state-of-the-art works, which reached at least 35%, 192%, and 68% area-time product improvements in terms of LUTs, BRAMs, and DSPs, respectively. Guangyan Li, Zewen Ye, Donald Donglong Chen, Wangchen Dai, Gaoyu Mao, Kejie Huang, Ray C. C. Cheung |
ACM Trans. Reconfigurable Technol. Syst. | 7 |
| 2024 | An Efficient FPGA-based Depthwise Separable Convolutional Neural Network Accelerator with Hardware PruningabstractConvolutional neural networks (CNNs) have been widely deployed in computer vision tasks. However, the computation and resource intensive characteristics of CNN bring obstacles to its application on embedded systems. This article proposes an efficient inference accelerator on Field Programmable Gate Array (FPGA) for CNNs with depthwise separable convolutions. To improve the accelerator efficiency, we make four contributions: (1) an efficient convolution engine with multiple strategies for exploiting parallelism and a configurable adder tree are designed to support three types of convolution operations; (2) a dedicated architecture combined with input buffers is designed for the bottleneck network structure to reduce data transmission time; (3) a hardware padding scheme to eliminate invalid padding operations is proposed; and (4) a hardware-assisted pruning method is developed to support online tradeoff between model accuracy and power consumption. Experimental results show that for MobileNetV2 the accelerator achieves 10× and 6× energy efficiency improvement over the CPU and GPU implementation, and 302.3 frames per second and 181.8 GOPS performance that is the best among several existing single-engine accelerators on FPGAs. The proposed hardware-assisted pruning method can effectively reduce 59.7% power consumption at the accuracy loss within 5%. Zhengyan Liu, Qiang Liu 0011, Shun Yan, Ray C. C. Cheung |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2023 | In-Network Aggregation with Transport Transparency for Distributed TrainingabstractRecent In-Network Aggregation (INA) solutions offload the all-reduce operation onto network switches to accelerate and scale distributed training (DT). On end hosts, these solutions build custom network stacks to replace the transport layer. The INA-oriented network stack cannot take advantage of the state-of-the-art performant transport layer implementation, and also causes complexity in system development and operation. Shuo Liu 0002, Qiaoling Wang, Junyi Zhang 0005, Wenfei Wu, Qinliang Lin, Yao Liu 0006, Marco Canini, Ray C. C. Cheung, Jianfei He |
ASPLOS (3) | 9 |
| 2023 | Bidirectionally Deformable Motion Modulation For Video-based Human Pose TransferabstractVideo-based human pose transfer is a video-to-video generation task that animates a plain source human image based on a series of target human poses. Considering the difficulties in transferring highly structural patterns on the garments and discontinuous poses, existing methods often generate unsatisfactory results such as distorted textures and flickering artifacts. To address these issues, we propose a novel Deformable Motion Modulation (DMM) that utilizes geometric kernel offset with adaptive weight modulation to simultaneously perform feature alignment and style transfer. Different from normal style modulation used in style transfer, the proposed modulation mechanism adaptively reconstructs smoothed frames from style codes according to the object shape through an irregular receptive field of view. To enhance the spatio-temporal consistency, we leverage bidirectional propagation to extract the hidden motion information from a warped image sequence generated by noisy poses. The proposed feature propagation significantly enhances the motion prediction ability by forward and backward propagation. Both quantitative and qualitative experimental results demonstrate superiority over the state-of-the-arts in terms of image fidelity and visual continuity. The source code is publicly available at github.com/rocketappslab/bdmm. Wing Yin Yu, Lai-Man Po, Ray C. C. Cheung, Yuzhi Zhao, Kun Li 0015 |
ICCV | 3 |
| 2023 | CO-Detector: Towards Complex Object Detection with Cross-Part Feature Learning in Remote SensingabstractObject detection in remote sensing imagery builds the essential foundation of aerial and satellite image understanding, being an important role in many common real-world tasks and attracting world-wide attention. In recent years, despite the great progress of common object detection in remote sensing and the proven success of deep learning in this field, yet complex object detection which consists of multiple objects with variable layouts in remote sensing (e.g., coal-fired power plant, airport, sewage treatment plant, etc.) is still challenging for complex composite spatial relationship, non-rigid boundaries, and complicated surrounding textures. These challenges necessitate developing specific complex object detection methods to learn inter-relationship and distinctive and discriminative features in complex objects. To address this problem, in this paper, we propose a method, i.e., CO-Detector, in an end-to-end manner, to achieve various complex composite object detection in remote sensing images with high accuracy and efficiency. The effectiveness of CO-Detector is built on three main parts: (a) First, as surrounding contexts are normally complicated and similar to complex objects, we propose a Tandem Attention Network (TAN), including a channel enhanced network and a spatial enhanced network, with a K-global max/average pooling, to restrain noise disturbance and highlight complex object features and boundaries. (b) Second, we design a Part Region Proposal Network (P-RPN) to learn the interrelationship between parts in one object, generating part proposals and locating discriminative and distinctive object parts finely. (c) Third, to detect the whole complex object as well as the parts, we propose a Part Detection Network (PDN) to detect the individual parts, and detect the whole object through multi-level fused features. We train our CO-Detector model with three selected categories (i.e., coal-fired power plant, airport, oil storage tank) in three datasets, and conduct comparative experiments to evaluate and verify the performance. The comprehensive experiment results show that our CO-Detector achieves a mAP of 80.23%, outperforming 4.17%-17.83% against other cutting-edge deep learning-based detection methods. The experiment results indicate our CO-Detector has promising performance and potential in various complex object detection in highresolution remote sensing images, pending to be utilized in real large-scale applications. Shuai Yuan 0005, Juepeng Zheng, Jierui Liu, Haohuan Fu, Ray C. C. Cheung |
IGARSS | 6 |
| 2023 | High-performance and Configurable SW/HW Co-design of Post-quantum Signature CRYSTALS-DilithiumabstractCRYSTALS-Dilithium is a lattice-based post-quantum digital signature scheme that is resistant to attacks by quantum computers and has been selected to be standardized in the NIST post-quantum cryptography (PQC) standardization process. However, the speed performance and design flexibility of the Dilithium still need to be evaluated. This article presents a high-performance software/hardware co-design of CRYSTALS-Dilithium based on the NIST PQC round-3 parameters. High-speed pipelined hardware modules for NTT/INTT, point-wise multiplication/addition, and for SHAKE are included in the design to accelerate the time-consuming operations in Dilithium. All hardware modules are parameterized, thus allowing full support of runtime configuration to increase versatility. Moreover, the proposed software/hardware architecture and tight operating workflows reduce the data transmission overhead between the processor and other hardware modules. The hardware accelerator is implemented with a reconfigurable logic on FPGA and is integrated with the high-performance ARM Cortex-A9 processor in the Xilinx Zynq Architecture. We measure the performance of the software/hardware system for Dilithium in NIST security levels 2, 3, and 5. Compared to pure software implementations, we achieve 8.7–12.5 times speedup in Key generation, 6.3–7.3 times speedup in Sign, and 9.1–12.2 times speedup in Verify operations. Gaoyu Mao, Donald Donglong Chen, Guangyan Li, Wangchen Dai, Abdurrashid Ibrahim Sanka, Çetin Kaya Koç, Ray C. C. Cheung |
ACM Trans. Reconfigurable Technol. Syst. | 7 |
| 2022 | A High-Performance FPGA Accelerator for CUR DecompositionabstractA matrix factorization is to decompose a matrix into a product of smaller matrices. It is widely used in machine learning algorithms. There are many matrix decomposition algorithms, and each has various applications. CUR matrix decomposition is a widely-used factorization tool that has been employed for dimension reduction and pattern recognition in many scientific and engineering applications, such as image processing, text mining, and wireless communications. In this paper we propose an efficient FPGA-based floating-point accelerator using high-level synthesis (HLS) for the CUR decomposition algorithm. Our experiment results demonstrate the better efficiency of our hardware design compared to the optimized CPU-based software solutions. The speedup of our FPGA-based architecture over the optimized software implementation ranges from 2.37 to 16.82 times for different dimensions of the data input matrix. We evaluated our design using large dimension matrices 1024 x 1024 and 2048 x 2048 and the experiment results demonstrated the efficiency of our design in terms of the utilized resources and latency. Finally, we have compared our design with other matrix decomposition algorithms such as SVD and QR decomposition, the experiment results demonstrated that CUR is more efficient than SVD and QR decomposition in terms of latency and required resources. Muhammad A. A. Abdelgawad, Ray C. C. Cheung, Hong Yan 0001 |
FPL | 2 |
| 2022 | PrefaceabstractPresents the introductory welcome message from the conference proceedings. May include the conference officers' congratulations to all involved with the conference event and publication of the proceedings record. Yun Liang 0001, Hiroki Nakahara, Wei Zhang 0012, Fubing Mao, Ray C. C. Cheung |
FPT | 5 |
| 2022 | Message from the General Chair and Program Co-ChairsabstractOn behalf of the FPT'22 Organizing Committee, we appreciate all of you for joining FPT'22 both in-person or virtually. We wished we could say “Welcome to Hong Kong” to all the attendees, but the ongoing Covid-19 travel restrictions still make overseas travel inconvenient for some of the attendees. Hence, we have worked very hard to give good support to both the in-person and virtual attendees. We bring a hybrid mode FPT'22 and hope to connect all the attendees together to enjoy the exciting program. Wei Zhang 0012, Ray C. C. Cheung, Yun Liang 0001, Hiroki Nakahara |
FPT | 2 |
| 2022 | Melting Glacier: A 37-Year (1984-2020) High-Resolution Glacier-Cover Record of MT. KilimanjaroabstractCommonly recognized as an important symbol of the tropics and global warming, the glacier loss on Mt. Kilimanjaro has received worldwide attention for decades. In this paper, we propose a high-resolution glacier-cover (GC) record of Mt. Kilimanjaro over the period from 1984 to 2020, using a novel deep learning-based semantic segmentation method and Google Earth images, as well as digital elevation model (DEM) and ERA5-Land (ERA5) for snowline and temperature variations analysis. Our method achieves an accuracy of 94.37%, which proves the model's capability to record the GC areas precisely. The results show that (1) the GC area dramatically decreases from 19.2 km2to 3.6 km2during 37 years, which decreases about 4% and 2% per year from 1984 to 2000 and from 2000 to 2020 respectively, (2) the snowline altitude rises from$4,651 m$to$5,088 m$by about$437 m$, and (3) the average$5,000 m$air temperature on Mt. Kilimanjaro increases from −2.1 °C to −1.1 °C by about 1 °C. This study indicates that there will be no GC within a few decades if the current loss continues. Shuai Yuan 0005, Juepeng Zheng, Lixian Zhang 0002, Runmin Dong, Yile Xing, Yuhan She, Haohuan Fu, Ray C. C. Cheung |
IGARSS | 8 |
| 2022 | Reconfigurable content-addressable memory (CAM) on FPGAs: A tutorial and survey
Muhammad Irfan 0004, Abdurrashid Ibrahim Sanka, Zahid Ullah 0001, Ray C. C. Cheung |
Future Gener. Comput. Syst. | 4 |
| 2022 | High Throughput Hardware/Software Heterogeneous System for RRPN-Based Scene Text DetectionabstractRotation Region Proposal Networks (RRPN) are used to generate rotated proposals with the information of text angle for arbitrary oriented scene text detection (STD). However, the computational complexity of RRPN inference is relatively high compared with other methods, which makes it difficult for massive deployment. In this paper, the first full-stack FPGA-CPU heterogeneous system design of RRPN-based STD algorithm is proposed. A hardware/software partition method is presented to analyze and split the tasks to enhance the computation efficiency of hardware. The fast 2D Winograd algorithm and block floating point are utilized to reduce computation complexity while maintaining a relatively high precision. The implementation results show that the peak performance of MAC arrays in the proposed architecture reaches 655.4 GOPS and the energy efficiency achieves 64.9 GOPS/W. By fully exploiting the parallel and pipelined merits in the algorithms, the first hardware architectures for skew non-maximum suppression (S-NMS) layer and rotation region-of-interest (RRoI) polling layer are proposed. The throughput of the proposed hardware/software heterogeneous system achieves 40 times and 1.4 times improvements compared with CPU and GPU, respectively. Moreover, the comprehensive operating expense ratio of pure CPU, GPU, and the proposed system is 80.7:2.5:1, which indicates that it is suitable for massive deployment. Yao Xin, Donald Donglong Chen, Chongyang Zeng, Yi Wang 0004, Ray C. C. Cheung |
IEEE Trans. Computers | 6 |
| 2021 | An FPGA-based MobileNet Accelerator Considering Network Structure CharacteristicsabstractConvolutional neural networks (CNNs) have been widely deployed in computer vision tasks. However, the computation and resource intensive characteristics of CNN bring obstacles to its application on embedded systems. MobileNet, as a representative of compact models, can reduce the amount of parameters and computation. A high-performance inference accelerator on FPGA for MobileNet is proposed in this paper. With respect to the three types of convolution operations, multiple parallel strategies are exploited and the corresponding hardware structures such as input buffer and configurable adder tree are designed. With respect to the bottleneck block, a dedicated architecture is proposed to reduce data transmission time. In addition, a hardware padding scheme to improve the efficiency of padding is proposed. The accelerator implemented on Virtex-7 FPGA reaches 70.8% Top-1 accuracy under 8-bit quantization. The accelerator achieves 302.3 FPS and 181.8 GOPS, which obtains 22.7x, 3.9x and 1.4x speedup compared to the implementations in Snapdragon 821 CPU, i7-6700HQ CPU and GTX 960M GPU, respectively. Shun Yan, Zhengyan Liu, Chenglong Zeng, Qiang Liu 0011, Bowen Cheng, Ray C. C. Cheung |
FPL | 7 |
| 2021 | Efficient High-Performance FPGA-Redis Hybrid NoSQL Caching System for Blockchain Scalability
Abdurrashid Ibrahim Sanka, Mehdi Hasan Chowdhury, Ray C. C. Cheung |
Comput. Commun. | 3 |
| 2021 | A survey of breakthrough in blockchain technology: Adoptions, applications, challenges and future research
Abdurrashid Ibrahim Sanka, Muhammad Irfan 0004, Ian Huang, Ray C. C. Cheung |
Comput. Commun. | 4 |
| 2021 | A systematic review of blockchain scalability: Issues, solutions, analysis and future research
Abdurrashid Ibrahim Sanka, Ray C. C. Cheung |
J. Netw. Comput. Appl. | 2 |
| 2021 | Scalable Fully Pipelined Hardware Architecture for In-Network Aggregated AllReduce CommunicationabstractThe Ring-AllReduce framework is currently the most popular solution to deploy industry-level distributed machine learning tasks. However, only about half of the maximum bandwidth can be achieved in the optimal condition. In recent years, several in-network aggregation frameworks have been proposed to overcome the drawback, but limited hardware information have been disclosed. In this paper, we propose a scalable fully-pipelined architecture that handles tasks like forwarding, aggregation and retransmission with no bandwidth loss. The architecture is implemented on a Xilinx Ultrascale FPGA that connects to 8 working servers with 10 Gb/s network adapters, and it is able to scale to more complicated scenarios involving more workers. Compared with Ring-AllReduce, using AllReduce-Switch improves the efficient bandwidth of AllReduce communication with a ratio of$1.75\times $. In image training tasks, the proposed hardware architecture helps to achieve up to$1.67\times $speedup to the training process. For computing-intensive models, the speedup from communication may be partially hidden by computing. In particular, for ResNet-50, AllReduce-Switch improves the training process with MPI and NCCL by$1.30\times $and$1.04\times $respectively. Yao Liu 0006, Junyi Zhang 0005, Shuo Liu 0002, Qiaoling Wang, Wangchen Dai, Ray C. C. Cheung |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2021 | Elastic Net Constraint-Based Tensor Model for High-Order Graph MatchingabstractThe procedure of establishing the correspondence between two sets of feature points is important in computer vision applications. In this article, an elastic net constraint-based tensor model is proposed for high-order graph matching. To control the tradeoff between the sparsity and the accuracy of the matching results, an elastic net constraint is introduced into the tensor-based graph matching model. Then, a nonmonotone spectral projected gradient (NSPG) method is derived to solve the proposed matching model. During the optimization of using NSPG, we propose an algorithm to calculate the projection on the feasible convex sets of elastic net constraint. Further, the global convergence of solving the proposed model using the NSPG method was proved. The superiority of the proposed method is verified through experiments on the synthetic data and natural images. Hu Zhu, Chunfeng Cui, Lizhen Deng, Ray C. C. Cheung, Hong Yan 0001 |
IEEE Trans. Cybern. | 4 |
| 2021 | An Efficient Parallel Processor for Dense Tensor ComputationabstractNowadays, many data are multidimensional, which are called tensors. Tensor computations have been applied in different fields and various software libraries have been developed. However, not much attention has been received for developing a hardware architecture to accelerate the tensor computations. In this article, an efficient and unified processing element (PE) array for the 3-D tensor computation is demonstrated. Our PE array is optimized for thin and tall tensor-matrix multiplication and two types of tensor times matrices chain (TTMc) operations. Our design is evaluated in three study cases and compared with the state-of-the-art design. By using computation partition and rearrangement, data movement between the field-programmable gate array (FPGA) and off-chip DDR memory can be reduced by O(I2), where I is the maximum range among all the dimensions of the data tensor. For TTMc implementation, clock frequency has been increased by 18% compared with the state-of-the-art implementation on the same FPGA chip. An experiment on 3-D volumetric data set rendering by tensor approximation method is conducted for demonstration. For the bricks reconstruction process, the runtime decreased by 50%, i.e., two times faster, on our FPGA implementation compared with that running on GPU. In CANDECOMP/PARAFAC decomposition, for one iteration, the runtime has been decreased by up to 93% compared with the programs implemented by Tensorly, which is a python library. Wei-pei Huang, Ray C. C. Cheung, Hong Yan 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2020 | Dynamic Sparse Training: Find Efficient Sparse Network From Scratch With Trainable Masked Layers
Zhe Xu 0008, Runbin Shi, Ray C. C. Cheung, Hayden Kwok-Hay So |
ICLR | 4 |
| 2020 | Binary convolutional neural network acceleration framework for rapid system prototyping
Zhe Xu 0008, Ray C. C. Cheung |
J. Syst. Archit. | 2 |
| 2020 | RPE-TCAM: Reconfigurable Power-Efficient Ternary Content-Addressable Memory on FPGAsabstractSoftware-defined networks (SDNs) are the future networks that enable the system to be more flexible and programmable using a centralized controller. Field-programmable gate arrays (FPGAs) serve as exemplary hardware to implement these adaptable networks. Ternary content-addressable memory (TCAM) is an essential part of every network to perform packet classification and forwarding, but they are missing in modern FPGAs. Researchers and FPGA vendors have proposed several designs to emulate TCAM using available memories on FPGA, but they are power inefficient due to the activating of entire circuitry in a single search operation. In this brief, we propose a novel power-aware reconfigurable FPGA-based TCAM architecture that enables only a portion of the hardware to perform the search operation. We performed an extensive design space exploration to find the optimal number of banks on Xilinx FPGAs, which provides the maximum power saving. Moreover, we propose a solution to bank overflow using backup CAM (BUC) to handle the overflowed CAM entries. The proposed TCAM improves the power consumption by 40% and maintains one-clock cycle update latency with no compromise on the throughput of the system compared with the state-of-the-art FPGA-based TCAM architectures. Muhammad Irfan 0004, Zahid Ullah 0001, Mehdi Hasan Chowdhury, Ray C. C. Cheung |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2019 | Bank-selective Strategy for Gate-based Ternary Content-addressable Memory on FPGAsabstractTernary content-addressable memory (CAM) is an associative memory which supports the storing of don't care ('X') bits. Field programmable gate arrays (FPGAs) are enriched with high speed hardware components such as memory elements, logical blocks, but do not have support for a TCAM. Researchers have emulated TCAM inside FPGAs using static random-access memory (SRAM) blocks, flip-flops, and lookup tables (LUTs). Power requirement of these emulated TCAMs is very high due to the parallel comparison of each bit of the search key with the stored words. In this paper, we present a bank-selective strategy to decrease the amount of power consumed by gate-based TCAM on FPGA. Instead of comparing the search key with all of the stored words, our proposed architecture compares the search key with a selected number of stored words based on classifier-bits in search key. A sample of 64×36 is implemented on Xilinx Virtex-6 FPGA with four banks, which reduced the power consumption by 44.8% compared to the state-of-the-art TCAM architecture. Hardware utilization of the proposed TCAM on FPGA with four banks is also reduced by 6% and 8% in the form of slice registers (SRs) and lookup tables (LUTs), respectively. Muhammad Irfan 0004, Ray C. C. Cheung, Zahid Ullah 0001 |
ASAP | 2 |
| 2019 | An Efficient Application Specific Instruction Set Processor (ASIP) for Tensor ComputationabstractIn the past decade, tensor computation is widely used in different areas. Various software toolbox have been released to assist tensor computation. However, there is still no hardware architecture to accelerate the tensor computation. This paper presents an efficient application specific instruction set processor (ASIP) for tensor computation. Different tensor computations are fully optimized in terms of resource usage and performance. We implement the ASIP on FPGA platform. We test our design by implementing the CANDECOMP/PARAFAC(CP) decomposition. Our design can achieve a low resource usage and run at 141 Mhz. Wei-pei Huang, Ray C. C. Cheung, Hong Yan 0001 |
ASAP | 2 |
| 2019 | Accurate and Compact Convolutional Neural Networks with Trained Binarization
Zhe Xu 0008, Ray C. C. Cheung |
BMVC | 2 |
| 2019 | High Performance Power-Efficient Gate-Based CAM for Reconfigurable ComputingabstractContent-addressable memory (CAM) is a high-speed lookup memory, which searches the entire memory in parallel and provides address of the input search word. Internet of things (IoT) technology needs access control list filters, deployed using CAM, to reduce the flow of invalid data through the network. Field-programmable gate arrays (FPGAs) play an important role to provide the computational power to IoT nodes using ultra low power modern devices. CAMs are emulated in FPGAs using different memories, i.e., block RAM, distributed RAM, and flip-flops (FFs). FPGA-based CAMs have large power consumption due to the involvement of each CAM cell in parallel comparison. This work presents a bank-selective strategy for gate-based binary CAM on FPGA, which reduces the amount of comparisons and reduces power consumption on target FPGA device. For every input search key, only one bank is activated using gated-clock to compare the input key to corresponding stored rules. A sample of 512 x 36 of the proposed architecture with four banks is implemented on Xilinx Virtex-6 FPGA that consumes 42% less dynamic power compared to the best available CAM designs. Hardware resources of the proposed CAM design on target FPGA is reduced in terms of slice registers (SRs) and lookup tables (LUTs) by 7.9% and 8.8%, respectively, compared with the latest prior work. Muhammad Irfan 0004, Zahid Ullah 0001, Ray C. C. Cheung |
MSN | 3 |
| 2019 | A robust background initialization algorithm with superpixel motion detection
Zhe Xu 0008, Biao Min, Ray C. C. Cheung |
Signal Process. Image Commun. | 3 |
| 2019 | Feature Selection Based on Tensor Decomposition and Object Proposal for Night-Time Multiclass Vehicle DetectionabstractNight-time vehicle detection is essential in building intelligent transportation systems (ITS) for road safety. Most of current night-time vehicle detection approaches focus on one or two classes of vehicles. In this paper, we present a novel multiclass vehicle detection system based on tensor decomposition and object proposal. Commonly used features such as histogram of oriented gradients and local binary pattern often produce useless image blocks (regions), which can result in unsatisfactory detection performance. Thus, we select blocks via feature ranking after tensor decomposition and only extract features from these selected blocks. To generate windows that contain all vehicles, we propose a novel object-proposal approach based on a state-of-the-art object-proposal method, local features, and image region similarity. The three terms are summed with learned weights to compute the reliability score of each proposal. A bio-inspired image enhancement method is used to enhance the brightness and contrast of input images. We have built a Hong Kong night-time multiclass vehicle dataset for evaluation. Our proposed vehicle detection approach can successfully detect four types of vehicles: 1) car; 2) taxi; 3) bus; and 4) minibus. Occluded vehicles and vehicles in the rain can also be detected. Our proposed method obtains 95.82% detection rate at 0.05 false positives per image, and it outperforms several state-of-the-art night-time vehicle detection approaches. Hulin Kuang, Long Chen 0005, Leanne Lai Chan, Ray C. C. Cheung, Hong Yan 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2018 | Lightweight Secure Processor Prototype on FPGAabstractLightweight devices usually come with an additional cryptographic co-processor for enabling the secrecy, in contrast, the master processor is typically a commercial processor where the required protection mechanism is missing. In this paper, an on-going effort in secured architecture named S-RISC-V based on RISC-V core is introduced. The mechanism of key generation used for memory protection is supported together with the joint efforts in the following perspectives, including ISA extension, compiler improvement, and hardware implementation. The architecture has been verified on Zedboard running at 25MHz, driven by the host ARM core. The area overhead is less than 10%, compared with the original RISC-V core. Yao Liu 0006, Ray C. C. Cheung, Hei Wong |
FPL | 2 |
| 2018 | ASIC Implementation of a Nonlinear Dynamical Model for Hippocampal ProsthesisabstractA hippocampal prosthesis is a very large scale integration (VLSI) biochip that needs to be implanted in the biological brain to solve a cognitive dysfunction. In this letter, we propose a novel low-complexity, small-area, and low-power programmable hippocampal neural network application-specific integrated circuit (ASIC) for a hippocampal prosthesis. It is based on the nonlinear dynamical model of the hippocampus: namely multi-input, multi-output (MIMO)-generalized Laguerre-Volterra model (GLVM). It can realize the real-time prediction of hippocampal neural activity. New hardware architecture, a storage space configuration scheme, low-power convolution, and gaussian random number generator modules are proposed. The ASIC is fabricated in 40 nm technology with a core area of 0.122 mm[Formula: see text] and test power of 84.4 [Formula: see text]W. Compared with the design based on the traditional architecture, experimental results show that the core area of the chip is reduced by 84.94% and the core power is reduced by 24.30%. Zhitong Qiao, Yan Han 0005, Xiangyu Li 0005, Dong Song, Theodore W. Berger, Ray C. C. Cheung |
Neural Comput. | 8 |
| 2018 | A fast inter CU decision algorithm for HEVC
Zhe Xu 0008, Biao Min, Ray C. C. Cheung |
Signal Process. Image Commun. | 3 |
| 2018 | FFT-Based McLaughlin's Montgomery Exponentiation without Conditional SelectionsabstractModular multiplication forms the basis of many cryptographic functions such as RSA, Diffie-Hellman key exchange, and ElGamal encryption. For large RSA moduli, combining the fast Fourier transform (FFT) with McLaughlin's Montgomery modular multiplication (MLM) has been validated to offer cost-effective implementation results. However, the conditional selections in McLaughlin's algorithm are considered to be inefficient and vulnerable to timing attacks, since extra long additions or subtractions may take place and the running time of MLM varies. In this work, we restrict the parameters of MLM by a set of new bounds and present a modified MLM algorithm involving no conditional selection. Compared to the original MLM algorithm, we inhibit extra operations caused by the conditional selections and accomplish constant running time for modular multiplications with different inputs. As a result, we improve both area-time efficiency and security against timing attacks. Based on the proposed algorithm, efficient FFT-based modular multiplication and exponentiation are derived. Exponentiation architectures with dual FFT-based multipliers are designed obtaining area-latency efficient solutions. The results show that our work offers a better efficiency compared to the state-of-the-art works from and above 2048-bit operand sizes. For single FFT-based modular multiplication, we have achieved constant running time and obtained area-latency efficiency improvements up to 24.3 percent for 1,024-bit and 35.5 percent for 4,096-bit operands, respectively. Wangchen Dai, Donald Donglong Chen, Ray C. C. Cheung, Çetin Kaya Koç |
IEEE Trans. Computers | 3 |
| 2017 | A low power V-band LC VCO with high Q varactor technique in 40 nm CMOS process
Yan Han 0005, Lu Jie 0001, Ray C. C. Cheung, Guangtao Feng |
Sci. China Inf. Sci. | 6 |
| 2017 | Area-Time Efficient Architecture of FFT-Based Montgomery MultiplicationabstractThe modular multiplication operation is the most time-consuming operation for number-theoretic cryptographic algorithms involving large integers, such as RSA and Diffie-Hellman. Implementations reveal that more than 75 percent of the time is spent in the modular multiplication function within the RSA for more than 1,024-bit moduli. There are fast multiplier architectures to minimize the delay and increase the throughput using parallelism and pipelining. However such designs are large in terms of area and low in efficiency. In this paper, we integrate the fast Fourier transform (FFT) method into the McLaughlin's framework, and present an improved FFT-based Montgomery modular multiplication (MMM) algorithm achieving high area-time efficiency. Compared to the previous FFT-based designs, we inhibit the zero-padding operation by computing the modular multiplication steps directly using cyclic and nega-cyclic convolutions. Thus, we reduce the convolution length by half. Furthermore, supported by the number-theoretic weighted transform, the FFT algorithm is used to provide fast convolution computation. We also introduce a general method for efficient parameter selection for the proposed algorithm. Architectures with single and double butterfly structures are designed obtaining low area-latency solutions, which we implemented on Xilinx Virtex-6 FPGAs. The results show that our work offers a better area-latency efficiency compared to the state-of-the-art FFT-based MMM architectures from and above 1,024-bit operand sizes. We have obtained area-latency efficiency improvements up to 50.9 percent for 1,024-bit, 41.9 percent for 2,048-bit, 37.8 percent for 4,096-bit and 103.2 percent for 7,680-bit operands. Furthermore, the operating latency is also outperformed with high clock frequency for length-64 transform and above. Wangchen Dai, Donald Donglong Chen, Ray C. C. Cheung, Çetin Kaya Koç |
IEEE Trans. Computers | 3 |
| 2017 | Area-Time Efficient Computation of Niederreiter Encryption on QC-MDPC Codes for Embedded HardwareabstractIn this paper, we present a fast implementation for QC-MDPC Niederreiter encryption. Existing high-speed implementations are considerably resource involving but the solution we propose here mitigates such situation while maintaining the high throughputs. In particular, new arithmetic for lightweight Hamming weight computation and a fast sorting network for MDPC decoding are proposed. A novel constant weight coding unit is proposed to enable standard asymmetric encryptions. For now, the design presented in this work is the fastest one of existing QC-MDPC code based encryptions in the public domain. The area-time product of this work drops by at least 53 percent compared to previous fast speed designs of QC-MDPC based encryptions. It is shown for instance that our implementation of encrypting engine can sign one encryption in 3.86 ms on a Xilinx Virtex-6 FPGA with 3371 slices. Our iterative decrypting engine can decrypt one ciphertext in 114.64 ms with 5271 slices and our faster non-iterative decrypting engine can decrypt in 65.76 ms with 8781 slices. Jingwei Hu 0001, Ray C. C. Cheung |
IEEE Trans. Computers | 2 |
| 2017 | A Fully Pipelined Hardware Architecture for Intra Prediction of HEVCabstractUltrahigh definition (UHD), such as 4K/8K, is becoming the mainstream of video resolution nowadays. High Efficiency Video Coding (HEVC) is the emerging video coding standard to process the encoding and decoding of UHD video. This paper first develops multiple techniques that allow the proposed hardware architecture for intra prediction of HEVC working in full pipeline. The proposed techniques include: 1) a novel buffer structure for reference samples; 2) a mode-dependent scanning order; and 3) an inverse method for reference sample extension. The size of the buffer is 3K b for luma component and 3K b for chroma components, providing sufficient accessing to the reference samples. Since the data dependency between two neighboring blocks is addressed by the modedependent scanning order, the proposed fully pipelined design can produce 4 pixels/clock cycle. As a result, the throughput of the proposed architecture is capable to support 3840×2160 videos at 30 frames/s. Biao Min, Zhe Xu 0008, Ray C. C. Cheung |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2016 | Parameter Space for the Architecture of FFT-Based Montgomery Modular MultiplicationabstractModular multiplication is the core operation in public-key cryptographic algorithms such as RSA and the Diffie-Hellman algorithm. The efficiency of the modular multiplier plays a crucial role in the performance of these cryptographic methods. In this paper, improvements to FFT-based Montgomery Modular Multiplication (FFTM3) using carry-save arithmetic and pre-computation techniques are presented. Moreover, pseudo-Fermat number transform is used to enrich the supported operand sizes for the FFTM3. The asymptotic complexity of our method is O(l log l log log l), which is the same as the Schonhage-Strassen multiplication algorithm (SSA). A systematic procedure to select suitable parameter set for the FFTM3is provided. Prototypes of the improved FFTM3multiplier with appropriate parameter sets are implemented on Xilinx Virtex-6 FPGA. Our method can perform 3,100-bit and 4,124-bit modular multiplications in 6.74 and 7.78 μs, respectively. It offers better computation latency and area-latency product compared to the state-of-the-art methods for operand size of 3,072-bit and above. Donald Donglong Chen, Gavin Xiaoxu Yao, Ray C. C. Cheung, Derek Chi-Wai Pao, Çetin Kaya Koç |
IEEE Trans. Computers | 3 |
| 2015 | Architecture Support for Task Out-of-Order Execution in MPSoCsabstractMulti-processor system on chip (MPSoC) has been widely applied in embedded systems in the past decades. However, it has posed great challenges to efficiently design and implement a rapid prototype for diverse applications due to heterogeneous instruction set architectures (ISA), programming interfaces and software tool chains. In order to solve the problem, this paper proposes a novel high level architecture support for automatic out-of-order (OoO) task execution on FPGA based heterogeneous MPSoCs. The architecture support is composed of a hierarchical middleware with an automatic task level OoO parallel execution engine. Incorporated with a hierarchical OoO layer model, the middleware is able to identify the parallel regions and generate the sources codes automatically. Besides, a runtime middleware Task-Scoreboarding analyzes the inter-task data dependencies and automatically schedules and dispatches the tasks with parameter renaming techniques. The middleware has been verified by the prototype built on FPGA platform. Examples and a JPEG case study demonstrate that our model can largely ease the burden of programmers as well as uncover the task level parallelism. Chao Wang 0003, Xi Li 0003, Junneng Zhang, Peng Chen 0004, Yunji Chen, Xuehai Zhou, Ray C. C. Cheung |
IEEE Trans. Computers | 7 |
| 2015 | An Application Specific Instruction Set Processor (ASIP) for Adaptive Filters in Neural ProstheticsabstractNeural coding is an essential process for neuroprosthetic design, in which adaptive filters have been widely utilized. In a practical application, it is needed to switch between different filters, which could be based on continuous observations or point process, when the neuron models, conditions, or system requirements have changed. As candidates of coding chip for neural prostheses, low-power general purpose processors are not computationally efficient especially for large scale neural population coding. Application specific integrated circuits (ASICs) do not have flexibility to switch between different adaptive filters while the cost for design and fabrication is formidable. In this research work, we explore an application specific instruction set processor (ASIP) for adaptive filters in neural decoding activity. The proposed architecture focuses on efficient computation for the most time-consuming matrix/vector operations among commonly used adaptive filters, being able to provide both flexibility and throughput. Evaluation and implementation results are provided to demonstrate that the proposed ASIP design is area-efficient while being competitive to commercial CPUs in computational performance. Yao Xin, Xiangyu Li 0005, Ray C. C. Cheung, Dong Song, Theodore W. Berger |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2015 | A Fast CU Size Decision Algorithm for the HEVC Intra EncoderabstractIntra coding plays a crucial role in the High Efficiency Video Coding (HEVC) standard. It provides the higher coding efficiency than the previous standard, H.264/Advanced Video Coding. The block partitioning in HEVC supports quad-tree-based coding unit (CU) structure from size 64×64 to 8×8. The new technique provides better performances on one hand, whereas on the other hand it also increases the coding complexity. In this paper, a novel fast algorithm is proposed for the CU size decision in intra coding. Both the global and local edge complexities in horizontal, vertical, 45° diagonal, and 135° diagonal directions are proposed and used to decide the partitioning of a CU. Coupled with handling its four sub-CUs in the same way, a CU is decided to be split, nonsplit, or undetermined for each depth. Compared with the reference software HM10.0, the encoding time is reduced by ~52% on average, with ~0.8% Bjontegaard Distortion-rate increasing and reasonable peak signal-to-noise ratio losses. Biao Min, Ray C. C. Cheung |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2015 | Z-TCAM: An SRAM-based Architecture for TCAMabstractTernary content addressable memories (TCAMs) perform high-speed lookup operation but when compared with static random access memories (SRAMs), TCAMs have certain limitations such as low storage density, relatively slow access time, low scalability, complex circuitry, and are very expensive. Thus, can we use the benefits of SRAM by configuring it (with additional logic) to enable it to behave like TCAM? This brief proposes a novel memory architecture, named Z-TCAM, which emulates the TCAM functionality with SRAM. Z-TCAM logically partitions the classical TCAM table along columns and rows into hybrid TCAM subtables, which are then processed to map on their corresponding memory blocks. Two example designs for Z-TCAM of sizes 512 × 36 and 64 × 32 have been implemented on Xilinx Virtex-7 field-programmable gate array. The design of 64 × 32 Z-TCAM has also been implemented using OSUcells library for 0.18 μm technology, which confirms the physical and technical feasibility of Z-TCAM. Search latency for each design is three clock cycles. The detailed implementation results and power measurements for each design have been reported thoroughly. Zahid Ullah 0001, Manish Kumar Jaiswal, Ray C. C. Cheung |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2014 | Big data genome sequencing on Zynq based clusters (abstract only)abstractNext-generation sequencing (NGS) problems have attracted many attentions of researchers in biological and medical computing domains. The current state-of-the-art NGS computing machines are dramatically lowering the cost and increasing the throughput of DNA sequencing. In this paper, we propose a practical study that uses Xilinx Zynq board to summarize acceleration engines using FPGA accelerators and ARM processors for the state-of-the-art short read mapping approaches. The heterogeneous processors and accelerators are coupled with each other using a general Hadoop distributed processing framework. First the reads are collected by the central server, and then distributed to multiple accelerators on the Zynq for hardware acceleration. Therefore, the combination of hardware acceleration and Map-Reduce execution flow could greatly accelerate the task of aligning short length reads to a known reference genome. Our approach is based on preprocessing the reference genomes and iterative jobs for aligning the continuous incoming reads. The hardware acceleration is based on the creditable read-mapping algorithm RMAP software approach. Furthermore, the speedup analysis on a Hadoop cluster, which concludes 8 development boards, is evaluated. Experimental results demonstrate that our proposed architecture and methods has the speedup of more than 112X, and is scalable with the number of accelerators. Finally, the Zynq based cluster has efficient potential to accelerate even general large scale big data applications. Chao Wang 0003, Xi Li 0003, Xuehai Zhou, Yunji Chen, Ray C. C. Cheung |
FPGA | 5 |
| 2014 | A complementary architecture for high-speed true random number generatorabstractIn this paper, we introduce a novel FPGA-based design for true random number generator (TRNG). It is able to harvest the timing difference caused by the nonuniformity of the Integrated Circuits (ICs) and use it to generate the randomness. Compared with the previous related work, this design uses a complementary scheme that leads to a doubled data rated output. The proposed complementary design has improved entropy and achieved higher throughput. The prototype design has been implemented and verified on a Xilinx Virtex-6 ML605 evaluation board. As a result, the generated random number stream is able to pass the statistical NIST and DIEHARD test suites showing a reliable performance. Meanwhile, it can approach the maximum data rate as 50 Mbps stably. Ray C. C. Cheung |
FPT | 2 |
| 2014 | A low-power inverter-based ΣΔ analog-to-digital converter for audio applications
Yan Han 0005, Ray C. C. Cheung, Tianlin Cao |
Sci. China Inf. Sci. | 5 |
| 2014 | GPU-based biclustering for microarray data analysis in neurocomputing
Benben Liu, Yao Xin, Ray C. C. Cheung, Hong Yan 0001 |
Neurocomputing | 3 |
| 2014 | Novel RNS Parameter Selection for Fast Modular MultiplicationabstractThe parameter selection of Residue Number Systems (RNS) has a great impact on its computational efficiency. This paper shows that a base extension, the most costly operation in RNS Montgomery multiplication, can be more efficient when the intervals between the RNS moduli are small. We propose a systematic RNS parameter selection procedure and two methods to select RNS moduli that lead to a reduced complexity. Our experimental results confirm the advantages of the selected moduli. Gavin Xiaoxu Yao, Junfeng Fan, Ray C. C. Cheung, Ingrid Verbauwhede |
IEEE Trans. Computers | 3 |
| 2014 | Design Exploration of Geometric Biclustering for Microarray Data Analysis in Data MiningabstractBiclustering is an important technique in data mining for searching similar patterns. Geometric biclustering (GBC) method is used to reduce the complexity of the NP-complete biclustering algorithm. This paper studies three commonly used modern platforms including multi-core CPU, GPU and FPGA to accelerate this GBC algorithm. By analyzing the parallelizing property of the GBC algorithm, we design 1) a multi-threaded software running on a server grade multi-core CPU system, 2) a CUDA program for GPU to accelerate the GBC algorithm, and 3) a novel parameterizable and scalable hardware architecture implemented on an FPGA. Genes microarray pattern analysis is employed as an example to demonstrate performance comparisons on different platforms. In particular, we compare the speed and energy efficiency of the three proposed methods. We found that 1) GPU achieves the highest average speedup of 48 × compared to single-threaded GBC program, 2) Our FPGA design can achieve higher speedup of 4 × for the computation for large microarray, and 3) FPGA consumes the least energy, which is about 3.53 × more efficient than the single-threaded GBC program. Benben Liu, Chi Wai Yu, Doris Z. Wang, Ray C. C. Cheung, Hong Yan 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2013 | Binding Hardware IPs to Specific FPGA Device via Inter-twining the PUF Response with the FSM of Sequential CircuitsabstractThe continuous growth in both capability and capacity for FPGA now requires significant resources invested in the hardware design, which results in two classes of main security issues: 1) the unauthorized use and piracy attacks including cloning, reverse engineering, tampering etc. 2) the licensing issue. Binding hardware IPs (HW-IPs) to specific FPGA devices can efficiently resolve these problems. However, previous binding techniques are all based on encryption and hence have three main drawbacks: 1) encryption-based proposals in commercial are limited to protect the single large FPGA configuration, 2) many encryption-based proposals depend on a trusted third party to involve the licensing protocol, and 3) the encryption-based binding methods use costly mechanisms such as secure ROM or flash memory to store FPGA specific cryptographic keys, which is not only expensive but also vulnerable to side-channel attacks, and the management and transport of secret keys became a practical issue. In this work, we propose a PUF-FSM binding technique completely different from the traditional encryption-based methods to address these shortcomings. Jiliang Zhang 0002, Yaping Lin, Yongqiang Lyu 0001, Ray C. C. Cheung, Wenjie Che, Qiang Zhou 0001, Jinian Bian |
FCCM | 4 |
| 2013 | Genome sequencing using mapreduce on FPGA with multiple hardware accelerators (abstract only)abstractThe genome sequencing problem with short reads is an emerging field with seemingly limitless possibilities for advances in numerous scientific research and application domains. It has been the hot topic during the past few years. Growing with the data population and the ease to access for personal users, how to shorten the response interval for short read mapping at a large scale computing domain is extremely important. In this paper we propose a novel FPGA-based acceleration solution with Map-Reduce framework on multiple hardware acceleration engines. The combination of hardware accelerators and Map-Reduce execution flow could greatly expedite the task of aligning short length reads to a known reference genome. Our approach is based on preprocessing the reference genomes and iterative jobs for aligning the continuous incoming reads. The read-mapping algorithm is modeled after the creditable RMAP software approach. Furthermore, theoretical speedup analysis on a MapReduce programming platform is presented, which demonstrates that our proposed architecture has efficient potential to reduce the average waiting time for large scale short reads applications. Chao Wang 0003, Xi Li 0003, Xuehai Zhou, Jim Martin 0001, Ray C. C. Cheung |
FPGA | 5 |
| 2013 | Design space explorations of Hybrid-Partitioned TCAM (HP-TCAM)abstractEven though, TCAM provides search operation in a constant time, when compared with Static Random Access Memories (SRAMs), TCAMs have certain limitations such as low storage density, relatively slow access time, low scalability, complex circuitry, and expensive costs. Hence, the need for a TCAM architecture arises that can use SRAM (with additional logic) to behave like TCAM. This paper presents the idea of Hybrid-Partitioned, SRAM-based architecture (HP-TCAM), which provides the same functionality as TCAM. We implemented and analyzed an example design of 512 × 36 HP-TCAM on Xilinx FPGAs with its different design parameters. Energy/bit/search, as an important metric, for the design is 47.13 fJ on Virtex-7 FPGA. Furthermore, we have provided in detail, all the implementation results and power consumption for our designs. Zahid Ullah 0001, Manish Kumar Jaiswal, Ray C. C. Cheung |
FPL | 3 |
| 2013 | FPGA IP protection by binding Finite State Machine to Physical Unclonable FunctionabstractIn this paper we propose a novel binding mechanism that can protect FPGA IP from being cloned, tampered, or misused; and facilitate the pay-per-use licensing to limit the FPGA IP's execution to specific FPGA devices only. In this mechanism, the FPGA vendors will provide each enrolled device with a Physical Unclonable Function (PUF) that can be deployed securely during fabrication process. The core vendor will embed an augmented Finite State Machine (FSM) into the original FSM structure of the hardware IP (HW-IP) to react on the PUF response to a given challenge. The proposed binding method does not need any Trusted Third Party (TTP) or block cipher for key management and exchange. We analyze several known attacks to hardware IP and show that our method is secure against these attacks. Experimental results on MCNC benchmarks show that the proposed method incurs small design overhead in terms of area, power and delay. Jiliang Zhang 0002, Yaping Lin, Yongqiang Lyu 0001, Gang Qu 0001, Ray C. C. Cheung, Wenjie Che, Qiang Zhou 0001, Jinian Bian |
FPL | 5 |
| 2013 | Fast simulation of Digital Spiking Silicon Neuron model employing reconfigurable dataflow computingabstractA new simulation scheme of the Digital Spiking Silicon Neuron (DSSN) model is proposed. This scheme is based on the reconfigurable dataflow computing paradigm and targets the Maxeler MaxWorkstation. Compared to the previous implementation of the DSSN network, the new scheme has the virtues of better flexibility and better programmability. More importantly, computing with dataflow cores takes good advantage of the intrinsic parallelism of the reconfigurable hardware and better pipelining is achievable. The proposed scheme has good potential of conducting large-scale and fast simulation of the DSSN-model-based network which is pivotal to future neuroscience research. Xiangyu Li 0005, Shridhar Choudhary, Ray C. C. Cheung, Takeshi Matsumoto, Masahiro Fujita 0004 |
FPT | 3 |
| 2013 | A reconfigurable architecture for real-time prediction of neural activityabstractIn this paper, we propose an FPGA-based hardware architecture for conducting real-time prediction of neural activity using a second-order generalized Laguerre-Volterra model (GLVM). This architecture serves as a rapid prototype of the prediction module of the future cognitive neural prosthetic device. We validate the functionality of the hardware model by utilizing the neuronal firing data of behaving rats trained to perform the delayed nonmatch-to-sample (DNMS) memory task. Xiangyu Li 0005, Ray C. C. Cheung, Rosa H. M. Chan, Dong Song, Theodore W. Berger |
ISCAS | 2 |
| 2013 | Design Automation Framework for Reconfigurable Interconnection NetworksabstractA reconfigurable interconnection network (RIN) is a custom-designed on-chip switching network yielding routing solutions for a pre-given set of applications. Like field programmable gate array (FPGA) routing networks, the RIN is used to make reconfigurable interconnections among functional blocks. Unlike FPGAs, the network topology of a RIN is irregular as it is designed for a given set of routing requirements and optimized for the area cost subject to given delay constraints. In this paper, we propose an automatic design scheme for RINs, including routing specification formulation, graph modelings, network topology designs, routing algorithms and multiplexer-based network circuit implementation. The choice of the design scheme is based on the existing routing network design practices and research, which give feasible solutions. Our scheme is to optimize the designs with the choice of design parameters. A computer-aided design (CAD) tool is developed based on the design scheme, which takes a set of routing requirements as input and produces the corresponding RIN network topology and network circuit in hardware description language format. We present the area costs of various RINs generated by the CAD tool subject to delay constraints, and illustrate the RIN design scheme with a reconfigurable multistream video system. Hongbing Fan, Yu-Liang Wu, Ray C. C. Cheung |
Comput. J. | 3 |
| 2013 | A memory-based NFA regular expression match engine for signature-based intrusion detection
Derek Chi-Wai Pao, Nga Lam Or, Ray C. C. Cheung |
Comput. Commun. | 3 |
| 2013 | HEALPIX DCT technique for compressing PCA-based illumination adjustable images
John Sum, Andrew Chi-Sing Leung, Ray C. C. Cheung, Tze-Yui Ho |
Neural Comput. Appl. | 3 |
| 2012 | Area-Efficient Architectures for Large Integer and Quadruple Precision Floating Point MultipliersabstractLarge integer multiplication and floating point multiplication are the two dominating operations for many scientific and cryptographic applications. Large integer multipliers generally have linearly but high area requirement according to a given bit-width. High precision requirements of a given application lead to the use of quadruple precision arithmetic, however its operation is dominated by large integer multiplication of the mantissa product. In this paper, we propose a hardware efficient approach for implementing a fully pipelined large integer multipliers, and further extending it to Quadruple Precision (QP) floating point multiplication. The proposed design uses less hardware resources in terms of DSP48 blocks and slices, while attaining high performance. Promising results are obtained when compared our designs with the best reported large integer multipliers and also QP floating point multiplier in literatures. For instance, our results have demonstrated a significant improvement for the proposed QP multiplier, for over 50% improvement in terms of the DSP48 block usage with a penalty of slight additional slices, when compared to the best result in the literature on a Virtex-4 device. Manish Kumar Jaiswal, Ray C. C. Cheung |
FCCM | 2 |
| 2012 | Low complexity and hardware-friendly spectral modular multiplicationabstractThe Schönhage-Strassen Algorithm (SSA) is an asymptotically fast multiplication algorithm with the complexity of O(l log l log log l) where l is the operand size. It outperforms other multiplication algorithms when l is large enough. One possible usage of such long integer multiplication is for cryptography. Innovated from SSA, the Interleaved Spectral Montgomery Modular Multiplication (ISM3) algorithm is proposed to accelerate the modular multiplication. ISM3 algorithm primarily interleaves the Montgomery modular multiplication algorithm between time and spectral (frequency) domain. We show that the tasks in each step of the proposed algorithm have little data dependency, and hence, extremely suitable for hardware implementation. We present the parallel ISM3architecture and implement it on Xilinx Virtex-II and Virtex-6 FPGAs. Experimental results show that our 3838-bit ISM3 is faster than the previous Montgomery multiplier. Moreover, our design can complete a 7678-bit modular multiplication in 3398 cycles in 17.98 μs on a Virtex-6 device. Donald Donglong Chen, Gavin Xiaoxu Yao, Çetin Kaya Koç, Ray C. C. Cheung |
FPT | 4 |
| 2012 | GPU-Based Biclustering for Neural Information Processing
Alan W. Y. Lo, Benben Liu, Ray C. C. Cheung |
ICONIP (5) | 3 |
| 2012 | An FPGA-based acceleration platform for auction algorithmabstractAuction algorithms have been applied in various linear network problems, such as assignment, transportation, max-flow and shortest path problem. The inherent parallel characteristics of these algorithms are well suited for FPGA hardware implementation. In this paper, we focus on the acceleration of auction algorithm to solve assignment problem. The main contribution is to set up a flexible platform to generate efficient and extendable application-based hardware acceleration. It aims at solving both symmetric and asymmetric assignment problem. Experimental results show that 10X speedup can be achieved using 128 Processing Elements for the problem size of 500. Ray C. C. Cheung, Bryan Hu |
ISCAS | 4 |
| 2012 | Faster Pairing Coprocessor Architecture
Gavin Xiaoxu Yao, Junfeng Fan, Ray C. C. Cheung, Ingrid Verbauwhede |
Pairing | 3 |
| 2012 | Hypergraph based geometric biclustering algorithm
Zhiguan Wang, Chi Wai Yu, Ray C. C. Cheung, Hong Yan 0001 |
Pattern Recognit. Lett. | 3 |
| 2011 | FPGA Implementation of Pairings Using Residue Number System and Lazy Reduction
Ray C. C. Cheung, Sylvain Duquesne, Junfeng Fan, Nicolas Guillermin, Ingrid Verbauwhede, Gavin Xiaoxu Yao |
CHES | 1 |
| 2011 | FPGA Architecture of Generalized Laguerre-Volterra MIMO Model for Neural Population Spiking ActivitiesabstractWe present a parallelized and pipelined architecture for a generalized Laguerre-Volterra MIMO system to identify the time-varying neural dynamics underlying spike activities. The proposed architecture consists of a first stage containing a vector convolution and MAC (Multiply and Accumulation) component, a second stage containing a pre-threshold potential updating unit with an error approximation function component, and a third stage consisting of a gradient calculation unit. A flexible and efficient architecture that can accommodate different design speed requirements are generated. Simulation results are rigorously analyzed. A hardware IP library for versatile models and applications is proposed. The design runs on a Xilinx Virtex-6 FPGA and the processing core produces data samples at a maximum clock rate of 357MHz, which is 4.37 × 105times faster than the corresponding software model running on an AMD Pheono 9750 Quad Core Processor. It occupies 216,766 LUTs, maximum 12 block-RAMs, and 2016 DSP-blocks. Xiangyu Li 0005, Ray C. C. Cheung, Wei Zhang 0044, Rosa H. M. Chan, Dong Song, Theodore W. Berger |
FCCM | 2 |
| 2011 | FPGA Architecture of Generalized Laguerre-Volterra MIMO Model for Neural Population ActivitiesabstractWe present a full-parallelized and pipelined architecture for a generalized Laguerre-Volterra MIMO system to identify the time-varying neural dynamics underlying spike activities. The proposed architecture consists of a first stage containing a vector convolution and MAC (Multiply and Accumulation) component, a second stage containing a pre-threshold potential updating unit with an error approximation function component, and a third stage consisting of a gradient calculation unit. A flexible and efficient architecture that can accommodate different design speed requirements is generated. The design runs on a Xilinx Virtex-6 FPGA and the processing core produces data samples at a speed of 1.33×106data frames/sec, which is 3.1×103times faster than the corresponding C model running on an Intel i7-860 Quad Core Processor. Xiangyu Li 0005, Rosa H. M. Chan, Wei Zhang 0044, C. W. Yu, Ray C. C. Cheung, Dong Song, Theodore W. Berger |
FPL | 5 |
| 2011 | Hydrate: Hybrid Reconfigurable Architecture ExpressionsabstractThis paper presents Hydrate (HYbriD Reconfigurable ArchiTecture Expressions), a generic architecture description language for exploring hybrid FPGA designs. Hydrate is based on the XML architecture description for the VPR tool. These expressions consist of variable, repeat and conditional statements to allow flexible, reusable and readable description for FPGA architectures, without modifying the VPR core. Two case studies, involving modern FPGA and floating-point applications, are used to illustrate our approach. Compared with VPR5.0 and VPR6.0, the architecture file format adopted by Hydrate is often much smaller and easier to understand, and over 90% of file size reduction for complex FPGAs can be achieved. Moreover, Hydrate is compatible with different versions of VPR, so that the powerful VPR tool flow can be used to explore future architectures. Chi Wai Yu, Fred Cox, Wayne Luk, Ray C. C. Cheung |
FPT | 4 |
| 2010 | Reconfigurable Number Theoretic Transform architectures for cryptographic applicationsabstractAs an important component of Spectral Modular Arithmetic (SMA) cryptographic co-processor, the efficient architectures of Number Theoretic Transforms (NTTs) on FPGA are discussed in this paper. We analyze characteristics of the NTTs for cryptographic applications, compare different arithmetic approaches, introduce an optimized solution for FPGA implementation, and developed several different architectures. Qualitative and quantitative analyses are provided to show the effectiveness of our proposed architectures. Gavin Xiaoxu Yao, Ray C. C. Cheung, Çetin Kaya Koç, Kim-Fung Man |
FPT | 2 |
| 2009 | A High-Performance Hardware Architecture for Spectral Hash AlgorithmabstractThe spectral hash algorithm is one of the round 1 candidates for the SHA-3 family, and is based on spectral arithmetic over a finite field, involving multidimensional discrete Fourier transformations over a finite field, data dependent permutations, rubic-type rotations, and affine and nonlinear functions. The underlying mathematical structures and operations pose interesting and challenging tasks for computer architects and hardware designers to create fast, efficient, and compact ASIC and FPGA realizations. In this paper, we present an efficient hardware architecture for the full 512-bit hash computation using the spectral hash algorithm. We have created a pipelined implementation on a Xilinx Virtex-4 XC4VLX200-11 FPGA which yields 100 MHz and occupies 38,328 slices, generating a throughput of 51.2 Gbps. Our fully parallel synthesized implementation shows that the spectral hash algorithm is about 100 times faster than the fastest SHA-1 implementation, while requiring only about 13 times as many logic slices. Ray C. C. Cheung, Çetin Kaya Koç, John D. Villasenor |
ASAP | 1 |
| 2009 | Hierarchical Segmentation for Hardware Function EvaluationabstractThis paper presents a method for evaluating functions based on piecewise polynomial approximations (splines) with a hierarchical segmentation scheme targeting hardware implementation. The methodology provides significant reduction in table size compared to traditional uniform segmentation approaches. The use of hierarchies involving uniform splines and splines with size varying by powers of two is particularly well suited for the coverage of nonlinear regions. The segmentation step is automated and supports user-supplied precision requirements and approximation method. Bit-widths of the coefficients and arithmetic operators are optimized to minimize circuit area and enable a guarantee of 1 unit in the last place (ulp) accuracy at the output. A coefficient transformation technique is also described, which significantly reduces the dynamic ranges of the fixed-point polynomial coefficients. The hierarchical segmentation method is illustrated using a set of functions including -(x/2) log2x, cos-1(x), radic(-ln(x)) , a high-degree rational function, ln(1+x), and 1/(1+x). Various degree-1 and degree-2 approximation results for precisions between 8 to 24 bits are given. Hardware realizations are demonstrated on a Xilinx Virtex-4 field-programmable gate array (FPGA). Dong-U Lee, Ray C. C. Cheung, Wayne Luk, John D. Villasenor |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2008 | Hardware Implementation Trade-Offs of Polynomial Approximations and InterpolationsabstractThis paper examines the hardware implementation tradeoffs when evaluating functions via piecewise polynomial approximations and interpolations for precisions up to 24 bits. In polynomial approximations, polynomials are evaluated using stored coefficients. Polynomial interpolations, however, require the coefficients to be computed on-the-fly using stored function values. Although it is known that interpolations require less memory than approximations at the expense of additional computation, the tradeoffs in memory, area, delay, and power consumption between the two approaches have not been examined in detail. This work quantitatively analyzes these tradeoffs for optimized approximations and interpolations across different functions and target precisions. Hardware architectures for degree-1 and degree-2 approximations and interpolations are described. The results show that the extent of memory savings realized by using interpolation is significantly lower than what is commonly believed. Furthermore, experimental results on a field-programmable gate array (FPGA) show that for high output precision, degree-1 interpolations offer considerable area and power savings over degree-1 approximations, but similar savings are not realized when degree-2 interpolations and approximations are compared. The availability of both interpolation-based and approximation-based designs offers a richer set of design tradeoffs than is available using either interpolation or approximation alone. Dong-U Lee, Ray C. C. Cheung, Wayne Luk, John D. Villasenor |
IEEE Trans. Computers | 2 |
| 2007 | Automatic Accuracy-Guaranteed Bit-Width Optimization for Fixed and Floating-Point SystemsabstractIn this paper we present Minibit+, an approach that optimizes the bit-widths of fixed-point and floating-point designs, while guaranteeing accuracy. Our approach adopts different levels of analysis giving the designer the opportunity to terminate it at any stage to obtain a result. Range analysis is achieved using a combined affine and interval arithmetic approach to reduce the number of bits. Precision analysis involves a coarse-grain and fine-grain analysis. The best representation, in fixed-point or floating-point, for the numbers is then chosen based on the range, precision and latency. Three case studies are used: discrete cosine transform, B-Splines and RGB to YCbCr color conversion. Our analysis can run over 200 times faster than current approaches to this problem while producing more accurate results, on average within 2-3% of an exhaustive search. William George Osborne, Ray C. C. Cheung, José Gabriel F. Coutinho, Wayne Luk, Oskar Mencer |
FPL | 2 |
| 2007 | Instrumented Multi-Stage Word-Length OptimizationabstractIn this paper we present a tool, LengthFinder, for optimizing word-lengths of hardware designs with fixed-point arithmetic based on analytical error models that guarantee accuracy. LengthFinder adopts a multi-stage approach, with four novel features. First, the code analysis stage selects loops to instrument, such that information about the number of iterations can be extracted to generate more accurate results. Second, aggressive heuristics are used to produce non-uniform word-lengths rapidly while meeting requirements from the guaranteed error functions. Third, a method capable of reducing the search space has been developed for data-partitioning with a variable word-length reduction. Fourth, a genetic algorithm with selective-crossover and high mutation probability is applied to obtain near-optimal results. The benefits of LengthFinder are illustrated with various case studies. We show that LengthFinder can run over 200 times faster than previous techniques (Lee et al., 2006), while producing more accurate results, relative to values obtained from integer linear programming. William George Osborne, José Gabriel F. Coutinho, Ray C. C. Cheung, Wayne Luk, Oskar Mencer |
FPT | 3 |
| 2007 | The exact channel density and compound design for generic universal switch blocksabstractA switch block of k sides W terminals on each side is said to be universal (a ( k , W )-USB) if it is routable for every set of 2-pin nets of channel density at most W . The generic optimum universal switch block design problem is to design a ( k , W )-USB with the minimum number of switches for every pair of ( k , W ). This problem was first proposed and solved for k =4 in Chang et al. [1996], and then solved for even W or for k ≤6 in Shuy et al. [2000] and Fan et al. [2002b]. No optimum ( k , W )-USB is known for k ≥7 and odd W ≥3. But it is already known that when W is a large odd number, a near-optimum ( k , W )-USB can be obtained by a disjoint union of ( W − f 2 ( k ))/2 copies of the optimum ( k , 2)-USB and a noncompound ( k , f 2 ( k ))-USB, where the value of f 2 ( k ) is unknown for k ≥8. In this article, we show that f 2 ( k ) = k +3− i /3, where 1≤ i ≤6 and i ≡ k (mod 6), and present an explicit design for the noncompound ( k , f 2 ( k ))-USB. Combining these two results we obtain the exact designs of ( k , W )-USBs for all k ≥7 and odd W ≥3. The new ( k , W )-USB designs also yield an efficient detailed routing algorithm. Hongbing Fan, Jiping Liu, Yu-Liang Wu, Ray C. C. Cheung |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2007 | Hardware Generation of Arbitrary Random Number Distributions From Uniform Distributions Via the Inversion MethodabstractWe present an automated methodology for producing hardware-based random number generator (RNG) designs for arbitrary distributions using the inverse cumulative distribution function (ICDF). The ICDF is evaluated via piecewise polynomial approximation with a hierarchical segmentation scheme that involves uniform segments and segments with size varying by powers of two which can adapt to local function nonlinearities. Analytical error analysis is used to guarantee accuracy to one unit in the last place (ulp). Compact and efficient RNGs that can reach arbitrary multiples of the standard deviation sigma can be generated. For instance, a Gaussian RNG based on our approach for a Xilinx Virtex-4 XC4VLX100-12 field-programmable gate array produces 16-bit random samples up to 8.2 sigma. It occupies 487 slices, 2 block-RAMs, and 2 DSP-blocks. The design is capable of running at 371 MHz and generates one sample every clock cycle. Ray C. C. Cheung, Dong-U Lee, Wayne Luk, John D. Villasenor |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2007 | A Flexible Architecture for Precise Gamma CorrectionabstractWe present a flexible hardware architecture for precise gamma correction via piece-wise linear polynomial approximations. Arbitrary gamma values, input bit widths, and output bit widths are supported. The gamma correction curve is segmented via a combination of uniform segments and segments whose sizes vary by powers of two. This segmentation method minimizes the number of segments required, while providing an efficient way for indexing the polynomial coefficients. The outputs are guaranteed to be accurate to one unit in the last place through an analytical bit-width analysis methodology. Hardware realizations of various gamma correction designs are demonstrated on a Xilinx Virtex-4 field-programmable gate array (FPGA). A pipelined 12-bit input/8-bit output design on an XC4VLX100-12 FPGA occupies 146 slices and one digital signal processing slice. It is capable of performing 378 million gamma correction operations per second. Dong-U Lee, Ray C. C. Cheung, John D. Villasenor |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2006 | Inversion-based hardware gaussian random number generator: A case study of function evaluation via hierarchical segmentationabstractWe present the design and implementation of a Gaussian random number generator (GRNG) via hierarchical segmentation. Gaussian samples are generated using the inversion method, which involves the evaluation of the inverse Gaussian cumulative distribution function (IGCDF). The IGCDF is highly nonlinear and is evaluated via piecewise polynomial approximations (splines) with a hierarchical segmentation scheme that involves uniform splines and splines with size varying by powers of two. This segmentation approach adapts the spline sizes according to the non-linearity of the function, allowing efficient evaluation of the IGCDF. Bit-widths of the fixed-point polynomial coefficients and arithmetic operators are optimized in an analytical manner to guarantee a precision accurate to one unit in the last place. Our architecture generates 16-bit Gaussian samples accurate to 8.2cr (standard deviations). A pipelined implementation on a Xilinx Virtex-4 XC4LX100-12 FPGA yields 371 MHz and occupies 543 slices, 2 block RAMs, and 2 DSP slices, generating one sample every clock cycle Dong-U Lee, Ray C. C. Cheung, John D. Villasenor, Wayne Luk |
FPT | 2 |
| 2006 | Decomposition Design Theory and Methodology for Arbitrary-Shaped Switch BoxesabstractWe consider the optimal design problem for arbitrary-shaped switch box, (r1...rk) which r, terminals are located on side i for i= 1...k and programmable switches are joining pairs of terminals from different sides. Previous investigations on switch box designs mainly focused on regular switch boxes in which all sides have the same number of terminals. By allowing different numbers of terminals on different sides, irregular switch boxes are more general and flexible for applications such as customized FPGAs and reconfigurable interconnection networks. The optimal switch box design problem is to design a switch box satisfying the given shape and routing capacity specifications with the minimum number of switches. We present a decomposition design method for a wide range of irregular switch boxes. The main idea of our method is to model a routing requirement as a nonnegative integer vector satisfying a system of linear equations and then derive a decomposition theory of routing requirements based on the theory of systems of linear Diophantine equations. The decomposition theory makes it possible to construct a large irregular switch box by combining small switch boxes of fixed sizes. Specifically, we can design a family of hyperuniversal (universal) (u-d + c)-SBs with B(h-) switches, where d and c are constant vectors and w is a scalar. We illustrate the design method by designing a class of optimal hyperuniversal irregular 3-sided switch boxes and a class of optimal rectangular universal switch boxes. Experimental results on the rectangular universal switch boxes with the VPR router show that the optimal design of irregular switch boxes does pay off. Hongbing Fan, Yu-Liang Wu, Ray C. C. Cheung, Jiping Liu |
IEEE Trans. Computers | 3 |
| 2006 | Accuracy-Guaranteed Bit-Width OptimizationabstractAn automated static approach for optimizing bit widths of fixed-point feedforward designs with guaranteed accuracy, called MiniBit, is presented. Methods to minimize both the integer and fraction parts of fixed-point signals with the aim of minimizing the circuit area are described. For range analysis, the technique in this paper identifies the number of integer bits necessary to meet range requirements. For precision analysis, a semianalytical approach with analytical error models in conjunction with adaptive simulated annealing is employed to optimize the number of fraction bits. The analytical models make it possible to guarantee overflow/underflow protection and numerical accuracy for all inputs over the user-specified input intervals. Using a stream compiler for field-programmable gate arrays (FPGAs), the approach in this paper is demonstrated with polynomial approximation, RGB-to-YCbCr conversion, matrix multiplication, B-splines, and discrete cosine transform placed and routed on a Xilinx Virtex-4 FPGA. Improvements for a given design reduce the area and the latency by up to 26% and 12%, respectively, over a design using optimum uniform fraction bit widths. Studies show that MiniBit-optimized designs are within 1% of the area produced from the integer linear programming approach Dong-U Lee, Altaf Abdul Gaffar, Ray C. C. Cheung, Oskar Mencer, Wayne Luk, George A. Constantinides |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2005 | Automating custom-precision function evaluation for embedded processorsabstractDue to resource and power constraints, embedded processors often cannot afford dedicated floating-point units. For instance, the IBM PowerPC processor embedded in Xilinx Virtex-II Pro FPGAs only supports emulated floating-point arithmetic, which leads to slow operation when floating-point arithmetic is desired. This paper presents a customizable mathematical library using fixed-point arithmetic for elementary function evaluation. We approximate functions via polynomial or rational approximations depending on the user-defined accuracy requirements. The data representation for the inputs and outputs are compatible with IEEE single-precision and double-precision floating-point formats. Results show that our 32-bit polynomial method achieves over 80 times speedup over the single-precision mathematical library from Xilinx, while our 64-bit polynomial method achieves over 30 times speedup. Ray C. C. Cheung, Dong-U Lee, Oskar Mencer, Wayne Luk, Peter Y. K. Cheung |
CASES | 1 |
| 2005 | Reconfigurable Elliptic Curve Cryptosystems on a ChipabstractThe paper presents a system-on-a-chip (SoC) architecture, which targets reconfigurable hardware, for elliptic curve cryptosystems (ECC). A four-level partitioning scheme is described for exploring the area and speed tradeoffs. A design generator is used to generate parameterisable building blocks for the configurable SoC architecture. A secure Web server, which runs on a reconfigurable soft-processor and an embedded hard-processor, shows over 2000 times speedup when computationally-intensive operations run on the customised building blocks. The embedded on-chip timer block gives accurate performance information. The design factors of configurable SoC architectures are also discussed and evaluated. Ray C. C. Cheung, Wayne Luk, Peter Y. K. Cheung |
DATE | 1 |
| 2005 | Ziggurat-based Hardware Gaussian Random Number GeneratorabstractAn architecture and implementation of a high performance Gaussian random number generator (GRNG) is described. The GRNG uses the Ziggurat algorithm which divides the area under the probability density function into three regions (rectangular, wedge and tail). The rejection method is then used and this amounts to determining whether a random point falls into one of the three regions. The vast majority of points lie in the rectangular region and are accepted to directly produce a random variate. For the nonrectangular regions, which occur 1.5% of the time, the exponential or logarithm functions must be computed and an iterative fixed point operation unit is used. Computation of the rectangular region is heavily pipelined and a buffering scheme is used to allow the processing of rectangular regions to continue to operate in parallel with evaluation of the wedge and tail computation. The resulting system can generate 169 million normally distributed random numbers per second on a Xilinx XC2VP3O-6 device. Guanglie Zhang, Philip H. W. Leong, Dong-U Lee, John D. Villasenor, Ray C. C. Cheung, Wayne Luk |
FPL | 5 |
| 2005 | Reconfigurable Acceleration for Monte Carlo Based Financial Simulation
Guanglie Zhang, Philip H. W. Leong, Chun Hok Ho, Kuen Hung Tsoi, Chris C. C. Cheung, Dong-U Lee, Ray C. C. Cheung, Wayne Luk |
FPT | 7 |
| 2005 | Customizable elliptic curve cryptosystemsabstractThis paper presents a method for producing hardware designs for elliptic curve cryptography (ECC) systems over the finite field GF(2/sup m/), using the optimal normal basis for the representation of numbers. Our field multiplier design is based on a parallel architecture containing multiple m-bit serial multipliers; by changing the number of such serial multipliers, designers can obtain implementations with different tradeoffs in speed, size and level of security. A design generator has been developed which can automatically produce a customised ECC hardware design that meets user-defined requirements. To facilitate performance characterization, we have developed a parametric model for estimating the number of cycles for our generic ECC architecture. The resulting hardware implementations are among the fastest reported: for a key size of 270 bits, a point multiplication in a Xilinx XC2V6000 FPGA at 35 MHz can run over 1000 times faster than a software implementation on a Xeon computer at 2.6 GHz. Ray C. C. Cheung, N. J. Telle, Wayne Luk, Peter Y. K. Cheung |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2004 | A System on Chip Design Framework for Prime Number Validation Using Reconfigurable Hardware
Ray C. C. Cheung |
FPL | 1 |
| 2004 | On Optimal Irregular Switch Box Designs
Hongbing Fan, Yu-Liang Wu, Ray C. C. Cheung, Jiping Liu |
FPL | 3 |
| 2004 | A scalable hardware architecture for prime number validationabstractThis work presents a scalable architecture for prime number validation which targets reconfigurable hardware. The primality test is crucial for security systems, especially for most public-key schemes. The Rabin-Miller Strong Pseudoprime Test has been mapped into hardware, which makes use of a circuit for computing Montgomery modular exponentiation to further speed up the validation and to reduce the hardware cost. A design generator has been developed to generate a variety of scalable and non-scalable Montgomery multipliers based on user-defined parameters. The performance and resource usage of our designs, implemented in Xilinx reconfigurable devices, have been explored using very large prime numbers. Our work demonstrates the flexibility and trade-offs in using reconfigurable platform for prototyping cryptographic hardware in embedded systems. It is shown that, for instance, a 1024-bit primality test can be completed in less than a second, and a low cost XC3S2000 FPGA chip can accommodate a 32k-bit scalable primality test with 64 parallel processing elements. Ray C. C. Cheung, Ashley Brown, Wayne Luk, Peter Y. K. Cheung |
FPT | 1 |
| 2003 | An FPGA-based re-configurable 24-bit 96kHz sigma-delta audio DACabstractThis paper presents a reconfigurable sigma-delta audio Digital-to-Analog Converter (DAC) which is suitable for embedded FPGA applications. The Sigma-Delta Modulator (SDM) design can be configured as a 3rd or 5th order SDM and allows different input word lengths. Different input sampling rates are also entertained by employing a programmable interpolator. The DAC accepts 16-/18-/20-/24-bit PCM data at sampling rates of 32/44.1/48/88.2/96 kHz for applications in CD, SACD and DVD audio. Ray C. C. Cheung, Kong-Pang Pun, Steve C. L. Yuen, Kuen Hung Tsoi, Philip H. W. Leong |
FPT | 1 |
| 2003 | On optimal hyperuniversal and rearrangeable switch box designsabstractThis paper explores theories on designing optimal multipoint interconnection structures, and proposes a simple switch box design scheme which can be directly applied to field programmable gate arrays (FPGAs), switch box designs, and communication switching network designs. We present a new hyperuniversal switch box designs with four sides and W terminals on each side, which is routable for every multipin net-routing requirement. This new design is proved to be optimum for W = 1, ..., 5 and close to optimum for W /spl ges/ 6 with 6.3 W switches. We also give a formal analysis and extensive benchmark experiments on routability comparisons between today's most well-known FPGA switch boxes like disjoint switch blocks (Xilinx XC4000 Type), Wilton's switch blocks, Universal switch blocks, and our Hyperuniversal switch boxes. We apply the design scheme to rearrangeable switching network designs targeting for applications of connecting multiple terminals (e.g., teleconferencing). Simply using a /spl kappa/-sided hyperuniversal switch block with a W /spl times/ W crossbar attached to each side, one can build a three-stage one-sided polygonal switching network capable of realizing every multipoint connection requirement on kW terminals. Besides, due to the fine-grained decomposition property of our design scheme, the new switch box designs are highly scalable and simple on physical layout and routing algorithm implementations. Hongbing Fan, Jiping Liu, Yu-Liang Wu, Ray C. C. Cheung |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2003 | Further improve circuit partitioning using GBAW logic perturbation techniquesabstractEfficient circuit partitioning is becoming more and more important as the size of modern circuits keeps increasing. Conventionally, circuit partitioning is solved without altering the circuit by modeling the circuit as a hypergraph for the ease of applying graph algorithms. However, there is room for further improvement on even optimal hypergraph partitioning results, if logic information can be applied for circuit perturbation. Such logic transformation based partitioning techniques are relatively less addressed. In this paper, we present a powerful multiway partitioning technique which applies efficient logic rewiring techniques for further improvement over already superior hypergraph partitioning results. The approach can integrate with any graph partitioner. We perform experiments on two-, three-, and four-way partitionings for MCNC benchmark circuits whose physical and logical information are both available. Our experimental results show that this partitioning approach is very powerful. For example, it can achieve a further 12.3% reduction in cut size upon already excellent pure graph partitioner (hMetis) results on two-way partitioning with an area penalty of only 0.34%. The outperforming results demonstrate the usefulness of this new partitioning technique. Yu-Liang Wu, Ray C. C. Cheung, David Ihsin Cheng, Hongbing Fan |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2002 | On Optimum Designs of Universal Switch Blocks
Hongbing Fan, Jiping Liu, Yu-Liang Wu, Ray C. C. Cheung |
FPL | 4 |
| 2001 | On Optimum Switch Box Designs for 2-D FPGAsabstractAll in-text\treferences\tunderlined\tin\tblue\tare\tlinked\tto\tpublications\ton\tResearchGate, letting you\taccess\tand\tread\tthem\timmediately. Hongbing Fan, Jiping Liu, Yu-Liang Wu, Ray C. C. Cheung |
DAC | 4 |
| 2001 | Further improve circuit partitioning using GBAW logic perturbation techniquesabstractEfficient circuit partitioning is gaining more importance with the increasing size of modern circuits. Conventionally, circuit partitioning is solved by modeling a circuit as a hypergraph for the ease of applying graph algorithms. However there exists room for further improvement on even optimum hypergraph partitioning results, if logic information can be applied for perturbation. In this paper we present a multi-way partitioning framework which can couple any excellent hypergraph partitioner and a noval logic perturbation based technique (GBAW) for further improvement over very excellent partitioning results. Our approach can integrate with any graph partitioner. We performed experiments on 2-, 3-, 4-, and 5-way partitionings for various circuits of different sizes from MCNC benchmarks. We have chosen the state-of-the-art hMetis-Kway to obtain high quality initial solutions for the experiments. Our experiments showed that this partitioning approach can achieve a further 15% reduction in cut size for 2-way partitioning with an area penalty of only 0.33%. The good results demonstrated the effectiveness of this new partitioning technique. Ray C. C. Cheung, Yu-Liang Wu, David Ihsin Cheng |
DATE | 1 |