EDBT 2026 Demo / reviewers in the wild / expert
Chiu-Wing Sham
dblp:74/6729
· DBLP profile ↗
52ranked-venue papers
9as first author
29since 2021 · last 2026
0000-0001-7007-6746ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 9 first-author · 4 since 2021Artificial intelligence and machine learning · 14 · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Computer networks · 5 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Memory-Efficient NTT Accelerator with a Highly Parallel Memory Mapping Scheme
Hengyu Ding, Houran Ji, Jinhang Chen, Chiu-Wing Sham, Yao Wang 0013 |
ISCAS | 5 |
| 2026 | Efficient FPGA Deployment of Power-of-Two Quantized Networks via Dynamic Hierarchical Base-Exponent Offset Encoding
Zongcheng Yue, Longyu Ma, Chiu-Wing Sham |
ISCAS | 4 |
| 2026 | Dual Student Discrepancy Correction for Semi-Supervised Medical Image SegmentationabstractABSTRACT Semi‐supervised medical image segmentation (SSMIS) has proven to be an effective solution that leverages limited labelled data and abundant unlabeled data, thereby significantly reducing the labour and cost associated with manual annotation. However, most of the existing teacher‐student frameworks are prone to suffer from confirmation bias during training, adversely affecting the performance of SSMIS. To address this challenge, we propose the Dual Student Discrepancy Correction framework (DSDC), which extends the Mean Teacher (MT) framework by incorporating an additional student model with identical architecture but independently updated parameters. This design mitigates the parameter coupling issue that may arise when updating the teacher model via Exponential Moving Average (EMA) in conventional single‐student paradigms. Moreover, the prediction discrepancy between the two student models is leveraged for error detection and correction, enabling the network to identify and rectify its own cognitive biases, ultimately enhancing segmentation accuracy. Comprehensive experiments on two public benchmarks, an MRI dataset (LA) and a CT dataset (Pancreas‐NIH), reveal that our DSDC framework surpasses current State‐of‐the‐Art (SOTA) approaches across all evaluation metrics. These findings substantiate the framework's effectiveness in SSMIS tasks. Code is accessible at https://github.com/Sangfugui/DSDC . Zhenfu Sang, Chong Fu 0001, Lin Cao 0003, Chiu-Wing Sham |
Expert Syst. J. Knowl. Eng. | 5 |
| 2026 | Enhancing medical image segmentation with collaborative and contrastive learning in mixed-domain settings
Haoming Yuan, Chong Fu 0001, Junxin Chen 0001, Xingwei Wang 0001, Chiu-Wing Sham |
Neurocomputing | 6 |
| 2026 | RedPIM: An Efficient PIM Accelerator Design with Reduced Analog-to-Digital ConversionsabstractReRAM-based Processing-In-Memory (PIM) architectures are compelling contenders for deep learning due to their ability to perform matrix-vector multiplications (MVMs) directly within the memory, significantly reducing data movement and enhancing computational efficiency. However, since MVMs occur in the analog domain, analog-to-digital converters (ADCs) dominate power consumption and area overhead in current implementations. To this end, we propose RedPIM, an efficient ReRAM-based PIM accelerator design for deep neural networks (DNNs) that reduces the number of analog-to-digital conversions. RedPIM exploits the fact that in ReRAM-based PIM accelerators, the overall energy consumption generally increases with the number of activated analog-to-digital conversions. Specifically, we introduce a novel training algorithm that is aware of the ADC overhead during activation value quantization and optimizes accuracy concurrently. From a hardware design perspective, we develop a lookup table (LUT)-based quantization module to enable efficient and low-cost activation value quantization. In addition, we propose an efficient adaptive operation unit (OU) size assignment scheme that further minimizes analog-to-digital conversions by considering activation sparsity and weight distribution. Extensive experimental results show our RedPIM reduces latency to 27.72% and energy consumption to 10.15% of the baseline, with minimal accuracy loss, making it a promising solution for enhancing DNN acceleration. The code for this project is available at: https://github.com/JialeLiLab/ADC_aware_Learning.git . Yulin Fu, Longyu Ma, Chiu-Wing Sham, Chong Fu 0001 |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2025 | LightFSA: A Lightweight Financial Sentiment Analysis Model
Chiu-Wing Sham, Longyu Ma, Chong Fu 0001 |
ICIC (15) | 3 |
| 2025 | A Novel Computing Paradigm for MobileNetV3 using MemristorabstractThe advancement in the field of machine learning is inextricably linked with the concurrent progress in domain-specific hardware accelerators such as GPUs and TPUs. However, the rapidly growing computational demands necessitated by larger models and increased data have become a primary bottleneck in further advancing machine learning, especially in mobile and edge devices. Currently, the neuromorphic computing paradigm based on memristors presents a promising solution. In this study, we introduce a memristor-based MobileNetV3 neural network computing paradigm and provide an end-to-end framework for validation. The results demonstrate that this computing paradigm achieves over 90% accuracy on the CIFAR10 dataset while saving inference time and reducing energy consumption. With the successful development and verification of MobileNetV3, the potential for realizing more memristor-based neural networks using this computing paradigm and open-source framework has significantly increased. This progress sets a groundbreaking pathway for future deployment initiatives. Longyu Ma, Chiu-Wing Sham, Chong Fu 0001 |
IJCNN | 4 |
| 2025 | LHA: Layer-wise Hardware Acceleration of Progressive Quantizing Inference through Partial Reconfiguration for Edge ComputingabstractAs the need for real-time, low-power deep learning at the edge increases, efficient hardware acceleration becomes crucial. Traditional edge hardware designs often scale to accommodate neural network sizes, which can degrade overall performance by taxing the hardware. To solve this, we propose a novel Layer-wise Hardware Acceleration (LHA) approach for Deep Neural Network (DNN) inference, leveraging progressive quantization and Partial Reconfiguration (PR). We first apply progressive quantization to systematically reduce the bit-width of network weights and activations, lowering computational and memory demands. Then, we utilize Field Programmable Gate Arrays (FPGAs) with PR capabilities to dynamically reconfigure hardware for each quantized network layer in sequence. This method optimizes FPGA resource usage, tailors to each layer’s needs, and reallocates freed resources to boost overall performance. Experiments show that LHA significantly enhances resource efficiency while maintaining inference performance on edge devices. Zongcheng Yue, Longyu Ma, Chiu-Wing Sham, Chong Fu 0001 |
IJCNN | 3 |
| 2025 | Joint Post-Training Pruning and Power-of-Two Quantization for Efficient Edge ComputingabstractRecent advancements in deep neural networks have created significant challenges for deploying these models on edge devices due to their computational and memory demands. We propose a novel integrated compression framework that combines nonlinear orthogonality-based channel pruning with progressive power-of-two (PoT) quantization to achieve efficient model compression for edge computing. Our framework first employs Radial Basis Function (RBF) kernel-based nonlinear orthogonality measurement to identify and remove redundant channels while preserving essential feature representations, then applies a layer-wise progressive power-of-two quantization scheme that enables efficient hardware implementation through bit-shift operations. Comprehensive experiments on CIFAR-10 and ImageNet demonstrate the effectiveness of our approach. On VGG16 with CIFAR-10, our method achieves 92.36% accuracy while reducing model size by 98.7% and computational complexity by 98.4%. On ResNet50 with ImageNet, we maintain 75.01% accuracy while achieving 95.93% model size reduction and 97.24% computational complexity reduction. Our framework significantly outperforms existing methods in terms of compression ratio and hardware efficiency while maintaining competitive accuracy. Zongcheng Yue, Longyu Ma, Chiu-Wing Sham, Chong Fu 0001 |
IJCNN | 3 |
| 2025 | Edge Priors Image Inpaintig With StyleGAN2abstractABSTRACT Image inpainting represents a fundamental task in computer vision, focusing primarily on the generation of missing content within an image to restore its integrity and aesthetics. Existing GAN‐based approaches often produce content with ambiguity and require a high training difficulties. Moreover, they tend to focus narrowly on damaged regions, leading to edge distortions that hinder generalisation. To address these challenges, we propose an algorithm that consist of two distinct networks. The first network, called Edge‐e4e, is designed for initial image restoration and integrates a pre‐trained StyleGAN2 as the generator to mitigate edge distortions. This network employs an encoder‐StyleGAN2 architecture, where only the encoder part is trained, thereby reducing training costs compared to traditional GAN methods. To resolve ambiguities in the restored content, we incorporate edge information into the damaged regions, guiding the network to generate content that is consistent with the original image. The second network, called Appending network, includes two style‐based encoders and a generator to improve the similarity between the images restored by Edge‐e4e and the original images. Specifically, we subtract the restored images from the input images in the channel dimension to obtain distortion maps, which serve as a prior to refine the restored images from Edge‐e4e. To further enhance the quality of refined images, we propose incorporating plugin and modulate plugin modules for style extraction and fusion. These modules utilise information from the input images and seamlessly integrate it into the style‐based generator. Experimental results demonstrate that our algorithm achieves high‐fidelity restoration and excellent generalisation, with optimal FID and Lpips metrics of 0.0631 and 0.875, respectively. The code is publicly available at: https://github.com/MengZhen‐Chi/Edge‐Pries‐Image‐Inpainting‐with‐StyleGAN2 . Mengzhen Chi, Chong Fu 0001, Xu Zheng 0002, Jialei Chen 0001, Chiu-Wing Sham |
Expert Syst. J. Knowl. Eng. | 6 |
| 2025 | Feature space expansion and compression with spatial-spectral augmentation for hyperspectral image Class-Incremental Learning
Ran Wu, Zongcheng Yue, Chiu-Wing Sham, Junbao Li |
Pattern Recognit. | 4 |
| 2025 | Multi-Scale Dynamic Sparse Attention UNet for Medical Image SegmentationabstractTransformers have recently gained significant attention in medical image segmentation due to their ability to capture long-range dependencies. However, the presence of excessive background noise in large regions of medical images introduces distractions and increases the computational burden on the fine-grained self-attention (SA) mechanism, which is a key component of the transformer model. Meanwhile, preserving fine-grained details is essential for accurately segmenting complex, blurred medical images with diverse shapes and sizes. Thus, we propose a novel Multi-scale Dynamic Sparse Attention (MDSA) module, which flexibly reduces computational costs while maintaining multi-scale fine-grained interactions with content awareness. Specifically, multi-scale aggregation is first applied to the feature maps to enrich the diversity of interaction information. Then, for each query, irrelevant key-value pairs are filtered out at a coarse-grained level. Finally, fine-grained SA is performed on the remaining key-value pairs. In addition, we design an enhanced downsampling merging (EDM) module and an enhanced upsampling fusion (EUF) module for building pyramid architectures. Using MDSA to construct the basic blocks, combined with EDMs and EUFs, we develop a UNet-like model named MDSA-UNet. Since MDSA-UNet dynamically processes only a small subset of relevant fine-grained features, it achieves strong segmentation performance with high computational efficiency. Extensive experiments on four datasets spanning three different types demonstrate that our MDSA-UNet, without using pre-training, significantly outperforms other non-pretrained methods and even competes with pre-trained models, achieving Dice scores of 82.10% on DDTI, 80.20% on TN3K, 90.75% on ISIC2018, and 91.05% on ACDC. Meanwhile, our model maintains lower complexity, with only 6.65 M parameters and 4.54 G FLOPs at a resolution of 224 × 224, ensuring both effectiveness and efficiency. Code is available at URL. Chong Fu 0001, Wenchao Zhang 0001, Junxin Chen 0001, Chiu-Wing Sham |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | Content-Aware Tunable Selective Encryption for HEVC Using Sine-Modular Chaotification ModelabstractExisting High Efficiency Video Coding (HEVC) selective encryption algorithms only consider the encoding characteristics of syntax elements to keep format compliance, but ignore the semantic features of video content, which may lead to unnecessary computational and bit rate costs. To tackle this problem, we present a content-aware tunable selective encryption (CATSE) scheme for HEVC. First, a deep hashing network is adopted to retrieve groups of pictures (GOPs) containing sensitive objects. Then, the retrieved sensitive GOPs and the remaining insensitive ones are encrypted with different encryption strengths. For the former, multiple syntax elements are encrypted to ensure security, whereas for the latter, only a few bypass-coded syntax elements are encrypted to improve the encryption efficiency and reduce the bit rate overhead. The keystream sequence used is extracted from the time series of a new improved logistic map with complex dynamic behavior, which is generated by our proposed sine-modular chaotification model. Finally, a reversible steganography is applied to embed the flag bits of the GOP type into the encrypted bitstream, so that the decoder can distinguish the encrypted syntax elements that need to be decrypted in different GOPs. Experimental results indicate that the proposed HEVC CATSE scheme not only provides high encryption speed and low bit rate overhead, but also has superior encryption strength than other state-of-the-art HEVC selective encryption algorithms. Qingxin Sheng, Chong Fu 0001, Zhaonan Lin, Junxin Chen 0001, Xingwei Wang 0001, Chiu-Wing Sham |
IEEE Trans. Multim. | 6 |
| 2025 | Content-Aware Selective Encryption for H.265/HEVC Using Deep Hashing Network and SteganographyabstractExisting selective encryption schemes for High Efficiency Video Coding (HEVC) only focus on the encoding characteristics of syntax elements in entropy coding and lack an understanding of the video content. Consequently, a large amount of unnecessary encryption operations are utilized to protect insensitive video frames, resulting in low encryption efficiency. In this article, we propose a content-aware selective encryption scheme for H.265/HEVC, which encrypts only the groups of pictures (GOPs) containing sensitive content and thus offers high efficiency. In our scheme, a deep hashing network is first adopted to retrieve video frames to determine the content-sensitive GOPs. Then, multiple prediction and residual syntax elements in sensitive GOPs are encrypted using a keystream sequence generated by the hyper-chaotic Lorenz system. In addition, the direct current coefficient of each \(4\times 4\) transform block is exchanged with a pseudo-randomly selected non-zero alternating current coefficient to further offer stronger visual distortion. Finally, the sign bits used for marking each GOP-type are reversibly embedded into the encrypted syntax elements to facilitate the decoder to distinguish the GOPs that need to be decrypted. Experimental results indicate that the proposed content-aware selective encryption scheme can efficiently protect sensitive content and is robust against all common attacks. Furthermore, it outperforms other state-of-the-art HEVC selective encryption algorithms in terms of security performance. Qingxin Sheng, Chong Fu 0001, Zhaonan Lin, Junxin Chen 0001, Xingwei Wang 0001, Chiu-Wing Sham |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2024 | Joint object contour points and semantics for instance segmentationabstractAbstract The edges of objects are of great significance to the task of instance segmentation. However, most of the current popular deep neural networks do not pay much attention to the object edge information. More importantly, using the down‐sampling pooling layer in the deep learning network, the edge detail information of the object will be lost. To address this issue, inspired by the manual annotation process, we propose Mask Point R‐CNN aiming at promoting the neural network's attention to the object boundary. Specifically, we introduce the auxiliary task of object contour point detection on the Mask R‐CNN framework, which can effectively improve the gradient flow between different tasks by multi‐task learning and repairing objects' boundary information via feature fusion. Consequently, the model can be more sensitive to the edges of the object and capture more geometric features. Quantitatively, the experimental results show that our Mask Point R‐CNN outperforms vanilla Mask R‐CNN by 3.8% on the Cityscapes dataset and 0.8% on the COCO dataset. Wenchao Zhang 0001, Chong Fu 0001, Mai Zhu, Lin Cao 0003, Ming Tie, Chiu-Wing Sham |
Expert Syst. J. Knowl. Eng. | 6 |
| 2024 | DMSA-UNet: Dual Multi-Scale Attention makes UNet more strong for medical image segmentation
Chong Fu 0001, Wenchao Zhang 0001, Chiu-Wing Sham, Junxin Chen 0001 |
Knowl. Based Syst. | 5 |
| 2024 | Attention-based deep supervised hashing for near duplicate video retrieval
Naifei Shi, Chong Fu 0001, Ming Tie, Wenchao Zhang 0001, Xingwei Wang 0001, Chiu-Wing Sham |
Neural Comput. Appl. | 6 |
| 2024 | Hyper-feature aggregation and relaxed distillation for class incremental learning
Ran Wu, Zongcheng Yue, Junbao Li, Chiu-Wing Sham |
Pattern Recognit. | 5 |
| 2024 | A Chaos-Based Tunable Selective Encryption Algorithm for H.265/HEVC With Semantic UnderstandingabstractExisting H.265/HEVC selective encryption (SE) schemes do not take into account the semantic features of input videos, nor do they adjust the encryption syntax elements according to the sensitivity of video content, which greatly limits their applicability. In this paper, we propose a chaos-based tunable H.265/HEVC SE scheme with semantic understanding. First, a deep hashing network is employed to identify content-sensitive videos by analyzing the semantic features of video sequences. Then, the non-sensitive videos and the retrieved sensitive ones are encrypted with different encryption strengths, respectively. Specifically, for non-sensitive videos, seven syntax elements with bypass-coded bins are selected for encryption at a constant bit rate. Hence, the encrypted bitstream keeps exactly the same compression ratio. To provide heavier visual distortion for content-sensitive videos, the regular-coded bins of four syntax elements and the intra prediction mode (IPM) are encrypted based on their corresponding encoding characteristics as well. Additionally, the selected syntax elements are all masked using a keystream generated by a chaotic system to ensure real-time constraints. Experimental results demonstrate that our suggested scheme offers format compatibility and is secure against all common attacks. Meanwhile, it outperforms state-of-the-art SE schemes in terms of security strength. Furthermore, the proposed scheme can be flexibly used in a wide range of applications according to the user’s requirements for encryption strength and bit rate. Qingxin Sheng, Chong Fu 0001, Ming Tie, Xingwei Wang 0001, Junxin Chen 0001, Chiu-Wing Sham |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | JPEG-compatible Joint Image Compression and Encryption Algorithm with File Size PreservationabstractJoint image compression and encryption algorithms are intensively investigated due to their powerful capability of simultaneous image data compression and sensitive information protection. Unfortunately, most of the existing algorithms suffered from either poor compression efficiency or weak encryption strength, making them vulnerable to cryptanalysis. To address these limitations, we propose a chaos-based JPEG-compatible joint image compression and encryption algorithm. We separate the luminance and chrominance coefficients to preserve file size and encrypt the discrete cosine transform (DCT) coefficients in parallel. The proposed inter-block DC encryption strategy achieves high encryption intensity based on the permutation-substitution structure. In addition, we apply both inter- and intra-block permutations to AC coefficients and strengthen the encryption using an inter-block substitution for non-zero AC coefficients. The results of security and performance analyses demonstrate that the proposed algorithm offers robust encryption of image data while maintaining compression efficiency for real-time transmission. Yuxiang Peng 0003, Chong Fu 0001, Guixing Cao, Junxin Chen 0001, Chiu-Wing Sham |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2024 | An efficient chaotic image encryption scheme using simultaneous permutation-diffusion operation
Qingxin Sheng, Chong Fu 0001, Zhaonan Lin, Junxin Chen 0001, Lin Cao 0003, Chiu-Wing Sham |
Vis. Comput. | 6 |
| 2023 | A one-time-pad-like chaotic image encryption scheme using data steganography
Qingxin Sheng, Chong Fu 0001, Zhaonan Lin, Ming Tie, Junxin Chen 0001, Chiu-Wing Sham |
J. Inf. Secur. Appl. | 6 |
| 2022 | CODH++: Macro-semantic differences oriented instance segmentation network
Wenchao Zhang 0001, Chong Fu 0001, Lin Cao 0003, Chiu-Wing Sham |
Expert Syst. Appl. | 4 |
| 2022 | A more compact object detector head network with feature enhancement and relational reasoning
Wenchao Zhang 0001, Chong Fu 0001, Xiang shi Chang, Teng fei Zhao, Chiu-Wing Sham |
Neurocomputing | 6 |
| 2022 | Protection of image ROI using chaos-based encryption and DCNN-based object detection
Chong Fu 0001, Yu Zheng 0021, Lin Cao 0003, Ming Tie, Chiu-Wing Sham |
Neural Comput. Appl. | 6 |
| 2022 | A fast parallel batch image encryption algorithm using intrinsic properties of chaos
Chong Fu 0001, Ming Tie, Chiu-Wing Sham, Hong-feng Ma |
Signal Process. Image Commun. | 4 |
| 2022 | Hardware Architecture of Layered Decoders for PLDPC-Hadamard CodesabstractProtograph-based low-density parity-check Hadamard codes (PLDPC-HCs) are a new type of ultimate-Shannon-limit-approaching codes. In this paper, we propose a hardware architecture for the PLDPC-HC layered decoders. The decoders consist mainly of random address memories, Hadamard sub-decoders and control logics. Two types of pipelined structures are presented and the latency and throughput of these two structures are derived. Implementation of the decoder design on an FPGA board shows that a throughput of 1.48 Gbps is achieved with a bit error rate (BER) of$10^{-5}$at around$E_{b}/N_{0}=-0.40$dB. The decoder can also achieve the same BER at$E_{b}/N_{0}=-1.14$dB with a reduced throughput of 0.20 Gbps. Pengwei Zhang 0001, Sheng Jiang 0001, Francis C. M. Lau 0002, Chiu-Wing Sham |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2022 | Spatially Coupled PLDPC-Hadamard Convolutional CodesabstractWe propose a new type of ultimate-Shannon-limit-approaching codes called spatially coupled protograph-based low-density parity-check Hadamard convolutional codes (SC-PLDPCH-CCs), which are constructed by spatially coupling PLDPC-Hadamard block codes. We develop an efficient decoding algorithm that combines pipeline decoding and layered scheduling for the decoding of SC-PLDPCH-CCs, and analyze the latency and complexity of the decoder. To estimate the decoding thresholds of SC-PLDPCH-CCs, we first propose a layered protograph extrinsic information transfer (PEXIT) algorithm to evaluate the thresholds of spatially coupled PLDPC-Hadamard terminated codes (SC-PLDPCH-TDCs) with a moderate coupling length. With the use of the proposed layered PEXIT method, we develop a genetic algorithm to find good SC-PLDPCH-TDCs in a systematic way. Then we extend the coupling length of these SC-PLDPCH-TDCs to form good SC-PLDPCH-CCs. Results show that our constructed SC-PLDPCH-CCs can achieve comparable thresholds to the block code counterparts. Simulations illustrate the superiority of the SC-PLDPCH-CCs over the block code counterparts and other state-of-the-art low-rate codes in terms of error performance. For the rate-0.00295 SC-PLDPCH-CC, a bit error rate of 10−5is achieved at$E_{b}/N_{0} = -1.465$dB, which is only 0.125 dB from the ultimate Shannon limit. Pengwei Zhang 0001, Francis C. M. Lau 0002, Chiu-Wing Sham |
IEEE Trans. Commun. | 3 |
| 2021 | Protograph-Based LDPC Hadamard CodesabstractIn this paper, we propose a new method to design low-density parity-check Hadamard (LDPC-Hadamard) codes — a type of ultimate-Shannon-limit approaching channel codes. The technique is based on applying Hadamard constraints to the check nodes in a generalized protograph-based LDPC code, followed by lifting the generalized protograph. We name the codes formedprotograph-basedLDPC Hadamard (PLDPC-Hadamard) codes. We also propose a modified Protograph Extrinsic Information Transfer (PEXIT) algorithm for analyzing and optimizing PLDPC-Hadamard code designs. The proposed algorithm further allows the analysis of PLDPC-Hadamard codes with degree- 1 and/or punctured nodes. We find codes with decoding thresholds ranging from −1.53 dB to −1.42 dB. At a BER of 10−5, the gaps of our codes to the ultimate-Shannon-limit range from 0.40 dB (for rate = 0.0494) to 0.16 dB (for rate = 0.003). Moreover, the error performance of our codes is comparable to that of the traditional LDPC-Hadamard codes. Finally, the BER performances of our codes after puncturing are simulated and compared. Pengwei Zhang 0001, Francis C. M. Lau 0002, Chiu-Wing Sham |
IEEE Trans. Commun. | 3 |
| 2020 | Protograph-based LDPC-Hadamard CodesabstractWe propose a novel type of ultimate-Shannon-limit-approaching codes, namely protograph-based low-density parity-check Hadamard (PLDPC-Hadamard) codes in this paper. We also propose a systematic way of analyzing such codes using Protograph EXtrinsic Information Transfer (PEXIT) charts. Using the analytical technique we have found a code of rate about 0.05 having a theoretical threshold of -1.42dB. At a BER of 10-5, the gaps of our code to the Shannon capacity for R=0.05 and to the ultimate Shannon limit are 0.25 dB and 0.40 dB, respectively. Pengwei Zhang 0001, Francis C. M. Lau 0002, Chiu-Wing Sham |
WCNC | 3 |
| 2018 | Revisiting Routability-Driven Placement for Analog and Mixed-Signal CircuitsabstractThe exponential increase in scale and complexity of very large-scale integrated circuits (VLSIs) poses a great challenge to current electronic design automation (EDA) techniques. As an essential step in the whole EDA layout synthesis, placement is attracting more and more attention, especially for analog and mixed-signal integrated circuits. Recently, experts in this field have observed a variety of analog-specific layout constraints to obtain high-performance placement solutions. These constraints include symmetry, alignment, boundary, preplace, abutment, range and maximum separation, and routability of the placement solutions. In this article, the effectiveness of slicing and nonslicing representation is investigated. Additionally, the technique of congestion-based virtual sizing is proposed. Experimental results show that the routability can be improved significantly by applying congestion-based virtual sizing. Results also show that the slicing representation can improve the regularity of the placement solutions and hence improve the routability with higher efficiency compared to the nonslicing representation. Hongxia Zhou, Chiu-Wing Sham, Hailong Yao 0002 |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2017 | Design and error performance of punctured hadamard codesabstractThis paper investigates the puncturing pattern design of Hadamard codes. When designing the puncturing patterns, the minimum Hamming distance property and cross-correlation property of the punctured Hadamard codes are considered. The theoretical upper bound of the minimum Hamming distance is derived for the punctured Hadamard codes and an algorithm is subsequently proposed for searching punctured codes with the minimum distance approaching the bound. Another approach to design the punctured codes is to minimize the cross-correlations between punctured codewords. A new metric is defined to measure the overall severity of the cross-correlation based on which punctured codes with the “lowest” cross-correlations are found. Simulation results show that both techniques are effective in searching for good punctured Hadamard codes. As expected, moreover, the punctured codes found by minimizing the cross-correlation outperform those optimized by maximizing minimum Hamming distance when a-posteriori-probability decoding is used. Sheng Jiang 0001, Francis C. M. Lau 0002, Wai M. Tam, Chiu-Wing Sham |
APCC | 4 |
| 2017 | Design of a high-throughput low-latency extended golay decoderabstractIn this paper, we propose a parallel architecture of the imperfect maximum likelihood decoding (IMLD) method, called PIMLD. It is further implemented onto an FPGA and applied to decode the (24,12,8) extended Golay code. Experimental results show that the proposed PIMLD decoder achieves 12.0 Gb/s throughput at 500 MHz frequency. Moreover, the latency for the decoder is only 5 clock cycles. Pengwei Zhang 0001, Francis C. M. Lau 0002, Chiu-Wing Sham |
APCC | 3 |
| 2016 | Rapid prototyping of multi-mode QC-LDPC decoder for 802.11n/ac standardabstractA multi-mode QC-LDPC decoder is proposed to satisfy the 802.11n/ac WiFi standard. With code-specific design, the overall performance of the decoder is enhanced while ensuring an on-the-fly reconfigurable ability. The proposed architecture has been synthesized using an FPGA for measurements. A state-of-art error rate and implementation complexity are reported. Meanwhile, the throughput has been increased to range from 382 MHz to 1852 MHz. Qing Lu 0003, Chiu-Wing Sham, Francis C. M. Lau 0002 |
ASP-DAC | 2 |
| 2015 | An architecture-algorithm co-design of artificial intelligence for Trax playerabstractTrax is a two-player game of simple rules but strategic depth. This article proposes an FPGA-based artificial intelligence for its endless version called Supertrax. An implementable algorithm is developed by combining several strategies and techniques functioning at various levels of software and hardware. These methods are developed using heuristics, multi-level pattern recognition, Monte-Carlo Tree Search, and path-based scheduling. A specific architecture has also been described to accommodate this algorithm. The proposal contributes a novel idea on this subject and its performance will be shown in the design competition to be held in FPT'15. Qing Lu 0003, Chiu-Wing Sham, Francis C. M. Lau 0002 |
FPT | 2 |
| 2015 | SIAR: Customized real-time interactive router for analog circuits
Hailong Yao 0002, Yici Cai, Qiang Zhou 0001, Chiu-Wing Sham |
Integr. | 5 |
| 2014 | Obstacle-avoiding rectilinear Steiner tree construction in sequential and parallel approach
Wing-Kai Chow, Evangeline F. Y. Young, Chiu-Wing Sham |
Integr. | 4 |
| 2012 | A new clock network synthesizer for modern VLSI designs
Jingwei Lu, Wing-Kai Chow, Chiu-Wing Sham |
Integr. | 3 |
| 2012 | Fast Power- and Slew-Aware Gated Clock Tree SynthesisabstractClock tree synthesis plays an important role on the total performance of chip. Gated clock tree is an effective approach to reduce the dynamic power usage. In this paper, two novel gated clock tree synthesizers, power-aware clock tree synthesizer (PACTS) and power- and slew-aware clock tree synthesizer (PSACTS), are proposed with zero skew achieved based on Elmore RC model. In PACTS, the topology of the clock tree is constructed with simultaneous buffer/gate insertion, which reduces the switched capacitance. In PSACTS, a more practical clock slew constraint is applied. Compared to previous works, clock tree synthesis is done first and followed by the insertions of clock gates. The clock slew changes a lot after the insertions of clock gates in real cases. In our work, the clock tree is constructed simultaneously with the insertions of clock gates. This ensures the limitation of the clock slew can be strictly satisfied while the limitation of the clock slew is always applied in the real design. The experimental results show that the power cost of our work is smaller and the runtime is reduced. The slew rate constraint is satisfied with a small clock skew from SPICE estimation. Generally, our work has better performance, improved efficiency and is more practical to be applied in the industry. Jingwei Lu, Wing-Kai Chow, Chiu-Wing Sham |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2011 | Efficient Decoding of QC-LDPC Codes Using GPUs
Yue Zhao 0011, Xu Chen 0018, Chiu-Wing Sham, Wai Man Tam, Francis C. M. Lau 0002 |
ICA3PP (1) | 3 |
| 2010 | A dual-MST approach for clock network synthesisabstractIn nanometer-scale VLSI physical design, clock network becomes a major concern on determining the total performance of digital circuit. Clock skew and PVT (process, voltage and temperature) variations contribute a lot to its behavior. Previous works mainly focused on skew and wirelength minimization. It may lead to negative influence towards these process variation factors. In this paper, a novel clock network synthesizer is proposed and several algorithms are introduced for performance improvement. A dual-MST (DMST) geometric matching approach is proposed for topology construction. It can help balancing the tree structure to reduce the variation effect. A recursive buffer insertion technique and a blockage handling method are also presented, and they are developed for proper distribution of buffers and saving of capacitance. Experimental results show that our matching approach is better than the traditional methods, and in particular our synthesizer has better performance compared to the results of the winner in the ISPD 2009 contest. Jingwei Lu, Wing-Kai Chow, Chiu-Wing Sham, Evangeline F. Y. Young |
ASP-DAC | 3 |
| 2009 | Block flipping and white space distribution for wirelength minimization
Chiu-Wing Sham, Evangeline F. Y. Young |
Integr. | 1 |
| 2009 | Congestion prediction in early stages of physical designabstractRoutability optimization has become a major concern in physical design of VLSI circuits. Due to the recent advances in VLSI technology, interconnect has become a dominant factor of the overall performance of a circuit. In order to optimize interconnect cost, we need a good congestion estimation method to predict routability in the early designing stages. Many congestion models have been proposed but there's still a lot of room for improvement. Besides, routers will perform rip-up and reroute operations to prevent overflow, but most models do not consider this case. The outcome is that the existing models will usually underestimate the routability. In this paper, we have a comprehensive study on our proposed congestion models. Results show that the estimation results of our approaches are always more accurate than the previous congestion models. Chiu-Wing Sham, Evangeline F. Y. Young, Jingwei Lu |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2008 | Optimizing wirelength and routability by searching alternative packings in floorplanningabstractRecent advances in VLSI technology have made optimization of the interconnect delay and routability of a circuit more important. We should consider interconnect planning as early as possible. We propose a postfloorplanning step to reduce the interconnect cost of a floorplan by searching alternative packings. If a packing contains a rectangular bounding box of a group of modules, we can rearrange the blocks in the bounding box to obtain a new floorplan with the same area, but possibly with a smaller interconnect cost. Experimental results show that we can reduce the interconnect cost of a packing without any penalty in area. Chiu-Wing Sham, Evangeline F. Y. Young, Hai Zhou 0001 |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2007 | Area reduction by deadspace utilization on interconnect optimized floorplanabstractInterconnect optimization has become the major concern in floorplanning. Many approaches would use simulated annealing (SA) with a cost function composed of a weighted sum of area, wirelength, and interconnect cost. These approaches can reduce the interconnect cost efficiently but the area penalty of the interconnect optimized floorplan is usually quite large. In this article, we propose an approach called deadspace utilization (DSU) to reclaim the unused area of an interconnect optimized floorplan by linear programming. Since modules are not necessarily rectangular in shape in floorplanning, some deadspace can be redistributed to the modules to increase the area occupied by each module. If the area of each module can be expanded by the same ratio, the whole floorplan can be compacted by that ratio to give a smaller floorplan. However, we will limit the compaction ratio to prevent overcongestion. Experiments show that we can apply this deadspace utilization technique to reduce the area and total wirelength of an interconnect optimized floorplan further while the routability can be maintained at the same time. Chiu-Wing Sham, Evangeline F. Y. Young |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2006 | Optimal cell flipping in placement and floorplanningabstractIn a placed circuit, there are a lot of movable cells that can be flipped to further reduce the total wirelength, without affecting the original placement solution. We aim at solving this flipping problem optimally. However, solving such a problem optimally is non-trivial given the gigantic sizes of modern circuits. We are able to identify a large portion of cells (about 75%) of which the orientation (flipped or not flipped) can be determined independent of the orientations of all the other cells. We have derived three non-trivial conditions to identify those so called independent cells, strictly solvable cells and conditionally solvable cells. In this way, we can greatly reduce the number of cells whose orientations are dependent on each other. Finally, the cell flipping problem of the remaining dependent cells can be formulated as a Mixed Integer Linear Programming (MILP) problem and solved optimally. However, this may still be too slow for extremely large circuits and we have applied two other methods, Linear Programming (LP) and Linear Programming followed by Mixed Integer Linear Programming (LP+MILP) to solve the problem. Experimental results show that by identifying those independent and solvable cells first and applying the LP+MILP technique, we can solve this flipping problem effectively and obtain results just 0.01% more than the optimal. In addition, we can improve the wirelength and number of overflow tiles by 5% and 9% respectively on the floorplanning benchmarks. Chiu-Wing Sham, Evangeline F. Y. Young, Chris C. N. Chu |
DAC | 1 |
| 2005 | Congestion prediction in floorplanningabstractRoutability optimization has become the major concern in floorplanning. In traditional floorplanners, area minimization is an important issue. Due to the recent advances in VLSI technology, interconnect has become a dominant factor to the overall performance of a circuit. Routability prediction is thus very important in the floorplanning stage. In this paper, we propose a new congestion model to predict the congestion after detailed routing which is not confined to the assumption of shortest Manhattan distance routes. We have compared our new models and some existing models with the actual congestion measures obtained by global routing some placement results (using the Capo placer [3]) with a publicly available maze router [2]. Results show that our models can make significant improvement in estimation accuracy over the other models. Chiu-Wing Sham, Evangeline F. Y. Young |
ASP-DAC | 1 |
| 2003 | Interconnect-driven floorplanning by searching alternative packingsabstractIn traditional floorplanners, area minimization is an important issue. Due to the recent advances in VLSI technology, the number of transistors in a design and their switching speeds are increasing rapidly. This results in the increasing importance of interconnect delay and routability of a circuit. We should consider interconnect planning and buffer planning as soon as possible. In this paper, we propose a method to reduce interconnect cost of a floorplan by searching alternative packings. We found that if a floorplan F contains some rectangular supermodules, we can rearrange the blocks in the supermodule to obtain a new floorplan with the same area as F but possibly with a smaller interconnect cost. Experimental results show that we can always reduce the interconnect cost of a floorplan without any penalty in area and runtime by using this method. Chiu-Wing Sham, Evangeline F. Y. Young, Hai Zhou 0001 |
ASP-DAC | 1 |
| 2003 | Routability-driven floorplanner with buffer block planningabstractIn traditional floorplanners, area minimization is an important issue. However, due to the recent advances in very large scale integration technology, the number of transistors in a design are increasing rapidly and so are their switching speeds. This has increased the importance of interconnect delay and routability in the overall performance of a circuit. We should consider interconnect planning, buffer planning, and routability as early as possible. In this paper, we study and implement a routability-driven floorplanner with congestion estimation and buffer planning. Our method is based on a simulated annealing approach that is divided into two phases: the area optimization and congestion optimization phases. In the area optimization phase, modules are roughly placed according to the total area and wirelength. In the congestion optimization phase, a floorplan is evaluated by its area, wirelength, congestion, and routability. We assume that buffers should be inserted at flexible intervals from each other for long enough wires and probabilistic analysis is performed to compute the congestion information taken into account the constraints in buffer locations. Our approach is able to reduce the average number of wires at the congested areas and allow more feasible insertions of buffers to satisfy the delay constraints without having much penalty in increasing the area of the floorplan. Chiu-Wing Sham, Evangeline F. Y. Young |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2002 | Congestion Estimation with Buffer Planning in Floorplan DesignabstractIn this paper, we study and implement a routability-driven floorplanner with buffer block planning. It evaluates the routability of a floorplan by computing the probability that a net will pass through each particular location of a floorplan taken into account buffer locations and routing blockages. Experimental results show that our congestion model can optimize congestion and delay (by successful buffer insertions) of a circuits better with only a slight penalty in area. Wai-Chiu Wong, Chiu-Wing Sham, Evangeline F. Y. Young |
DATE | 2 |
| 2002 | Routability driven floorplanner with buffer block planningabstractIn traditional floorplanners, area minimization is an important issue. However, due to the recent advances in VLSI technology, the number of transistors in a design are increasing rapidly and so are their switching speeds. This has increased the importance of interconnect delay and routability in the overall performance of a circuit. We should consider interconnect planning, buffer planning and routability as early as possible. In this paper, we study and implement a routability-driven floorplanner with congestion estimation and buffer planning. Our method is based on a simulated annealing approach that is divided into two phases: the area optimization phase and the congestion optimization phase. In the area optimization phase, modules are roughly placed according to the total area and wirelength. In the congestion optimization phase, a floorplan will be evaluated by its area, wirelength, congestion and routability. We assume that every buffer should be inserted at a flexible interval from each other for long enough wires and probabilistic analysis is performed to compute the congestion information taken into accounts the constraints in buffer locations. Our approach is able to reduce the average number of wires at the congested areas and allow more feasible insertions of buffers to satisfy the delay constraints without having much penalty in increasing the area of the floorplan. Chiu-Wing Sham, Evangeline F. Y. Young |
ISPD | 1 |
| 2001 | A bitstream reconfigurable FPGA implementation of the WSAT algorithmabstractA field programmable gate array (FPGA) implementation of a coprocessor which uses the WSAT algorithm to solve Boolean satisfiability problems is presented. The input is a SAT problem description file from which a software program directly generates a problem-specific circuit design which can be downloaded to a Xilinx Virtex FPGA device and executed to find a solution. On an XCV300, problems of 50 variables and 170 clauses can be solved. Compared with previous approaches, it avoids the need for resynthesis, placement, and routing for different constraints. Our coprocessor is eminently suitable for embedded applications where energy, weight and real-time response are of concern. Philip H. W. Leong, Chiu-Wing Sham, H. Y. Wong, Wing Seung Yuen, Monk-Ping Leong |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |