EDBT 2026 Demo / reviewers in the wild / expert
Wan-Chi Siu
dblp:61/5203
· DBLP profile ↗
254ranked-venue papers
8as first author
14since 2021 · last 2026
0000-0001-8280-0367ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 182 · 8 first-author · 11 since 2021Artificial intelligence and machine learning · 34 · 2 since 2021Systems, architecture and hardware · 29 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 since 2021Computer networks · 3Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Quality-Assisted Domain Transfer for Fast Face Super-ResolutionabstractFace super-resolution (FSR) is a challenging task, especially when the low-resolution (LR) image does not have sufficient information to build a high-resolution (HR) image. In this paper, we propose a quality-assisted domain transfer framework that uses a two-stage pipeline for efficient face SR. Instead of treating all LR images uniformly, we use a quality evaluator to classify images in different LR domains with distinct quality levels and determine the starting points for domain transfer adaptively. We further design a new residual domain learning strategy that integrates the LR domain with learned residual information. This approach allows us to avoid the correction formulation in diffusion models and the momentum correction in the originally proposed domain transfer approach and enables the model to learn the transformation from the LR domain to the HR domain gradually while preserving content information. The experimental results demonstrate competitive performance with the best inference speed among all approaches. Yi-Hao Cheng, Wan-Chi Siu, S. C. Chan 0001 |
IEEE Signal Process. Lett. | 2 |
| 2026 | Domain Fusion in Latent Space for Face Video Super-Resolution
Shaohua Jia, Pengyu Liu 0001, Kebin Jia, Anthony H. Chan, Wan-Chi Siu |
IEEE Signal Process. Lett. | 5 |
| 2025 | PUMPS: Skeleton-Agnostic Point-Based Universal Motion Pre-Training for Synthesis in Human Motion Tasks
Clinton Mo, Kun Hu 0008, Chengjiang Long, Dong Yuan 0001, Wan-Chi Siu, Zhiyong Wang 0001 |
ICCV | 5 |
| 2025 | Domain Transfer Generative Model for New Face GenerationabstractWe propose a novel fast approach for face image generation based on a domain transfer method. This approach incorporates variance injection within each gradual transfer process, which is facilitated by a conditioned U-Net and a novel pattern similarity loss. The variance injection enables the exploration of unseen regions within each domain, resulting in the generation of diverse, photorealistic images that extend beyond the training data distribution. Our approach offers a promising alternative to GAN-based and diffusion models, providing enhanced controllability on the generation process. This work reveals the potential of domain transfer as a new type of generative model, though further research is needed to fully explore and refine this nascent framework. Results of extensive experimental work have shown that this new fast approach can successfully generate quality photorealistic faces with diverse face features. It verifies that this is a new and attractive approach for new data synthesis with high-quality. Chun-Chuen Hui, Wan-Chi Siu, H. Anthony Chan |
ICIP | 2 |
| 2025 | DTLS-Inpaint: Yet Another Efficient Image Inpainting with Domain TransferabstractImage inpainting, a process of reconstructing missing or corrupted regions of an image, has evolved significantly with models such as LaMa, AOT GAN, and RePaint with Diffusion models (DMs). These are successful models. However, there is room for improvement, especially on high computation requirement, stability of training and mode collapse. In this paper, we introduce the Domain Transfer in Latent Space (DTLS-Inpaint) model for Inpainting tasks. DTLS-Inpaint adopts a progressive domain transfer strategy, where masked areas are gradually restored by transitioning through increasingly realistic inpainting domains, bearing similarity of DMs but operating without noise-adding and denoising processes. Instead, it utilizes latent space transformations at each timestep which will determine an appropriate amount of inpainting in the new domain, and produce progressively improved intermediate results. Experimental results show that this new approach is extremely fast and able to produce inpainted images comparable or even better than the state-of-the art approaches. Hon Man Hammond Lee, Wan-Chi Siu |
ICIP | 2 |
| 2025 | Unveiling image source: Instance-level camera device linking via context-aware deep Siamese network
Mingjie Zheng 0002, Ngai-Fong Law, Wan-Chi Siu |
Expert Syst. Appl. | 3 |
| 2024 | Video Assisted Face Recognition in Smart ClassroomabstractStudent recognition in smart classrooms is challenging due to low resolution, occlusions, and various face orientations. We propose a video-based approach for robust face recognition. It firstly collects a continuous sequence of face images through inter-frame analysis. Then, it aggregates face features, via statistical elimination of outliers. We introduce a new consistency loss to address varying face resolutions, enabling the learning of scale-robust features. Experimental results demonstrate that our proposed approach improves recognition performance under various conditions compared to previous methods. In our experiments, it achieved the highest recognition accuracy (83.57%) for 8×8 resolution cases. Li-Wen Wang, Wan-Chi Siu, Yi-Hao Cheng, H. Anthony Chan |
ISCAS | 2 |
| 2024 | Edge fusion back projection GAN for large scale face super resolutionabstractFace super-resolution is an important low-level vision task that has wide applications. Existing deep-learning-based face super-resolution (SR) methods often optimize the image super-resolution network by directly minimizing the pixel or feature level distance between the synthetic low-resolution face and the ground truth face image. These methods usually generate blur or over smooth results and lack of high-frequency face details. These are especially true for making high super-resolution of faces. Say for example, a super-resolution of 16x, only 0.4 % of the reference data points are available. To address the problem, we propose a novel network with edge fusion, back projection, and GAN prior (EFBPGAN) which can significantly improve the visual quality and generate realistic faces. To further make use of the spatial information and keep the structural consistency, we have developed new edge fusion and spatial fusion modules. We also propose a back projection extensive based coarse to fine SR pipeline to suppress the distortion and artifacts caused by GAN. Much experimental work has been done, results of which show that our proposed EFBPGAN can outperform the state-of-the-art approaches not only on numerical metrics but also on subjective visual evaluations. Wan-Chi Siu |
J. Vis. Commun. Image Represent. | 2 |
| 2024 | AnlightenDiff: Anchoring Diffusion Probabilistic Model on Low Light Image EnhancementabstractLow-light image enhancement aims to improve the visual quality of images captured under poor illumination. However, enhancing low-light images often introduces image artifacts, color bias, and low SNR. In this work, we propose AnlightenDiff, an anchoring diffusion model for low light image enhancement. Diffusion models can enhance the low light image to well-exposed image by iterative refinement, but require anchoring to ensure that enhanced results remain faithful to the input. We propose a Dynamical Regulated Diffusion Anchoring mechanism and Sampler to anchor the enhancement process. We also propose a Diffusion Feature Perceptual Loss tailored for diffusion based model to utilize different loss functions in image domain. AnlightenDiff demonstrates the effect of diffusion models for low-light enhancement and achieving high perceptual quality results. Our techniques show a promising future direction for applying diffusion models to image enhancement. Cheuk-Yiu Chan, Wan-Chi Siu, Yuk-Hee Chan, H. Anthony Chan |
IEEE Trans. Image Process. | 2 |
| 2023 | Intelligent Painter: Picture Composition with Resampling Diffusion ModelabstractHave you ever thought that you can be an intelligent painter? This means that you can paint a picture with a few expected objects in mind, or with a desirable scene. This is different from normal inpainting approaches for which the location of specific objects cannot be determined. In this paper, we present an intelligent painter that generate a person's imaginary scene in one go, given explicit hints. We propose a resampling strategy for Denoising Diffusion Probabilistic Model (DDPM) to intelligently compose unconditional harmonized pictures according to the input subjects at specific locations. By exploiting the diffusion property, we resample efficiently to produce realistic pictures. Experimental results show that our resampling method favors the semantic meaning of the generated output efficiently and generates less blurry output. Quantitative analysis of image quality assessment shows that our method produces higher perceptual quality images compared with the state-of-the-art methods. Wing-Fung Ku, Wan-Chi Siu, H. Anthony Chan |
ICIP | 2 |
| 2022 | See360: Novel Panoramic View InterpolationabstractWe present See360, which is a versatile and efficient framework for 360° panoramic view interpolation using latent space viewpoint estimation. Most of the existing view rendering approaches only focus on indoor or synthetic 3D environments and render new views of small objects. In contrast, we suggest to tackle camera-centered view synthesis as a 2D affine transformation without using point clouds or depth maps, which enables an effective 360° panoramic scene exploration. Given a pair of reference images, the See360 model learns to render novel views by a proposed novel Multi-Scale Affine Transformer (MSAT), enabling the coarse-to-fine feature rendering. We also propose a Conditional Latent space AutoEncoder (C-LAE) to achieve view interpolation at any arbitrary angle. To show the versatility of our method, we introduce four training datasets, namely UrbanCity360, Archinterior360, HungHom360 and Lab360, which are collected from indoor and outdoor environments for both real and synthetic rendering. Experimental results show that the proposed method is generic enough to achieve real-time rendering of arbitrary views for all four datasets. In addition, our See360 model can be applied to view synthesis in the wild: with only a short extra training time (approximately 10 mins), and is able to render unknown real-world scenes. The superior performance of See360 opens up a promising direction for camera-centered view rendering and 360° panoramic view interpolation. Marie-Paule Cani, Wan-Chi Siu |
IEEE Trans. Image Process. | 3 |
| 2021 | Photo-Realistic Image Super-Resolution via Variational AutoencodersabstractThere is a great leap in objective accuracy on image super-resolution, which recently brings a new challenge on image super-resolution with larger up-scaling (e.g. 4×) using pixel based distortion for measurement. This causes over-smooth effect which cannot grasp well the perceptual similarity. The advent of generative adversarial networks makes it possible super-resolve a low-resolution image to generate photo-realistic images sharing distribution with the high-resolution images. However, generative networks suffer from problems of mode-collapse and unrealistic sample generation. We propose to perform Image Super-Resolution via Variational AutoEncoders (SR-VAE) learning according to the conditional distribution of the high-resolution images induced by the low-resolution images. Given that the Conditional Variational Autoencoders tend to generate blur images, we add the conditional sampling mechanism to narrow down the latent subspace for reconstruction. To evaluate the model generalization, we use KL loss to measure the divergence between latent vectors and standard Gaussian distribution. Eventually, in order to balance the trade-off between super-resolution distortion and perception, not only that we use pixel based loss, we also use the modified deep feature loss between SR and HR images to estimate the reconstruction. In experiments, we evaluated a large number of datasets to make comparison with other state-of-the-art super-resolution approaches. Results on both objective and subjective measurements show that our proposed SR-VAE can achieve good photo-realistic perceptual quality closer to the natural image manifold while maintain low distortion. Wan-Chi Siu, Yui-Lam Chan |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Features Guided Face Super-Resolution via Hybrid Model of Deep Learning and Random ForestsabstractFace hallucination or super-resolution is a practical application of general image super-resolution which has been recently studied by many researchers. The challenge of good face hallucination comes from a variety of poses, illuminations, facial expressions, and other degradations. In many proposed methods, researchers resolve it by using a generative neural network to reduce the perceptual loss so we can generate a photo-realistic image. The problem is that researchers usually overlook the fidelity of the super-resolved image which could affect further facial image processing. Meanwhile, many CNN based approaches cascade multiple networks to extract facial prior information to improve super-resolution quality. Because of the end-to-end design, the details are missing for investigation. In this paper, we combine new techniques in convolutional neural network and random forests to a Hierarchical CNN based Random Forests (HCRF) approach for face super-resolution in a coarse-to-fine manner. In the proposed approach, we focus on a general approach that can handle facial images with various conditions without pre-processing. To the best of our knowledge, this is the first paper that combines the advantages of deep learning with random forests for face super-resolution. To achieve superior performance, we propose two novel CNN models for coarse facial image super-resolution and segmentation and then apply new random forests to target on local facial features refinement making use of the segmentation results. Extensive benchmark experiments on subjective and objective evaluation show that HCRF can achieve comparable speed and competitive performance compared with state-of-the-art super-resolution approaches for very low-resolution images. Wan-Chi Siu, Yui-Lam Chan |
IEEE Trans. Image Process. | 2 |
| 2021 | Fast Monocular Visual Place Recognition for Non-Uniform Vehicle Speed and Varying Lighting EnvironmentabstractThis paper presents a novel Fast Monocular Visual Place Recognition (FMPR) with a shallow path-oriented offline learning stage and an online place recognition and tracking stage. FMPR uses a tube of frames with a humanlike key frame recognition to solve place recognition for situations with varying speeds and changing lighting conditions, which are two most commonly encountered situations in real life. We propose an offline learning to analyze the correlation of all video frames in a reference path and to extract effective feature patches of key frames with an offline feature-shifts approach to achieve real-time place recognition. Our recognition results are on the basis of both the instant feature matching of frames and the historical recognition results which impose temporal logic constraints on the movement of a vehicle. Experimental results demonstrate that our proposed method can achieve comparable or even better performance compared with the state-of-the-art methods on different challenging datasets, especially for the case which requires a trade-off between the performance and the processing time. We believe that our FMPR offers a useful alternative to computationally expensive deep learning-based methods especially for applications with battery-powered or resource-limited devices. Chu-Tak Li, Wan-Chi Siu |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2020 | Video Lightening with Dedicated CNN ArchitectureabstractDarkness brings us uncertainty, worry and low confidence. This is a problem not only applicable to us walking in a dark evening but also for drivers driving a car on the road with very dim or even without lighting condition. To address this problem, we propose a new CNN structure named as Video Lightening Network (VLN) that regards the low-light enhancement as a residual learning task, which is useful as reference to indirectly lightening the environment, or for vision-based application systems, such as driving assistant systems. The VLN consists of several Lightening Back-Projection (LBP) and Temporal Aggregation (TA) blocks. Each LBP block enhances the low-light frame by domain transfer learning that iteratively maps the frame between the low- and normal-light domains. A TA block handles the motion among neighboring frames by investigating the spatial and temporal relationships. Several TAs work in a multi-scale way, which compensates the motions at different levels. The proposed architecture has a consistent enhancement for different levels of illuminations, which significantly increases the visual quality even in the extremely dark environment. Extensive experimental results show that the proposed approach outperforms other methods under both objective and subjective metrics. Li-Wen Wang, Wan-Chi Siu, Chu-Tak Li, Daniel Pak-Kong Lun |
ICPR | 2 |
| 2020 | Deep Lightening Network for Low-Light Image EnhancementabstractWe propose a Deep Lightening Network (DLN) for low-light image enhancement. Inspire by the domain transfer study, we propose a novel cycle learning structure to learn the mapping relationship between low- and normal-light images. Each DLN consists of several Lightening Back-Projection (LBP) blocks that learn the residual between low- and normal-light images. To efficiently estimate the local and global information, we fuse the features from different LBP results. Experimental results on different datasets show that our proposed DLN approach outperforms other approaches in all objective and subjective measures. Li-Wen Wang, Wan-Chi Siu, Daniel Pak-Kong Lun |
ISCAS | 3 |
| 2020 | Machine Learning-Based Fast Intra Mode Decision for HEVC Screen Content Coding via Decision TreesabstractThe screen content coding (SCC) extension of high efficiency video coding (HEVC) improves coding gain for screen content videos by introducing two new coding modes, namely, intra block copy (IBC) and palette (PLT) modes. However, the coding gain is achieved at the increased cost of computational complexity. In this paper, we propose a decision tree-based framework for fast intra mode decision by investigating various features in the training sets. To avoid the exhaustive mode searching process, a sequential arrangement of decision trees is proposed to check each mode separately by inserting a classifier before checking a mode. As compared with the previous approaches where both IBC and PLT modes are checked for screen content blocks (SCBs), the proposed coding framework is more flexible which facilitates either the IBC or PLT mode to be checked for SCBs such that computational complexity is further reduced. To enhance the accuracy of decision trees, dynamic features are introduced, which reveal the unique intermediate coding information of a coding unit (CU). Then, if all the modes are decided to be skipped for a CU at the last depth level, at least one possible mode is assigned by a CU-type decision tree. Furthermore, a decision tree constraint technique is developed to reduce the rate-distortion performance loss. Compared with the HEVC-SCC reference software SCM-8.3, the proposed algorithm reduces computational complexity by 47.62% on average with a negligible Bjøntegaard delta bitrate (BDBR) increase of 1.42% under all-intra (AI) configurations, which outperforms all the state-of-the-art algorithms in the literature. Wei Kuang, Yui-Lam Chan, Sik-Ho Tsang, Wan-Chi Siu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | DeepSCC: Deep Learning-Based Fast Prediction Network for Screen Content CodingabstractScreen content coding (SCC) is an extension of high efficiency video coding (HEVC), and it is developed to improve the coding efficiency of screen content videos by adopting two new coding modes: Intra Block Copy (IBC) and Palette (PLT). However, the flexible quadtree-based coding tree unit (CTU) partitioning structure and various mode candidates make the fast algorithms of the SCC extremely challenging. To efficiently reduce the computational complexity of SCC, we propose a deep learning-based fast prediction network DeepSCC that contains two parts: DeepSCC-I and DeepSCC-II. Before feeding to DeepSCC, incoming coding units (CUs) are divided into two categories: dynamic CTUs and stationary CTUs. For dynamic CTUs having different content as their collocated CTUs, DeepSCC-I takes raw sample values as the input to make fast predictions. For stationary CTUs having the same content as their collocated CTUs, DeepSCC-II additionally utilizes the optimal mode maps of the stationary CTU to further reduce the computational complexity. Compared with the HEVC-SCC reference software SCM-8.3, the proposed DeepSCC reduces the encoding time by 48.81% on average with a negligible Bjøntegaard delta bitrate increase of 1.18% under all-intra configuration. Wei Kuang, Yui-Lam Chan, Sik-Ho Tsang, Wan-Chi Siu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Online-Learning-Based Bayesian Decision Rule for Fast Intra Mode and CU Partitioning Algorithm in HEVC Screen Content CodingabstractScreen content coding (SCC) is an extension of high efficiency video coding by adopting new coding modes to improve the coding efficiency of SCC at the expense of increased complexity. This paper proposes an online-learning approach for fast mode decision and coding unit (CU) size decision in SCC. To make a fast mode decision, the corner point is first extracted as a unique feature in screen content, which is an essential pre-processing step to guide Bayesian decision modeling. Second, the distinct color number in a CU is derived as another unique feature in screen content to build the precise model using online-learning for skipping unnecessary modes. Third, the correlation of the modes among spatial neighboring CUs is analyzed to further eliminate unnecessary mode candidates. Finally, the Bayesian decision rule using online-learning is applied again to make a fast CU size decision. To ensure the accuracy of the Bayesian decision models, new scene change detection is designed to update the models. Results show that the proposed algorithm achieves 36.69% encoding time reduction with 1.08% Bjøntegaard delta bitrate (BDBR) increment under all intra configuration. By integrating into the existing fast SCC approach, the proposed algorithm reduces 48.83% encoding time with a 1.78% increase in BDBR. Wei Kuang, Yui-Lam Chan, Sik-Ho Tsang, Wan-Chi Siu |
IEEE Trans. Image Process. | 4 |
| 2020 | Lightening Network for Low-Light Image EnhancementabstractLow-light image enhancement is a challenging task that has attracted considerable attention. Pictures taken in low-light conditions often have bad visual quality. To address the problem, we regard the low-light enhancement as a residual learning problem that is to estimate the residual between low- and normal-light images. In this paper, we propose a novel Deep Lightening Network (DLN) that benefits from the recent development of Convolutional Neural Networks (CNNs). The proposed DLN consists of several Lightening Back-Projection (LBP) blocks. The LBPs perform lightening and darkening processes iteratively to learn the residual for normal-light estimations. To effectively utilize the local and global features, we also propose a Feature Aggregation (FA) block that adaptively fuses the results of different LBPs. We evaluate the proposed method on different datasets. Numerical results show that our proposed DLN approach outperforms other methods under both objective and subjective metrics. Li-Wen Wang, Wan-Chi Siu, Daniel Pak-Kong Lun |
IEEE Trans. Image Process. | 3 |
| 2019 | Semi-Supervised Deep Vision-Based Localization Using Temporal Correlation Between Consecutive FramesabstractVision-based localization is a temporal informative task in which we can obtain information about the ego-motion of a vehicle from the historical information via examining consecutive frames. Sufficient temporal information helps to reduce the search space of the next location. Hence, both efficiency and accuracy of the localization system can be enhanced. This paper presents a semi-supervised deep vision-based localization algorithm, using a novel tubing strategy to find the starting location of a vehicle. We group different number of consecutive frames as sets of tubes based on their temporal correlation to achieve pair searching with variable tube sizes. We also enhance an off-the-shelf network model with our modified training data generation method to improve the discrimination power of the features given by the model. Experimental results show that our proposed temporal correlation based initialization module can confidently localize the starting location of a vehicle (for a certain journey), and achieve 40% precision improvement over that of the conventional CNN approaches. Chu-Tak Li, Wan-Chi Siu, Daniel Pak-Kong Lun |
ICIP | 2 |
| 2019 | Clustering-Based Compression for Population DNA SequencesabstractDue to the advancement of DNA sequencing techniques, the number of sequenced individual genomes has experienced an exponential growth. Thus, effective compression of this kind of sequences is highly desired. In this work, we present a novel compression algorithm called Reference-based Compression algorithm using the concept of Clustering (RCC). The rationale behind RCC is based on the observation about the existence of substructures within the population sequences. To utilize these substructures, k-means clustering is employed to partition sequences into clusters for better compression. A reference sequence is then constructed for each cluster so that sequences in that cluster can be compressed by referring to this reference sequence. The reference sequence of each cluster is also compressed with reference to a sequence which is derived from all the reference sequences. Experiments show that RCC can further reduce the compressed size by up to 91.0 percent when compared with state-of-the-art compression approaches. There is a compromise between compressed size and processing time. The current implementation in Matlab has time complexity in a factor of thousands higher than the existing algorithms implemented in C/C++. Further investigation is required to improve processing time in future. Kin-On Cheng, Ngai-Fong Law, Wan-Chi Siu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2019 | Reduced-Complexity Intra Block Copy (IntraBC) Mode With Early CU Splitting and Pruning for HEVC Screen Content CodingabstractA screen content coding (SCC) extension to high efficiency video coding has been developed to incorporate many new coding tools in order to achieve better coding efficiency for videos mixed with camera-captured content and graphics/text/animation. For instance, the Intra Block Copy (IntraBC) mode helps to encode repeating patterns within the same frame while the Palette mode aims at encoding screen content with a few major colors. However, the IntraBC mode brings along high computational complexity due to the exhaustive block matching within the same frame though there are already some constraints and fast approaches applied to the IntraBC mode to reduce its complexity. Thus, we propose a fast intracoding scheme to reduce the complexity of using the IntraBC mode in SCC. Screen content always contains no sensor noise resulting in the characteristics with pixel exactness along both horizontal and vertical directions. These characteristics pave the way for mode skipping and early coding unit (CU) splitting. Besides, early CU pruning and early termination are proposed based on the rate distortion cost to further reduce encoder complexity. Moreover, we also propose reducing the complexity of the IntraBC mode by checking the hash value of each block candidate and the current block during block matching. With our proposed scheme, the encoding time is reduced compared with the SCC while the coding efficiency can still be maintained with a minor increase in the bjontegaard delta bitrate. Sik-Ho Tsang, Yui-Lam Chan, Wei Kuang, Wan-Chi Siu |
IEEE Trans. Multim. | 4 |
| 2018 | Cascaded Random Forests for Fast Image Super-ResolutionabstractDue to the development of deep learning, image super- resolution has achieved huge improvement on both subjective and objective qualities. However, the computation is still a problem for real-time applications. In this paper, we propose a Cascaded Random Forest for Image Super-Resolution (CRFSR) which screens sufficient simple features to train a much robust and efficient model for image super-resolution. To further boost up the super-resolution performance, an extra Gaussian Mixture Model (GMM) based layer is added as the final refinement. Extensive experimental results show that the cascaded decision trees continue performing better when more features are selected for refinement. The analysis on both computation time and reconstruction fidelity indicates the superior performance of our proposed CRFSR and CRFSR+ with extra GMM-based layer on natural images. Wan-Chi Siu |
ICIP | 2 |
| 2018 | Fast HEVC to SCC Transcoding Based on Decision TreesabstractScreen Content Coding (SCC) is an extension of the High-Efficiency Video Coding (HEVC) for encoding screen content videos. However, there are many legacy screen content videos already encoded by HEVC. To efficiently migrate screen content videos from the existing HEVC to the emerging SCC, a machine learning based fast transcoding algorithm is proposed by using decision trees in this paper. To speed up the transcoding process, the intermediate data from both the HEVC decoder side and the SCC encoder side are jointly analyzed. Then the optimal coding unit (CU) sizes are mapped from HEVC to SCC while the mode candidates are adaptively checked according to the decision tree outcomes in the re-encoding process. Experimental results show that an average of 48.20% re-encoding time reduction is achieved with only 1.47% Bjontegaard delta bitrate loss using All Intra (AI) configuration. Wei Kuang, Yui-Lam Chan, Sik-Ho Tsang, Wan-Chi Siu |
ICME | 4 |
| 2018 | Probability-Based Depth Intra-Mode Skipping Strategy and Novel VSO Metric for DMM Decision in 3D-HEVCabstractMultiview video plus depth format has been adopted as the emerging 3D video representation recently. It includes a limited number of textures and depth maps to synthesize additional virtual views. Since the quality of depth maps influences the view synthesis process, their sharp edges should be well preserved to avoid mixing foreground with background. To address this issue, 3D-High Efficiency Video Coding (HEVC) introduces new coding tools, a partition-based intra mode [depth modeling mode (DMM)], a residual description technique [segmentwise depth coding (SDC)], and a more complex rate-distortion (RD) evaluation with view synthesis optimization (VSO), to provide more accurate predictions and achieve higher compression rate. However, these new techniques introduce a lot of possible candidates, and each of them requires complicated RD calculation in the process of intra-mode decision. They lead to unacceptable computational burden in a 3D-HEVC encoder. Therefore, in this paper, we raise two efficient techniques for depth intra-mode decision. First, by investigating the statistical characteristics of variance distributions in the two partitions of DMM, a simple but efficient criterion based on the squared Euclidean distance of variances (SEDV) is suggested to evaluate RD costs of the DMM candidates instead of the time-consuming VSO process. Second, a probability-based early depth intra-mode decision is proposed to select only the most promising mode and make the early determination of using SDC based on the low-complexity RD cost in rough mode decision. Experimental results show that the proposed algorithm with these two new techniques provides 33%-48% time reduction with little drop in the coding performance compared with the state-of-the-art algorithms. Hongbin Zhang 0005, Chang-Hong Fu 0002, Yui-Lam Chan, Sik-Ho Tsang, Wan-Chi Siu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2017 | Depth modelling mode decision for depth intra coding via good featureabstractThe depth modelling modes (DMM) and 35 conventional intra modes (CHIMs) introduced in 3D-HEVC results in unacceptable huge complexity of depth intra coding. However, some redundancy between DMM and CHIMs could be avoided to accelerate the process. In this paper, a good feature-corner point (CP) is proposed to evaluate the orientation of edge in a given prediction unit (PU), by which a binary classifier is created. We further investigate the probability distribution of DMM, which is selected as the optimal intra mode in each category. According to the statistical analysis, the skipping of DMM decision is proposed to eliminate the cases which have been predicted well by CHIMs. The experimental results show that, compared with the test model HTM-13.0 of 3D-HEVC, the proposed algorithm can yield about 17% time reduction for depth intra coding with almost no degradation in coding performance. Chang-Hong Fu 0002, Ya-Wen Zhao, Hongbin Zhang 0005, Yui-Lam Chan, Wan-Chi Siu |
ICIP | 5 |
| 2017 | Fast mode decision algorithm for HEVC screen content intra codingabstractScreen Content coding (SCC) is one of an extension to High Efficiency Video Coding (HEVC) developed by the Joint Collaborative Team on Video Coding (JCT-VC). It adopts two new coding tools, intra block copy (IBC) and palette (PLT) modes, to improve the compression performance for intra coding. Nevertheless, mode selection causes a substantial increase in encoding complexity. In this paper, a fast mode decision algorithm, which makes use of early mode skip decision based on the Bayesian decision rule using online learning, is proposed. The proposed algorithm is implemented in the SCC reference software SCM-7.0. Experimental results show that the proposed algorithm can achieve 23.2% complexity reduction on average with only 0.58% Bjontegaard delta bitrate loss in All Intra (AI) configurations. Wei Kuang, Sik-Ho Tsang, Yui-Lam Chan, Wan-Chi Siu |
ICIP | 4 |
| 2017 | Decoder side merge mode and AMVP in HEVC screen content codingabstractIntra Block Copy (IBC) mode in a screen content coding (SCC) extension in High Efficiency Video Coding (HEVC) provides high coding gain by performing motion estimation (ME) and motion compensation (MC) to find the repetitive patterns within the same frame. Merge mode and Advanced Motion Vector Prediction (AMVP), which are originally used for inter mode, are also applied to the IBC mode. However, there are redundant coding bits when they are applied to IBC. Therefore, we propose decoder-side merge mode and AMVP for IBC in SCC so as to remove the redundancy. Experimental shows that the proposed method can achieve up to 0.25% Bjontegaard delta bitrate (BD-rate) reduction compared to the conventional SCC with negligible impact to encoding and decoding complexity. Sik-Ho Tsang, Wei Kuang, Yui-Lam Chan, Wan-Chi Siu |
ICIP | 4 |
| 2017 | Adaptive search range by depth variant decaying weights for HEVC inter texture codingabstractEmerging high-efficiency video coding (HEVC) outperforms H.264 by a gain of 50% bitrate reduction while maintaining almost the same perceptual quality. However, it induces higher coding complexity due to its adoption of recursive block partitioning mechanism in motion estimation (ME) with a fixed search range. For an objective of reducing the computational burden in HEVC, this paper proposes an adaptive search range algorithm by using depth map information. With the aid of depth intensity variations among neighboring blocks, associated weights to the neighboring blocks are derived. The proposed weighted sum of the motions from the neighboring blocks is formulated to provide a suitable search range for each block. The simulation results demonstrated that proposed adaptive search range is compatible to not only full-search (FS) but also fast Test Zone Search (TZS) in HEVC. The proposed algorithm could reduce significant coding time on average with negligible rate-distortion degradation. Tsz-Kwan Lee, Yui-Lam Chan, Wan-Chi Siu |
ICME | 3 |
| 2017 | Depth-projected determination for adaptive search range in motion estimation for HEVCabstractHigh Efficient Video Coding (HEVC) improves coding efficiency but suffers from high computational complexity due to its quad-tree partitioning structure in motion estimation (ME). In recent development of 3D video technology, depth map from the 3D video provides an intimation of the objects' distance from the projected screen in a 3D scene, which inspires the authors to explore the adaptive search range determination for complexity reduction in HEVC. The proposed algorithm exploits the high temporal correlation between the depth map and the motion in texture. By utilizing this correlation and the potential impact of 3D-to-2D projection, a depth/motion relationship is built for a tailor-made search range with a depth-projected scale factor to skip unnecessary search points in ME. Besides, the proposed ASR algorithm can work well with other fast ME algorithms with up to 53% of average coding time reduction whereas the coding efficiency can be maintained. Tsz-Kwan Lee, Yui-Lam Chan, Wan-Chi Siu |
ISCAS | 3 |
| 2017 | Fast image super-resolution via Randomized Multi-split ForestsabstractThis paper proposes a novel learning-based image Super-Resolution via a Randomized Multi-split Forests model (SRRMF). The proposed method uses the LR-HR training patch pairs to model the nonlinear patch manifold into a pairs of linear subspaces. The key idea of this approach is to use several decision trees split randomly the training data into different classes. A linear regression model is learnt to map the relationship between LR and HR patches at the end of the leaf nodes. In order to make full use of the generalization ability of the random forests, we randomize the grow of the decision tree to cover more possibilities. Furthermore, we modify the splitting function by using Multi-Split Binary Test (MSBT) function so that we can use more feature information to derive more accurate classification result to match patch subspace. Extended experimental results show that image super-resolution using our proposed method can achieve the state-of-the-art super-resolution performance with reduced computation time. Wan-Chi Siu, Yui-Lam Chan |
ISCAS | 2 |
| 2017 | Segment-based view synthesis optimization scheme in 3D-HEVC
Huan Dou, Yui-Lam Chan, Kebin Jia, Wan-Chi Siu |
J. Vis. Commun. Image Represent. | 4 |
| 2017 | Learning Hierarchical Decision Trees for Single-Image Super-ResolutionabstractSparse representation has been extensively studied for image super-resolution (SR), and it achieved great improvement. Deep-learning-based SR methods have also emerged in the literature to pursue better SR results. In this paper, we propose to use a set of decision tree strategies for fast and high-quality image SR. Our proposed SR using decision tree (SRDT) method takes the divide-and-conquer strategy, which performs a few simple binary tests to classify an input low-resolution (LR) patch into one of the leaf nodes and directly multiplies this LR patch with the regression model at that leaf node for regression. Both the classification process and the regression process take an extremely small amount of computation. To further boost the SR results, we introduce a SR using hierarchical decision trees (SRHDT) method, which cascades multiple layers of decision trees for SR and progressively refines the estimated high-resolution image. Inspired by the random forests approach, which combines regression models from an ensemble of decision trees, we propose to fuse regression models from relevant leaf nodes within the same decision tree to form a more robust approach. The SRHDT method with fused regression model (SRHDT_f) improves further the SRHDT method by 0.1-dB in PNSR. Our experimental results show that our initial approach, the SRDT method, achieves SR results comparable to those of the sparse-representation-based method and the deep-learning-based method, but our method is much faster. Furthermore, our enhanced version, the SRHDT_f method, achieves more than 0.3-dB higher PSNR than that of the A+ method, which is the state-of-the-art method in SR. Junjie Huang 0002, Wan-Chi Siu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Adaptive Search Range for HEVC Motion Estimation Based on Depth InformationabstractHigh Efficiency Video Coding achieves twofold coding efficiency improvement compared with its predecessor H.264/MPEG-4 Advanced Video Coding. However, it suffers from high computational complexity due to its quad-tree structure in motion estimation (ME). This paper exposes the use of depth maps in the multiview video plus depth format for relieving the computational burden. The depth map provides an intimation of the objects' distance from the projected screen in a 3D scene, which is explored in adaptive search range determination in this paper. The proposed algorithm exploits the high temporal correlation between the depth map and the motion in texture. By utilizing this correlation, a depth/motion relationship map is built for a mapping process. For each block, this forms a tailor-made search range with a motion-aware asymmetric shape to skip unnecessary search points in ME. The obtained search range can be further adjusted by taking the influence of 3D-to-2D projection into consideration. Simulation results reveal that, compared to the full search approach, the proposed algorithm can reduce the complexity by 93% on average, whereas the coding efficiency can be maintained. Besides, the proposed search range determination can work well with other fast search ME algorithms in the literature. Tsz-Kwan Lee, Yui-Lam Chan, Wan-Chi Siu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2016 | Quadtree decision for depth intra coding in 3D-HEVC by good featureabstract3D-HEVC is a good coding solution for multi-view video plus depth data. It achieves good coding performance of synthesized views. However, depth intra coding brings unbearable complexity, which is the most urgent issue to be solved for the practical applications. Typically, depth maps have a good feature of structure or less texture compared with natural videos. Therefore, in this paper, a fast depth intra coding algorithm is proposed to speed up the quadtree decision by the good feature-corner point (CP). The proposed algorithm can adaptively extract CPs and preallocate the depth level of coding quadtree. The large size of coding units (CUs) can be skipped for blocks, which have higher predicted depth level. On the contrary, the blocks, with lower predicted depth level, do not check the smaller size of CUs. Simulation results show that the proposed algorithm can provide about 41% time reduction while maintaining the BD performance. Hongbin Zhang 0005, Yui-Lam Chan, Chang-Hong Fu 0002, Sik-Ho Tsang, Wan-Chi Siu |
ICASSP | 5 |
| 2015 | View synthesis optimization based on texture smoothness for 3D-HEVCabstractThis paper presents view synthesis optimization for 3D-HEVC based on a new texture smoothness process. In the original method, all pixels are exhaustively rendered to get distortions from synthesized views. Since not all pixels from the distorted depth map may cause distortions in the synthesized view, it brings unnecessary coding complexity. In this paper, lines of pixels are skipped based on the analysis of pixel regularity from smooth texture regions. It is due to the fact that the distorted disparity may not have much effect on the synthesized view in smooth texture regions. The proposed method can reduce the coding complexity of view synthesis optimization without significant performance loss. Huan Dou, Yui-Lam Chan, Kebin Jia, Wan-Chi Siu |
ICASSP | 4 |
| 2015 | Fast image interpolation with decision treeabstractThis paper proposes a fast image interpolation method using decision tree. This new fast image interpolation with decision tree (FIDT) method can achieve state-of-the-art image interpolation performance and requires only 10% computational time of the soft adaptive interpolation (SAI) method. During training, the proposed method recursively divides the training data at a non-leaf node into two child nodes according to the binary test which can maximize the information gain of a division. At the end, for each of the leaf node, a linear regression model is learned according to the training data at that leaf node. In the image interpolation phase, input image patches are passed into the learned decision tree. According to the stored binary test at each non-leaf node, each input image patch will be classified into its left or right child node until a leaf node is reached. The high-resolution image patch of the input image patch can then be predicted efficiently using the learned linear regression model at the leaf node. Junjie Huang 0002, Wan-Chi Siu |
ICASSP | 2 |
| 2015 | Fast and efficient intra coding techniques for smooth regions in screen content coding based on boundary prediction samplesabstractThis paper presents fast and efficient intra prediction algorithms for screen content coding (SCC). The proposed algorithms focus on smooth regions frequently appeared in screen content videos, which have the characteristics of noiselessness. All the samples in a noiseless smooth region exhibit exactly the same pixel value. We then propose two intra coding techniques for noiseless smooth regions in SCC based on the smoothness of the boundary samples which are used for intra prediction. Our proposed algorithm can reduce computational complexity by at most 26.7% while keeping nearly the same video quality. Moreover, by removing the redundant coding bits for intra prediction modes, computational complexity can be further reduced to at most 53.3% in terms of encoding time with bitrate reduction up to 1.2%. Sik-Ho Tsang, Yui-Lam Chan, Wan-Chi Siu |
ICASSP | 3 |
| 2015 | Real time railway extraction by angle alignment measureabstractRail extraction is a fundamental and important step in railway Driver Assistant System, which is now an important application of image processing. The task is challenging as the railway is exposed to different environments. This paper proposes a railway extraction scheme, using a novel connectivity measure method named Angle Alignment Measure. The proposed scheme is robust to luminance and color variation, without edge extraction process. Railways with different lengths and patterns can be extracted under various lighting and weather conditions. More importantly, the computation complexity of the proposed scheme is very low, requiring only on average 25ms to process a frame on smart phone and 5ms on desktop computer, which are significantly better than algorithms in the literature. Wan-Chi Siu |
ICIP | 2 |
| 2015 | Efficient depth intra mode decision by reference pixels classification in 3D-HEVCabstractThe uniform intra prediction increases the intra prediction modes up to 35 and brings better coding efficiency in HEVC. Besides, depth modelling modes (DMMs) are introduced in depth intra coding of 3D-HEVC to preserve sharp edges and avoid ringing artifacts in a synthesized view. Meanwhile, the encoding time of depth intra coding rapidly increases due to a huge number of intra mode candidates. Based on the spatial correlation of a depth map, we find that not all of the intra modes are necessary to be considered in most cases. Hence, a fast content-dependent depth intra mode decision algorithm is raised in this paper by classifying the spatial distribution of the reference pixels. Simulation results show that the proposed adaptive fast algorithm can save 21%-35% time of the depth coding with the insignificant bit rate increase compared with the state-of-the-art algorithm. Hongbin Zhang 0005, Chang-Hong Fu 0002, Yui-Lam Chan, Sik-Ho Tsang, Wan-Chi Siu |
ICIP | 5 |
| 2015 | Practical application of random forests for super-resolution imagingabstractIn this paper, a novel learning-based single image super-resolution method using random forest is proposed. Different from example-based super-resolution methods which search for similar image patches from an external database or the input image, and the sparse representation model based methods which rely on the sparse representation, this proposed super-resolution with random forest (SRRF) method takes the divide-and-conquer strategy. Random forest is applied to classify the training LR-HR patch pairs into a number of classes. Within every class, a simple linear regression model is used to model the relationship between the LR image patches and their corresponding HR image patches. Experimental results show that the proposed SRRF method can generate the state-of-the-art super-resolved images with near real-time performance. Junjie Huang 0002, Wan-Chi Siu |
ISCAS | 2 |
| 2015 | Learning-based image interpolation via robust k-NN searching for coherent AR parameters estimation
Kwok-Wai Hung, Wan-Chi Siu |
J. Vis. Commun. Image Represent. | 2 |
| 2015 | Single-image super-resolution using iterative Wiener filter based on nonlocal means
Kwok-Wai Hung, Wan-Chi Siu |
Signal Process. Image Commun. | 2 |
| 2015 | Compression of Multiple DNA Sequences Using Intra-Sequence and Inter-Sequence SimilaritiesabstractTraditionally, intra-sequence similarity is exploited for compressing a single DNA sequence. Recently, remarkable compression performance of individual DNA sequence from the same population is achieved by encoding its difference with a nearly identical reference sequence. Nevertheless, there is lack of general algorithms that also allow less similar reference sequences. In this work, we extend the intra-sequence to the inter-sequence similarity in that approximate matches of subsequences are found between the DNA sequence and a set of reference sequences. Hence, a set of nearly identical DNA sequences from the same population or a set of partially similar DNA sequences like chromosome sequences and DNA sequences of related species can be compressed together. For practical compressors, the compressed size is usually influenced by the compression order of sequences. Fast search algorithms for the optimal compression order are thus developed for multiple sequences compression. Experimental results on artificial and real datasets demonstrate that our proposed multiple sequences compression methods with fast compression order search are able to achieve good compression performance under different levels of similarity in the multiple DNA sequences. Kin-On Cheng, Paula Wu, Ngai-Fong Law, Wan-Chi Siu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2015 | Fast Image Interpolation via Random ForestsabstractThis paper proposes a two-stage framework for fast image interpolation via random forests (FIRF). The proposed FIRF method gives high accuracy, as well as requires low computation. The underlying idea of this proposed work is to apply random forests to classify the natural image patch space into numerous subspaces and learn a linear regression model for each subspace to map the low-resolution image patch to high-resolution image patch. The FIRF framework consists of two stages. Stage 1 of the framework removes most of the ringing and aliasing artifacts in the initial bicubic interpolated image, while Stage 2 further refines the Stage 1 interpolated image. By varying the number of decision trees in the random forests and the number of stages applied, the proposed FIRF method can realize computationally scalable image interpolation. Extensive experimental results show that the proposed FIRF(3, 2) method achieves more than 0.3 dB improvement in peak signal-to-noise ratio over the state-of-the-art nonlocal autoregressive modeling (NARM) method. Moreover, the proposed FIRF(1, 1) obtains similar or better results as NARM while only takes its 0.3% computational time. Junjie Huang 0002, Wan-Chi Siu |
IEEE Trans. Image Process. | 2 |
| 2014 | Block-adaptive DCT-Wiener image up-samplingabstractDCT-Wiener image up-sampling scheme is highly desirable since it makes use of the advantages of the information in both spatial and DCT domains. The idea is to combine the observed low-frequency DCT coefficients with the estimated high-frequency DCT coefficients obtained by the Wiener filters in the spatial domain. However, the available 1-D and 2-D Wiener filters that were proposed for high-frequency DCT coefficients estimation are block non-adaptive, mainly due to the limited information from the observed image. In this paper, we propose a block-adaptive Wiener filter by utilizing the information from external training data. During the online estimation, for each image block, the k-nearest relevant DCT LR-HR block pairs are searched from the training data, in order to estimate the coefficients of the Wiener filter. Experimental results show that the proposed block-adaptive Wiener filter improves the PSNR value of the DCT-Wiener scheme by 1.5 dB compared with that using non-adaptive 1-D Wiener filter. Kwok-Wai Hung, Wan-Chi Siu |
ICASSP | 2 |
| 2014 | Using large color LBP in generalized hough transformabstractA new color based generalized Hough transform algorithm which initially generates a similar color map between the prototype object and a test image and uses a novel Large Color Local Binary Pattern (LCLBP) descriptor on the similar color map for feature extraction and object detection is proposed. The novel LCLBP descriptor can efficiently capture the local structure of target object on the similar color map. According to the experiment results, the proposed algorithm can provide comparable accuracy with the state-of-the-art Hough algorithm, but ours is 27 times faster. Junjie Huang 0002, Wan-Chi Siu |
ICIP | 2 |
| 2014 | Patch based image denoising using the finite ridgelet transform for less artifacts
Yunxia Liu 0001, Ngai-Fong Law, Wan-Chi Siu |
J. Vis. Commun. Image Represent. | 3 |
| 2014 | Establishment of linkages across GOP boundaries for reverse playback on compressed video
Chang-Hong Fu 0002, Yui-Lam Chan, Wan-Chi Siu |
Signal Process. Image Commun. | 3 |
| 2014 | Novel DCT-Based Image Up-Sampling Using Learning-Based Adaptive k-NN MMSE EstimationabstractImage up-sampling in the discrete cosine transform (DCT) domain is a challenging problem because DCT coefficients are de-correlated, such that it is nontrivial to estimate directly high-frequency DCT coefficients from observed low-frequency DCT coefficients. In the literature, DCT-based up-sampling algorithms usually pad zeros as high-frequency DCT coefficients or estimate such coefficients with limited success mainly due to the nonadaptive estimator and restricted information from a single observed image. In this paper, we tackle the problem of estimating high-frequency DCT coefficients in the spatial domain by proposing a learning-based scheme using an adaptive k-nearest neighbor weighted minimum mean squares error (MMSE) estimation framework. Our proposed scheme makes use of the information from precomputed dictionaries to formulate an adaptive linear MMSE estimator for each DCT block. The scheme is able to estimate high-frequency DCT coefficients with very successful results. Experimental results show that the proposed up-sampling scheme produces the minimal ringing and blocking effects, and significantly better results compared with the state-of-the-art algorithms in terms of peak signal-to-noise ratio (more than 1 dB), structural similarity, and subjective quality measurements. Kwok-Wai Hung, Wan-Chi Siu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2013 | Hybrid DCT-Wiener-based interpolation via learnt Wiener filterabstractThe hybrid DCT-Wiener-based (DCT-WB) interpolation scheme provides a powerful framework to interpolate an image by utilizing the information in both spatial and DCT domain. In this paper, we investigate the bottleneck of this hybrid scheme and propose a 2D non-separable block-based Wiener filter for the hybrid scheme. The Wiener filter is learnt using training image pairs through the minimum mean squares error estimation. The proposed Wiener filter resolves the quarter-pixel shift issue and provides much better performance over the original 1D 6-tap pixel-based Wiener filter. Experimental results show that incorporating the proposed Wiener filter into the hybrid scheme improves the PSNR (0.44 dB), SSIM and subjective quality for our extensive experimental work on testing images with various contents. Kwok-Wai Hung, Wan-Chi Siu |
ICASSP | 2 |
| 2013 | Region-based weighted prediction algorithm for H.264/AVC video codingabstractThis paper proposes a novel region-based weighted prediction (WP) algorithm to encode scenes with complex brightness variations. It facilitates the use of multiple WP parameter sets in a single reference frame by utilizing the framework of multiple reference frame motion estimation (MRF-ME). With this arrangement, different macroblocks in the current frame can use different WP parameter sets even when they are predicted from the same reference frame. To support this, a region partitioning process is designed to divide the current frame into different regions where each one has some degree of uniformity in its brightness variation. Multiple sets of region-based WP parameters can then be estimated accurately. Consequently, the proposed algorithm can improve prediction in scenes with different degrees of brightness variations in different regions of the same picture. Results show that the region-based algorithm can achieve significant coding gains of scenes with complex brightness variations. Sik-Ho Tsang, Tsz-Kwan Lee, Yui-Lam Chan, Wan-Chi Siu |
ISCAS | 4 |
| 2013 | Improved hierarchial intra prediction based on adaptive interpolation filtering for lossless compressionabstractOften the adaptive interpolation filter provides a more accurate interpolation technique in the formulation of the sub-pixel reference frame for motion estimation and compensation. Inspired by its possible advantages, we make use of the adaptive interpolation filter for efficient lossless intra prediction in H.264/AVC. Specifically, four subblocks are firstly formed by sampling pixels in one macroblock/block, and then a hierarchical intra prediction is performed on these subblocks. The first subblock is predicted based on the intra spatial prediction method in H.264/AVC, and subsequently it is encoded and reconstructed. The remaining three subblocks are then predicted based on adaptive interpolation filters by using the neighboring reconstructed samples. The residual block is encoded by rate optimization. Experimental results show that the proposed algorithm can improve the compression efficiency of lossless intra coding. Wan-Chi Siu |
ISCAS | 2 |
| 2013 | Improved algorithm for detecting zero-quantised discrete cosine transform coefficients in H.264/AVC (revised version)abstractIn this study, an efficient approach for detecting zero‐quantised discrete cosine transform (DCT) coefficients for video coding is developed. Compared with conventional detection methods of zero‐quantised DCT coefficients used in H.264/AVC, the proposed algorithm has two major features. First, a new classification of patterns for DCT, quantisation, inverse quantisation and inverse discrete cosine transform processes is proposed. By taking a zigzag scanning order into the classification, the quantised DCT coefficients can be coded efficiently. Second, the thresholds for detecting zero‐quantised DCT coefficients are determined by combining the Gaussian distribution with a theoretical analysis of the DCT and quantisation in H.264/AVC. Experimental results show that the proposed algorithm achieves an average timesaving of more than 40% compared with the algorithm in reference software JM12.2 of H.264/AVC. When compared with other algorithms in the literature, it also gives the best performance in terms of both rate‐distortion and time saving. Wan-Chi Siu |
IET Image Process. | 2 |
| 2013 | Motion estimation in low-delay hierarchical p-frame coding using motion vector composition
Tsz-Kwan Lee, Yui-Lam Chan, Wan-Chi Siu |
J. Vis. Commun. Image Represent. | 3 |
| 2013 | Region-Based Weighted Prediction for Coding Video With Local Brightness VariationsabstractThis paper presents a new region-based scheme for the estimation of weighted prediction (WP) parameter sets for encoders of the H.264/MPEG-4 AVC standard. The proposed scheme is specifically designed for handling local brightness variations (LBVs) in video scenes. It is achieved by making use of multiple WP parameter sets for various regions and assigning them to the same reference frame. An accurate estimation of multiple WP parameter sets is accomplished by: 1) partitioning regions with a simple WP parameter estimator; 2) selecting regions where WP should be applied; and 3) estimating accurate WP parameter sets with a quasioptimal WP parameter estimator. The multiple WP parameter sets of different regions are encoded using the framework of multiple reference frames in the H.264/MPEG-4 AVC standard. With this arrangement, the proposed scheme is compliant with the H.264/MPEG-4 AVC standard. To reduce the implementation cost, a reduction of the memory requirement is realized via look-up tables (LUTs). Experimental results show that the region-based scheme can efficiently handle scenes with global and LBVs and achieve significant coding gain over other WP schemes. Furthermore, our scheme with LUTs can reduce the memory requirement by about 80% while keeping the same coding efficiency as that without LUTs. Sik-Ho Tsang, Yui-Lam Chan, Wan-Chi Siu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2013 | Novel Adaptive Algorithm for Intra Prediction With Compromised Modes Skipping and Signaling Processes in HEVCabstractUp to 35 intra prediction modes are available for each Luma prediction unit in the coming HEVC standard. This can provide more accurate predictions and thereby improve the compression efficiency of intra coding. However, the encoding complexity is thus increased dramatically due to a large number of modes involved in the intra mode decision process. In addition, more overhead bits should be assigned to signal the mode index. Intuitively, it is not necessary for all modes to be checked and signaled all the time. Therefore, a novel adaptive modes skipping algorithm for mode decision and signaling processing is presented in this paper. More specifically, three optimized candidate sets with 1, 19, and 35 intra prediction modes are initiated for each prediction unit in the proposed algorithm. Based on the statistical properties of the neighboring reference samples used for intra prediction, the proposed algorithm is able to adaptively select the optimal set from the three candidates for each prediction unit preceding the mode decision and signaling processing. As a result, the mode decision process can be speeded up due to some modes skipping in the first two sets, and importantly less bits are required to signal the mode index. Experimental results show that, compared to the test model HM7.0 of HEVC, BD-Rate savings of 0.18% and as well as 0.18% on average are achieved for AI-Main and AI-HE10 cases for low-bitrate ranges, and the average encoding times can also be reduced by 8%-38% and 8%-34% for AI-Main and AI-HE10 cases in low-bitrate ranges, respectively. Wan-Chi Siu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2012 | A two dimensional camera identification method based on image sensor noiseabstractIn this paper, we propose a two-dimensional digital camera identification method based on the photo-response non-uniformity (PRNU). The traditional identification method is based on a correlation estimator which calculates the correlation between the reference PRNU and the PRNU extracted from the testing image. However, the correlation calculated greatly depends on the image content. To reduce the image content effect in classification, a correlation predictor is trained based on different types of image features. By using the predicted correlation and the actual correlation, a 2D classifier using support vector machine is proposed in this paper. Experimental results show that the proposed method can have a more flexible threshold setting which gives a better identification results as compared to the traditional identification method. Lit-Hung Chan, Ngai-Fong Law, Wan-Chi Siu |
ICASSP | 3 |
| 2012 | Single image super-resolution using iterative Wiener filterabstractIn this paper, we propose an iterative Wiener filter which can simultaneously perform interpolation and restoration by using non-local means to directly model the correlation between the desired high-resolution image and observed low-resolution image. A novel mechanism is proposed to control the decay speed of the correlation function while iteratively updating both estimated correlation and high-resolution image. During the iterations, the image is decomposed into patches with similar intensities at initial iterations and the patches are connected naturally with good convergence. Experimental results show that the proposed algorithm is able to produce natural image structures, and provides better PSNR and visual quality than the state-of-the-art algorithms using the sparse representation and natural image priors. Kwok-Wai Hung, Wan-Chi Siu |
ICASSP | 2 |
| 2012 | Depth-assisted nonlocal means hole filling for novel view synthesisabstractIn novel view synthesis using the depth image based rendering, there exist some unknown pixel intensities (holes) due to the unexposed area. In this paper, we propose a depth-assisted nonlocal means algorithm to fill the holes using the information in the current frame and other frames in the synthesized video. The nonlocal means has been successfully applied to the video super-resolution applications. In hole filling, the challenge is the irregular interpolation. However, there can be a relatively reliable depth map, such that we are able combine the irregular intensity information and the reliable depth information, and make use of the formulation of nonlocal means to fill the holes. Experimental results show that the proposed algorithm outperforms the conventional spatial and temporal approaches in terms of visual quality. Kwok-Wai Hung, Wan-Chi Siu |
ICIP | 2 |
| 2012 | Keyframe selection for motion capture using motion activity analysisabstractMotion capture data acquired from high definition cameras creates accurate human motion representation but introduces many redundant frames which pose a problem in data storage and motion retrieval purposes. In this paper, a keyframing approach is proposed to reduce the motion data by extracting keyframes using motion analysis approach in sampling windows. Motion changes in sampling windows for original motion without frame skipping and with frame skipping are computed. The difference in the motion changes is the main aspect in deciding whether the frames in sampling windows are possible candidates for keyframe selection. Simulation results showed that the proposed method is able to achieve an overall good visual quality for different types of motion. It also gives an improvement of up to 52% in terms of mean square error measurement, as compared to the existing keyframe extraction method, which is curve simplification method. Ming-Hwa Kim, Lap-Pui Chau, Wan-Chi Siu |
ISCAS | 3 |
| 2012 | A new 3-phase design exploration methodology for video processor designabstractWhen making video processor design, conventional design exploration methodologies take extremely long time in parameter optimization but the final design may not necessarily meet the application requirements since the architecture cannot deviate too much from the initial design. To speed up the design process, statistical performance models were used to guide the simulation; however their accuracy is questionable. In this paper, a new 3-phase design exploration methodology for video processor is proposed. It makes use of an almost cycle-accurate performance model to provide information for refining the processor architecture. It can derive the optimal architecture in a much shorter period of time than the conventional methods. We successfully implemented a few video coding/decoding applications on the video processor derived from the proposed methodology. Simulation results show that it outperforms other video processors in both cost and performance perspectives. Wing-Yee Lo, Daniel Pak-Kong Lun, Wan-Chi Siu |
ISCAS | 3 |
| 2012 | Iterative search strategy with selective bi-directional prediction for low complexity multiview video coding
Zhipin Deng, Yui-Lam Chan, Kebin Jia, Chang-Hong Fu 0002, Wan-Chi Siu |
J. Vis. Commun. Image Represent. | 5 |
| 2012 | Flash scene video coding using weighted prediction
Sik-Ho Tsang, Yui-Lam Chan, Wan-Chi Siu |
J. Vis. Commun. Image Represent. | 3 |
| 2012 | Iterative bicluster-based least square framework for estimation of missing values in microarray gene expression data
Kin-On Cheng, Ngai-Fong Law, Wan-Chi Siu |
Pattern Recognit. | 3 |
| 2012 | Hybrid motion estimation scheme for secondary SP-frame coding using inter-frame correlation and FMO
Ki-Kit Lai, Yui-Lam Chan, Chang-Hong Fu 0002, Wan-Chi Siu |
Signal Process. Image Commun. | 4 |
| 2012 | Robust Soft-Decision Interpolation Using Weighted Least SquaresabstractSoft-decision adaptive interpolation (SAI) provides a powerful framework for image interpolation. The robustness of SAI can be further improved by using weighted least-squares estimation, instead of least-squares estimation in both of the parameter estimation and data estimation steps. To address the mismatch issue of "geometric duality" during parameter estimation, the residuals (prediction errors) are weighted according to the geometric similarity between the pixel of interest and the residuals. The robustness of data estimation can be improved by modeling the weights of residuals with the well-known bilateral filter. Experimental results show that there is a 0.25-dB increase in peak signal-to-noise ratio (PSNR) for a sample set of natural images after the suggested improvements are incorporated into the original SAI. The proposed algorithm produces the highest quality in terms of PSNR and subjective quality among sophisticated algorithms in the literature. Kwok-Wai Hung, Wan-Chi Siu |
IEEE Trans. Image Process. | 2 |
| 2011 | Single image super-resolution using Gaussian process regressionabstractIn this paper we address the problem of producing a high-resolution image from a single low-resolution image without any external training set. We propose a framework for both magnification and deblurring using only the original low-resolution image and its blurred version. In our method, each pixel is predicted by its neighbors through the Gaussian process regression. We show that when using a proper covariance function, the Gaussian process regression can perform soft clustering of pixels based on their local structures. We further demonstrate that our algorithm can extract adequate information contained in a single low-resolution image to generate a high-resolution image with sharp edges, which is comparable to or even superior in quality to the performance of other edge-directed and example-based super-resolution algorithms. Experimental results also show that our approach maintains high-quality performance at large magnifications. He He 0001, Wan-Chi Siu |
CVPR | 2 |
| 2011 | Fast video interpolation/upsampling using linear motion modelabstractRecently, the probabilistic motion field was proposed for super-resolution reconstruction (SRR). In the interpolation step of SRR, a missing pixel can be estimated by the weighted average of neighboring pixels, which are weighted by the errors with the missing pixel. However, the errors are far from true values due to the approximated missing pixel in calculating the errors. Hence, in this paper, we propose a linear motion model to better approximate the errors, which results in a better interpolation quality. Experimental results show that a gain of 0.6 dB in PSNR is achievable using this linear motion model, and only a small number of neighboring pixels have to be used for fast interpolation/upsampling. Kwok-Wai Hung, Wan-Chi Siu |
ICIP | 2 |
| 2011 | Improved lossless coding algorithm in H.264/AVC based on hierarchical intra predictionabstractIn this paper, an improved lossless intra prediction algorithm based on H.264/AVC framework is proposed. In the new algorithm, samples in a macroblock/block are hierarchically predicted, instead of using a block-based prediction as a whole. More specifically, four groups are extracted from the samples in the MB/block. Samples in the first group are firstly predicted based on directional intra prediction method, and then samples in other groups are predicted using the samples in the first group as the reference. As a result, the information left in the residual block can be reduced since the samples can be predicted accurately by using nearer references. Experimental results show that compared with other methods in the literature the proposed algorithm gives a much better compression performance with the lowest encoding complexity. Wan-Chi Siu |
ICIP | 2 |
| 2011 | Fast iterative search for motion and disparity estimation in stereoscopic video codingabstractIn this paper, a fast algorithm is proposed to speed up the motion and disparity estimation in stereoscopic video coding. Based on the stereo-motion consistency constraint, an iterative search strategy is suggested to get the motion and disparity vectors simultaneously. A credible base vector selection scheme and an adaptive search range adjustment technique are designed to further strengthen the iterative search. Results show that the complexity can be significantly reduced compared to the JMVM full search with a negligible quality drop. Zhipin Deng, Kebin Jia, Yui-Lam Chan, Chang-Hong Fu 0002, Wan-Chi Siu |
VCIP | 5 |
| 2011 | Low-complexity intra prediction algorithm for video down-sizing transcoderabstractThis paper presents an efficient intra mode decision scheme for down-sizing video transcoding in H.264 using support vector machines (SVMs). In order to reduce the high computational complexity of intra prediction in the H.264 re- encoder, the proposed scheme uses SVMs to exploit the correlation between coding information extracted from the input high-resolution bit-stream and the coding modes of macro- blocks in down-sized video, and turns the intra mode decision problem into pattern classification problem to achieve low complexity video transcoding framework. After the SVMs classifier, improbable modes are eliminated and only a small number of candidate modes are carried out using the RDO operations. Hence, remarkable computing time of up to 65.3% can be saved with trivial bit gain and quality loss. Zhuo-Yi Lu, Kebin Jia, Wan-Chi Siu |
VCIP | 3 |
| 2011 | A fast stereoscopic video coding algorithm based on JMVM
Zhipin Deng, Kebin Jia, Yui-Lam Chan, Chang-Hong Fu 0002, Wan-Chi Siu |
Sci. China Inf. Sci. | 5 |
| 2011 | Improved SIMD Architecture for High Performance Video ProcessorsabstractSingle instruction multiple data (SIMD) execution is in no doubt an efficient way to exploit the data level parallelism in image and video applications. However, SIMD execution bottlenecks must be tackled in order to achieve high execution efficiency. We first analyze in this paper the implementation of two major kernel functions of H.264/AVC namely, SATD and subpel interpolation, in conventional SIMD architectures to identify the bottlenecks in traditional approaches. Based on the analysis results, we propose a new SIMD architecture with two novel features: 1) parallel memory structure with variable block size and word length support, and 2) configurable SIMD structure. The proposed parallel memory structure allows great flexibility for programmers to perform data access of different block sizes and different word lengths. The configurable SIMD structure allows almost “random” register file access and slightly different operations in ALUs inside SIMD. The new features greatly benefit the realization of H.264/AVC kernel functions. For instance, the fractional motion estimation, particularly the half to quarter pixel interpolation, can now be executed with minimal or no additional memory access. When comparing with the conventional SIMD systems, the proposed SIMD architecture can have a further speedup of 2.1X to 4.6X when implementing H.264/AVC kernel functions. Based on Amdahl's law, the overall speedup of H.264/AVC encoding application can be projected to be 2.46X. We expect significant improvement can also be achieved when applying the proposed architecture to other image and video processing applications. Wing-Yee Lo, Daniel Pak-Kong Lun, Wan-Chi Siu, Jiqiang Song |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2011 | Transform Kernel Selection Strategy for the H.264/AVC and Future Video Coding StandardsabstractIn this paper, we propose a new discrete cosine transform (DCT)-like kernel IK(5, 7, 3) and revitalize another DCT-like kernel IK(13, 17, 7) for the transform coding process of hybrid video coding. Making use one of these kernels together with the H.264/AVC kernel IK(1, 2, 1), we are able to design new multiple-kernel schemes which give better coding performance over that of the conventional approaches. All these schemes make use of the adaptive kernel mechanism at macroblock-level (MB-AKM), which requires heavy computation during the encoding process. We subsequently discovered that a rate-distortion feature extracted from a pair of kernels gives an intrinsic property that can be used to select a better kernel for a two-kernel MB-AKM system. This is a powerful tool with theoretical interest and practical uses. In order to reduce computation substantially, we make use of this tool to make an analysis and design of a frame-level adaptive kernel mechanism and come up with a simple solution that the kernel IK(1, 2, 1) be used for I-frames and P-frames and the kernel IK(5, 7, 3) be used for B-frames coding. This proposed frame-based AKM gives similar, or even better, performance as the proposed macroblock-based AKM. Furthermore, it substantially reduces computation and certainly gives a good improvement in terms of the PSNR and bitrate compared to those obtained from the H.264/AVC default scheme and other MB-AKM schemes available in the literature. Chau-Wai Wong, Wan-Chi Siu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2011 | Comments on "2-D Order-16 Integer Transforms for HD Video Coding"abstractIn a recent paper, Dong proposed a set of order-16 nonorthogonal integer cosine transforms (NICTs). They proved that the reconstruction error caused by the nonorthogonality is negligible as compared to the error caused by the quantization. However, we would like to point out three problems found in derivations and also give two comments. Nevertheless, the problems are defects only, hence do not affect the overall justifications to the proposed NICT. This letter is to enhance and clarify the proof of Dong 's work. Chau-Wai Wong, Wan-Chi Siu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2011 | Analysis of Dyadic Approximation Error for Hybrid Video Codecs With Integer TransformsabstractIn this paper, we present an analysis of the dyadic approximation error introduced by the integerization of transform coding in H.264/AVC-like codecs. We derive the analytical formulations for dyadic approximation error and nonorthogonality error. We further classify the dyadic approximation error into a "system error" and a "nonflat error," and proposed two models for them. We found that the "nonflat error" has a substantial impact on video quality if the number of shifting bits at decoder side (DQ_BITS) is small. We also give a theoretical justification on why scaling factors at encoder side are better to be adapted to the rescaling factors at decoder side in H.264/AVC-like codecs. Chau-Wai Wong, Wan-Chi Siu |
IEEE Trans. Image Process. | 2 |
| 2010 | Adaptive Directional Window Selection for Edge-Directed InterpolationabstractIn this paper, we present an adaptive directional window selection for the edge-directed interpolation. The new window selection can solve the problem of covariance mismatch in high frequency and texture regions. It makes use of a practical directional elliptic window which works according to the edge direction sliding along an edge and then subsequently chooses the best window evaluated by choosing the elliptic window which has the lowest Means Square Error (MSE). Experimental results show that by the proposed technique can generate a high quality interpolated image which is better than other edge directed interpolation approaches. Experimental results also provided on different images to justify the value of this approach at the end of the paper. Chi-Shing Wong, Wan-Chi Siu |
ICCCN | 2 |
| 2010 | Improved image interpolation using bilateral filter for weighted least square estimationabstractNew edge-directed interpolation (NEDI) consists two steps. The two steps are parameter and data estimation. The second step can be replaced by a recently proposed technique called soft-decision to consider the consistency of image structure during this data estimation. The original idea of both steps is to assume equal variances for all estimation errors, such that an ordinary least squares (OLS) estimator can be used. Due to the existence of noise, different object layers, changing in image structures, different spatial distance to the missing data, etc, we observe that the estimation errors of data samples have unequal variances. Hence, a weighted least square (WLS) estimator should be used for both steps. The bilateral filter, which can accurately remove noise and preserve image structure, has been used to model successfully the weights of squared residuals, such that we can apply it to both steps of the estimation. Experimental results show that the average PSNR of this improved interpolation method is 0.47 dB and 0.23 dB higher than two similar approaches, NEDI and Soft-decision Adaptive Interpolation (SAI) using 24 natural images from Kodak. The subjective results show improvement as well. Kwok-Wai Hung, Wan-Chi Siu |
ICIP | 2 |
| 2010 | Message from the general chairabstractOn behalf of the ICIP 2010 Organizing Committee, I extend to you a heartfelt welcome to Hong Kong. This is a unique city where the eastern and western cultures meet. If it were square-shaped, the size of Hong Kong would be no more that 20×20 square miles. Despite its small area, Hong Kong is one of the most popular tourist destinations in Asia. In 2008, a total of 29.5 million people visited Hong Kong, a 4.7% increase since the year 2007. Hong Kong, better known as Pearl of the Orient, is famous for its spectacular harbour views. September is one of the best times to visit Hong Kong. It is very pleasant and there is plenty of sunshine. The average temperature is 23°C and the average humidity is 84%. There could be typhoons during September. However the chance of typhoon hitting is less at the end of September. Wan-Chi Siu |
ICIP | 1 |
| 2010 | Fast video object detection via multiple background modelingabstractIn this paper, a robust background extraction and novel object detection are proposed, which comprise of filtering operations to detect non background objects in a monitoring scene. Conventionally, a statistical background model is extracted by using a training sequence without foreground objects and the background model parameters are being updated continuously to adapt changes in the scene. However, it is not possible to require a monitoring scene to be static. Furthermore, static objects in the scene could be adapted into the background. Problems arise when static objects start to move again. The convention method would produce false alarms in the detection process. In our proposed algorithm, two background models are constructed by using N-bins histogram method to indicate short term and long term changes of the monitoring scene. We then apply background subtractions to the current frame to obtain two error frames, which are combined for objects detection and classification. Extensive experimental work has been done, results of which show that the present approach provides a better solution compared with the conventional approach, including to resolve the problem of re-active objects. Kin-Yi Yam, Wan-Chi Siu, Ngai-Fong Law, Chok-Ki Chan |
ICIP | 2 |
| 2010 | Local affine motion prediction for H.264 without extra overheadabstractConventional coding system using translation only motion estimation and compensation system cannot efficiently handle complex inter-frame motion activity including scaling, rotation and various forms of distortion. In this paper, we propose a local affine motion prediction method, which manages to enhance the inter-frame image prediction quality using the conventional motion vectors. Our method is characterized with the fact that no extra bit has to be sent to the decoder for proper decoding. Experimental results show that our method manages to achieve a maximum average bit rate reduction of 1.7% compared to the conventional inter-frame prediction method using translation only motion compensation techniques, at no cost on quality degradation. Hoi-Kok Cheung, Wan-Chi Siu |
ISCAS | 2 |
| 2010 | A new motion vector composition algorithm for fast-forward video playback in H.264abstractWith the rapid growth of streaming digital videos, it is desirable to access video segments of interest by searching through the video contents with a faster speed than a normal playback. Fast-forward playback is the key function that enables quick browsing of videos. It can be realized by a frame-skipping transcoder which transcodes only the frames required for playback at the desired fast speed. Various motion vector (MV) composition algorithms aim at reducing the computational complexity of the transcoder. They only perform fairly in limited skipped frames scenarios. In this paper, a new vector selection algorithm is proposed to compose a new motion vector (MV) from a set of candidate MVs for minimizing prediction errors due to a larger frame-skipping factor. Experimental results show that the proposed algorithm can deliver a remarkable improvement on the rate-distortion performance over other algorithms. Tsz-Kwan Lee, Chang-Hong Fu 0002, Yui-Lam Chan, Wan-Chi Siu |
ISCAS | 4 |
| 2010 | Improved hybrid coding scheme for intra 4×4 residual block produced by H.264/AVCabstractIn this paper, the intra residual macroblock produced by H.264 is investigated. Based on its characteristics, an efficient two-layer coding scheme for the intra residual macroblock is developed. The rate-distortion performance for the proposed coding scheme is evaluated. Experimental results show that our algorithm can achieve better coding performance. Wan-Chi Siu |
ISCAS | 2 |
| 2010 | Fast motion and disparity estimation for multiview video coding
Zhipin Deng, Kebin Jia, Yui-Lam Chan, Chang-Hong Fu 0002, Wan-Chi Siu |
Frontiers Comput. Sci. China | 5 |
| 2010 | An efficient motion vector composition algorithm for fast-forward playback in a video streaming system
Chang-Hong Fu 0002, Tsz-Kwan Lee, Yui-Lam Chan, Wan-Chi Siu |
J. Vis. Commun. Image Represent. | 4 |
| 2010 | Fast extraction of wavelet-based features from JPEG images for joint retrieval with JPEG2000 images
Kin-On Cheng, Ngai-Fong Law, Wan-Chi Siu |
Pattern Recognit. | 3 |
| 2010 | An efficient retinex-like brightness normalization method for coding camera flashes and strong brightness variation in videos
Hoi-Kok Cheung, Wan-Chi Siu, David Dagan Feng, Zhiyong Wang 0001 |
Signal Process. Image Commun. | 2 |
| 2010 | Quantized Transform-Domain Motion Estimation for SP-Frame Coding in Viewpoint Switching of Multiview VideoabstractThe brand-new SP-frame in H.264 facilitates drift-free bitstream switching. Notwithstanding the guarantee of seamless switching, the cost is the bulky size of secondary SP-frames. This induces a significant amount of additional space or bandwidth for storage or transmission. In this paper, our investigation reveals that the size of secondary SP-frames is more severe when the correlation between the two bitstreams becomes smaller. Examples include viewpoint switching in multiview video and bitstream switching in single-view video with complex motion. For this reason, a new motion estimation and compensation technique, which is operated in the quantized-transform (QDCT) domain, is designed for coding secondary SP-frames. Our proposed work aims at keeping the secondary SP-frames as small as possible without affecting the size of primary SP-frames by incorporating QDCT-domain motion estimation and compensation in the secondary SP-frame coding. Simulation results show our proposed scheme overwhelmingly outperforms the conventional pixel-domain motion estimation technique. As a consequence, the size of secondary SP-frames can be reduced remarkably, especially in multiview video and single-view video with complex motion. Ki-Kit Lai, Yui-Lam Chan, Wan-Chi Siu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2009 | H.264 fast intra mode selection algorithm based on direction difference measure in the pixel domainabstractIn this paper, a fast mode decision algorithm for Intra prediction in the H.264/AVC is proposed. We use the characteristics of each directional prediction mode to compute the strength of directional differences in the original pixel domain to find the minimal direction error. This is the first time reported in the literature that the intrinsic differences between the real-data and the predictors of modes are used to form an algorithm for mode decision. The approach allows us to select several better candidate modes for evaluation instead of using the full search. Experimental results show that the proposed method can achieve more than 80% reduction in computation with negligible degradation in rate-distortion performance, and the results are better than other algorithms available in the literature. Wan-Chi Siu |
ICASSP | 2 |
| 2009 | New motion compensation model via frequency classification for fast video super-resolutionabstractA typical dynamic reconstruction-based super-resolution video involves three independent processes: registration, fusion and restoration. Fast video super-resolution systems apply translational motion compensation model for registration with low computational cost. Traditional motion compensation model assumes that the whole spectrum of pixels is consistent between frames. In reality, the low frequency component of pixels often varies significantly. We propose a translational motion compensation model via frequency classification for video super-resolution systems. A novel idea to implement motion compensation by combining the up-sampled current frame and the high frequency part of the previous frame through the SAD framework is presented. Experimental results show that the new motion compensation model via frequency classification has an advantage of 2dB gain on average over that of the traditional motion compensation model. The SR quality has 0.25dB gain on average after the fusion process which is to minimize error by making use of the new motion compensated frame. Kwok-Wai Hung, Wan-Chi Siu |
ICIP | 2 |
| 2009 | Kurtosis-based super-resolution algorithmabstractA kurtosis-based super-resolution image reconstruction algorithm is proposed in this paper. Firstly, we give the definition of the kurtosis image and analyze its two properties: (i) the kurtosis image is Gaussian noise invariant, and (ii) the absolute value of a kurtosis image becomes smaller as the the image gets smoother. Then we build a constrained absolute local kurtosis maximization function to estimate the high-resolution image by fusing multiple blurred low-resolution images corrupted by intensive white Gaussian noise. The Lagrange multiplier is used to solve the combinatorial optimization problem. Experimental results demonstrate that the proposed method is better than the conventional algorithms in terms of visual inspection and robustness, using both synthetic and real world examples under severe noise background. It has an improvement of 0.5 to 2.0 dB in PSNR over other approaches. Jianping Qiao, Xiangzeng Meng, Wan-Chi Siu |
ICME | 4 |
| 2009 | Efficient Inter Mode Decision for H.263 to H.264 Video Transcoding using Support Vector MachinesabstractThis paper presents an efficient mode decision algorithm for H.263 to H.264 interframe transcoding. The proposed scheme uses a support vector machines (SVMs) approach to investigate the relationship between data extracted from H.263 decoding stage and the optimal coding mode in H.264 re-encoding process. Based on the SVM classifier, the transcoder only enables a subset of candidate modes for each macroblock in the rate-distortion-optimized mode decision in H.264. The objective is to eliminate unlikely modes in earlier stages in order to achieve computation saving. Simulation results show that the proposed method can reduce the computational complexity of interframe transcoding by up to 77% while maintaining similar rate-distortion performance. Xuan Jing, Wan-Chi Siu, Lap-Pui Chau, Anthony G. Constantinides |
ISCAS | 2 |
| 2009 | Energy-based adaptive transform scheme in the DPRT domain and its application to image denoising
Yunxia Liu 0001, Yuhua Peng, Wan-Chi Siu |
Signal Process. | 3 |
| 2009 | Compressed-Domain Techniques for Error-Resilient Video Transcoding Using RPSabstractIn video applications where video sequences are compressed and stored in a storage device for future delivery, the encoding process is typically carried out without enough prior knowledge about the channel characteristics of a network. Error-resilient transcoding plays an important role to provide an addition of resilience to the video data, where or whenever it is needed. Recently, a reference picture selection (RPS) scheme has been adopted in an error-resilient transcoder in order to reduce error effects for the already encoded video bitstream. In this approach, the transcoder learns through a feedback channel about the damaged parts of a previously coded frame and then decides to code the next P-frame not relative to the most recent, but to an older, reference picture, which is known to be error-free in the decoder. One straightforward approach of adopting RPS in error-resilient transcoding is to decode all the P-frames from the previously nearest I-frame to the current transmitted frame which is then re-encoded with a new reference frame; this can create undesirable complexity in the transcoder as well as introduce re-encoding errors. In this paper, some novel techniques are suggested for an effective implementation of RPS in the error-resilient transcoder with the minimum requirement on its complexity. All the proposed techniques will manipulate video data in the compressed domain such that the computational loading of the transcoder is greatly reduced. By utilizing these new compressed-domain techniques, we develop a new structure to handle various types of macroblocks in the transcoder which re-uses motion vectors and prediction errors from the encoded bitstream. Experimental results demonstrate that significant improvements in terms of transcoder complexity and quality of reconstructed video can be achieved by employing our compressed-domain techniques. Yui-Lam Chan, Hoi-Kin Cheung, Wan-Chi Siu |
IEEE Trans. Image Process. | 3 |
| 2008 | Retinex based motion estimation for sequences with brightness variations and its application to H.264abstractConventional motion estimation does not take inter-frame brightness variations into consideration, which causes inefficient video coding for sequences involving brightness variations. H.264 provides a specific mode called weighted prediction targeting to improve the coding efficiency for this case. In this paper, we propose a Retinex based motion estimation scheme which effectively removes the inter-frame de-correlation factor resulting from brightness variations. We also propose to use some DCT techniques to generate the Retinex images for both current and reference images and apply conventional motion estimation and compensation procedures for coding. We applied the scheme to the H.264 testing the efficiency in the multiple reference frame motion compensation environment. Experimental results show that our proposed scheme outperforms the H.264 system with weighted prediction enabled. It allows the system to use a smaller number of reference frames for coding, e.g. 2, to achieve a similar (or slightly better) compression efficiency of the H.264 system using 5 reference frames. Hoi-Kok Cheung, Wan-Chi Siu, David Dagan Feng, Zhiyong Wang 0001 |
ICASSP | 2 |
| 2008 | Windowing technique for the DCT based retinex algorithm to handle videos with brightness variations coded using the H.264abstractConventional block based motion estimators assume constant inter-frame object brightness. Pixel discrepancy is resulted primarily from motion factor without considering the influence of brightness changes. In this paper, we propose a simple and efficient windowing technique using the Hamming window and integrate it to our previously proposed algorithm. The algorithm is based on retinex approach using the DCT technique and designed to handle brightness variations. The new technique manages to greatly reduce the influence of ripple effect and further increase the compression efficiency without adding any extra overhead bits to the bit-stream. We applied the scheme to H.264 for testing. Experimental results show that the retinex based approach is an effective technique to handle inter-frame brightness variations and outperforms the H.264 system with weighted prediction enabled for sequence involving brightness variations. With our proposed windowing technique using the Hamming window, the coding efficiency can be further improved by a maximum of 0.17dB. Hoi-Kok Cheung, Wan-Chi Siu, David Dagan Feng, Zhiyong Wang 0001 |
ICIP | 2 |
| 2008 | Viewpoint switching in multiview videos using SP-framesabstractThe distinguishing feature of multiview video lies in the interactivity, which allows users to select their favourite viewpoint. It switches bitstream at a particular view when necessary instead of transmitting all the views. The new SP-frame in H.264 is originally developed for multiple bit-rate streaming with the support of seamless switching. The SP-frame can also be directly employed in the viewpoint switching of multiview videos. Notwithstanding the guarantee of seamless switching using SP-frames, the cost is the bulky size of secondary SP-frames. This induces a significant amount of additional space or bandwidth for storage or transmission, especially for the multiview scenario. For this reason, a new motion estimation and compensation technique operating in the quantized transform (QDCT) domain is designed for coding secondary SP-frame in this paper. Our proposed work aims at keeping the secondary SP-frames as small as possible without affecting the size of primary SP-frames by incorporating QDCT-domain motion estimation and compensation in the secondary SP-frame coding. Simulation results show that the size of secondary SP-frames can be reduced remarkably in viewpoint switching. Ki-Kit Lai, Yui-Lam Chan, Chang-Hong Fu 0002, Wan-Chi Siu |
ICIP | 4 |
| 2008 | Efficient and low-complexity image coding with the lifting scheme and modified SPIHTabstractIn this paper, we propose an efficient and low complexity image coding algorithm based on the lifting wavelet transform and listless modified SPIHT (LWT-LMSPIHT). LWT-LMSPIHT jointly considers the advantages of progressive transmission and spatial scalability that were not fully provided by the SPIHT algorithm, thus it outperforms the SPIHT at low bit rates coding. The coding efficiency of LWT-LMSPIHT comes from three aspects. The lifting scheme lowers the number of arithmetic operations of the wavelet transform. Moreover, a significance reordering of the modified SPIHT ensures that it codes more significant information earlier in the bit stream belonging to the lower frequency bands than SPIHT to better exploit the energy compaction of the wavelet coefficients. Finally, a listless structure further reduces the amount of memory and improves the speed of compression by more than 47% for a 512×512 image, as compared with the SPIHT algorithm. Hong Pan 0001, Wan-Chi Siu, Ngai-Fong Law |
IJCNN | 2 |
| 2008 | Image annotation with parametric mixture model based multi-class multi-labelingabstractImage annotation, which labels an image with a set of semantic terms so as to bridge the semantic gap between low level features and high level semantics in visual information retrieval, is generally posed as a classification problem. Recently, multi-label classification has been investigated for image annotation since an image presents rich contents and can be associated with multiple concepts (i.e. labels). In this paper, a parametric mixture model based multi-class multi-labeling approach is proposed to tackle image annotation. Instead of building classifiers to learn individual labels exclusively, we model images with parametric mixture models so that the mixture characteristics of labels can be simultaneously exploited in both training and annotation processes. Our proposed method has been benchmarked with several state-of-the-art methods and achieved promising results. Zhiyong Wang 0001, Wan-Chi Siu, David Dagan Feng |
MMSP | 2 |
| 2008 | Identification of coherent patterns in gene expression data using an efficient biclustering algorithm and parallel coordinate visualizationabstractBACKGROUND: The DNA microarray technology allows the measurement of expression levels of thousands of genes under tens/hundreds of different conditions. In microarray data, genes with similar functions usually co-express under certain conditions only 1. Thus, biclustering which clusters genes and conditions simultaneously is preferred over the traditional clustering technique in discovering these coherent genes. Various biclustering algorithms have been developed using different bicluster formulations. Unfortunately, many useful formulations result in NP-complete problems. In this article, we investigate an efficient method for identifying a popular type of biclusters called additive model. Furthermore, parallel coordinate (PC) plots are used for bicluster visualization and analysis. RESULTS: We develop a novel and efficient biclustering algorithm which can be regarded as a greedy version of an existing algorithm known as pCluster algorithm. By relaxing the constraint in homogeneity, the proposed algorithm has polynomial-time complexity in the worst case instead of exponential-time complexity as in the pCluster algorithm. Experiments on artificial datasets verify that our algorithm can identify both additive-related and multiplicative-related biclusters in the presence of overlap and noise. Biologically significant biclusters have been validated on the yeast cell-cycle expression dataset using Gene Ontology annotations. Comparative study shows that the proposed approach outperforms several existing biclustering algorithms. We also provide an interactive exploratory tool based on PC plot visualization for determining the parameters of our biclustering algorithm. CONCLUSION: We have proposed a novel biclustering algorithm which works with PC plots for an interactive exploratory analysis of gene expression data. Experiments show that the biclustering algorithm is efficient and is capable of detecting co-regulated genes. The interactive analysis enables an optimum parameter determination in the biclustering algorithm so as to achieve the best result. In future, we will modify the proposed algorithm for other bicluster models such as the coherent evolution model. Kin-On Cheng, Ngai-Fong Law, Wan-Chi Siu, Alan Wee-Chung Liew |
BMC Bioinform. | 3 |
| 2008 | Peak constrained two-dimensional quadrantally symmetric eigenfilter design without transition band specification
Chi-Wah Kok, Wan-Chi Siu, Ying-Man Law |
Signal Process. | 2 |
| 2008 | Peak constrained least-squares QMF banks
Chi-Wah Kok, Wan-Chi Siu, Ying-Man Law |
Signal Process. | 2 |
| 2008 | A fast and low memory image coding algorithm based on lifting wavelet transform and modified SPIHT
Hong Pan 0001, Wan-Chi Siu, Ngai-Fong Law |
Signal Process. Image Commun. | 2 |
| 2008 | Frame Complexity-Based Rate-Quantization Model for H.264/AVC Intraframe Rate ControlabstractIn this letter, we present an adaptive intraframe rate-quantization (R-Q) model for H.264/AVC video coding. The proposed method aims at selecting accurate quantization parameters (QP) for intra-coded frames according to the target bit rate. By taking gradient-based frame complexity measure into consideration, the model parameters can be adaptively updated. Experimental results show that when employing our proposed R-Q model, the intraframe target bits mismatch ratio can be reduced by up to 75% as compared to the traditional Cauchy-density-based model. Hence, this is extremely useful for H.264/AVC rate control applications. Xuan Jing, Lap-Pui Chau, Wan-Chi Siu |
IEEE Signal Process. Lett. | 3 |
| 2008 | New Block-Based Motion Estimation for Sequences with Brightness Variation and Its Application to Static Sprite Generation for Video CompressionabstractIn this brief, a new local motion estimator is proposed which can accurately estimate motion activities under varying strong brightness conditions. The proposed estimator makes use of a new block division technique which manages practically to get rid of the adverse influence caused by brightness changes between frames. We also propose a new static sprite coding system using the proposed local motion estimator. The system is characterized not only with the features of accurate motion estimation under varying brightness conditions, but also possesses the capability of coding the brightness variability of the background scene using a single layered sprite image. Experimental results show that the resulting static sprite coding system improves the PSNR by 6.32 dB as compared with the conventional static sprite coding system when the background scenes of the video sequences involve strong brightness variations in the spatial and time domains. Hoi-Kok Cheung, Wan-Chi Siu, David Dagan Feng, Tom Weidong Cai |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2007 | Efficient Motion Estimation in H.264 Reverse TranscodingabstractIn this paper, we propose a fast reverse motion estimation algorithm with efficient mode decision for reverse transcoding of H.264 bitstream. By analyzing the motion vectors and modes decoded from the forward bitstream, the best mode and motion vector for each backward transcoded macroblock are estimated. A remarkable reduction of computational complexity involved in reverse motion estimation can be achieved by the proposed algorithm with only negligible impact on the rate-distortion performance. Chang-Hong Fu 0002, Yui-Lam Chan, Wan-Chi Siu |
ICIP (5) | 3 |
| 2007 | A Simplified Dual-Bitstream MPEG Video Streaming System with VCR FunctionalitiesabstractNowadays, video playback devices have only limited fast-forward/backward playback and even they cannot provide backward playback. The limitation is due to the motion compensated prediction technique adopted in the MPEG standards. One possible way to support browsing functionalities is to store an additional reverse-encoded bitstream into the server. However, this additional bitstream increases the storage requirement of the video server significantly. In this paper, we exploit the redundancy inherent between the forward and reverse-encoded bitstreams in order to achieve a substantial reduction on the size of the reverse-encoded bitstream. The server accesses and manipulates various macroblocks from the forward and reverse-encoded bistreams to facilitate various browsing operations. Experimental results show that, as compared to the conventional dual-bitstream scheme, the new scheme significantly reduces the storage requirement due to the additional reverse-encoded bitstream. Tak-Piu Ip, Yui-Lam Chan, Chang-Hong Fu 0002, Wan-Chi Siu |
ICIP (6) | 4 |
| 2007 | BiVisu: software tool for bicluster detection and visualizationabstractUNLABELLED: BiVisu is an open-source software tool for detecting and visualizing biclusters embedded in a gene expression matrix. Through the use of appropriate coherence relations, BiVisu can detect constant, constant-row, constant-column, additive-related as well as multiplicative-related biclusters. The biclustering results are then visualized under a 2D setting for easy inspection. In particular, parallel coordinate (PC) plots for each bicluster are displayed, from which objective and subjective cluster quality evaluation can be performed. AVAILABILITY: BiVisu has been developed in Matlab and is available at http://www.eie.polyu.edu.hk/~nflaw/Biclustering/. Kin-On Cheng, Ngai-Fong Law, Wan-Chi Siu, T. H. Lau |
Bioinform. | 3 |
| 2007 | Unified feature analysis in JPEG and JPEG 2000-compressed domains
K. M. Au, Ngai-Fong Law, Wan-Chi Siu |
Pattern Recognit. | 3 |
| 2007 | Multiscale directional filter bank with applications to structured and random texture retrieval
Kin-On Cheng, Ngai-Fong Law, Wan-Chi Siu |
Pattern Recognit. | 3 |
| 2007 | A filter design strategy for binary field wavelet transform using the perpendicular constraint
Ngai-Fong Law, Wan-Chi Siu |
Signal Process. | 2 |
| 2007 | A general contrast function based blind source separation method for convolutively mixed independent sources
Chi-Tat Leung, Wan-Chi Siu |
Signal Process. | 2 |
| 2007 | A Novel Fast and Reduced Redundancy Structure for Multiscale Directional Filter BanksabstractThe multiscale directional filter bank (MDFB) improves the radial frequency resolution of the contourlet transform by introducing an additional decomposition in the high-frequency band. The increase in frequency resolution is particularly useful for texture description because of the quasi-periodic property of textures. However, the MDFB needs an extra set of scale and directional decomposition, which is performed on the full image size. The rise in computational complexity is, thus, prominent. In this paper, we develop an efficient implementation framework for the MDFB. In the new framework, directional decomposition on the first two scales is performed prior to the scale decomposition. This allows sharing of directional decomposition among the two scales and, hence, reduces the computational complexity significantly. Based on this framework, two fast implementations of the MDFB are proposed. The first one can maintain the same flexibility in directional selectivity in the first two scales while the other has the same redundancy ratio as the contourlet transform. Experimental results show that the first and the second schemes can reduce the computational time by 33.3%-34.6% and 37.1%-37.5%, respectively, compared to the original MDFB algorithm. Meanwhile, the texture retrieval performance of the proposed algorithms is more or less the same as the original MDFB approach which outperforms the steerable pyramid and the contourlet transform approaches. Kin-On Cheng, Ngai-Fong Law, Wan-Chi Siu |
IEEE Trans. Image Process. | 3 |
| 2007 | New Architecture for MPEG Video Streaming System With Backward Playback SupportabstractMPEG digital video is becoming ubiquitous for video storage and communications. It is often desirable to perform various video cassette recording (VCR) functions such as backward playback in MPEG videos. However, the predictive processing techniques employed in MPEG severely complicate the backward-play operation. A straightforward implementation of backward playback is to transmit and decode the whole group-of-picture (GOP), store all the decoded frames in the decoder buffer, and play the decoded frames in reverse order. This approach requires a significant buffer in the decoder, which depends on the GOP size, to store the decoded frames. This approach could not be possible in a severely constrained memory requirement. Another alternative is to decode the GOP up to the current frame to be displayed, and then go back to decode the GOP again up to the next frame to be displayed. This approach does not need the huge buffer, but requires much higher bandwidth of the network and complexity of the decoder. In this paper, we propose a macroblock-based algorithm for an efficient implementation of the MPEG video streaming system to provide backward playback over a network with the minimal requirements on the network bandwidth and the decoder complexity. The proposed algorithm classifies macroblocks in the requested frame into backward macroblocks (BMBs) and forward/backward macroblocks (FBMBs). Two macroblock-based techniques are used to manipulate different types of macroblocks in the compressed domain and the server then sends the processed macroblocks to the client machine. For BMBs, a VLC-domain technique is adopted to reduce the number of macroblocks that need to be decoded by the decoder and the number of bits that need to be sent over the network in the backward-play operation. We then propose a newly mixed VLC/DCT-domain technique to handle FBMBs in order to further reduce the computational complexity of the decoder. With these compressed-domain techniques, the proposed architecture only manipulates macroblocks either in the VLC domain or the quantized DCT domain resulting in low server complexity. Experimental results show that, as compared to the conventional system, the new streaming system reduces the required network bandwidth and the decoder complexity significantly. Chang-Hong Fu 0002, Yui-Lam Chan, Tak-Piu Ip, Wan-Chi Siu |
IEEE Trans. Image Process. | 4 |
| 2007 | Extended Analysis of Motion-Compensated Frame Difference for Block-Based Motion Prediction ErrorabstractIn the past, most design and optimization work on hybrid video codecs relied mainly on experimental evidence. A proper theoretical model is always desirable, since this allows us to explain the phenomena of existing codecs and to design better ones. In this paper, we make use of the first-order Markov model to derive an approximated separable autocorrelation model for the block-based motion compensation frame difference (MCFD) signal. A major assumption of our derivation is that the net deformation of pixels is directional, in general, rather than a uniform error distribution in a block. We have also shown that the imperfect block-based motion compensation is significant to the theoretical study and the behavior of motion-compensated codecs. Results of our experimental work show that the derived model can describe the statistical characteristics of the MCFD signals accurately. The model also shows that the imperfectly formulated block-based motion compensation can result in an incorrect MCFD autocorrelation function while, conversely, it can form a better block-based motion compensation scheme. Ko-Cheung Hui, Wan-Chi Siu |
IEEE Trans. Image Process. | 2 |
| 2007 | On Transcoding a B-Frame to a P-Frame in the Compressed DomainabstractOnly a limited number of methods have been proposed to realize heterogeneous transcoding, for example from MPEG-2 to H.263, or from H.264 to H.263. The major difficulties of transcoding a B-picture to a P-picture are that the incoming discrete cosine transform (DCT) coefficients of the B-frame are prediction errors arising from both forward and backward predictions, whilst the prediction errors in the DCT domain arising from the prediction using the previous frame alone are not available. The required new prediction errors need to be re-estimated in the pixel domain. This process involves highly complex computation and introduces re-encoding errors. We propose a new approach to convert a B-picture into a P-picture by making use of some properties of motion compensation in the DCT domain and the direct addition of DCT coefficients. We derive a set of equations and formulate the problem of how to obtain the DCT coefficients. One difficulty is that the last P-frame inside a GOP with an IBBP structure, for example, needs to be transcoded to become the last P-frame in the IPPP structure, and it has to be linked to the previous reconstructed P-frame instead of to the I-frame. We increased the speed of the transcoding process by making use of the motion activity which is expressed in terms of the correlation between pictures. The whole transcoding process is done in the transform domain, hence re-encoding errors are completely avoided. Results from our experimental work show that the proposed video transcoder not only achieves a speed-up of two to six times that of the conventional video transcoder, but it also substantially improves the quality of the video. Wan-Chi Siu, Yui-Lam Chan, Kai-Tat Fung |
IEEE Trans. Multim. | 1 |
| 2006 | Adopting SP/SI-Frames in Dual-Bitstream Video Streaming with VCR SupportabstractDigital video cassette recording (VCR) operations such as fast-forward and fast-reverse playbacks enable quick and user-friendly browsing of video. However, the predictive techniques adopted in current video standards severely complicate these operations. One approach to implement the fast-forward/reverse playback is to store an additional reverse-encoded bitstream into the server. Once a client requests a fast-forward/reverse operation, the server can select an appropriate frame for the client from either the forward or reverse-encoded bitstreams to reduce the network traffic and the decoder complexity. Unfortunately, the forward and reverse-encoded bitstreams are encoded separately. The frame that has previously decoded by the client may not be exactly identical to the reference of the current selected frame and the mismatch problem occurs frequently. In this paper, a novel H.264 dual-bitstream scheme aiming at providing fast-forward/reverse playback based on SP/SI-frames is proposed to eliminate mismatch errors during switching between the forward and reverse-encoded bitstreams. As a result, the proposed scheme enhances the performance of the conventional dual-bitstream scheme. Tak-Piu Ip, Yui-Lam Chan, Wan-Chi Siu |
ICASSP (2) | 3 |
| 2006 | Co-occurrence features of multi-scale directional filter bank for texture characterizationabstractIn this paper, we propose to use co-occurrence features computed from multi-scale directional filter bank (MDFB) for texture characterization. As the filter band coefficients are localized frequency components, features from co-occurrence matrices of filter bands can characterize structures of textures by describing correlation among coefficients. Our experiments show that the co-occurrence features outperform energy features considerably in texture retrieval. In particular, they significantly improve the retrieval rate for textures with weak directionality and periodicity while still maintains a high retrieval rate for regular textures as the energy features Kin-On Cheng, Ngai-Fong Law, Wan-Chi Siu |
ISCAS | 3 |
| 2006 | New results on exhaustive search algorithm for motion estimation using adaptive partial distortion search and successive elimination algorithmabstractMotion estimation is one of the most computational-intensive tasks in video compression. In order to reduce the amount of computation, various fast motion estimation algorithms have been developed. These fast algorithms can be classified into two groups. One is the lossy motion-estimation approach, which may have some degradation of predicted images, and the other is lossless, which means that the quality of the predicted images is exactly the same as those obtained by the conventional full search algorithm. The partial distortion search and successive elimination algorithm are two well-known techniques belonging to the second kind of approach. These two algorithms use different checking criteria to eliminate as much redundant computations as possible. Actually, the working principles of these methods are independent to each others and it is possible to apply them sequentially in order to achieve greater saving in computation. In this paper, we propose a new fast full-search motion estimation algorithm which can exploit fully the advantages of adaptive partial distortion search and successive elimination algorithm. Experimental results show that this proposed algorithm has an average speed-up of 13.31 as compared with the full search algorithm in terms of computational efficiency. This result is much better than the method simply combing both partial distortion search and successive elimination algorithm, which has an average computational speed-up of 10.60. For a practical realization using a PC, the average execution time speed-up for our algorithm is 4.96, which is also the best performance among all algorithms tested. Man-Yau Chiu, Wan-Chi Siu |
ISCAS | 2 |
| 2006 | On re-composition of motion compensated macroblocks for DCT-based video transcoding
Kai-Tat Fung, Wan-Chi Siu |
Signal Process. Image Commun. | 2 |
| 2006 | Marker-based image segmentation relying on disjoint set union
Hai Gao, Weisi Lin, Ping Xue 0001, Wan-Chi Siu |
Signal Process. Image Commun. | 4 |
| 2006 | Efficient reverse-play algorithms for MPEG video with VCR supportabstractReverse playback is the most common video cassette recording (VCR) function in digital video players and it involves playing video frames in reverse order. However, the predictive processing techniques employed in MPEG severely complicate the reverse-play operation. For displaying single frame during reverse playback, all frames from the previous I-frame to the requested frame must be sent by the server and decoded by the client machine. It requires much higher bandwidth of the network and complexity of the decoder. In this paper, we propose a compressed-domain approach for an efficient implementation of the MPEG video streaming system to provide reverse playback over a network with the minimal requirements on the network bandwidth and the decoder complexity. In the proposed video streaming server, it classifies macroblocks in the requested frame into two categories—backward macroblocks (BMBs) and forward macroblock (FMBs). Two novel MB-based techniques are used to manipulate the necessary MBs in the compressed domain and the server then sends the processed MBs to the client machine. For BMBs, we propose a sign inversion technique, which is operated in the variable length coding (VLC) domain, to reduce the number of MBs that need to be decoded by the decoder and the number of bits that need to be sent over the network in the reverse-play operation. The server also identifies the previous related MBs of FMBs and those related maroblocks coded without motion compensation are then processed by a technique of direction addition of discrete cosine transform (DCT) coefficients to further reduce the computational complexity of the client decoder. With the sign inversion and direct addition of DCT coefficients, the proposed architecture only manipulates MBs either on the VLC domain or DCT domain to achieve the server with low complexity. Experimental results show that, as compared to the conventional system, the new streaming system reduces the required network bandwidth and the decoder complexity significantly. Chang-Hong Fu 0002, Yui-Lam Chan, Wan-Chi Siu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2006 | DCT-based video downscaling transcoder using split and merge techniqueabstractFor a conventional downscaling video transcoder, a video server has firstly to decompress the video, perform downscaling operations in the pixel domain, and then recompress it. This is computationally intensive. However, it is difficult to perform video downscaling in the discrete cosine transform (DCT)- domain since the prediction errors of each frame are computed from its immediate past higher resolution frames. Recently, a fast algorithm for DCT domain image downsampling has been proposed to obtain the downsampled version of DCT coefficients with low computational complexity. However, there is a mismatch between the downsampled version of DCT coefficients and the resampled motion vectors. In other words, significant quality degradation is introduced when the derivation of the original motion vectors and the resampled motion vector is large. In this paper, we propose a new architecture to obtain resampled DCT coefficients in the DCT domain by using the split and merge technique. Using our proposed video transcoder architecture, a macroblock is splitted into two regions: dominant region and the boundary region. The dominant region of the macroblock can be transcoded in the DCT domain with low computational complexity and re-encoding error can be avoided. By transcoding the boundary region adaptively, low computational complexity can also be achieved. More importantly, the re-encoding error introduced in the boundary region can be controlled more dynamically. Experimental results show that our proposed video downscaling transcoder can lead to significant computational savings as well as videos with high quality as compared with the conventional approach. The proposed video transcoder is useful for video servers that provide quality service in real-time for heterogeneous clients. Kai-Tat Fung, Wan-Chi Siu |
IEEE Trans. Image Process. | 2 |
| 2005 | Diversity and importance measures for video downscalingabstractIn video downscaling, simply reusing the motion vectors extracted from an incoming video bitstream may not result in good quality pictures. In this paper, we propose an adaptive motion vector re-composition algorithm using two new measures: the diversity and importance measures of motion vectors. Using the importance measure, our proposed scheme manages to differentiate the most representative motion vector as a consideration to recompose a new motion vector. In addition, the diversity measure provides information for the video transcoder controlling the size of the refinement window to achieve a significant reduction of computational complexity. Experimental results show that our proposed adaptive motion vector re-composition scheme provides a high coding efficiency in terms of both quality and complexity. Kai-Tat Fung, Wan-Chi Siu |
ICASSP (2) | 2 |
| 2005 | New Pixel-DCT Domain Coding Technique for Object Based and Frame Based Prediction ErrorabstractThe discrete cosine transform (DCT) is widely used in modern video compression standards, such as the ITU-T H.263 and the ISO MPEG-4, to achieve high compression efficiency. A major merit of the DCT is its capability in high energy compaction for natural images. However, the motion prediction error frame is not a natural image but synthetically generated by the process of motion compensation. This process degrades the energy compaction efficiency of the DCT. We study the spatial distribution of the prediction errors resulting from either the full-search motion estimation or other fast search algorithms in order to improve the coding efficiency of the DCT. Subsequently, a mixed spatial-DCT-based coding scheme is proposed for coding the prediction errors. Our experimental results show that this coding scheme can successfully improve the compression performance of the traditional DCT-based video coder with block based motion compensation for arbitrary shaped video objects and, video sequences which contain moderate to high motion activities. Ko-Cheung Hui, Wan-Chi Siu |
ICASSP (2) | 2 |
| 2005 | Direct image retrieval in JPEG and JPEG2000abstractImages are often compressed using JPEG or JPEG2000. Many retrieval systems operated in either uncompressed or compressed domains have been proposed. However, retrieving in multiple domains typically involves full decompression for feature analysis in spatial domain. Common features in different domains are thus worth of investigation for direct image indexing. By employing a common subband filtering model, outputs from JPEG and JPEG2000 can be compared directly without having a full decompression. Despite of high compression, similar translation and rotation invariant features can be extracted from these two domains. Simulation results reveal that JPEG and JPEG2000 compressed images can be searched from one another irrespective of the compression ratio. Our experimental studies confirm that retrieval in multiple domains is possible without a full decompression. Ka Man Au, Ngai-Fong Law, Wan-Chi Siu |
ICIP (1) | 3 |
| 2005 | Reference picture selection in an already MPEG encoded bitstreamabstractReference picture selection (RPS) is the most common error resilience method for robust transmission over lossy networks. However, RPS has been studied for use in real-time encoding, but has not been examined in transmitting an already encoded MPEG bitstream. In this paper, we propose a compressed-domain approach for the efficient implementation of RPS in the pre-encoded MPEG bitstream with minimum requirement on the server complexity. In the proposed algorithm, a novel macroblock-based algorithm is used to adaptively select the necessary macroblocks, manipulate them in compressed-domain and send the processed macroblocks to the receiver. Experimental results show that, as compared to the original RPS, the new algorithm reduces the required server complexity significantly. Hoi-Kin Cheung, Yui-Lam Chan, Wan-Chi Siu |
ICIP (1) | 3 |
| 2005 | Conversion between DCT coefficients and it coefficients in the compressed domain for H.263 to H.264 video transcodingabstractFor transcoding a video sequence from the H.263 format to the H.264 format, it is beneficial to reuse as much information as possible in the original sequence. However, given the significant differences between the H.263 and the H.264 algorithms, transcoding is much more complex. Motivated by this, a set of operators is derived for converting the DCT coefficients to integer transform (IT) coefficients in the compressed domain. Using these operators, the proposed architecture is able to transcode the DCT coefficients to the integer transform coefficients directly. In other words, no complete decoding and reencoding processes are required. To further speed up the transcoding process, an approximation form of the DCT coefficients is proposed. Experimental results show that the proposed video transcoder can reduce substantially the amount of computation and provide transcoded videos with better quality as compared with those obtained from the conventional cascaded video transcoder. The proposed video transcoder can combine any existing fast searching algorithm for multiple reference frames and variable block sizes estimation to speed up the transcoding process. Kai-Tat Fung, Wan-Chi Siu |
ICIP (3) | 2 |
| 2005 | New adaptive partial distortion search using clustered pixel matching error CharacteristicabstractIn order to reduce the computation load, many conventional fast block-matching algorithms have been developed to reduce the set of possible searching points in the search window. All of these algorithms produce some quality degradation of a predicted image. Alternatively, another kind of fast block-matching algorithms which do not introduce any prediction error as compared with the full-search algorithm is to reduce the number of necessary matching evaluations for every searching point in the search window. The partial distortion search (PDS) is a well-known technique of the second kind of algorithms. In the literature, many researches tried to improve both lossy and lossless block-matching algorithms by making use of an assumption that pixels with larger gradient magnitudes have larger matching errors on average. Based on a simple analysis, it is found that, on average, pixel matching errors with similar magnitudes tend to appear in clusters for natural video sequences. By using this clustering characteristic, we propose an adaptive PDS algorithm which significantly improves the computation efficiency of the original PDS. This approach is much better than other algorithms which make use of the pixel gradients. Furthermore, the proposed algorithm is most suitable for motion estimation of both opaque and boundary macroblocks of an arbitrary-shaped object in MPEG-4 coding. Ko-Cheung Hui, Wan-Chi Siu, Yui-Lam Chan |
IEEE Trans. Image Process. | 2 |
| 2004 | Improved macroblock-based reverse play algorithm for MPEG video streaming
Chang-Hong Fu 0002, Yui-Lam Chan, Wan-Chi Siu |
ICIP | 3 |
| 2004 | A compressed-domain heterogeneous video transcoder
Wan-Chi Siu, Kai-Tat Fung, Yui-Lam Chan |
ICIP | 1 |
| 2004 | Adaptive dual-point Hough transform for object recognition
Chun-Pong Chau, Wan-Chi Siu |
Comput. Vis. Image Underst. | 2 |
| 2004 | Generalized Hough Transform Using Regions with Homogeneous Color
Chun-Pong Chau, Wan-Chi Siu |
Int. J. Comput. Vis. | 2 |
| 2004 | Adaptive partial distortion search for block motion estimation
Yui-Lam Chan, Ko-Cheung Hui, Wan-Chi Siu |
J. Vis. Commun. Image Represent. | 3 |
| 2004 | An efficient recursive shortest spanning tree algorithm using linking propertiesabstractSpeed is a great concern in the recursive shortest spanning tree (RSST) algorithm as its applications are focused on image segmentation and video coding, in which a large amount of data is processed. Several efficient RSST algorithms have been proposed in the literature, but the linking properties are not properly addressed and used in these algorithms and they are intended to produce a truncated RSST. This paper categorizes the linking process into three classes based on link weights. These linking processes are defined as the linking process for link weight equal to zero (LPLW-Z), the linking process for link weight equal to one (LPLW-O), and the linking process for link weight equal to real number (LPLW-R). We study these linking properties and apply them to an efficient RSST algorithm. The proposed efficient RSST algorithm is novel, as it makes use of linking properties, and its resulting shortest spanning tree is truly identical to that produced by the conventional algorithm. Our experimental results show that the percentages of links for the three classes are 17%, 27%, and 58%, respectively. This paper proposes a prediction method for LPLW-O, as a result of which the vertex weight of the next region can be determined by comparing sizes of the merging regions. It is also demonstrated that the proposed LPLW-O with prediction approach is applicable to the multiple-stage merging. Our experimental results show that the proposed algorithm has a substantial improvement over the conventional RSST algorithm. Sai Ho Kwok, Anthony G. Constantinides, Wan-Chi Siu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2004 | Low-complexity and high-quality frame-skipping transcoder for continuous presence multipoint video conferencingabstractThis paper presents a new frame-skipping transcoding approach for video combiners in multipoint video conferencing. Transcoding is regarded as a process of converting a previously compressed video bitstream into a lower bitrate bitstream. A high transcoding ratio may result in an unacceptable picture quality when the incoming video bitstream is transcoded with the full frame rate. Frame skipping is often used as an efficient scheme to allocate more bits to representative frames, so that an acceptable quality for each frame can be maintained. However, the skipped frame must be decompressed completely, and should act as the reference frame to the nonskipped frame for reconstruction. The newly quantized DCT coefficients of prediction error need to be recomputed for the nonskipped frame with reference to the previous nonskipped frame; this can create an undesirable complexity in the real time application as well as introduce re-encoding error. A new frame-skipping transcoding architecture for improved picture quality and reduced complexity is proposed. The proposed architecture is mainly performed on the discrete cosine transform (DCT) domain to achieve a low complexity transcoder. It is observed that the re-encoding error is avoided at the frame-skipping transcoder when the strategy of direct summation of DCT coefficients is employed. By using the proposed frame-skipping transcoder and dynamically allocating more frames to the active participants in video combining, we are able to make more uniform peak signal-to-noise ratio (PSNR) performance of the subsequences and the video qualities of the active subsequences can be improved significantly. Kai-Tat Fung, Yui-Lam Chan, Wan-Chi Siu |
IEEE Trans. Multim. | 3 |
| 2003 | An adaptive partial distortion search for block motion estimationabstractFast search algorithms for block motion estimation reduce the set of possible displacements for locating the motion vector. All algorithms produce some quality degradation of the predicted image. To reduce the computational complexity of the full search algorithm without introducing any loss in the predicted image, we propose a Hilbert-grouped partial distortion search algorithm (HGPDS) by grouping the representative pixels based on pixel activities in the Hilbert scan. By using the grouped information and computing the accumulated partial distortion of the representative pixels before that of other pixels, impossible candidates can be rejected sooner and the remaining computation involved in the matching criterion can be reduced remarkably. In addition, we also suggest a smart search strategy which is an excellent complement of the HGPDS to form an efficient partial distortion search algorithm. The new search strategy rearranges the search order such that the most possible candidates are searched first and this rearrangement will increase the probability of early rejection of impossible motion vectors. Simulation results show that the proposed algorithm has a significant computational speed-up and is the fastest when compared to the conventional partial distortion search algorithms. Yui-Lam Chan, Wan-Chi Siu |
ICASSP (3) | 2 |
| 2003 | Maximal disk based histogram for shape retrievalabstractWe propose a robust and efficient representation scheme for shape retrieval, which is based on the normalized maximal disks used to represent the shape of an object. The maximal disks are extracted by means of a fast skeletonization technique with a pruning algorithm. The logarithm of the radii of the normalized maximal disks is used to construct a histogram to represent the shape. The retrieval performance of this maximal disk based histogram approach is compared to other methods, including moment invariants, Zernike moments, and curvature scale-space. Experimental results show that our proposed representation scheme outperforms the other methods under affine transformation and different noise levels. Wai-Pak Choi, Kin-Man Lam 0001, Wan-Chi Siu |
ICASSP (3) | 3 |
| 2003 | A fast and efficient computational structure for the 2D over-complete wavelet transformabstractWe studied the computational complexity of the over-complete wavelet representation for the commonly used Spline wavelet family with an arbitrary order. By deriving a general expression for the complexity, it is shown that the inverse transform is significantly more costly in computation than the forward transform. In order to reduce the computational complexity, a new spatial implementation is proposed. This new implementation exploits the redundancy between the lowpass and the bandpass outputs that is inherent to the over-complete wavelet scheme. It is shown that the new implementation can greatly simplify the computations, give an efficient inverse structure and allow the use of an arbitrary boundary extension method without affecting the ease of the inverse transform. Ngai-Fong Law, Wan-Chi Siu |
ICASSP (3) | 2 |
| 2003 | New DCT-domain transcoding using split and merge techniqueabstractFor the conventional downscaling video transcoder, the video server would be to first decompress the video, perform the downscaling operation in the pixel domain, and then recompress it. This is computational intensive. However, it is difficult to perform video downscaling in the DCT-domain since the prediction errors of each frame are computed from its immediate past higher resolution frames. Besides, the motion vector need to resample due to the size of the video is changed. Due to the mismatch of the resampled motion vector with the incoming DCT coefficients, the video transcoder need to recalculate the new DCT coefficient with lower resolution in pixel domain; this can create undesirable complexity as well as introduce re-encoding error. In this paper, we propose a new architecture to obtain the new DCT coefficients and the new motion vector by reuse the incoming motion vector and DCT coefficients. Since our proposed transcoder is mainly performing in DCT domain, low computational complexity can be achieved as well as the re-encoding can be reduced. Experimental result show that our proposed video downscaling transcoder can lead to significant computational savings as well as provide a high video quality compared to the conventional approach. Kai-Tat Fung, Wan-Chi Siu |
ICIP (1) | 2 |
| 2003 | Efficient Learning in Adaptive Processing of Data Structures
Siu-Yeung Cho, Zheru Chi, Zhiyong Wang 0001, Wan-Chi Siu |
Neural Process. Lett. | 4 |
| 2003 | Hierarchical content classification and script determination for automatic document image processing
Zheru Chi, Qing Wang 0006, Wan-Chi Siu |
Pattern Recognit. | 3 |
| 2003 | Extraction of the Euclidean skeleton based on a connectivity criterion
Wai-Pak Choi, Kin-Man Lam 0001, Wan-Chi Siu |
Pattern Recognit. | 3 |
| 2003 | Spatially eigen-weighted Hausdorff distances for human face recognition
Kwan-Ho Lin, Kin-Man Lam 0001, Wan-Chi Siu |
Pattern Recognit. | 3 |
| 2003 | Human face recognition based on spatially weighted Hausdorff distance
Baofeng Guo, Kin-Man Lam 0001, Kwan-Ho Lin, Wan-Chi Siu |
Pattern Recognit. Lett. | 4 |
| 2003 | Efficient multiplier structure for realization of the discrete cosine transform
Lap-Pui Chau, Wan-Chi Siu |
Signal Process. Image Commun. | 2 |
| 2003 | Fast motion estimation of arbitrarily shaped video objects in MPEG-4
Ko-Cheung Hui, Wan-Chi Siu, Yui-Lam Chan |
Signal Process. Image Commun. | 2 |
| 2003 | A robust scheme for live detection of human faces in color images
Kwok-Wai Wong, Kin-Man Lam 0001, Wan-Chi Siu |
Signal Process. Image Commun. | 3 |
| 2003 | A fast fractal image coding based on kick-out and zero contrast conditionsabstractA fast algorithm for fractal image coding based on a single kick-out condition and the zero contrast prediction is proposed in this paper. The single kick-out condition can avoid a large number of range-domain block matches when finding the best matched domain block. An efficient method for zero contrast prediction is also proposed, which can determine whether the contrast factor for a domain block is zero or not, and compute the corresponding difference between the range block and the transformed domain block efficiently and exactly. The proposed algorithm can achieve the same reconstructed image quality as the exhaustive search, and can greatly reduce the required computation or runtime. In addition, this algorithm does not need any pre-processing step or additional memory for its implementation, and can combine with other fast fractal algorithms to further improve the speed. Experimental results show that the runtime is reduced by about 50% of that of the exhaustive search method. When combined with the DCT Inner Product algorithm, the required runtime for the algorithm can be further reduced by about 50%. The proposed algorithm was also compared to two other fast fractal algorithms. Experimental results also show that our algorithm achieves a better efficiency and requires a much smaller amount of memory for implementation. Cheung-Ming Lai, Kin-Man Lam 0001, Wan-Chi Siu |
IEEE Trans. Image Process. | 3 |
| 2003 | An improved algorithm for learning long-term dependency problems in adaptive processing of data structuresabstractMany researchers have explored the use of neural-network representations for the adaptive processing of data structures. One of the most popular learning formulations of data structure processing is backpropagation through structure (BPTS). The BPTS algorithm has been successful applied to a number of learning tasks that involve structural patterns such as logo and natural scene classification. The main limitations of the BPTS algorithm are attributed to slow convergence speed and the long-term dependency problem for the adaptive processing of data structures. In this paper, an improved algorithm is proposed to solve these problems. The idea of this algorithm is to optimize the free learning parameters of the neural network in the node representation by using least-squares-based optimization methods in a layer-by-layer fashion. Not only can fast convergence speed be achieved, but the long-term dependency problem can also be overcome since the vanishing of gradient information is avoided when our approach is applied to very deep tree structures. Siu-Yeung Cho, Zheru Chi, Wan-Chi Siu, Ah Chung Tsoi |
IEEE Trans. Neural Networks | 3 |
| 2002 | A new approach using modified Hausdorff distances with eigenface for human face recognitionabstractHausdorff distance is an efficient measure of the similarity of two point sets. In this paper, we propose two new spatially weighted Hausdorff distance measures for human face recognition, namely, spatially eigen-weighted Hausdorff distance (SEWHD) and spatially eigen-weighted 'doubly' Hausdorff distance (SEW2HD). These new Hausdorff distances incorporate the information about the location of important facial features so that distances at those regions will be emphasized. The weighting function used in the Hausdorff distance measure is based on an eigenface, which has a large value at locations of important facial features and can reflect the face structure more effectively. Experimental results based on a combination of the ORL, MIT, and Yale face databases show that SEW2HD can achieve recognition rates of 83%, 90% and 92% for the first one, the first three and the first five likely matched faces, respectively, while the corresponding recognition rates of SEWHD are 80%, 83% and 88%, respectively. Kwan-Ho Lin, Kin-Man Lam 0001, Wan-Chi Siu |
ICARCV | 3 |
| 2002 | An efficient algorithm for the extraction of a Euclidean skeletonabstractThe skeleton is essential for general shape representation but the discrete representation of an image presents a lot of problems that may influence the process of skeleton extraction. Some of the methods are memory-intensive and computationally intensive, and require a complex data structure. In this paper, we propose a fast, efficient and accurate skeletonization method for the extraction of a well-connected Euclidean skeleton based on a signed sequential Euclidean distance map. A connectivity criterion that can be used to determine whether a given pixel inside an object is a skeleton point is proposed. The criterion is based on a set of points along the object boundary, which are the nearest contour points to the pixel under consideration and its 8 neighbors. The extracted skeleton is of single-pixel width without requiring a linking algorithm or iteration process. Experiments show that the runtime of our algorithm is faster than. those of using the distance transformation and is linearly proportional to the number of pixels of an image. Wai-Pak Choi, Kin-Man Lam 0001, Wan-Chi Siu |
ICASSP | 3 |
| 2002 | Block motion estimation using adaptive partial distortion searchabstractThe conventional search algorithms for block motion estimation reduce the set of possible displacements for locating the motion vector. All of these algorithms produce some quality degradation of the predicted image. To reduce the computational complexity of the full search algorithm without introducing any loss in the predicted image, we propose an adaptive partial distortion search algorithm (APDS) by selecting the most representative pixels with high activities, such as edges and texture which contribute most to the matching criterion. The APDS algorithm groups the representative pixels based on the pixel activities in the Hilbert scan. By using the grouped information and computing the accumulated partial distortion of the representative pixels before that of the other pixels, impossible candidates can be rejected sooner and the remaining computation involved in the matching criterion can be reduced remarkably. Simulation results show that the proposed APDS algorithm has a significant computational speed-up and is the fastest when compared to the conventional partial distortion search algorithms. Yui-Lam Chan, Wan-Chi Siu, Ko-Cheung Hui |
ICME (1) | 2 |
| 2002 | A dynamic video combiner for multipoint video conferencing using wavelet transformabstractA new architecture of video combiner for multipoint video conferencing is proposed. The proposed video combiner is wavelet-based which extracts the motion activities information from the video bitstreams produced by a wavelet-based video coder. Using the progressive properties of wavelet transform, the encoded bitstream become scalable. Hence, the video quality of inactive sub-sequences can be easily adjusted in the video combiner by discarding the fine detail information bitstreams. In other words, more bits can be reallocated to the active sub-sequences to achieve a good visual quality with smooth motion. In addition, the video coder is region-based so that different wavelet kernels can be used for the foreground and the background. This setting can on one hand reduce the computational complexity significantly. On the other hand, by considering the unequal importance of various regions, a high video quality in foreground can always be guaranteed and an acceptable quality in background can be maintained even under low bitrate environments. Since the video combiner only needs to rearrange the video quality level according to their motion activities, no re-encoding process is required. Therefore, a significant computational complexity saving can be achieved as compared to the conventional video combiner using a transcoding approach. The new video combiner is then used to realize a multipoint video conferencing and some results are presented to show the improvement in performance due to our proposed architecture. Kai-Tat Fung, Wan-Chi Siu, Ngai-Fong Law |
ICME (2) | 2 |
| 2002 | Fast over-complete wavelet implementation for spline familyabstractWe have studied the computational complexity of the over-complete wavelet representation for the commonly used spline wavelet family with an arbitrary order. By deriving a general expression for the complexity, it is shown that the inverse transform is nearly three times more costly in computation than the forward transform. In order to reduce the computational complexity, a new spatial implementation is proposed. This new implementation is based on the exploitation of redundancy between the lowpass and the bandpass outputs that is inherent to the over-complete wavelet scheme. It is shown that the new implementation can greatly simplify computation and give an efficient inverse structure. Ngai-Fong Law, Wan-Chi Siu |
ICME (1) | 2 |
| 2002 | Robust Hausdorff distance for shape matching
Wai-Pak Choi, Kin-Man Lam 0001, Wan-Chi Siu |
VCIP | 3 |
| 2002 | An independent component analysis based weight initialization method for multilayer perceptrons
Yat-Fung Yam, Chi-Tat Leung, Peter Kwong-Shun Tam, Wan-Chi Siu |
Neurocomputing | 4 |
| 2002 | Extraction and Optimization of B-Spline PBD Templates for Recognition of Connected Handwritten Digit StringsabstractThe recognition of connected handwritten digit strings is a challenging task due mainly to two problems: poor character segmentation and unreliable isolated character recognition. The authors first present a rational B-spline representation of digit templates based on Pixel-to-Boundary Distance (PBD) maps. We then present a neural network approach to extract B-spline PBD templates and an evolutionary algorithm to optimize these templates. In total, 1000 templates (100 templates for each of 10 classes) were extracted from and optimized on 10426 training samples from the NIST Special Database 3. By using these templates, a nearest neighbor classifier can successfully reject 90.7 percent of nondigit patterns while achieving a 96.4 percent correct classification of isolated test digits. When our classifier is applied to the recognition of 4958 connected handwritten digit strings (4555 2-digit, 355 3-digit, and 48 4-digit strings) from the NIST Special Database 3 with a dynamic programming approach, it has a correct classification rate of 82.4 percent with a rejection rate of as low as 0.85 percent. Our classifier compares favorably in terms of correct classification rate and robustness with other classifiers that are tested. Zhongkang Lu, Zheru Chi, Wan-Chi Siu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2002 | New architecture for dynamic frame-skipping transcoderabstractTranscoding is a key technique for reducing the bit rate of a previously compressed video signal. A high transcoding ratio may result in an unacceptable picture quality when the full frame rate of the incoming video bitstream is used. Frame skipping is often used as an efficient scheme to allocate more bits to the representative frames, so that an acceptable quality for each frame can be maintained. However, the skipped frame must be decompressed completely, which might act as a reference frame to nonskipped frames for reconstruction. The newly quantized discrete cosine transform (DCT) coefficients of the prediction errors need to be re-computed for the nonskipped frame with reference to the previous nonskipped frame; this can create undesirable complexity as well as introduce re-encoding errors. In this paper, we propose new algorithms and a novel architecture for frame-rate reduction to improve picture quality and to reduce complexity. The proposed architecture is mainly performed on the DCT domain to achieve a transcoder with low complexity. With the direct addition of DCT coefficients and an error compensation feedback loop, re-encoding errors are reduced significantly. Furthermore, we propose a frame-rate control scheme which can dynamically adjust the number of skipped frames according to the incoming motion vectors and re-encoding errors due to transcoding such that the decoded sequence can have a smooth motion as well as better transcoded pictures. Experimental results show that, as compared to the conventional transcoder, the new architecture for frame-skipping transcoder is more robust, produces fewer requantization errors, and has reduced computational complexity. Kai-Tat Fung, Yui-Lam Chan, Wan-Chi Siu |
IEEE Trans. Image Process. | 3 |
| 2001 | Issues of filter design for binary wavelet transformabstractWavelet decomposition has recently been generalized to the binary field in which the arithmetic is performed wholly in GF(2). In order to maintain an invertible binary wavelet transform with desirable multiresolution properties, the bandwidth, the perfect reconstruction and the vanishing moment constraints are placed on the binary filters. While they guarantee an invertible transform, the transform becomes non-orthogonal and non-biorthogonal in which the inverse filters could be signal-length dependent. We propose to apply the perpendicular constraint on the binary filters to make them length-independent. A filter design strategy is outlined in which a filter design for a length of eight is given. We also propose an efficient implementation structure for the binary filters that saves memory space and reduces the computational complexity. Ngai-Fong Law, Wan-Chi Siu |
ICASSP | 2 |
| 2001 | Dynamic frame skipping for high-performance transcodingabstractTranscoding is a process of converting a previously compressed video bitstream into a lower bit-rate bitstream. When some incoming frames are dropped for the frame-rate conversion in transcoding, the newly quantized DCT coefficients of prediction error need to be re-computed, which can create an undesirable complexity as well as introduce re-encoding error. We propose a new architecture for a frame-skipping transcoder to improve picture quality and to reduce complexity. It is observed that re-encoding error is reduced significantly when the strategy of direct summation of DCT coefficients and the error compensation feedback loop are employed. Furthermore, we propose a frame-rate control scheme which can dynamically adjust the number of skipped frames according to the incoming motion vectors and the re-encoding error due to transcoding such that the decoded sequence can have smooth motion as well as better transcoded pictures. Experimental results show that, as compared to the conventional transcoder, the new frame-skipping transcoder is more robust, produces smaller requantization errors, and has simple computational complexity. Kai-Tat Fung, Yui-Lam Chan, Wan-Chi Siu |
ICIP (1) | 3 |
| 2001 | Priority search technique for MPEG-4 motion estimation of arbitrarily shaped video objectabstractOne of the main differences between MPEG-4 and previously standardized video coding schemes is the support of arbitrarily shaped video objects, for which most of the existing fast motion estimation algorithms are not suitable. The conventional fast motion estimation algorithm works well for opaque macroblocks, but not in the case of a boundary macroblock which contains a large number of local minima on its error surface. We propose a fast search algorithm which incorporates the binary alpha-plane to predict accurately the motion vectors of boundary macroblocks. Besides, these accurate motion vectors can be used to develop a novel priority search algorithm which is an efficient search strategy for the remaining opaque macroblocks. Experimental results show that, compared to the conventional methods, our approach requires a low computational complexity and provides a significant improvement in terms of accuracy in motion-compensated video object planes. Ko-Cheung Hui, Yui-Lam Chan, Wan-Chi Siu |
ICIP (3) | 3 |
| 2001 | Fast algorithm for binary field wavelet transform for image processingabstractThe lifting scheme for the real field wavelet transform has provided a new insight into its practical implementation. This paper shows that a similar scheme can be developed for the binary field wavelet transform. In particular, by using the Euclidean algorithm, the binary filters can be decomposed into a finite sequence of simple lifting steps over the binary field. This provides an alternative method for the implementations of the binary field wavelet transform for image processing applications. It is found that the new implementation can reduce the number of arithmetic operations involved in the transform and allow an efficient in-place implementation structure. Ngai-Fong Law, Alan Wee-Chung Liew, Wan-Chi Siu |
ICIP (2) | 3 |
| 2001 | Locating the human eye using fractal dimensionsabstractA new method for locating eye pairs based on valley field detection and measurement of fractal dimensions is proposed. Fractal dimension is an efficient representation of the texture of facial features. Possible eye candidates to an image with a complex background are identified by valley field detection. The eye candidates are then grouped to form eye pairs if their local properties for eyes are satisfied. Two eyes are matched if they have similar roughness and orientation as represented by fractal dimensions. We propose a modified approach to estimate the fractal dimensions that are less sensitive to lighting conditions and provide information about the orientation of an image under consideration. Possible eye pairs are further verified by comparing the fractal dimensions of the eye-pair window and the corresponding face region with the respective means of the fractal dimensions of the eye-pair windows and the face regions. The means of the fractal dimensions are obtained based on a number of facial images in a database. Experiments have shown that this approach is fast and reliable. This shows that the texture of the eyes can be represented very well by fractal surfaces. Kwan-Ho Lin, Kin-Man Lam 0001, Wan-Chi Siu |
ICIP (3) | 3 |
| 2001 | An adaptive active contour model for highly irregular boundaries
Wai-Pak Choi, Kin-Man Lam 0001, Wan-Chi Siu |
Pattern Recognit. | 3 |
| 2001 | A novel approach for human face detection from color images under complex background
Kwok-Wai Wong, Kin-Man Lam 0001, Wan-Chi Siu |
Pattern Recognit. | 3 |
| 2001 | Document image template matching based on component block list
Hanchuan Peng, Fuhui Long, Zheru Chi, Wan-Chi Siu |
Pattern Recognit. Lett. | 4 |
| 2001 | Successive structural analysis using wavelet transform for blocking artifacts suppression
Ngai-Fong Law, Wan-Chi Siu |
Signal Process. | 2 |
| 2001 | A new fast algorithm for computing prime-Length DCT through cyclic convolutions
Rui-Xiang Yin, Wan-Chi Siu |
Signal Process. | 2 |
| 2001 | Improved techniques for automatic image segmentationabstractMathematical morphology is very attractive for automatic image segmentation because it efficiently deals with geometrical descriptions such as size, area, shape, or connectivity that can be considered as segmentation-oriented features. This paper presents an image-segmentation system based on some well-known strategies. The segmentation process is divided into three basic steps, namely: simplification, marker extraction, and boundary decision. Simplification, which makes use of area morphology, removes unnecessary information from the image to make it easy to segment. Marker extraction identifies the presence of homogeneous regions. A new marker extraction design is proposed in this paper. It is based on both luminance and color information. The goal of boundary decision is to precisely locate the boundary of regions detected by the marker extraction. This decision is based on a region-growing algorithm which is a modified watershed algorithm. A new color distance is also defined for this algorithm. In both marker extraction and boundary decision, color measurement is used to replace grayscale measurement and L*a*b* color space is used to replace the more straightforward spaces such as the RGB color space and YUV color space. Hai Gao, Wan-Chi Siu, Chao-Huan Hou |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2001 | A robust model generation technique for model-based video codingabstractIn conventional model-based coding schemes, predefined static models are generally used. These models cannot adapt to new situations, and hence, they have to be very specific and cannot be generated from a single generic model even though they are very similar. We present a model-generation technique that can gradually build a model and dynamically modify it according to new video frames scanned. The proposed technique is robust to the object's orientation in the view and can be efficiently implemented with a parallel processing technique. As a result, the proposed technique is more attractive to the practical use of model-based coding techniques in real applications. Manson Siu, Yuk-Hee Chan, Wan-Chi Siu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2001 | An efficient low bit-rate video-coding algorithm focusing on moving regionsabstractBlock-based motion estimation and compensation are the most popular techniques for video coding. However, as the shape and the structure of an object in a picture are arbitrary, the performance of such conventional block-based methods may not be satisfactory. In this paper, a very low bit-rate video coding algorithm that focuses on moving regions is proposed. The aim is to improve the coding performance, which gives better subjective and objective quality than that of the conventional coding methods at the same bit rate. Eight patterns are pre-defined to approximate the moving regions in a macroblock. The patterns are then used for motion estimation and compensation to reduce the prediction errors. Furthermore, in order to increase the compression performance, the residual errors of a macroblock are rearranged into a block with no significant increase of high-order DCT coefficients. As a result, both the prediction efficiency and the compression efficiency are improved. Kwok-Wai Wong, Kin-Man Lam 0001, Wan-Chi Siu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2001 | An efficient search strategy for block motion estimation using image featuresabstractBlock motion estimation using the exhaustive full search is computationally intensive. Fast search algorithms offered in the past tend to reduce the amount of computation by limiting the number of locations to be searched. Nearly all of these algorithms rely on this assumption: the mean absolute difference (MAD) distortion function increases monotonically as the search location moves away from the global minimum. Essentially, this assumption requires that the MAD error surface be unimodal over the search window. Unfortunately, this is usually not true in real-world video signals. However, we can reasonably assume that it is monotonic in a small neighborhood around the global minimum. Consequently, one simple strategy, but perhaps the most efficient and reliable, is to place the checking point as close as possible to the global minimum. In this paper, some image features are suggested to locate the initial search points. Such a guided scheme is based on the location of certain feature points. After applying a feature detecting process to each frame to extract a set of feature points as matching primitives, we have extensively studied the statistical behavior of these matching primitives, and found that they are highly correlated with the MAD error surface of real-world motion vectors. These correlation characteristics are extremely useful for fast search algorithms. The results are robust and the implementation could be very efficient. A beautiful point of our approach is that the proposed search algorithm can work together with other block motion estimation algorithms. Results of our experiment on applying the present approach to the block-based gradient descent search algorithm (BBGDS), the diamond search algorithm (DS) and our previously proposed edge-oriented block motion estimation show that the proposed search strategy is able to strengthen these searching algorithms. As compared to the conventional approach, the new algorithm, through the extraction of image features, is more robust, produces smaller motion compensation errors, and has a simple computational complexity. Yui-Lam Chan, Wan-Chi Siu |
IEEE Trans. Image Process. | 2 |
| 2000 | Document Image Matching Based on Component BlocksabstractDocument image matching is the key technique for document registration and retrieval. In this paper, a new matching algorithm based on document component block list and component block tree is proposed. Our method can effectively make use of the local information of each page block and the global information of page layout, while it is also robust to image distortion, filled-in text, and noises. This algorithm is then refined and applied to automatic data extraction of column forms. A demonstrating software package has been developed. Hanchuan Peng, Fuhui Long, Wan-Chi Siu, Zheru Chi, David Dagan Feng |
ICIP | 3 |
| 2000 | Multimodal Interface Techniques in Content-Based Multimedia Retrieval
Jinchang Ren, Rongchun Zhao, David Dagan Feng, Wan-Chi Siu |
ICMI | 4 |
| 2000 | An Efficient and Accurate Algorithm for Extracting a SkeletonabstractIn this paper, a non-iterative method is proposed, which is fast, efficient, and more importantly, is robust to boundary noise and rotation. Unnecessary branches and hairs can be reduced by the adjustment of the residual distance and the skeleton can be represented in a hierarchical manner. The reconstruction error can also be estimated using the residual distance. A new definition of a skeleton and the criteria of being a skeleton point are introduced. The effect of boundary noise and curved boundary are investigated and compared to other skeletonization algorithms. Finally, the reconstruction of an object using its skeleton and the associated radii of the maximal disks, is illustrated. Wai-Pak Choi, Kin-Man Lam 0001, Wan-Chi Siu |
ICPR | 3 |
| 2000 | Recursive algorithm for the realization of the discrete cosine transformabstractRecursive algorithms have been found very effective for realization using software and VLSI techniques. Recently, some recursive algorithms have been proposed for the realization of the discrete cosine transform (DCT). In this paper, an efficient recursive algorithm for the computation of the DCT is proposed. By using some appropriate iterative techniques, the formulation of the arbitrary length forward DCT (FDCT) and inverse DCT (IDCT) can be effectively implemented by recursive equations, and the hardware complexity is further reduced as compared to approaches in the literature. The proposed algorithm is suitable for both software and VLSI implementations. Lap-Pui Chau, Wan-Chi Siu |
ISCAS | 2 |
| 2000 | A Semi-Parametric Hybrid Neural Model for Nonlinear Blind Signal SeparationabstractNonlinear blind signal separation is an important but rather difficult problem. Any general nonlinear independent component analysis algorithm for such a problem should specify which solution it tries to find. Several recent neural networks for separating the post nonlinear blind mixtures are limited to the diagonal nonlinearity, where there is no cross-channel nonlinearity. In this paper, a new semi-parametric hybrid neural network is proposed to separate the post nonlinearly mixed blind signals where cross-channel disturbance is included. This hybrid network consists of two cascading modules, which are a neural nonlinear module for approximating the post nonlinearity and a linear module for separating the predicted linear blind mixtures. The nonlinear module is a semi-parametric expansion made up of two sub-networks, one of which is a linear model and the other of which is a three-layer perceptron. These two sub-networks together produce a "weak" nonlinear operator and can approach relatively strong nonlinearity by tuning parameters. A batch learning algorithm based on the entropy maximization and the gradient descent method is deduced. This model is successfully applied to a blind signal separation problem with two sources. Our simulation results indicate that this hybrid model can effectively approach the cross-channel post nonlinearity and achieve a good visual quality as well as a high signal-to-noise ratio in some cases. Hanchuan Peng, Zheru Chi, Wan-Chi Siu |
Int. J. Neural Syst. | 3 |
| 2000 | Efficient recursive algorithm for the inverse discrete cosine transformabstractRecursive algorithms have been found very effective for realization using software and very large scale integrated circuit (VLSI) techniques. Previously, some recursive algorithms have been proposed for the realization of the inverse discrete cosine transform (IDCT). In this paper, an efficient recursive algorithm for the IDCT with arbitrary length is presented. By using some appropriate iterative techniques, the formulation of the IDCT can be implemented effectively using recursive equations, and the hardware complexity is further reduced as compared with the approaches in the literature. Lap-Pui Chau, Wan-Chi Siu |
IEEE Signal Process. Lett. | 2 |
| 2000 | Speech enhancement using the constrained-optimization techniqueabstractWe address a problem of speech enhancement: recovering a speech source from a mixture of its delayed versions and additive noise. By using the constrained-optimization technique, the second order statistics based algorithm is developed. The new proposed algorithm requires no strong limitations to the speech signal and the noise. Simulation results show that our algorithm achieves a better performance as compared to other algorithms. Wei Li 0012, Wan-Chi Siu |
IEEE Signal Process. Lett. | 2 |
| 2000 | A study of the Lamarckian evolution of recurrent neural networksabstractTraining neural networks by evolutionary search can require a long computation time. In certain situations, using Lamarckian evolution, local search and evolutionary search can complement each other to yield a better training algorithm. This paper demonstrates the potential of this evolutionary-learning synergy by applying it to train recurrent neural networks in an attempt to resolve a long-term dependency problem and the inverted pendulum problem. This work also aims at investigating the interaction between local search and evolutionary search when they are combined; it is found that the combinations are particularly efficient when the local search is simple. In the case where no teacher signal is available for the local search to learn the desired task directly, the paper proposes a related local task for the local search to learn, and finds that this approach is able to reduce the training time considerably. Kim W. C. Ku, Man-Wai Mak, Wan-Chi Siu |
IEEE Trans. Evol. Comput. | 3 |
| 1999 | Reliable search strategy for block motion estimation by measuring the error surfaceabstractThe conventional search algorithms for block matching motion estimation reduce the set of possible displacements for locating the motion vector. Nearly all of these algorithms rely on the assumption: the distortion function increases monotonically as the search location moves away from the global minimum. Obviously, this assumption essentially requires that the error surface be unimodal over the search window. Unfortunately, this is usually not true in real-world video signals. We formulate a criterion to check the confidence of unimodal error surface over the search window. The proposed confidence measure of error surface, CMES, would be a good measure for identifying whether the searching should continue or not. It is found that this proposed measure is able to strengthen the conventional fast search algorithms for block matching motion estimation. Experimental results show that, as compared to the conventional approach, the new algorithm through the CMES is more robust, produces smaller motion compensation errors, and requires simple computational complexity. Yui-Lam Chan, Wan-Chi Siu |
ICASSP | 2 |
| 1999 | A Feature-Assisted Search Strategy for Block Motion EstimationabstractBlock motion estimation using the exhaustive full search is computationally intensive. Previous fast search algorithms tend to reduce the computation by limiting the number of locations to be searched. Nearly all of these algorithms rely on the assumption: the mean absolute distortion (MAD) function increases monotonically as the search location moves away from the global minimum. Unfortunately, this is usually not true in real-world video signals. However, we can reasonably assume that it is monotonic in a small neighbourhood around the global minimum. Consequently, one simple, but perhaps the most efficient and reliable strategy, is to put the checking point as close as possible to the global minimum. In this paper, some image features are suggested to locate the initial search points. Such a guided scheme is based on the location of some feature points. After a feature detecting process was applied to each frame to extract a set of feature points as matching primitives, we studied extensively the statistical behaviour of these matching primitives and found that they are highly correlated with the MAD error surface of real-world motion vectors. These correlation characteristics are extremely useful for fast search algorithms. The results are robust and the implementation could be very efficient. Yui-Lam Chan, Wan-Chi Siu |
ICIP (2) | 2 |
| 1999 | Generalized Dual-Point Hough Transform for Object RecognitionabstractIn this paper, a modified version of dual-point generalized Hough transform has been proposed. Inspired by the result of an analysis of the shapes of objects, we are able to improve the efficiency of the dual-point generalized Hough transform. A characteristic angle, which is governed by the points selected in the recognition process, will be determined for the generation of the R-table. This angle is based on making the number of index entries in the R-table as large as possible, the number of entries per indexes as small as possible, and the number of access per point as small as possible. Experimental results show that it can improve the transform by increasing the detection accuracy and speed after the modification is made into the transform. Chun-Pong Chau, Wan-Chi Siu |
ICIP (1) | 2 |
| 1999 | Progressive Image Coding Based on Visually Important FeaturesabstractA novel scalable coding scheme is proposed in which wavelet transform is used to induce a multiple resolution framework for progressive coding. The proposed algorithm corresponds well to human visual characteristics. In particular, our proposed algorithm would display visually important information, such as edges first with visually not-so-important information displayed progressively. A common problem with edge-based coding is the high bit rare. We propose a way to make a compromise between the high bit rate and the skeleton coding idea. As simulation results show, the shape of an image can be recognized with only one or two decoded skeleton frames, i.e., bit rates under 0.05. This is important for image browsing application as images can be recognized with a low bit rate. Also, the accurate coding of edges facilitates the motion estimation in video coding. Therefore, the proposed algorithm has great potential in both image and video coding. Ngai-Fong Law, Wan-Chi Siu |
ICIP (2) | 2 |
| 1999 | Adaptive Shrinkage Algorithm for Ringing Suppression with Smoothness ConstraintabstractA non-iterative wavelet-based algorithm was proposed to reduce the ringing artifacts associated with a lossy compressed image. The proposed algorithm is based on the fact that an increase in the magnitude of the quantized wavelet coefficients leads to a decrease in visual smoothness. Thus a shrinkage algorithm is applied to maintain visual smoothness. The proposed algorithm is, however, adaptive in nature and the amount of shrinkage depends on edge strength, region activity, and compression ratio. Experimental results have confirmed that the adaptive algorithm could suppress the ringing artifacts and improve visual smoothness, especially around edges where ringing is severe. Ngai-Fong Law, Wan-Chi Siu |
ICIP (3) | 2 |
| 1999 | Genetic Algorithm for the Extraction of Nonanalytic Objects from Multiple Dimensional Parameter Space
Pui-Kin Ser, Clifford S. T. Choy, Wan-Chi Siu |
Comput. Vis. Image Underst. | 3 |
| 1999 | A background-thinning-based approach for separating and recognizing connected handwritten digit strings
Zhongkang Lu, Zheru Chi, Wan-Chi Siu |
Pattern Recognit. | 3 |
| 1999 | Reliable block motion estimation through the confidence measure of error surface
Yui-Lam Chan, Wan-Chi Siu |
Signal Process. | 2 |
| 1999 | Recovery of single source signal from noisy and reverberant environments using second-order statistics
Wei Li 0012, J. C. H. Poon, Wan-Chi Siu |
Signal Process. | 3 |
| 1999 | Adding learning to cellular genetic algorithms for training recurrent neural networksabstractThis paper proposes a hybrid optimization algorithm which combines the efforts of local search (individual learning) and cellular genetic algorithms (GA's) for training recurrent neural networks (RNN's). Each weight of an RNN is encoded as a floating point number, and a concatenation of the numbers forms a chromosome. Reproduction takes place locally in a square grid with each grid point representing a chromosome. Two approaches, Lamarckian and Baldwinian mechanisms, for combining cellular GA's and learning have been compared. Different hill-climbing algorithms are incorporated into the cellular GA's as learning methods. These include the real-time recurrent learning (RTRL) and its simplified versions, and the delta rule. The RTRL algorithm has been successively simplified by freezing some of the weights to form simplified versions. The delta rule, which is the simplest form of learning, has been implemented by considering the RNN's as feedforward networks during learning. The hybrid algorithms are used to train the RNN's to solve a long-term dependency problem. The results show that Baldwinian learning is inefficient in assisting the cellular GA. It is conjectured that the more difficult it is for genetic operations to produce the genotypic changes that match the phenotypic changes due to learning, the poorer is the convergence of Baldwinian learning. Most of the combinations using the Lamarckian mechanism show an improvement in reducing the number of generations required for an optimum network; however, only a few can reduce the actual time taken. Embedding the delta rule in the cellular GA's has been found to be the fastest method. It is also concluded that learning should not be too extensive if the hybrid algorithm is to be benefit from learning. Kim W. C. Ku, Man-Wai Mak, Wan-Chi Siu |
IEEE Trans. Neural Networks | 3 |
| 1998 | On Block Motion Estimation Using a Novel Search Strategy for an Improved Adaptive Pixel Decimation
Yui-Lam Chan, Wan-Chi Siu |
J. Vis. Commun. Image Represent. | 2 |
| 1998 | On the efficient computation of 2-d image moments using the discrete radon transform
Tak-Wai Shen, Daniel Pak-Kong Lun, Wan-Chi Siu |
Pattern Recognit. | 3 |
| 1998 | A new approach for real-time reduction of blocking effect
Sung-Wai Hong, Yuk-Hee Chan, Wan-Chi Siu |
Signal Process. | 3 |
| 1998 | Fast sequential implementation of "neural-gas" network for vector quantizationabstractAlthough the "neural-gas" network proposed by Martinetz et al. in 1993 has been proven for its optimality in vector quantizer design and has been demonstrated to have good performance in time-series prediction, its high computational complexity (NlogN) makes it a slow sequential algorithm. We suggest two ideas to speedup its sequential realization: (1) using a truncated exponential function as its neighborhood function and (2) applying a new extension of the partial distance elimination method (PDE). This fast realization is compared with the original version of the neural-gas network for codebook design in image vector quantization. The comparison indicates that a speedup of five times is possible, while the quality of the resulting codebook is almost the same as that of the straightforward realization. Clifford S. T. Choy, Wan-Chi Siu |
IEEE Trans. Commun. | 2 |
| 1998 | A practical postprocessing technique for real-time block-based coding systemabstractA noniterative postprocessing method is proposed to restore the images encoded with block-based compression standards such as Joint Photographic Experts Group (JPEG). This method classifies small local boundary regions according to their intensity distribution and then selects appropriate linear predictors to estimate the corresponding boundary pixels in the regions. This approach is easy to implement and we found in our simulations that its restoration performance was very respectable compared with the reported postprocessing methods. Yuk-Hee Chan, Sung-Wai Hong, Wan-Chi Siu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 1998 | Adaptive temporal decimation algorithm with dynamic time windowabstractDecimation approaches for image processing have been widely used for various applications. For video processing, decimation refers to sampling the frame rate in order to reduce the number of processing frames. Most of these temporal decimation methods discard whole frames; as a result, some high-speed motions could be completely eliminated while some redundant frames might remain in the processing frames. An adaptive temporal decimation approach has been successfully developed by Olstad (1993) to take both spatial and temporal information into consideration and is fully compatible with some existing discrete cosine transform (DCT)-based standards, such as MPEG and H.261. Moreover, it theoretically preserves all high activity motions and discards all low activity motions. However, we found that it is still not fully adaptive due to the confinement of the size of the time window. The discontinuity detection method is quite complex and, more importantly, the efficiency of coding block position maps is fairly low. We propose to resolve the problem of the time window by a dynamic time window approach. By using variable sizes of the time window, the optimal number of remaining frames could be produced. It also enhances the visual quality of the resulting video while the compression is comparable with the conventional approach. Based on our proposed algorithm, a simple but efficient quantization process has been used to replace the highly complex temporal discontinuity detection. The conventional adaptive temporal decimation algorithm operates on the basis of block sequences, but our dynamic approach which can retain all high activity blocks operates on the spatio-temporal domain. This approach can reduce redundant planes with slow activity and give higher precision for blocks with high activity. Experimental results show that the proposed algorithm achieves the optimal number of remaining frames. Sai Ho Kwok, Wan-Chi Siu, Anthony G. Constantinides |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1998 | A deblocking technique for block-transform compressed image using wavelet transform modulus maximaabstractIn this work, we introduce a deblocking algorithm for Joint Photographic Experts Group (JPEG) decoded images using the wavelet transform modulus maxima (WTMM) representation. Under the WTMM representation, we can characterize the blocking effect of a JPEG decoded image as: 1) small modulus maxima at block boundaries oversmooth regions; 2) noise or irregular structures near strong edges; and 3) corrupted edges across block boundaries. The WTMM representation not only provides characterization of the blocking effect, but also enables simple and local operations to reduce the adverse effect due to this problem. The proposed algorithm first performs a segmentation on a JPEG decoded image to identify the texture regions by noting that their WTMM have small variation in regularity. We do not process the modulus maxima of these regions, to avoid the image texture being "oversmoothed"by the algorithm. Then, the singularities in the remaining regions of the blocky image and the small modulus maxima at block boundaries are removed. We link up the corrupted edges, and regularize the phase of modulus maxima as well as the magnitude of strong edges. Finally,the image is reconstructed using the projection onto convex set (POCS)technique on the processed WTMM of that JPEG decoded image.This simple algorithm improves the quality of a JPEG decoded image inthe senses of signal-to-noise ratio (SNR) as well as visual quality. We also compare the performance of our algorithm to the previous approaches,such as CLS and POCS methods. The most remarkable advantage of the WTMM deblocking algorithm is that we can directly process the edges and texture of an image using its WTMM representation. Richard T. C. Hsung, Daniel Pak-Kong Lun, Wan-Chi Siu |
IEEE Trans. Image Process. | 3 |
| 1998 | Minimum Dynamic SPECT Image Acquisition Time Required for T1-201 Tracer Kinetic Modelling
Chi-Hoi Lau, Stefan Eberl, David Dagan Feng, Hidehiro Iida, Daniel Pak-Kong Lun, Wan-Chi Siu, Yoshikazu Tamura, George J. Bautovich, Yukihiko Ono |
IEEE Trans. Medical Imaging | 6 |
| 1998 | Dynamic Imaging and Tracer Kinetic Modeling for Emission Tomography Using Rotating DetectorsabstractWhen performing dynamic studies using emission tomography the tracer distribution changes during acquisition of a single set of projections. This is particularly true for some positron emission tomography (PET) systems which, like single photon emission computed tomography (SPECT), acquire data over a limited angle at any time, with full projections obtained by rotation of the detectors. In this paper, an approach is proposed for processing data from these systems, applicable to either PET or SPECT. A method of interpolation, based on overlapped parabolas, is used to obtain an estimate of the total counts in each pixel of the projections for each required frame-interval, which is the total time to acquire a single complete set of projections necessary for reconstruction. The resultant projections are reconstructed using traditional filtered backprojection (FBP) and tracer kinetic parameters are estimated using a method which relies on counts integrated over the frame-interval rather than instantaneous values. Simulated data were used to illustrate the technique's capabilities with noise levels typical of those encountered in either PET or SPECT. Dynamic datasets were constructed, based on kinetic parameters for fluoro-deoxy-glucose (FDG) and use of either a full ring detector or rotating detector acquisition. For the rotating detector, use of the interpolation scheme provided reconstructed dynamic images with reduced artefacts compared to unprocessed data or use of linear interpolation. Estimates for the metabolic rate of glucose had similar bias to those obtained from a full ring detector. Chi-Hoi Lau, David Dagan Feng, Brian F. Hutton, Daniel Pak-Kong Lun, Wan-Chi Siu |
IEEE Trans. Medical Imaging | 5 |
| 1998 | A class of competitive learning models which avoids neuron underutilization problemabstractIn this paper, we study a qualitative property of a class of competitive learning (CL) models, which is called the multiplicatively biased competitive learning (MBCL) model, namely that it avoids neuron underutilization with probability one as time goes to infinity. In the MBCL, the competition among neurons is biased by a multiplicative term, while only one weight vector is updated per learning step. This is of practical interest since its instances have computational complexities among the lowest in existing CL models. In addition, in applications like classification, vector quantizer design and probability density function estimation, a necessary condition for optimal performance is to avoid neuron underutilization. Hence, it is possible to define instances of MBCL to achieve optimal performance in these applications. Clifford S. T. Choy, Wan-Chi Siu |
IEEE Trans. Neural Networks | 2 |
| 1997 | Distortion sensitive competitive learning for vector quantizer designabstractWe propose the distortion sensitive competitive learning (DSCL) algorithm for codebook design in image vector quantization. The algorithm is based on the equidistortion principle for an asymptotically optimal vector quantizer after Gersho (1979) and from Ueda and Nakano (1994). The DSCL is simple and efficient in that a single weight vector update is performed per training vector, and the processing speed of the DSCL in a sequential or multiprocessor environment can further be improved by applying a modified partial distance elimination (MPDE) method. Simulations indicate that the DSCL outperforms some previously proposed neural algorithms, including the "neural-gas" from Martinetz et al. (1993) and the DEFCL from Butler and Jiang (1996). In combining with the MPDE, the DSCL is faster than the "neural-gas" up to a factor of 45 times on a sequential machine, and yet arrives at better codebooks with the same number of iterations. Clifford S. T. Choy, Wan-Chi Siu |
ICASSP | 2 |
| 1997 | A New Approach for Restoring Block-Transform Coded Images with Estimation of Correlation MatricesabstractThis paper presents a new restoration approach to reduce coding artifacts in block-transform image coding. Different from conventional restoration techniques, the proposed one is non-iterative and requires a low computational cost, yet can reconstruct objectively and subjectively better images. This good performance is achieved because of the following advantages the proposed approach has: (i) efficient incorporation of the solution bound into restoration; and (ii) effective exploitation of local image properties and statistical knowledge about the quantizers used. Steven Sheung-On Choy, Yuk-Hee Chan, Wan-Chi Siu |
ICIP (2) | 3 |
| 1997 | A Scalable and Adaptive Temporal Segmentation Algorithm for Video Coding
Sai Ho Kwok, Wan-Chi Siu, Anthony G. Constantinides |
CVGIP Graph. Model. Image Process. | 2 |
| 1997 | Reduction of block-transform image coding artifacts by using local statistics of transform coefficientsabstractThis letter presents a new approach to reduce coding artifacts in transform image coding. We approach the problem in an estimation of each transform coefficient from its quantized version with its local mean and variance. The proposed method can significantly reduce coding artifacts of low bit-rate coded images, and at the same time guarantee that the resulting images satisfies the quantization error constraint. Steven Sheung-On Choy, Yuk-Hee Chan, Wan-Chi Siu |
IEEE Signal Process. Lett. | 3 |
| 1997 | Variable temporal-length 3-D discrete cosine transform codingabstractThree-dimensional discrete cosine transform (3-D DCT) coding has the advantage of reducing the interframe redundancy among a number of consecutive frames, while the motion compensation technique can only reduce the redundancy of at most two frames. However, the performance of the 3-D DCT coding will be degraded for complex scenes with a greater amount of motion. This paper presents a 3-D DCT coding with a variable temporal length that is determined by the scene change detector. Our idea is to let the motion activity in each block be very low, while the efficiency of the 3-D DCT coding could be increased. Experimental results show that this technique is indeed very efficient. The present approach has substantial improvement over the conventional fixed-length 3-D DCT coding and is also better than that of the Moving Picture Expert Group (MPEG) coding. Yui-Lam Chan, Wan-Chi Siu |
IEEE Trans. Image Process. | 2 |
| 1997 | A technique for extracting physiological parameters and the required input function simultaneously from PET image measurements: theory and simulation studyabstractPositron emission tomography (PET) is an important tool for enabling quantification of human brain function. However, quantitative studies using tracer kinetic modeling require the measurement of the tracer time-activity curve in plasma (PTAC) as the model input function. It is widely believed that the insertion of arterial lines and the subsequent collection and processing of the biomedical signal sampled from the arterial blood are not compatible with the practice of clinical PET, as it is invasive and exposes personnel to the risks associated with the handling of patient blood and radiation dose. Therefore, it is of interest to develop practical noninvasive measurement techniques for tracer kinetic modeling with PET. In this paper, a technique is proposed to extract the input function together with the physiological parameters from the brain dynamic images alone. The identifiability of this method is tested rigorously by using Monte Carlo simulation. The results show that the proposed method is able to quantify all the required parameters by using the information obtained from two or more regions of interest (ROI's) with very different dynamics in the PET dynamic images. There is no significant improvement in parameter estimation for the local cerebral metabolic rate of glucose (LCMRGlc) if the number of ROI's are more than three. The proposed method can provide very reliable estimation of LCMRGlc, which is our primary interest in this study. David Dagan Feng, Koon-Pong Wong, Chi-Ming Wu, Wan-Chi Siu |
IEEE Trans. Inf. Technol. Biomed. | 4 |
| 1996 | An improved quantitative measure of image restoration qualityabstractThe field of image restoration lacks promising comparison vehicle for judging the effectiveness of competing algorithms. By far the most widely adopted quantitative measure of image restoration quality is the SNR improvement. However, we find that the SNR improvement is of low precision, which will adversely hinder it from being a reliable measure. It is also noted that another limitation of the SNR improvement is that it cannot reveal clearly the extent to which the image quality is improved. We devise an alternative measure for quantitative evaluation of image restoration quality. The proposed measure is much more precise than the SNR improvement. Moreover, the proposed measure contains finite and meaningful reference points in its measurements, to provide us with a better insight into the effectiveness of restoration algorithms that under study. Steven Sheung-On Choy, Yuk-Hee Chan, Wan-Chi Siu |
ICASSP | 3 |
| 1996 | Fast algorithm for 2-D image moments via the Radon transformabstractA fast algorithm for the computation of the two-dimensional image moments is proposed. In our approach, a new discrete Radon transform (DRT) is used for the major part of the algorithm. The new DRT preserves an important property of the continuous Radon transform that the regular or geometric moments can be directly obtained from the projection data. With this property, the computation of a two-dimensional (2-D) image moments can be decomposed as a number of one-dimensional (1-D) ones, hence greatly reducing the computational complexity. Comparisons of the computational complexity and performance with some known methods are also given. It is shown that the proposed algorithm significantly reduces the complexity and computation time. Tak-Wai Shen, Daniel Pak-Kong Lun, Wan-Chi Siu |
ICASSP | 3 |
| 1996 | A practical real-time post-processing technique for block effect eliminationabstractIn this paper, a non-iterative post-processing method is proposed to restore the images encoded with block-based compression standards such as JPEG. This method classifies small local boundary regions according to their intensity distribution and then select appropriate linear predictors to estimate the corresponding boundary pixels in the regions. This approach is easy to implement and we found in our simulations that its restoration performance was very respectable compared with the reported post-processing methods. Sung-Wai Hong, Yuk-Hee Chan, Wan-Chi Siu |
ICIP (2) | 3 |
| 1996 | A deblocking technique for JPEG decoded image using wavelet transform modulus maxima representationabstractIn this paper, we introduce a local deblocking algorithm for JPEG decoded images using the wavelet transform modulus maxima (WTMM) representation. Under the WTMM representation, we can characterize the blocking effect as: 1) small modulus maxima at block boundaries over smooth regions; 2) noises or irregular structures near strong edges; 3) corrupted edges across block boundaries. The WTMM representation not only provides characterization of the blocking effect, but also enables simple and local operations on those singularities. The proposed algorithm first performs a segmentation to discriminate the texture regions of an image based on the WTMM local regularity variance. We then keep the modulus maxima of these regions, which are of low regularity variation, unchange to avoid the image texture being "over-smoothed" by the algorithm. Then, the singularities on the remaining regions of the blocky image and small modulus maxima at block boundaries are removed. Then, we link up the corrupted edges and regularize the phase of modulus maxima as well as the amplitude of strong edges. Finally, the image is reconstructed using the projection onto convex sets (POCS) technique on the processed WTMM of the JPEG decoded image. This simple algorithm improves the quality of JPEG decoded image in the sense of signal to noise ratio as well as visual quality. We also compare the performance of our algorithm with the previous approaches and show the superiority over them. The most remarkable advantage of the WTMM deblocking algorithm is that it incorporates direct edges and texture operations into the WTMM representation. Richard T. C. Hsung, Daniel Pak-Kong Lun, Wan-Chi Siu |
ICIP (2) | 3 |
| 1996 | New adaptive pixel decimation for block motion vector estimationabstractA new adaptive technique based on pixel decimation for the estimation of motion vector is presented. In a traditional approach, a uniform pixel decimation is used. Since part of the pixels in each block do not enter into the matching criterion, this approach limits the accuracy of the motion vector. In this paper, we select the most representative pixels based on image content in each block for the matching criterion. This is due to the fact that high activity in the luminance signal such as edges and texture mainly contributes to the matching criterion. Our approach can compensate the drawback in standard pixel decimation techniques. Computer simulations show that this technique is close to the performance of the exhaustive search with significant computational reduction. Yui-Lam Chan, Wan-Chi Siu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1995 | A new block motion vector estimation using adaptive pixel decimationabstractBlock motion estimation is being widely used in video coding. A new adaptive technique based on pixel decimation for estimating motion vector is presented. In the traditional approach, a uniform pixel decimation is used. Since some pixels in each block do not enter into the matching criterion, this approach might limit the accuracy of the motion vector. We select the most representative pixels based on the image content in each block for the matching criterion. This is due to the fact that high activity in the luminance signal such as edges and texture contributes mainly to the matching criterion. Our approach can compensate the drawback in standard pixel decimation techniques. Computer simulations show that this technique is close to the performance of the exhaustive search method with a significant reduction in computational complexity. Yui-Lam Chan, Wan-Chi Siu |
ICASSP | 2 |
| 1995 | Algorithm for solving bipartite subgraph problem with probabilistic self-organizing learningabstractSelf-organizing model has been successfully applied to solving some combinatorial optimization problems, including the travelling salesman problem, the routing problem and the cell-placement problem, but there has not much work reported on its application to solving the graph partitioning problem. We propose a novel mapping which has not been proposed before, with some changes to the original Kohonen's (1982) algorithm so as to enable it to solve a partitioning problem-the bipartite subgraph problem. This new approach is compared with to the maximum neural network for solving the same problem, showing that the performance of our new approach is superior to that of the maximum neural network. Clifford S. T. Choy, Wan-Chi Siu |
ICASSP | 2 |
| 1995 | New 2n discrete cosine transform algorithm using recursive filter structureabstractThe discrete cosine transform (DCT) is widely used in digital signal processing. It is always desirable to look for more efficient algorithms for the realization of the DCT. We generalize a formulation for converting a length-2/sup n/ DCT into n groups of equations, then apply a novel technique for its implementation. The sizes of the groups are 2/sup m/, for m=n-1,...,0. While their structures are extremely regular. The realization can then be converted into the simplest recursive filter form, which is of particularly simple for practical implementation. The filter structure is numerically stable, since it involves no division at all. Wan-Chi Siu, Yuk-Hee Chan, Lap-Pui Chau |
ICASSP | 1 |
| 1995 | Highly efficient coding schemes for contour line drawingsabstractIn this paper, adaptive coding schemes for contour line drawings based on chain code representation is presented. In this scheme, the chain code or the chain-difference code of a contour is modeled as an n-order Markov sequence and then coded with an arithmetic coding scheme adaptively. Experimental result shows that the proposed approach is better than some other conventional approaches. Yuk-Hee Chan, Wan-Chi Siu |
ICIP (3) | 2 |
| 1995 | Subband adaptive regularization method for removing blocking effectabstractThis paper presents two new approaches to remove blocking effect in low-bit rate transform coded images by using subband decomposition/reconstruction technique. They are designed to act as a supplementary post-processing step of the JPEG standard. Both approaches make use of the noise characteristic of each subband to bound the maximum tolerable error and the smoothness of the restored images in restoring subband images with regularization. One of them will also utilize the spatial activity of the restoring images to tighten the bounds. Computer simulations showed that the new adaptive objective functions could achieve a better restoration performance in terms of both subjective and objective measures than did other conventional objective functions. Sung-Wai Hong, Yuk-Hee Chan, Wan-Chi Siu |
ICIP | 3 |
| 1995 | Fast Interframe Transfrom Coding Based on Characteristics of Transform Coefficients and Frame DifferenceabstractThe interframe transform coding has been seldom used in practice because of the considerable computational complexity. To reduce the computational complexity, a fast algorithm is proposed which reduces the number of operations by limiting the calculation of transform coefficients without significant quality degradation. Different modes of transformation are performed according to frame difference. This proposed fast interframe coding, like the MPEG, has an asymmetric property with decoding being much faster than encoding. Computer simulations show that this fast algorithm can significantly reduce the computational burden in interframe transform coding and is even faster than MPEG-like coding. Yui-Lam Chan, Wan-Chi Siu |
ISCAS | 2 |
| 1995 | Peak Detection in Hough Transform Via Self-Organizing LearningabstractIn this paper, we suggest a novel concept of applying the self-organizing map (SOM) in the Hough domain for a significant reduction of the Hough space. By using the SOM as the output space of the generalized Hough transform, the conventional 4-D Hough domain is replaced by a 10/spl times/10 map, organized in a rectangular grid. Experimental results indicate high accuracy in voting is attainable despite its small memory requirement. Clifford S. T. Choy, Pui-Kin Ser, Wan-Chi Siu |
ISCAS | 3 |
| 1995 | An Adaptive Constrained Least Square Approach for Removing Blocking EffectabstractThis paper presents a new adaptive objective function based on the regularized iterative block reduction technique for low-bit rate transform coded images. Also, a better initial estimate for the regularization approach is presented. Two types of prior knowledge are used: the first type bounds the maximum tolerable error (roughness), and the second type restricts the high-frequency content (smoothness) of the restorated images. Computer simulations showed that the new adaptive objective function with the proposed initial estimate performed better on both subjective and objective measures than did a previously proposed objective function. Sung-Wai Hong, Yuk-Hee Chan, Wan-Chi Siu |
ISCAS | 3 |
| 1995 | On the Convolution Property of a New Discrete Radon Transform and its Efficient Inversion Algorithm
Daniel Pak-Kong Lun, Richard T. C. Hsung, Wan-Chi Siu |
ISCAS | 3 |
| 1995 | New Single-Pass Algorithm for Parallel Thinning
Steven Sheung-On Choy, Clifford S. T. Choy, Wan-Chi Siu |
Comput. Vis. Image Underst. | 3 |
| 1995 | A New Generalized Hough Transform for the Detection of Irregular Objects
Pui-Kin Ser, Wan-Chi Siu |
J. Vis. Commun. Image Represent. | 2 |
| 1995 | Memory compression for straight line recognition using the Hough transform
Pui-Kin Ser, Wan-Chi Siu |
Pattern Recognit. Lett. | 2 |
| 1995 | Generation of moment invariants and their uses for character recognition
Wai-Hong Wong, Wan-Chi Siu, Kin-Man Lam 0001 |
Pattern Recognit. Lett. | 2 |
| 1995 | In search of the optimal searching sequence for VQ encodingabstractThe codeword searching sequence is sometimes very vital to the efficiency of a vector quantization (VQ) encoding algorithm. In this paper, we evaluate some necessary criteria for the derivation of an optimal searching sequence and derive the optimal searching sequence based on such criteria. Yuk-Hee Chan, Wan-Chi Siu |
IEEE Trans. Commun. | 2 |
| 1994 | Efficient formulation for the realization of discrete cosine transform using recursive structureabstractEffective formulations for the conversion of the discrete Fourier transform (DFT) into a recursive structure are available and have been very effective for the realization using software, hardware and VLSI techniques. Little research work has been reported on the effective way to convert the discrete cosine transform (DCT) into a recursive form and the related realization. We propose a new method to convert a prime length DCT into a recursive structure. A trivial approach is to convert the DCT into the DFT and to apply Goertzel's (1958) algorithm for the rest of the realization. However, this method is inefficient. In our approach, we convert a prime length DCT into suitable transforms with half of the original length to effect fast realization. The number of operations is greatly reduced and the structure is extremely regular.> Lap-Pui Chau, Wan-Chi Siu |
ICASSP (3) | 2 |
| 1994 | Approach of using a density equalizing function to self-organizing learning for solving travelling salesman problemabstractProposes a new approach which requires neither neuron addition nor deletion, and at the same time, N neurons are sufficient to solve an N-city travelling salesman problem. the authors begin with a description of their model, and then results for applying the model to solve the 30-city problem from Hopfield are presented. Results of practical testing show that the present approach always converges. It has the highest chance to achieve the optimal solution, and gives the best most probable solution, as compared to other self-organizing algorithms.> Clifford S. T. Choy, Wan-Chi Siu |
ICASSP (2) | 2 |
| 1994 | A New Adaptive Interframe Transform Coding using Directional ClassificationabstractInterframe transform coding is affected not only by the statistics of spatial details within a frame, but also by the variation of the amount of movement and other temporal activities in different regions of the image sequence. Therefore, adaptive techniques have to be used in order to achieve good image quality. In this paper, we propose a new version of the adaptive interframe coding method, namely directional classification, which is based on image sequence statistics. Blocks with different perceptual features such as edges and high motion activity are categorised to different classes. Then, a new adaptive quantization, associated with appropriate scanning and Huffman coding, are employed based on the classification map. Coding tests using computer simulation show this technique is indeed very efficient.> Yui-Lam Chan, Wan-Chi Siu |
ICIP (2) | 2 |
| 1994 | New Adaptive Iterative Image Restoration AlgorithmabstractIt has been shown in the literature that adaptive regularized image restoration is superior to the non-adaptive case. However, the adaptivity introduced in most proposed iterative algorithms is based only on the application of the space-variant smoothing operator. It is found that these adaptive algorithms suffer from insufficient smoothing of the flat image regions. In this paper, an adaptive iterative image restoration algorithm, which applies both techniques of space-variant smoothing and space-variant restoration, is proposed to overcome the stated problem. It is shown by experiments that the restored images obtained by the proposed algorithm are better in terms of both numerical measurement and visual quality.> Steven Sheung-On Choy, Yuk-Hee Chan, Wan-Chi Siu |
ICIP (2) | 3 |
| 1994 | The Determination of the Searching Sequence for VQ EncodingabstractMost VQ encoding algorithms are proposed to reject codewords or stop searching codewords as soon as possible by using various criteria. This paper proposes an algorithm to dynamically determine the codeword searching sequence for a given input vector, which can further improve the codeword elimination efficiency. An encoding algorithm is also proposed based on this idea.> Yuk-Hee Chan, Wan-Chi Siu |
ISCAS | 2 |
| 1994 | Object Recognition with a 2-D Hough DomainabstractIn this paper, we report a new Generalized Hough Transform (GHT) for the recognition of non-analytic objects. Our proposed GHT algorithm is able to replace the conventional 4-D Hough space with a 2-D one. The significant reduction of memory requirement leads to a great simplification of the peak searching time for the recognition. By an analysis of the angular deviation of the edge operator, we also derive a new profile for the modeling of peaks, which is extremely helpful for our peak searching and verification processes.> Pui-Kin Ser, Wan-Chi Siu |
ISCAS | 2 |
| 1994 | An Approach to Subband DCT Image Coding
Yuk-Hee Chan, Wan-Chi Siu |
J. Vis. Commun. Image Represent. | 2 |
| 1994 | General approach for the realization of DCT/IDCT using convolutions
Yuk-Hee Chan, Wan-Chi Siu |
Signal Process. | 2 |
| 1994 | A Pipeline Design for the Realization of the Prime Factor Algorithm Using the Extended Diagonal StructureabstractIn this brief contribution, an efficient pipeline architecture is proposed for the realization of the Prime Factor Algorithm (PFA) for digital signal processing. By using the extended diagonal feature of the Chinese Remainder Theorem (CRT) mapping, we show that the input data sequence can be directly loaded into a multidimensional array for the PFA computation without any permutation. Short length modules are modified such that an in-place and in-order computation is allowed. The computed results can then be directly restored back to the memory array without the need for further reordering. More importantly, the CRT mapping can also be used to represent the output data, hence we can utilize the extended diagonal feature of the CRT mapping to directly send the computed results to the outside world. As compared to the previous approaches, the present approach requires no shifting or rotation during the data loading and retrieval processes. In the case of multidimensional PFA computation, it does not require the computation to be split up into a number of two-dimensional computations. Hence, the overhead required for data loading and retrieval in each two-dimensional stage can be saved. Daniel Pak-Kong Lun, Wan-Chi Siu |
IEEE Trans. Computers | 2 |
| 1994 | Efficient implementation of discrete cosine transform using recursive filter structureabstractWe generalize a formulation for converting a length-2/sup n/ discrete cosine transform into n groups of equations, then apply a novel technique for its implementation. The sizes of the groups are 2/sup n-1/, 2/sup n-2/, ...2/sup 0/ respectively, while their structures are extremely regular. The realization can then be converted into recursive filter form, which is particularly simple for practical implementation.> Yuk-Hee Chan, Lap-Pui Chau, Wan-Chi Siu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 1993 | Generalized approach for the realization of discrete cosine transform using cyclic convolutions
Yuk-Hee Chan, Wan-Chi Siu |
ICASSP (3) | 2 |
| 1993 | Invariant Hough transform with matching technique for the recognition of non-analytic objects
Pui-Kin Ser, Wan-Chi Siu |
ICASSP (5) | 2 |
| 1993 | On the theoretical lower bound of the multiplicative complexity for DCT
Yuk-Hee Chan, Wan-Chi Siu |
ISCAS | 2 |
| 1993 | Generation of Chain-coded Contours and Contours Inclusion Relationship Under Multiprocessor Environment
Clifford S. T. Choy, Wan-Chi Siu |
ISCAS | 2 |
| 1993 | Transform-based fine-corase vector quantization
Kin-Man Lam 0001, Wan-Chi Siu, Kai-Ming Tse |
ISCAS | 2 |
| 1993 | Improved formulations for fast polynomial transform
Albert Ming Loh, Wan-Chi Siu |
ISCAS | 2 |
| 1993 | Automatic generation of moment invariants and the use of higher order moments for character recognition
Wai-Hong Wong, Wan-Chi Siu, Kin-Man Lam 0001 |
ISCAS | 2 |
| 1992 | On inherent in-place and in-order features of the prime factor algorithmabstractFor the computation of the prime factor algorithm (PFA), an in-place and in-order approach is always desirable because it reduces the memory requirement for the storage of the temporary results, and the computation time which is required to unscramble the output sequence to a proper order. In fact, the processing time required for this unscrambling process can take up as much as 50% of the overall computation time. It is shown that the PFA has an intrinsic property that allows it to be easily realized in an in-place and in-order form. No extra operation is required as in the previous propositions. Nevertheless, the sequence length of the PFA computation must be carefully selected. The conditions under which a particular sequence length is possible for a natural in-place and in-order PFA computation are analyzed. The result is useful to both the hardware and software realization of the PFA.> Daniel Pak-Kong Lun, Wan-Chi Siu |
ICASSP | 2 |
| 1992 | Efficient mapping scheme for the prime factor discrete Hartley transformabstractCompared to the complexity for realizing the prime factor discrete Fourier transform (DFT), the prime factor discrete Hartley transform requires some extra arithmetic operations for the realization of the prime factor mapping. These extra arithmetic operations can take up as much as 40% of the total arithmetic operations required. A new prime factor mapping scheme which requires no extra arithmetic operations is proposed for the computation of the discrete Hartley transform. It is achieved by embedding all the extra arithmetic operations into the subsequent short length computations, whereas the arithmetic complexities of these embedded short length modules remain unchanged.> Wan-Chi Siu, Daniel Pak-Kong Lun |
ICASSP | 1 |
| 1991 | Data Routing Networks for Systolic/Pipeline Realization of Prime Factor MappingabstractIt is pointed out that transformed data computed by systolic/pipeline processors using the data shuffling network recently proposed by T.K. Troung et al. (ibid., vol.37, p.266-73, Mar. 1988) cannot be unscrambled by simply reversing the cyclic row and cyclic column shufflings. This can be amended by the proposed restoration scheme. In addition, efficient architectures for the data routing networks with low circuit complexities are proposed. These form useful building blocks for very-high-throughput hardware realizations.> Kar-Lik Wong, Wan-Chi Siu |
IEEE Trans. Computers | 2 |
| 1990 | Fast detection of ellipses using chord bisectorsabstractThe horizontal and vertical chord bisectors of objects are used to detect elliptic objects in a binary image. In 3-D applications, a circular plane viewed from different viewing angles will always be projected as a pseudoellipse. Hence, the detection of ellipses, which are parameterized with five parameters (namely x0, y0, a, b and theta ), is much more general and useful for both 2-D and 3-D image recognition tasks. For any general ellipse, the loci of its chord bisectors and chord lengths possess various properties which can be utilized to extract the center (0, y0) and to facilitate discrimination of elliptic objects from others. Different pairs of parallel strips are used to calculate the hypothesized values of two newly defined parameters of an ellipse. The statistical modes for these two parameters are extracted for the computation of the remaining parameters theta , a, and b. This two-stage algorithm is fast due to the fact that locating bisecting points solves only very simple additions and shifting operations, and the second stage involves very few computations for each object.> Wan-Chi Siu |
ICASSP | 2 |
| 1990 | Yet a faster address generation scheme for the computation of prime factor algorithmsabstractAn in-place, in-order address generation scheme is proposed for the realization of prime factor mapping (PFM). The new approach has the characteristic of forming systematic and regular structures. Hence it is suitable for realizations using both high-level and low-level languages. Furthermore, it requires very few modulo operations and no modulo inverse for its computation; such inverses often take up memory space for their storage and/or extra time for the computation in other address generation algorithms. The approach has been implemented using Fortran 77 and the assembly language of the 320C25 DSP. It shows that a maximum of 86% saving in address generation time can be achieved as compared to the conventional approach.> Daniel Pak-Kong Lun, Wan-Chi Siu |
ICASSP | 3 |
| 1989 | Fast address generation for the computation of prime factor algorithmsabstractThe authors propose an address generation scheme for in-place in-order prime factor algorithms. This scheme achieves high efficiency by using simple indirect addressing techniques to replace complicated modulo operations of previous methods. It is shown that the scheme is most suitable for software realizations using assembly languages. A hardware address generator based on this mapping scheme is also suggested. Its architectural simplicity makes it suitable for fabrication as a single-chip peripheral to add onto current digital signal processors or to be integrated into future DSP (digital signal processor) architectures. In both cases, a 30% to 50% reduction in computation time is achievable in the realization of prime factor algorithms using digital signal processors.> Kar-Lik Wong, Wan-Chi Siu |
ICASSP | 2 |
| 1988 | A nesting algorithm for very fast discrete Fourier transformsabstractThe use of a nesting discrete Fourier transform technique to compute discrete Fourier transform is proposed. This technique only relies on two primitive modules and other modules are generated by a standard nesting procedure. The speed of computation of this approach is comparable to the speed of computation of the WFTA, whereas the program size of the present approach is smaller than that of the WFTA. This approach is most suitable for cases where there are restrictions on memory size.> Wan-Chi Siu |
ICASSP | 1 |
| 1986 | A hardware efficient realisation of number theoretic convolversabstractIn this paper, we propose hardware realisations of Number Theoretic Transforms that are based on the transformation of their fundamental relationships into recursive filter forms with single integer poles. Furthermore use is made of Read-Only- Memory(ROM) to effect the multiplications by the root of unity, α Suitable NTTs are then suggested for the fast computation of cyclic convolutions using multi-dimensional and multi-modular techniques. The required ROM size in the proposed realisations is small and the control of data flow is simple and straightforward. This new class of Number Theoretic Transforms can relax considerably the normal sequence length and wordlength constraints for the NTT. Wan-Chi Siu, Anthony G. Constantinides |
ICASSP | 1 |
| 1984 | Hardware realization of Mersenne number transforms for fast digital convolutionabstractIn this paper we convert the Mersenne-prime and Mersenne-composite Number Transforms into recursive filter form and propose simple hardware structures to carry out the fast implementation of circular convolutions. We also present the results of our study employing efficient methods to complute long circular convolutions using the Mersenne Number Transforms. Wan-Chi Siu, Anthony G. Constantinides |
ICASSP | 1 |