EDBT 2026 Demo / reviewers in the wild / expert
Yui-Lam Chan
dblp:04/2096
· DBLP profile ↗
85ranked-venue papers
15as first author
13since 2021 · last 2026
0000-0002-1473-094XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 73 · 13 first-author · 11 since 2021Systems, architecture and hardware · 7 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-Branch Aesthetic and Technical Perspectives With Cross Tri-Fusion Attention for No-Reference Audio-Visual Quality Assessment
Ngai-Wing Kwong, Yui-Lam Chan, Ziyin Huang, Sik-Ho Tsang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Long Short-Term Fusion by Multi-Scale Distillation for Screen Content Video Quality EnhancementabstractDifferent from natural videos, where artifacts distributed evenly, the artifacts of compressed screen content videos mainly occur in the edge areas. Besides, these videos often exhibit abrupt scene switches, resulting in noticeable distortions in video reconstruction. Existing multiple-frame models using a fixed range of neighbor frames face challenges in effectively enhancing frames during scene switches and lack efficiency in reconstructing high-frequency details. To address these limitations, we propose a novel method that effectively handles scene switches and reconstructs high-frequency information. In the feature extraction part, we develop long-term and short-term feature extraction streams, in which the long-term feature extraction stream learns the contextual information, and the short-term feature extraction stream extracts more related information from shorter input to assist the long-term stream to handle fast motion and scene switches. To further enhance the frame quality during scene switches, we incorporate a similarity-based neighbor frame selector before feeding frames into the short-term stream. This selector identifies relevant neighbor frames, aiding in the efficient handling of scene switches. To dynamically fuse the short-term feature and long-term features, the muti-scale feature distillation focuses on adaptively recalibrating channel-wise feature responses to achieve effective feature distillation. In the reconstruction part, a high-frequency reconstruction block is proposed for guiding the model to restore the high-frequency components. Experimental results demonstrate the significant advancements achieved by our proposed Long Short-term Fusion by Multi-Scale Distillation (LSFMD) method in enhancing the quality of compressed screen content videos, surpassing the current state-of-the-art methods. Ziyin Huang, Yui-Lam Chan, Ngai-Wing Kwong, Sik-Ho Tsang, Kin-Man Lam 0001, Bingo Wing-Kuen Ling |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Multi-Frame Spatiotemporal Feature and Hierarchical Learning Approach for No-Reference Screen Content Video Quality AssessmentabstractThe rapid adoption of remote work, online conferencing, and shared-screen collaboration has significantly increased the usage of screen content videos (SCVs), creating a growing need for reliable quality assessment to maintain excellent quality of service. While several full-reference SCV quality assessment (SCVQA) methods have been proposed, their practical application is often limited by the unavailability of reference videos. Existing no-reference SCVQA (NR-SCVQA) methods rely on handcrafted features and focus solely on specific distortions and features, potentially limiting their generalization ability. Moreover, they fail to explore the underlying spatiotemporal information of SCVs, which could hinder their performance. In this work, we propose a novel deep learning-based NR-SCVQA model specifically tailored to capture the comprehensive spatiotemporal features of SCVs to overcome these issues and challenges posed by the SCVQA task. Our approach incorporates a dual-channel spatiotemporal convolutional neural network (DCST-CNN) module to extract both content-aware and edge-aware spatiotemporal quality features, which enables an effective spatiotemporal quality feature representation learning for the downstream SCVQA task. Building upon the DCST-CNN, we further propose a Temporal Pyramid Transformer (TPT) module to fuse spatiotemporal features across multiple temporal scales, enabling the model to capture both short-term and long-term temporal dependencies within an SCV for hierarchical learning. The proposed DCST-CNN and TPT modules work together to provide a robust and accurate NR-SCVQA framework. We conduct experiments on SCVQA databases to validate the effectiveness of our model, which outperforms existing state-of-the-art NR-SCVQA method. The results demonstrate the strength and applicability of our approach in real-world SCVQA tasks. Ngai-Wing Kwong, Yui-Lam Chan, Sik-Ho Tsang, Ziyin Huang, Kin-Man Lam 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Frame Similarity-Based Screen Content Video Quality Enhancement via Adaptive Long Short-Term FusionabstractCompressed screen content videos often exhibit artifacts in edge areas and suffer from distortions during scene switches, where content abruptly changes between frames. Existing multi-frame models, which use a fixed range of neighbor frames, struggle with these switches. To address this, we propose a novel method that effectively handles scene switches. Our approach utilizes Long-term Feature Extraction (LFE) to capture contextual information, while the Frame Similarity-based Short-term Feature Extraction (FSFE) focuses on texture information to manage fast motion and scene switches. In FSFE, a Similarity-based Neighbor Frame Selector (SNFS) is designed to choose relevant neighbor frames for the short-term stream, enhancing the quality of scene switch frames. To fuse short-term and long-term features adaptively, we introduce a local-spatial and global-channel attention module, which recalibrates spatial and channel-wise feature responses. Experimental results show that our Frame Similarity-Based via Adaptive Long Short-Term Fusion (FSLST) method significantly improves the quality of compressed videos, outperforming current state-of-the-art methods. Ziyin Huang, Yui-Lam Chan, Ngai-Wing Kwong, Sik-Ho Tsang, Kin-Man Lam 0001, Bingo Wing-Kuen Ling |
VCIP | 2 |
| 2024 | Solving the imbalanced dataset problem in surveillance image blur classification
Yikun Pan, Sik-Ho Tsang, Tom Tak-Lam Chan, Yui-Lam Chan, Daniel Pak-Kong Lun |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | Spatio-temporal feature learning for enhancing video quality based on screen content characteristics
Ziyin Huang, Yui-Lam Chan, Sik-Ho Tsang, Ngai-Wing Kwong, Kin-Man Lam 0001, Bingo Wing-Kuen Ling |
J. Vis. Commun. Image Represent. | 2 |
| 2024 | Spatiotemporal feature learning for no-reference gaming content video quality assessment
Ngai-Wing Kwong, Yui-Lam Chan, Sik-Ho Tsang, Ziyin Huang, Kin-Man Lam 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2024 | Structured Adversarial Self-Supervised Learning for Robust Object Detection in Remote Sensing ImagesabstractObject detection plays a crucial role in scene understanding and has extensive practical applications. In the field of remote sensing object detection, both detection accuracy and robustness are of significant concern. Existing methods heavily rely on sophisticated adversarial training strategies that tend to improve robustness at the expense of accuracy. However, detection robustness is not always indicative of improved accuracy. Therefore, in this paper, we research how to enhance robustness, while still preserving high accuracy, or even improve both simultaneously, with simple vanilla adversarial training or even in the absence thereof. In pursuit of a solution, we first conduct an exploratory investigation by shifting our attention from adversarial training, referred to as adversarial fine-tuning, to adversarial pretraining. Specifically, we propose a novel pretraining paradigm, namely structured adversarial self-supervised (SASS) pretraining, to strengthen both clean accuracy and adversarial robustness for object detection in remote sensing images. At a high level, SASS pretraining aims to unify adversarial learning and self-supervised learning into pretraining and encode structured knowledge into pretrained representations for powerful transferability to downstream detection. Moreover, to fully explore the inherent robustness of vision Transformers and facilitate their pretraining efficiency, by leveraging the recent masked image modeling (MIM) as the pretext task, we further instantiate SASS pretraining into a concise end-to-end framework, named structured adversarial MIM (SA-MIM). SA-MIM consists of two pivotal components, structured adversarial attack and structured MIM (S-MIM). The former establishes structured adversaries for the context of adversarial pretraining, while the latter introduces a structured local-sampling global-masking strategy to adapt to hierarchical encoder architectures. Comprehensive experiments on three different datasets have demonstrated the significant superiority of the proposed pretraining paradigm over previous counterparts for remote sensing object detection. More importantly, regardless of with or without adversarial fine-tuning, it enables simultaneous improvements on detection accuracy and robustness as expected, promisingly alleviating the dependence on complicated adversarial fine-tuning. Kin-Man Lam 0001, Tianshan Liu, Yui-Lam Chan, Qi Wang 0009 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Optimized Quality Feature Learning for Video Quality AssessmentabstractRecently, some transfer learning-based methods have been adopted in video quality assessment (VQA) to compensate for the lack of enormous training samples and human annotation labels. But these methods induce a domain gap between source and target domains, resulting in a sub-optimal feature representation that deteriorates the accuracy. This paper proposes the optimized quality feature learning via a multi-channel convolutional neural network (CNN) with the gated recurrent unit (GRU) for no-reference (NR) VQA. First, the multi-channel CNN is pre-trained on the image quality assessment (IQA) domain using non-human annotation labels, which is inspired by self-supervised learning. Then, semi-supervised learning is used to fine-tune CNN and transfer the knowledge from IQA to VQA while considering motion information for optimized quality feature learning. Finally, all frame quality features are extracted as the input of GRU to obtain the final video quality. Experimental results demonstrate that our model achieves better performance than state-of-the-art VQA approaches. Ngai-Wing Kwong, Yui-Lam Chan, Sik-Ho Tsang, Daniel Pak-Kong Lun |
ICASSP | 2 |
| 2023 | Image super resolution via combination of two dimensional quaternion valued singular spectrum analysis based denoising, empirical mode decomposition based denoising and discrete cosine transform based denoising methods
Yingdan Cheng, Bingo Wing-Kuen Ling, Ziyin Huang, Yui-Lam Chan |
Multim. Tools Appl. | 5 |
| 2021 | Photo-Realistic Image Super-Resolution via Variational AutoencodersabstractThere is a great leap in objective accuracy on image super-resolution, which recently brings a new challenge on image super-resolution with larger up-scaling (e.g. 4×) using pixel based distortion for measurement. This causes over-smooth effect which cannot grasp well the perceptual similarity. The advent of generative adversarial networks makes it possible super-resolve a low-resolution image to generate photo-realistic images sharing distribution with the high-resolution images. However, generative networks suffer from problems of mode-collapse and unrealistic sample generation. We propose to perform Image Super-Resolution via Variational AutoEncoders (SR-VAE) learning according to the conditional distribution of the high-resolution images induced by the low-resolution images. Given that the Conditional Variational Autoencoders tend to generate blur images, we add the conditional sampling mechanism to narrow down the latent subspace for reconstruction. To evaluate the model generalization, we use KL loss to measure the divergence between latent vectors and standard Gaussian distribution. Eventually, in order to balance the trade-off between super-resolution distortion and perception, not only that we use pixel based loss, we also use the modified deep feature loss between SR and HR images to estimate the reconstruction. In experiments, we evaluated a large number of datasets to make comparison with other state-of-the-art super-resolution approaches. Results on both objective and subjective measurements show that our proposed SR-VAE can achieve good photo-realistic perceptual quality closer to the natural image manifold while maintain low distortion. Wan-Chi Siu, Yui-Lam Chan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Efficient Depth Intra Frame Coding in 3D-HEVC by Corner PointsabstractTo improve the coding performance of depth maps, 3D-HEVC includes several new depth intra coding tools at the expense of increased complexity due to a flexible quadtree Coding Unit/Prediction Unit (CU/PU) partitioning structure and a huge number of intra mode candidates. Compared to natural images, depth maps contain large plain regions surrounded by sharp edges at the object boundaries. Our observation finds that the features proposed in the literature either speed up the CU/PU size decision or intra mode decision and they are also difficult to make proper predictions for CUs/PUs with the multi-directional edges in depth maps. In this work, we reveal that the CUs with multi-directional edges are highly correlated with the distribution of corner points (CPs) in the depth map. CP is proposed as a good feature that can guide to split the CUs with multi-directional edges into smaller units until only single directional edge remains. This smaller unit can then be well predicted by the conventional intra mode. Besides, a fast intra mode decision is also proposed for non-CP PUs, which prunes the conventional HEVC intra modes, skips the depth modeling mode decision, and early determines segment-wise depth coding. Furthermore, a two-step adaptive corner point selection technique is designed to make the proposed algorithm adaptive to frame content and quantization parameters, with the capability of providing the flexible tradeoff between the synthesized view quality and complexity. Simulation results show that the proposed algorithm can provide about 66% time reduction of the 3D-HEVC intra encoder without incurring noticeable performance degradation for synthesized views and it also outperforms the previous state-of-the-art algorithms in term of time reduction and ∆ BDBR. Chang-Hong Fu 0002, Yui-Lam Chan, Hongbin Zhang 0005, Sik-Ho Tsang, Mengting Xu |
IEEE Trans. Image Process. | 2 |
| 2021 | Features Guided Face Super-Resolution via Hybrid Model of Deep Learning and Random ForestsabstractFace hallucination or super-resolution is a practical application of general image super-resolution which has been recently studied by many researchers. The challenge of good face hallucination comes from a variety of poses, illuminations, facial expressions, and other degradations. In many proposed methods, researchers resolve it by using a generative neural network to reduce the perceptual loss so we can generate a photo-realistic image. The problem is that researchers usually overlook the fidelity of the super-resolved image which could affect further facial image processing. Meanwhile, many CNN based approaches cascade multiple networks to extract facial prior information to improve super-resolution quality. Because of the end-to-end design, the details are missing for investigation. In this paper, we combine new techniques in convolutional neural network and random forests to a Hierarchical CNN based Random Forests (HCRF) approach for face super-resolution in a coarse-to-fine manner. In the proposed approach, we focus on a general approach that can handle facial images with various conditions without pre-processing. To the best of our knowledge, this is the first paper that combines the advantages of deep learning with random forests for face super-resolution. To achieve superior performance, we propose two novel CNN models for coarse facial image super-resolution and segmentation and then apply new random forests to target on local facial features refinement making use of the segmentation results. Extensive benchmark experiments on subjective and objective evaluation show that HCRF can achieve comparable speed and competitive performance compared with state-of-the-art super-resolution approaches for very low-resolution images. Wan-Chi Siu, Yui-Lam Chan |
IEEE Trans. Image Process. | 3 |
| 2020 | Low-Complexity Intra Prediction for Screen Content Coding by Convolutional Neural NetworkabstractScreen content coding (SCC) is developed to encode screen content videos, and it is an extension of High Efficiency Video Coding (HEVC). Since screen content videos contain computer-generated content that shows special characteristics, SCC adopts the new Intra Block Copy mode and Palette mode besides the HEVC based Intra mode to improve the coding efficiency. However, the exhaustive mode searching process makes the SCC encoder computational expensive. In this paper, a low-complexity intra prediction algorithm is proposed by the convolutional neural network (CNN). The proposed network skips unnecessary coding units (CUs) and mode candidates by imitating the behavior of the original SCC encoder. The network first decides if a CU size should be checked by analyzing global features, and it decides which mode should be checked by analyzing the local features. Experimental results show that the proposed algorithm achieves 53.44% computational complexity reduction on average with 1.94% Bjentegaard delta bitrate loss under All Intra configuration. Wei Kuang, Yui-Lam Chan, Sik-Ho Tsang |
ISCAS | 2 |
| 2020 | 360-Degree Intra Coding Mode for Equirectangular Projection Format VideosabstractRecent advances in display, networking, and computing technologies have resulted in changing industry focus towards 360-degree/omnidirectional images as witnessed by increased interest in virtual reality (VR) and augmented reality (AR). Numerous curves are generated for 360-degree images due to the lens curvature and projection format. However, in High Efficiency Video Coding (HEVC), the conventional angular intra modes cannot handle them well since only straight lines can be predicted. More enhanced 360-degree image coding is essential for higher efficiency of storage and transmission. Therefore, we propose a new coding mode, called 360-degree intra mode, for predicting the coding units with curves. Experimental results show that our 360-degree intra mode improves the coding efficiency of HEVC by 0.32% on average and up to 0.74% Bjontegaard delta bit rate reduction. Sik-Ho Tsang, Yui-Lam Chan |
ISCAS | 2 |
| 2020 | FastSCCNet: Fast Mode Decision in VVC Screen Content Coding via Fully Convolutional NetworkabstractScreen content coding have been supported recently in Versatile Video Coding (VVC) to improve the coding efficiency of screen content videos by adopting new coding modes which are dedicated to screen content video compression. Two new coding modes called Intra Block Copy (IBC) and Palette (PLT) are introduced. However, the flexible quad-tree plus multi-type tree (QTMT) coding structure for coding unit (CU) partitioning in VVC makes the fast algorithm of the SCC particularly challenging. To efficiently reduce the computational complexity of SCC in VVC, we propose a deep learning based fast prediction network, namely FastSCCNet, where a fully convolutional network (FCN) is designed. CUs are classified into natural content block (NCB) and screen content block (SCB). With the use of FCN, only one shot inference is needed to classify the block types of the current CU and all corresponding sub-CUs. After block classification, different subsets of coding modes are assigned according to the block type, to accelerate the encoding process. Compared with the conventional SCC in VVC, our proposed FastSCCNet reduced the encoding time by 29.88% on average, with negligible bitrate increase under all-intra configuration. To the best of our knowledge, it is the first approach to tackle the computational complexity reduction for SCC in VVC. Sik-Ho Tsang, Ngai-Wing Kwong, Yui-Lam Chan |
VCIP | 3 |
| 2020 | Overview of current development in depth map coding of 3D video and its futureabstract3D videos have attracted attention from academia and industry after great success in the film industry. Multiview video plus depth (MVD) is the most popular 3D video format to provide vivid 3D feeling and has been adopted as an international 3D video coding standard, namely 3D extension of high efficiency video coding (HEVC). MVD includes a limited number of textures and depth maps to synthesise virtual views. In MVD, depth samples describe the distance between a camera and an actual object as a grey‐level image. The characteristics of depth maps are quite different from texture images. Consequently, new coding tools are designed for depth maps in 3D‐HEVC to improve the coding efficiency at the expense of high‐computational complexity, which faces great challenges in coding systems. Depth map coding is also an important technique in immersive media to support three degrees of freedom 3DoF+/6DoF applications such as virtual reality/augmented reality. The study starts with an overview of what has been done over the last decade in 3D‐HEVC, especially depth map coding, regarding theories, methodologies, current research and state‐of‐the‐art fast approaches. Following this, a comprehensive comparison of the reviewed techniques is presented, and an outlook on its future trends is provided. Yui-Lam Chan, Chang-Hong Fu 0002, Hao Chen 0043, Sik-Ho Tsang |
IET Signal Process. | 1 |
| 2020 | Early termination for fast intra mode decision in depth map coding using DIS-inheritance
Chang-Hong Fu 0002, Hao Chen 0043, Yui-Lam Chan, Sik-Ho Tsang, Xiaohua Zhu 0001 |
Signal Process. Image Commun. | 3 |
| 2020 | Machine Learning-Based Fast Intra Mode Decision for HEVC Screen Content Coding via Decision TreesabstractThe screen content coding (SCC) extension of high efficiency video coding (HEVC) improves coding gain for screen content videos by introducing two new coding modes, namely, intra block copy (IBC) and palette (PLT) modes. However, the coding gain is achieved at the increased cost of computational complexity. In this paper, we propose a decision tree-based framework for fast intra mode decision by investigating various features in the training sets. To avoid the exhaustive mode searching process, a sequential arrangement of decision trees is proposed to check each mode separately by inserting a classifier before checking a mode. As compared with the previous approaches where both IBC and PLT modes are checked for screen content blocks (SCBs), the proposed coding framework is more flexible which facilitates either the IBC or PLT mode to be checked for SCBs such that computational complexity is further reduced. To enhance the accuracy of decision trees, dynamic features are introduced, which reveal the unique intermediate coding information of a coding unit (CU). Then, if all the modes are decided to be skipped for a CU at the last depth level, at least one possible mode is assigned by a CU-type decision tree. Furthermore, a decision tree constraint technique is developed to reduce the rate-distortion performance loss. Compared with the HEVC-SCC reference software SCM-8.3, the proposed algorithm reduces computational complexity by 47.62% on average with a negligible Bjøntegaard delta bitrate (BDBR) increase of 1.42% under all-intra (AI) configurations, which outperforms all the state-of-the-art algorithms in the literature. Wei Kuang, Yui-Lam Chan, Sik-Ho Tsang, Wan-Chi Siu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | DeepSCC: Deep Learning-Based Fast Prediction Network for Screen Content CodingabstractScreen content coding (SCC) is an extension of high efficiency video coding (HEVC), and it is developed to improve the coding efficiency of screen content videos by adopting two new coding modes: Intra Block Copy (IBC) and Palette (PLT). However, the flexible quadtree-based coding tree unit (CTU) partitioning structure and various mode candidates make the fast algorithms of the SCC extremely challenging. To efficiently reduce the computational complexity of SCC, we propose a deep learning-based fast prediction network DeepSCC that contains two parts: DeepSCC-I and DeepSCC-II. Before feeding to DeepSCC, incoming coding units (CUs) are divided into two categories: dynamic CTUs and stationary CTUs. For dynamic CTUs having different content as their collocated CTUs, DeepSCC-I takes raw sample values as the input to make fast predictions. For stationary CTUs having the same content as their collocated CTUs, DeepSCC-II additionally utilizes the optimal mode maps of the stationary CTU to further reduce the computational complexity. Compared with the HEVC-SCC reference software SCM-8.3, the proposed DeepSCC reduces the encoding time by 48.81% on average with a negligible Bjøntegaard delta bitrate increase of 1.18% under all-intra configuration. Wei Kuang, Yui-Lam Chan, Sik-Ho Tsang, Wan-Chi Siu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Online-Learning-Based Bayesian Decision Rule for Fast Intra Mode and CU Partitioning Algorithm in HEVC Screen Content CodingabstractScreen content coding (SCC) is an extension of high efficiency video coding by adopting new coding modes to improve the coding efficiency of SCC at the expense of increased complexity. This paper proposes an online-learning approach for fast mode decision and coding unit (CU) size decision in SCC. To make a fast mode decision, the corner point is first extracted as a unique feature in screen content, which is an essential pre-processing step to guide Bayesian decision modeling. Second, the distinct color number in a CU is derived as another unique feature in screen content to build the precise model using online-learning for skipping unnecessary modes. Third, the correlation of the modes among spatial neighboring CUs is analyzed to further eliminate unnecessary mode candidates. Finally, the Bayesian decision rule using online-learning is applied again to make a fast CU size decision. To ensure the accuracy of the Bayesian decision models, new scene change detection is designed to update the models. Results show that the proposed algorithm achieves 36.69% encoding time reduction with 1.08% Bjøntegaard delta bitrate (BDBR) increment under all intra configuration. By integrating into the existing fast SCC approach, the proposed algorithm reduces 48.83% encoding time with a 1.78% increase in BDBR. Wei Kuang, Yui-Lam Chan, Sik-Ho Tsang, Wan-Chi Siu |
IEEE Trans. Image Process. | 2 |
| 2019 | Mode Skipping for HEVC Screen Content Coding via Random ForestabstractScreen content coding (SCC) is the extension to high-efficiency video coding (HEVC) for compressing screen content videos. New coding tools, intrablock copy (IBC), and palette (PLT) modes, are introduced to encode screen content (SC) such as texts and graphics. The IBC mode is used for encoding repeating patterns by performing block matching within the same frame, while the PLT mode is designed for SC with few distinct colors by coding the major colors and their corresponding locations using an index map. However, the use of IBC and PLT modes increases the encoder complexity remarkably though coding efficiency can be improved. Therefore, we propose to have a mode skipping approach to reduce the encoder complexity of SCC by making use of SC characteristics, neighbor coding unit (CU) correlations, and intermediate cost information via random forest (RF). Detailed feature analyses and sample preparation are also described. A novel hyperparameter tuning approach with the consideration of coding bitrate and encoding time is proposed for RFs at each CU size to further boost the encoding process. Experimental results show that our proposed approach can obtain 45.06% average encoding time reduction with only a 1.08% increase in Bjøntegaard delta bitrate. Average encoding time can even be reduced to 58.57% by regulating the hyperparameters. Sik-Ho Tsang, Yui-Lam Chan, Wei Kuang |
IEEE Trans. Multim. | 2 |
| 2019 | Reduced-Complexity Intra Block Copy (IntraBC) Mode With Early CU Splitting and Pruning for HEVC Screen Content CodingabstractA screen content coding (SCC) extension to high efficiency video coding has been developed to incorporate many new coding tools in order to achieve better coding efficiency for videos mixed with camera-captured content and graphics/text/animation. For instance, the Intra Block Copy (IntraBC) mode helps to encode repeating patterns within the same frame while the Palette mode aims at encoding screen content with a few major colors. However, the IntraBC mode brings along high computational complexity due to the exhaustive block matching within the same frame though there are already some constraints and fast approaches applied to the IntraBC mode to reduce its complexity. Thus, we propose a fast intracoding scheme to reduce the complexity of using the IntraBC mode in SCC. Screen content always contains no sensor noise resulting in the characteristics with pixel exactness along both horizontal and vertical directions. These characteristics pave the way for mode skipping and early coding unit (CU) splitting. Besides, early CU pruning and early termination are proposed based on the rate distortion cost to further reduce encoder complexity. Moreover, we also propose reducing the complexity of the IntraBC mode by checking the hash value of each block candidate and the current block during block matching. With our proposed scheme, the encoding time is reduced compared with the SCC while the coding efficiency can still be maintained with a minor increase in the bjontegaard delta bitrate. Sik-Ho Tsang, Yui-Lam Chan, Wei Kuang, Wan-Chi Siu |
IEEE Trans. Multim. | 2 |
| 2018 | Early Intra Block Partition Decision for Depth Maps in 3D-HEVCabstractIn the three-dimensional extension of the high efficiency video coding standard (3D-HEVC), the optimal depth intra coding block partition structure with smallest rate-distortion (RD) cost is decided after recursively checking all possible partition levels (0-3). The block partition method, together with all intra modes for depth maps at each partition level, dramatically increases the computational complexity of depth intra coding. In this paper, we propose an early termination method for intra block partitioning, in which the termination decision could be decided in a simplified form associated with the current block and its first sub-block. Simulation results demonstrate that the proposed algorithm could save 45.10% of depth coding time with almost no increase in BDBR. Hao Chen 0043, Chang-Hong Fu 0002, Yui-Lam Chan, Xiaohua Zhu 0001 |
ICIP | 3 |
| 2018 | Fast HEVC to SCC Transcoding Based on Decision TreesabstractScreen Content Coding (SCC) is an extension of the High-Efficiency Video Coding (HEVC) for encoding screen content videos. However, there are many legacy screen content videos already encoded by HEVC. To efficiently migrate screen content videos from the existing HEVC to the emerging SCC, a machine learning based fast transcoding algorithm is proposed by using decision trees in this paper. To speed up the transcoding process, the intermediate data from both the HEVC decoder side and the SCC encoder side are jointly analyzed. Then the optimal coding unit (CU) sizes are mapped from HEVC to SCC while the mode candidates are adaptively checked according to the decision tree outcomes in the re-encoding process. Experimental results show that an average of 48.20% re-encoding time reduction is achieved with only 1.47% Bjontegaard delta bitrate loss using All Intra (AI) configuration. Wei Kuang, Yui-Lam Chan, Sik-Ho Tsang, Wan-Chi Siu |
ICME | 2 |
| 2018 | Probability-Based Depth Intra-Mode Skipping Strategy and Novel VSO Metric for DMM Decision in 3D-HEVCabstractMultiview video plus depth format has been adopted as the emerging 3D video representation recently. It includes a limited number of textures and depth maps to synthesize additional virtual views. Since the quality of depth maps influences the view synthesis process, their sharp edges should be well preserved to avoid mixing foreground with background. To address this issue, 3D-High Efficiency Video Coding (HEVC) introduces new coding tools, a partition-based intra mode [depth modeling mode (DMM)], a residual description technique [segmentwise depth coding (SDC)], and a more complex rate-distortion (RD) evaluation with view synthesis optimization (VSO), to provide more accurate predictions and achieve higher compression rate. However, these new techniques introduce a lot of possible candidates, and each of them requires complicated RD calculation in the process of intra-mode decision. They lead to unacceptable computational burden in a 3D-HEVC encoder. Therefore, in this paper, we raise two efficient techniques for depth intra-mode decision. First, by investigating the statistical characteristics of variance distributions in the two partitions of DMM, a simple but efficient criterion based on the squared Euclidean distance of variances (SEDV) is suggested to evaluate RD costs of the DMM candidates instead of the time-consuming VSO process. Second, a probability-based early depth intra-mode decision is proposed to select only the most promising mode and make the early determination of using SDC based on the low-complexity RD cost in rough mode decision. Experimental results show that the proposed algorithm with these two new techniques provides 33%-48% time reduction with little drop in the coding performance compared with the state-of-the-art algorithms. Hongbin Zhang 0005, Chang-Hong Fu 0002, Yui-Lam Chan, Sik-Ho Tsang, Wan-Chi Siu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | Depth modelling mode decision for depth intra coding via good featureabstractThe depth modelling modes (DMM) and 35 conventional intra modes (CHIMs) introduced in 3D-HEVC results in unacceptable huge complexity of depth intra coding. However, some redundancy between DMM and CHIMs could be avoided to accelerate the process. In this paper, a good feature-corner point (CP) is proposed to evaluate the orientation of edge in a given prediction unit (PU), by which a binary classifier is created. We further investigate the probability distribution of DMM, which is selected as the optimal intra mode in each category. According to the statistical analysis, the skipping of DMM decision is proposed to eliminate the cases which have been predicted well by CHIMs. The experimental results show that, compared with the test model HTM-13.0 of 3D-HEVC, the proposed algorithm can yield about 17% time reduction for depth intra coding with almost no degradation in coding performance. Chang-Hong Fu 0002, Ya-Wen Zhao, Hongbin Zhang 0005, Yui-Lam Chan, Wan-Chi Siu |
ICIP | 4 |
| 2017 | Fast mode decision algorithm for HEVC screen content intra codingabstractScreen Content coding (SCC) is one of an extension to High Efficiency Video Coding (HEVC) developed by the Joint Collaborative Team on Video Coding (JCT-VC). It adopts two new coding tools, intra block copy (IBC) and palette (PLT) modes, to improve the compression performance for intra coding. Nevertheless, mode selection causes a substantial increase in encoding complexity. In this paper, a fast mode decision algorithm, which makes use of early mode skip decision based on the Bayesian decision rule using online learning, is proposed. The proposed algorithm is implemented in the SCC reference software SCM-7.0. Experimental results show that the proposed algorithm can achieve 23.2% complexity reduction on average with only 0.58% Bjontegaard delta bitrate loss in All Intra (AI) configurations. Wei Kuang, Sik-Ho Tsang, Yui-Lam Chan, Wan-Chi Siu |
ICIP | 3 |
| 2017 | Decoder side merge mode and AMVP in HEVC screen content codingabstractIntra Block Copy (IBC) mode in a screen content coding (SCC) extension in High Efficiency Video Coding (HEVC) provides high coding gain by performing motion estimation (ME) and motion compensation (MC) to find the repetitive patterns within the same frame. Merge mode and Advanced Motion Vector Prediction (AMVP), which are originally used for inter mode, are also applied to the IBC mode. However, there are redundant coding bits when they are applied to IBC. Therefore, we propose decoder-side merge mode and AMVP for IBC in SCC so as to remove the redundancy. Experimental shows that the proposed method can achieve up to 0.25% Bjontegaard delta bitrate (BD-rate) reduction compared to the conventional SCC with negligible impact to encoding and decoding complexity. Sik-Ho Tsang, Wei Kuang, Yui-Lam Chan, Wan-Chi Siu |
ICIP | 3 |
| 2017 | Adaptive search range by depth variant decaying weights for HEVC inter texture codingabstractEmerging high-efficiency video coding (HEVC) outperforms H.264 by a gain of 50% bitrate reduction while maintaining almost the same perceptual quality. However, it induces higher coding complexity due to its adoption of recursive block partitioning mechanism in motion estimation (ME) with a fixed search range. For an objective of reducing the computational burden in HEVC, this paper proposes an adaptive search range algorithm by using depth map information. With the aid of depth intensity variations among neighboring blocks, associated weights to the neighboring blocks are derived. The proposed weighted sum of the motions from the neighboring blocks is formulated to provide a suitable search range for each block. The simulation results demonstrated that proposed adaptive search range is compatible to not only full-search (FS) but also fast Test Zone Search (TZS) in HEVC. The proposed algorithm could reduce significant coding time on average with negligible rate-distortion degradation. Tsz-Kwan Lee, Yui-Lam Chan, Wan-Chi Siu |
ICME | 2 |
| 2017 | Depth-projected determination for adaptive search range in motion estimation for HEVCabstractHigh Efficient Video Coding (HEVC) improves coding efficiency but suffers from high computational complexity due to its quad-tree partitioning structure in motion estimation (ME). In recent development of 3D video technology, depth map from the 3D video provides an intimation of the objects' distance from the projected screen in a 3D scene, which inspires the authors to explore the adaptive search range determination for complexity reduction in HEVC. The proposed algorithm exploits the high temporal correlation between the depth map and the motion in texture. By utilizing this correlation and the potential impact of 3D-to-2D projection, a depth/motion relationship is built for a tailor-made search range with a depth-projected scale factor to skip unnecessary search points in ME. Besides, the proposed ASR algorithm can work well with other fast ME algorithms with up to 53% of average coding time reduction whereas the coding efficiency can be maintained. Tsz-Kwan Lee, Yui-Lam Chan, Wan-Chi Siu |
ISCAS | 2 |
| 2017 | Fast image super-resolution via Randomized Multi-split ForestsabstractThis paper proposes a novel learning-based image Super-Resolution via a Randomized Multi-split Forests model (SRRMF). The proposed method uses the LR-HR training patch pairs to model the nonlinear patch manifold into a pairs of linear subspaces. The key idea of this approach is to use several decision trees split randomly the training data into different classes. A linear regression model is learnt to map the relationship between LR and HR patches at the end of the leaf nodes. In order to make full use of the generalization ability of the random forests, we randomize the grow of the decision tree to cover more possibilities. Furthermore, we modify the splitting function by using Multi-Split Binary Test (MSBT) function so that we can use more feature information to derive more accurate classification result to match patch subspace. Extended experimental results show that image super-resolution using our proposed method can achieve the state-of-the-art super-resolution performance with reduced computation time. Wan-Chi Siu, Yui-Lam Chan |
ISCAS | 3 |
| 2017 | Segment-based view synthesis optimization scheme in 3D-HEVC
Huan Dou, Yui-Lam Chan, Kebin Jia, Wan-Chi Siu |
J. Vis. Commun. Image Represent. | 2 |
| 2017 | Adaptive Search Range for HEVC Motion Estimation Based on Depth InformationabstractHigh Efficiency Video Coding achieves twofold coding efficiency improvement compared with its predecessor H.264/MPEG-4 Advanced Video Coding. However, it suffers from high computational complexity due to its quad-tree structure in motion estimation (ME). This paper exposes the use of depth maps in the multiview video plus depth format for relieving the computational burden. The depth map provides an intimation of the objects' distance from the projected screen in a 3D scene, which is explored in adaptive search range determination in this paper. The proposed algorithm exploits the high temporal correlation between the depth map and the motion in texture. By utilizing this correlation, a depth/motion relationship map is built for a mapping process. For each block, this forms a tailor-made search range with a motion-aware asymmetric shape to skip unnecessary search points in ME. The obtained search range can be further adjusted by taking the influence of 3D-to-2D projection into consideration. Simulation results reveal that, compared to the full search approach, the proposed algorithm can reduce the complexity by 93% on average, whereas the coding efficiency can be maintained. Besides, the proposed search range determination can work well with other fast search ME algorithms in the literature. Tsz-Kwan Lee, Yui-Lam Chan, Wan-Chi Siu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | Quadtree decision for depth intra coding in 3D-HEVC by good featureabstract3D-HEVC is a good coding solution for multi-view video plus depth data. It achieves good coding performance of synthesized views. However, depth intra coding brings unbearable complexity, which is the most urgent issue to be solved for the practical applications. Typically, depth maps have a good feature of structure or less texture compared with natural videos. Therefore, in this paper, a fast depth intra coding algorithm is proposed to speed up the quadtree decision by the good feature-corner point (CP). The proposed algorithm can adaptively extract CPs and preallocate the depth level of coding quadtree. The large size of coding units (CUs) can be skipped for blocks, which have higher predicted depth level. On the contrary, the blocks, with lower predicted depth level, do not check the smaller size of CUs. Simulation results show that the proposed algorithm can provide about 41% time reduction while maintaining the BD performance. Hongbin Zhang 0005, Yui-Lam Chan, Chang-Hong Fu 0002, Sik-Ho Tsang, Wan-Chi Siu |
ICASSP | 2 |
| 2015 | View synthesis optimization based on texture smoothness for 3D-HEVCabstractThis paper presents view synthesis optimization for 3D-HEVC based on a new texture smoothness process. In the original method, all pixels are exhaustively rendered to get distortions from synthesized views. Since not all pixels from the distorted depth map may cause distortions in the synthesized view, it brings unnecessary coding complexity. In this paper, lines of pixels are skipped based on the analysis of pixel regularity from smooth texture regions. It is due to the fact that the distorted disparity may not have much effect on the synthesized view in smooth texture regions. The proposed method can reduce the coding complexity of view synthesis optimization without significant performance loss. Huan Dou, Yui-Lam Chan, Kebin Jia, Wan-Chi Siu |
ICASSP | 2 |
| 2015 | Fast and efficient intra coding techniques for smooth regions in screen content coding based on boundary prediction samplesabstractThis paper presents fast and efficient intra prediction algorithms for screen content coding (SCC). The proposed algorithms focus on smooth regions frequently appeared in screen content videos, which have the characteristics of noiselessness. All the samples in a noiseless smooth region exhibit exactly the same pixel value. We then propose two intra coding techniques for noiseless smooth regions in SCC based on the smoothness of the boundary samples which are used for intra prediction. Our proposed algorithm can reduce computational complexity by at most 26.7% while keeping nearly the same video quality. Moreover, by removing the redundant coding bits for intra prediction modes, computational complexity can be further reduced to at most 53.3% in terms of encoding time with bitrate reduction up to 1.2%. Sik-Ho Tsang, Yui-Lam Chan, Wan-Chi Siu |
ICASSP | 2 |
| 2015 | Efficient depth intra mode decision by reference pixels classification in 3D-HEVCabstractThe uniform intra prediction increases the intra prediction modes up to 35 and brings better coding efficiency in HEVC. Besides, depth modelling modes (DMMs) are introduced in depth intra coding of 3D-HEVC to preserve sharp edges and avoid ringing artifacts in a synthesized view. Meanwhile, the encoding time of depth intra coding rapidly increases due to a huge number of intra mode candidates. Based on the spatial correlation of a depth map, we find that not all of the intra modes are necessary to be considered in most cases. Hence, a fast content-dependent depth intra mode decision algorithm is raised in this paper by classifying the spatial distribution of the reference pixels. Simulation results show that the proposed adaptive fast algorithm can save 21%-35% time of the depth coding with the insignificant bit rate increase compared with the state-of-the-art algorithm. Hongbin Zhang 0005, Chang-Hong Fu 0002, Yui-Lam Chan, Sik-Ho Tsang, Wan-Chi Siu |
ICIP | 3 |
| 2014 | Establishment of linkages across GOP boundaries for reverse playback on compressed video
Chang-Hong Fu 0002, Yui-Lam Chan, Wan-Chi Siu |
Signal Process. Image Commun. | 2 |
| 2013 | Region-based weighted prediction algorithm for H.264/AVC video codingabstractThis paper proposes a novel region-based weighted prediction (WP) algorithm to encode scenes with complex brightness variations. It facilitates the use of multiple WP parameter sets in a single reference frame by utilizing the framework of multiple reference frame motion estimation (MRF-ME). With this arrangement, different macroblocks in the current frame can use different WP parameter sets even when they are predicted from the same reference frame. To support this, a region partitioning process is designed to divide the current frame into different regions where each one has some degree of uniformity in its brightness variation. Multiple sets of region-based WP parameters can then be estimated accurately. Consequently, the proposed algorithm can improve prediction in scenes with different degrees of brightness variations in different regions of the same picture. Results show that the region-based algorithm can achieve significant coding gains of scenes with complex brightness variations. Sik-Ho Tsang, Tsz-Kwan Lee, Yui-Lam Chan, Wan-Chi Siu |
ISCAS | 3 |
| 2013 | Motion estimation in low-delay hierarchical p-frame coding using motion vector composition
Tsz-Kwan Lee, Yui-Lam Chan, Wan-Chi Siu |
J. Vis. Commun. Image Represent. | 2 |
| 2013 | Region-Based Weighted Prediction for Coding Video With Local Brightness VariationsabstractThis paper presents a new region-based scheme for the estimation of weighted prediction (WP) parameter sets for encoders of the H.264/MPEG-4 AVC standard. The proposed scheme is specifically designed for handling local brightness variations (LBVs) in video scenes. It is achieved by making use of multiple WP parameter sets for various regions and assigning them to the same reference frame. An accurate estimation of multiple WP parameter sets is accomplished by: 1) partitioning regions with a simple WP parameter estimator; 2) selecting regions where WP should be applied; and 3) estimating accurate WP parameter sets with a quasioptimal WP parameter estimator. The multiple WP parameter sets of different regions are encoded using the framework of multiple reference frames in the H.264/MPEG-4 AVC standard. With this arrangement, the proposed scheme is compliant with the H.264/MPEG-4 AVC standard. To reduce the implementation cost, a reduction of the memory requirement is realized via look-up tables (LUTs). Experimental results show that the region-based scheme can efficiently handle scenes with global and LBVs and achieve significant coding gain over other WP schemes. Furthermore, our scheme with LUTs can reduce the memory requirement by about 80% while keeping the same coding efficiency as that without LUTs. Sik-Ho Tsang, Yui-Lam Chan, Wan-Chi Siu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2012 | Iterative search strategy with selective bi-directional prediction for low complexity multiview video coding
Zhipin Deng, Yui-Lam Chan, Kebin Jia, Chang-Hong Fu 0002, Wan-Chi Siu |
J. Vis. Commun. Image Represent. | 2 |
| 2012 | Flash scene video coding using weighted prediction
Sik-Ho Tsang, Yui-Lam Chan, Wan-Chi Siu |
J. Vis. Commun. Image Represent. | 2 |
| 2012 | Hybrid motion estimation scheme for secondary SP-frame coding using inter-frame correlation and FMO
Ki-Kit Lai, Yui-Lam Chan, Chang-Hong Fu 0002, Wan-Chi Siu |
Signal Process. Image Commun. | 2 |
| 2011 | Fast iterative search for motion and disparity estimation in stereoscopic video codingabstractIn this paper, a fast algorithm is proposed to speed up the motion and disparity estimation in stereoscopic video coding. Based on the stereo-motion consistency constraint, an iterative search strategy is suggested to get the motion and disparity vectors simultaneously. A credible base vector selection scheme and an adaptive search range adjustment technique are designed to further strengthen the iterative search. Results show that the complexity can be significantly reduced compared to the JMVM full search with a negligible quality drop. Zhipin Deng, Kebin Jia, Yui-Lam Chan, Chang-Hong Fu 0002, Wan-Chi Siu |
VCIP | 3 |
| 2011 | A fast stereoscopic video coding algorithm based on JMVM
Zhipin Deng, Kebin Jia, Yui-Lam Chan, Chang-Hong Fu 0002, Wan-Chi Siu |
Sci. China Inf. Sci. | 3 |
| 2010 | New frame type for view access in MVCabstractThis paper proposes a new frame-type, PSB-frame, for view access in multi-view video coding (MVC). With this new frame-type, a corresponding SI-frame that has been supported by the H.264 extended profile can work with the PSB-frame to provide an immediate access point on B-frames. This new PSB-frame facilities some recent applications such as view access of an multi-view bitstream. In this paper, we provide the arrangements for encoding and decoding mechanisms of a PSB-frame and describe how PSB-frames can be used in MVC bitstreams for view access. Simulation results show that the solution with the new frame-type outperforms the existing approaches. Chang-Hong Fu 0002, Yui-Lam Chan, Ki-Kit Lai |
ICIP | 2 |
| 2010 | H.264 video coding with multiple weighted prediction modelsabstractWeighted prediction is a video coding tool to encode scenes with brightness variations. However, no single WP model works well for all types of brightness variations. In this paper, a novel single reference frame multiple WP models (SRefMWP) scheme is proposed to facilitate the use of multiple WP models in different macroblocks of the current frame even when they are predicted from the same reference. It provides this feature by making a new arrangement of the multiple frame buffers in multiple reference frame motion estimation. Experimental results show that the proposed SRefMWP can improve prediction in scenes with different types of brightness variations, and even benefit to scenes that contain local brightness variation. Sik-Ho Tsang, Yui-Lam Chan |
ICIP | 2 |
| 2010 | A new motion vector composition algorithm for fast-forward video playback in H.264abstractWith the rapid growth of streaming digital videos, it is desirable to access video segments of interest by searching through the video contents with a faster speed than a normal playback. Fast-forward playback is the key function that enables quick browsing of videos. It can be realized by a frame-skipping transcoder which transcodes only the frames required for playback at the desired fast speed. Various motion vector (MV) composition algorithms aim at reducing the computational complexity of the transcoder. They only perform fairly in limited skipped frames scenarios. In this paper, a new vector selection algorithm is proposed to compose a new motion vector (MV) from a set of candidate MVs for minimizing prediction errors due to a larger frame-skipping factor. Experimental results show that the proposed algorithm can deliver a remarkable improvement on the rate-distortion performance over other algorithms. Tsz-Kwan Lee, Chang-Hong Fu 0002, Yui-Lam Chan, Wan-Chi Siu |
ISCAS | 3 |
| 2010 | Fast motion and disparity estimation for multiview video coding
Zhipin Deng, Kebin Jia, Yui-Lam Chan, Chang-Hong Fu 0002, Wan-Chi Siu |
Frontiers Comput. Sci. China | 3 |
| 2010 | An efficient motion vector composition algorithm for fast-forward playback in a video streaming system
Chang-Hong Fu 0002, Tsz-Kwan Lee, Yui-Lam Chan, Wan-Chi Siu |
J. Vis. Commun. Image Represent. | 3 |
| 2010 | Quantized Transform-Domain Motion Estimation for SP-Frame Coding in Viewpoint Switching of Multiview VideoabstractThe brand-new SP-frame in H.264 facilitates drift-free bitstream switching. Notwithstanding the guarantee of seamless switching, the cost is the bulky size of secondary SP-frames. This induces a significant amount of additional space or bandwidth for storage or transmission. In this paper, our investigation reveals that the size of secondary SP-frames is more severe when the correlation between the two bitstreams becomes smaller. Examples include viewpoint switching in multiview video and bitstream switching in single-view video with complex motion. For this reason, a new motion estimation and compensation technique, which is operated in the quantized-transform (QDCT) domain, is designed for coding secondary SP-frames. Our proposed work aims at keeping the secondary SP-frames as small as possible without affecting the size of primary SP-frames by incorporating QDCT-domain motion estimation and compensation in the secondary SP-frame coding. Simulation results show our proposed scheme overwhelmingly outperforms the conventional pixel-domain motion estimation technique. As a consequence, the size of secondary SP-frames can be reduced remarkably, especially in multiview video and single-view video with complex motion. Ki-Kit Lai, Yui-Lam Chan, Wan-Chi Siu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2009 | Compressed-Domain Techniques for Error-Resilient Video Transcoding Using RPSabstractIn video applications where video sequences are compressed and stored in a storage device for future delivery, the encoding process is typically carried out without enough prior knowledge about the channel characteristics of a network. Error-resilient transcoding plays an important role to provide an addition of resilience to the video data, where or whenever it is needed. Recently, a reference picture selection (RPS) scheme has been adopted in an error-resilient transcoder in order to reduce error effects for the already encoded video bitstream. In this approach, the transcoder learns through a feedback channel about the damaged parts of a previously coded frame and then decides to code the next P-frame not relative to the most recent, but to an older, reference picture, which is known to be error-free in the decoder. One straightforward approach of adopting RPS in error-resilient transcoding is to decode all the P-frames from the previously nearest I-frame to the current transmitted frame which is then re-encoded with a new reference frame; this can create undesirable complexity in the transcoder as well as introduce re-encoding errors. In this paper, some novel techniques are suggested for an effective implementation of RPS in the error-resilient transcoder with the minimum requirement on its complexity. All the proposed techniques will manipulate video data in the compressed domain such that the computational loading of the transcoder is greatly reduced. By utilizing these new compressed-domain techniques, we develop a new structure to handle various types of macroblocks in the transcoder which re-uses motion vectors and prediction errors from the encoded bitstream. Experimental results demonstrate that significant improvements in terms of transcoder complexity and quality of reconstructed video can be achieved by employing our compressed-domain techniques. Yui-Lam Chan, Hoi-Kin Cheung, Wan-Chi Siu |
IEEE Trans. Image Process. | 1 |
| 2008 | Viewpoint switching in multiview videos using SP-framesabstractThe distinguishing feature of multiview video lies in the interactivity, which allows users to select their favourite viewpoint. It switches bitstream at a particular view when necessary instead of transmitting all the views. The new SP-frame in H.264 is originally developed for multiple bit-rate streaming with the support of seamless switching. The SP-frame can also be directly employed in the viewpoint switching of multiview videos. Notwithstanding the guarantee of seamless switching using SP-frames, the cost is the bulky size of secondary SP-frames. This induces a significant amount of additional space or bandwidth for storage or transmission, especially for the multiview scenario. For this reason, a new motion estimation and compensation technique operating in the quantized transform (QDCT) domain is designed for coding secondary SP-frame in this paper. Our proposed work aims at keeping the secondary SP-frames as small as possible without affecting the size of primary SP-frames by incorporating QDCT-domain motion estimation and compensation in the secondary SP-frame coding. Simulation results show that the size of secondary SP-frames can be reduced remarkably in viewpoint switching. Ki-Kit Lai, Yui-Lam Chan, Chang-Hong Fu 0002, Wan-Chi Siu |
ICIP | 2 |
| 2007 | A High Performance Lossless Bayer Image Compression SchemeabstractDemosaicing and compression are generally performed sequentially in most digital cameras. Recent reports show that the compression-first scheme outperforms the conventional demosaicing-first scheme in terms of image quality and complexity. In this paper, an efficient lossless compression scheme for Bayer images is presented. It exploits a context matching technique to rank the neighboring pixels for predicting a pixel. Besides, an adaptive color difference estimation scheme is also proposed to remove the spectral redundancy. Simulation results show that the proposed algorithm can achieve a better compression performance as compared with the existing lossless CFA image coding methods. King-Hong Chung, Yuk-Hee Chan, Chang-Hong Fu 0002, Yui-Lam Chan |
ICIP (2) | 4 |
| 2007 | An Efficient Combined Demosaicing and Zooming Algorithm for Digital CameraabstractColor demosaicing and digital zooming are common processes in digital cameras and they often employ similar interpolation concepts based on the information extracted from the raw sensor data. Realizing them independently is not efficient as separate extraction processes are required. It may also cause inconsistent utilization of the raw sensor data in different stages. This paper presents a low-complexity combined algorithm which directly extracts edge information from raw sensor data and exploits it consistently and efficiently in both demosaicing and zooming. The proposed algorithm can produce zoomed full-color images and zoomed CFA images with outstanding performance as compared with conventional approaches. King-Hong Chung, Yuk-Hee Chan, Chang-Hong Fu 0002, Yui-Lam Chan |
ICIP (4) | 4 |
| 2007 | Efficient Motion Estimation in H.264 Reverse TranscodingabstractIn this paper, we propose a fast reverse motion estimation algorithm with efficient mode decision for reverse transcoding of H.264 bitstream. By analyzing the motion vectors and modes decoded from the forward bitstream, the best mode and motion vector for each backward transcoded macroblock are estimated. A remarkable reduction of computational complexity involved in reverse motion estimation can be achieved by the proposed algorithm with only negligible impact on the rate-distortion performance. Chang-Hong Fu 0002, Yui-Lam Chan, Wan-Chi Siu |
ICIP (5) | 2 |
| 2007 | A Simplified Dual-Bitstream MPEG Video Streaming System with VCR FunctionalitiesabstractNowadays, video playback devices have only limited fast-forward/backward playback and even they cannot provide backward playback. The limitation is due to the motion compensated prediction technique adopted in the MPEG standards. One possible way to support browsing functionalities is to store an additional reverse-encoded bitstream into the server. However, this additional bitstream increases the storage requirement of the video server significantly. In this paper, we exploit the redundancy inherent between the forward and reverse-encoded bitstreams in order to achieve a substantial reduction on the size of the reverse-encoded bitstream. The server accesses and manipulates various macroblocks from the forward and reverse-encoded bistreams to facilitate various browsing operations. Experimental results show that, as compared to the conventional dual-bitstream scheme, the new scheme significantly reduces the storage requirement due to the additional reverse-encoded bitstream. Tak-Piu Ip, Yui-Lam Chan, Chang-Hong Fu 0002, Wan-Chi Siu |
ICIP (6) | 2 |
| 2007 | New Architecture for MPEG Video Streaming System With Backward Playback SupportabstractMPEG digital video is becoming ubiquitous for video storage and communications. It is often desirable to perform various video cassette recording (VCR) functions such as backward playback in MPEG videos. However, the predictive processing techniques employed in MPEG severely complicate the backward-play operation. A straightforward implementation of backward playback is to transmit and decode the whole group-of-picture (GOP), store all the decoded frames in the decoder buffer, and play the decoded frames in reverse order. This approach requires a significant buffer in the decoder, which depends on the GOP size, to store the decoded frames. This approach could not be possible in a severely constrained memory requirement. Another alternative is to decode the GOP up to the current frame to be displayed, and then go back to decode the GOP again up to the next frame to be displayed. This approach does not need the huge buffer, but requires much higher bandwidth of the network and complexity of the decoder. In this paper, we propose a macroblock-based algorithm for an efficient implementation of the MPEG video streaming system to provide backward playback over a network with the minimal requirements on the network bandwidth and the decoder complexity. The proposed algorithm classifies macroblocks in the requested frame into backward macroblocks (BMBs) and forward/backward macroblocks (FBMBs). Two macroblock-based techniques are used to manipulate different types of macroblocks in the compressed domain and the server then sends the processed macroblocks to the client machine. For BMBs, a VLC-domain technique is adopted to reduce the number of macroblocks that need to be decoded by the decoder and the number of bits that need to be sent over the network in the backward-play operation. We then propose a newly mixed VLC/DCT-domain technique to handle FBMBs in order to further reduce the computational complexity of the decoder. With these compressed-domain techniques, the proposed architecture only manipulates macroblocks either in the VLC domain or the quantized DCT domain resulting in low server complexity. Experimental results show that, as compared to the conventional system, the new streaming system reduces the required network bandwidth and the decoder complexity significantly. Chang-Hong Fu 0002, Yui-Lam Chan, Tak-Piu Ip, Wan-Chi Siu |
IEEE Trans. Image Process. | 2 |
| 2007 | On Transcoding a B-Frame to a P-Frame in the Compressed DomainabstractOnly a limited number of methods have been proposed to realize heterogeneous transcoding, for example from MPEG-2 to H.263, or from H.264 to H.263. The major difficulties of transcoding a B-picture to a P-picture are that the incoming discrete cosine transform (DCT) coefficients of the B-frame are prediction errors arising from both forward and backward predictions, whilst the prediction errors in the DCT domain arising from the prediction using the previous frame alone are not available. The required new prediction errors need to be re-estimated in the pixel domain. This process involves highly complex computation and introduces re-encoding errors. We propose a new approach to convert a B-picture into a P-picture by making use of some properties of motion compensation in the DCT domain and the direct addition of DCT coefficients. We derive a set of equations and formulate the problem of how to obtain the DCT coefficients. One difficulty is that the last P-frame inside a GOP with an IBBP structure, for example, needs to be transcoded to become the last P-frame in the IPPP structure, and it has to be linked to the previous reconstructed P-frame instead of to the I-frame. We increased the speed of the transcoding process by making use of the motion activity which is expressed in terms of the correlation between pictures. The whole transcoding process is done in the transform domain, hence re-encoding errors are completely avoided. Results from our experimental work show that the proposed video transcoder not only achieves a speed-up of two to six times that of the conventional video transcoder, but it also substantially improves the quality of the video. Wan-Chi Siu, Yui-Lam Chan, Kai-Tat Fung |
IEEE Trans. Multim. | 2 |
| 2006 | Adopting SP/SI-Frames in Dual-Bitstream Video Streaming with VCR SupportabstractDigital video cassette recording (VCR) operations such as fast-forward and fast-reverse playbacks enable quick and user-friendly browsing of video. However, the predictive techniques adopted in current video standards severely complicate these operations. One approach to implement the fast-forward/reverse playback is to store an additional reverse-encoded bitstream into the server. Once a client requests a fast-forward/reverse operation, the server can select an appropriate frame for the client from either the forward or reverse-encoded bitstreams to reduce the network traffic and the decoder complexity. Unfortunately, the forward and reverse-encoded bitstreams are encoded separately. The frame that has previously decoded by the client may not be exactly identical to the reference of the current selected frame and the mismatch problem occurs frequently. In this paper, a novel H.264 dual-bitstream scheme aiming at providing fast-forward/reverse playback based on SP/SI-frames is proposed to eliminate mismatch errors during switching between the forward and reverse-encoded bitstreams. As a result, the proposed scheme enhances the performance of the conventional dual-bitstream scheme. Tak-Piu Ip, Yui-Lam Chan, Wan-Chi Siu |
ICASSP (2) | 2 |
| 2006 | Efficient reverse-play algorithms for MPEG video with VCR supportabstractReverse playback is the most common video cassette recording (VCR) function in digital video players and it involves playing video frames in reverse order. However, the predictive processing techniques employed in MPEG severely complicate the reverse-play operation. For displaying single frame during reverse playback, all frames from the previous I-frame to the requested frame must be sent by the server and decoded by the client machine. It requires much higher bandwidth of the network and complexity of the decoder. In this paper, we propose a compressed-domain approach for an efficient implementation of the MPEG video streaming system to provide reverse playback over a network with the minimal requirements on the network bandwidth and the decoder complexity. In the proposed video streaming server, it classifies macroblocks in the requested frame into two categories—backward macroblocks (BMBs) and forward macroblock (FMBs). Two novel MB-based techniques are used to manipulate the necessary MBs in the compressed domain and the server then sends the processed MBs to the client machine. For BMBs, we propose a sign inversion technique, which is operated in the variable length coding (VLC) domain, to reduce the number of MBs that need to be decoded by the decoder and the number of bits that need to be sent over the network in the reverse-play operation. The server also identifies the previous related MBs of FMBs and those related maroblocks coded without motion compensation are then processed by a technique of direction addition of discrete cosine transform (DCT) coefficients to further reduce the computational complexity of the client decoder. With the sign inversion and direct addition of DCT coefficients, the proposed architecture only manipulates MBs either on the VLC domain or DCT domain to achieve the server with low complexity. Experimental results show that, as compared to the conventional system, the new streaming system reduces the required network bandwidth and the decoder complexity significantly. Chang-Hong Fu 0002, Yui-Lam Chan, Wan-Chi Siu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2005 | Reference picture selection in an already MPEG encoded bitstreamabstractReference picture selection (RPS) is the most common error resilience method for robust transmission over lossy networks. However, RPS has been studied for use in real-time encoding, but has not been examined in transmitting an already encoded MPEG bitstream. In this paper, we propose a compressed-domain approach for the efficient implementation of RPS in the pre-encoded MPEG bitstream with minimum requirement on the server complexity. In the proposed algorithm, a novel macroblock-based algorithm is used to adaptively select the necessary macroblocks, manipulate them in compressed-domain and send the processed macroblocks to the receiver. Experimental results show that, as compared to the original RPS, the new algorithm reduces the required server complexity significantly. Hoi-Kin Cheung, Yui-Lam Chan, Wan-Chi Siu |
ICIP (1) | 2 |
| 2005 | New adaptive partial distortion search using clustered pixel matching error CharacteristicabstractIn order to reduce the computation load, many conventional fast block-matching algorithms have been developed to reduce the set of possible searching points in the search window. All of these algorithms produce some quality degradation of a predicted image. Alternatively, another kind of fast block-matching algorithms which do not introduce any prediction error as compared with the full-search algorithm is to reduce the number of necessary matching evaluations for every searching point in the search window. The partial distortion search (PDS) is a well-known technique of the second kind of algorithms. In the literature, many researches tried to improve both lossy and lossless block-matching algorithms by making use of an assumption that pixels with larger gradient magnitudes have larger matching errors on average. Based on a simple analysis, it is found that, on average, pixel matching errors with similar magnitudes tend to appear in clusters for natural video sequences. By using this clustering characteristic, we propose an adaptive PDS algorithm which significantly improves the computation efficiency of the original PDS. This approach is much better than other algorithms which make use of the pixel gradients. Furthermore, the proposed algorithm is most suitable for motion estimation of both opaque and boundary macroblocks of an arbitrary-shaped object in MPEG-4 coding. Ko-Cheung Hui, Wan-Chi Siu, Yui-Lam Chan |
IEEE Trans. Image Process. | 3 |
| 2004 | Improved macroblock-based reverse play algorithm for MPEG video streaming
Chang-Hong Fu 0002, Yui-Lam Chan, Wan-Chi Siu |
ICIP | 2 |
| 2004 | A compressed-domain heterogeneous video transcoder
Wan-Chi Siu, Kai-Tat Fung, Yui-Lam Chan |
ICIP | 3 |
| 2004 | Adaptive partial distortion search for block motion estimation
Yui-Lam Chan, Ko-Cheung Hui, Wan-Chi Siu |
J. Vis. Commun. Image Represent. | 1 |
| 2004 | Low-complexity and high-quality frame-skipping transcoder for continuous presence multipoint video conferencingabstractThis paper presents a new frame-skipping transcoding approach for video combiners in multipoint video conferencing. Transcoding is regarded as a process of converting a previously compressed video bitstream into a lower bitrate bitstream. A high transcoding ratio may result in an unacceptable picture quality when the incoming video bitstream is transcoded with the full frame rate. Frame skipping is often used as an efficient scheme to allocate more bits to representative frames, so that an acceptable quality for each frame can be maintained. However, the skipped frame must be decompressed completely, and should act as the reference frame to the nonskipped frame for reconstruction. The newly quantized DCT coefficients of prediction error need to be recomputed for the nonskipped frame with reference to the previous nonskipped frame; this can create an undesirable complexity in the real time application as well as introduce re-encoding error. A new frame-skipping transcoding architecture for improved picture quality and reduced complexity is proposed. The proposed architecture is mainly performed on the discrete cosine transform (DCT) domain to achieve a low complexity transcoder. It is observed that the re-encoding error is avoided at the frame-skipping transcoder when the strategy of direct summation of DCT coefficients is employed. By using the proposed frame-skipping transcoder and dynamically allocating more frames to the active participants in video combining, we are able to make more uniform peak signal-to-noise ratio (PSNR) performance of the subsequences and the video qualities of the active subsequences can be improved significantly. Kai-Tat Fung, Yui-Lam Chan, Wan-Chi Siu |
IEEE Trans. Multim. | 2 |
| 2003 | An adaptive partial distortion search for block motion estimationabstractFast search algorithms for block motion estimation reduce the set of possible displacements for locating the motion vector. All algorithms produce some quality degradation of the predicted image. To reduce the computational complexity of the full search algorithm without introducing any loss in the predicted image, we propose a Hilbert-grouped partial distortion search algorithm (HGPDS) by grouping the representative pixels based on pixel activities in the Hilbert scan. By using the grouped information and computing the accumulated partial distortion of the representative pixels before that of other pixels, impossible candidates can be rejected sooner and the remaining computation involved in the matching criterion can be reduced remarkably. In addition, we also suggest a smart search strategy which is an excellent complement of the HGPDS to form an efficient partial distortion search algorithm. The new search strategy rearranges the search order such that the most possible candidates are searched first and this rearrangement will increase the probability of early rejection of impossible motion vectors. Simulation results show that the proposed algorithm has a significant computational speed-up and is the fastest when compared to the conventional partial distortion search algorithms. Yui-Lam Chan, Wan-Chi Siu |
ICASSP (3) | 1 |
| 2003 | Fast motion estimation of arbitrarily shaped video objects in MPEG-4
Ko-Cheung Hui, Wan-Chi Siu, Yui-Lam Chan |
Signal Process. Image Commun. | 3 |
| 2002 | Block motion estimation using adaptive partial distortion searchabstractThe conventional search algorithms for block motion estimation reduce the set of possible displacements for locating the motion vector. All of these algorithms produce some quality degradation of the predicted image. To reduce the computational complexity of the full search algorithm without introducing any loss in the predicted image, we propose an adaptive partial distortion search algorithm (APDS) by selecting the most representative pixels with high activities, such as edges and texture which contribute most to the matching criterion. The APDS algorithm groups the representative pixels based on the pixel activities in the Hilbert scan. By using the grouped information and computing the accumulated partial distortion of the representative pixels before that of the other pixels, impossible candidates can be rejected sooner and the remaining computation involved in the matching criterion can be reduced remarkably. Simulation results show that the proposed APDS algorithm has a significant computational speed-up and is the fastest when compared to the conventional partial distortion search algorithms. Yui-Lam Chan, Wan-Chi Siu, Ko-Cheung Hui |
ICME (1) | 1 |
| 2002 | New architecture for dynamic frame-skipping transcoderabstractTranscoding is a key technique for reducing the bit rate of a previously compressed video signal. A high transcoding ratio may result in an unacceptable picture quality when the full frame rate of the incoming video bitstream is used. Frame skipping is often used as an efficient scheme to allocate more bits to the representative frames, so that an acceptable quality for each frame can be maintained. However, the skipped frame must be decompressed completely, which might act as a reference frame to nonskipped frames for reconstruction. The newly quantized discrete cosine transform (DCT) coefficients of the prediction errors need to be re-computed for the nonskipped frame with reference to the previous nonskipped frame; this can create undesirable complexity as well as introduce re-encoding errors. In this paper, we propose new algorithms and a novel architecture for frame-rate reduction to improve picture quality and to reduce complexity. The proposed architecture is mainly performed on the DCT domain to achieve a transcoder with low complexity. With the direct addition of DCT coefficients and an error compensation feedback loop, re-encoding errors are reduced significantly. Furthermore, we propose a frame-rate control scheme which can dynamically adjust the number of skipped frames according to the incoming motion vectors and re-encoding errors due to transcoding such that the decoded sequence can have a smooth motion as well as better transcoded pictures. Experimental results show that, as compared to the conventional transcoder, the new architecture for frame-skipping transcoder is more robust, produces fewer requantization errors, and has reduced computational complexity. Kai-Tat Fung, Yui-Lam Chan, Wan-Chi Siu |
IEEE Trans. Image Process. | 2 |
| 2001 | Dynamic frame skipping for high-performance transcodingabstractTranscoding is a process of converting a previously compressed video bitstream into a lower bit-rate bitstream. When some incoming frames are dropped for the frame-rate conversion in transcoding, the newly quantized DCT coefficients of prediction error need to be re-computed, which can create an undesirable complexity as well as introduce re-encoding error. We propose a new architecture for a frame-skipping transcoder to improve picture quality and to reduce complexity. It is observed that re-encoding error is reduced significantly when the strategy of direct summation of DCT coefficients and the error compensation feedback loop are employed. Furthermore, we propose a frame-rate control scheme which can dynamically adjust the number of skipped frames according to the incoming motion vectors and the re-encoding error due to transcoding such that the decoded sequence can have smooth motion as well as better transcoded pictures. Experimental results show that, as compared to the conventional transcoder, the new frame-skipping transcoder is more robust, produces smaller requantization errors, and has simple computational complexity. Kai-Tat Fung, Yui-Lam Chan, Wan-Chi Siu |
ICIP (1) | 2 |
| 2001 | Priority search technique for MPEG-4 motion estimation of arbitrarily shaped video objectabstractOne of the main differences between MPEG-4 and previously standardized video coding schemes is the support of arbitrarily shaped video objects, for which most of the existing fast motion estimation algorithms are not suitable. The conventional fast motion estimation algorithm works well for opaque macroblocks, but not in the case of a boundary macroblock which contains a large number of local minima on its error surface. We propose a fast search algorithm which incorporates the binary alpha-plane to predict accurately the motion vectors of boundary macroblocks. Besides, these accurate motion vectors can be used to develop a novel priority search algorithm which is an efficient search strategy for the remaining opaque macroblocks. Experimental results show that, compared to the conventional methods, our approach requires a low computational complexity and provides a significant improvement in terms of accuracy in motion-compensated video object planes. Ko-Cheung Hui, Yui-Lam Chan, Wan-Chi Siu |
ICIP (3) | 2 |
| 2001 | An efficient search strategy for block motion estimation using image featuresabstractBlock motion estimation using the exhaustive full search is computationally intensive. Fast search algorithms offered in the past tend to reduce the amount of computation by limiting the number of locations to be searched. Nearly all of these algorithms rely on this assumption: the mean absolute difference (MAD) distortion function increases monotonically as the search location moves away from the global minimum. Essentially, this assumption requires that the MAD error surface be unimodal over the search window. Unfortunately, this is usually not true in real-world video signals. However, we can reasonably assume that it is monotonic in a small neighborhood around the global minimum. Consequently, one simple strategy, but perhaps the most efficient and reliable, is to place the checking point as close as possible to the global minimum. In this paper, some image features are suggested to locate the initial search points. Such a guided scheme is based on the location of certain feature points. After applying a feature detecting process to each frame to extract a set of feature points as matching primitives, we have extensively studied the statistical behavior of these matching primitives, and found that they are highly correlated with the MAD error surface of real-world motion vectors. These correlation characteristics are extremely useful for fast search algorithms. The results are robust and the implementation could be very efficient. A beautiful point of our approach is that the proposed search algorithm can work together with other block motion estimation algorithms. Results of our experiment on applying the present approach to the block-based gradient descent search algorithm (BBGDS), the diamond search algorithm (DS) and our previously proposed edge-oriented block motion estimation show that the proposed search strategy is able to strengthen these searching algorithms. As compared to the conventional approach, the new algorithm, through the extraction of image features, is more robust, produces smaller motion compensation errors, and has a simple computational complexity. Yui-Lam Chan, Wan-Chi Siu |
IEEE Trans. Image Process. | 1 |
| 1999 | Reliable search strategy for block motion estimation by measuring the error surfaceabstractThe conventional search algorithms for block matching motion estimation reduce the set of possible displacements for locating the motion vector. Nearly all of these algorithms rely on the assumption: the distortion function increases monotonically as the search location moves away from the global minimum. Obviously, this assumption essentially requires that the error surface be unimodal over the search window. Unfortunately, this is usually not true in real-world video signals. We formulate a criterion to check the confidence of unimodal error surface over the search window. The proposed confidence measure of error surface, CMES, would be a good measure for identifying whether the searching should continue or not. It is found that this proposed measure is able to strengthen the conventional fast search algorithms for block matching motion estimation. Experimental results show that, as compared to the conventional approach, the new algorithm through the CMES is more robust, produces smaller motion compensation errors, and requires simple computational complexity. Yui-Lam Chan, Wan-Chi Siu |
ICASSP | 1 |
| 1999 | A Feature-Assisted Search Strategy for Block Motion EstimationabstractBlock motion estimation using the exhaustive full search is computationally intensive. Previous fast search algorithms tend to reduce the computation by limiting the number of locations to be searched. Nearly all of these algorithms rely on the assumption: the mean absolute distortion (MAD) function increases monotonically as the search location moves away from the global minimum. Unfortunately, this is usually not true in real-world video signals. However, we can reasonably assume that it is monotonic in a small neighbourhood around the global minimum. Consequently, one simple, but perhaps the most efficient and reliable strategy, is to put the checking point as close as possible to the global minimum. In this paper, some image features are suggested to locate the initial search points. Such a guided scheme is based on the location of some feature points. After a feature detecting process was applied to each frame to extract a set of feature points as matching primitives, we studied extensively the statistical behaviour of these matching primitives and found that they are highly correlated with the MAD error surface of real-world motion vectors. These correlation characteristics are extremely useful for fast search algorithms. The results are robust and the implementation could be very efficient. Yui-Lam Chan, Wan-Chi Siu |
ICIP (2) | 1 |
| 1999 | Reliable block motion estimation through the confidence measure of error surface
Yui-Lam Chan, Wan-Chi Siu |
Signal Process. | 1 |
| 1998 | On Block Motion Estimation Using a Novel Search Strategy for an Improved Adaptive Pixel Decimation
Yui-Lam Chan, Wan-Chi Siu |
J. Vis. Commun. Image Represent. | 1 |
| 1997 | Variable temporal-length 3-D discrete cosine transform codingabstractThree-dimensional discrete cosine transform (3-D DCT) coding has the advantage of reducing the interframe redundancy among a number of consecutive frames, while the motion compensation technique can only reduce the redundancy of at most two frames. However, the performance of the 3-D DCT coding will be degraded for complex scenes with a greater amount of motion. This paper presents a 3-D DCT coding with a variable temporal length that is determined by the scene change detector. Our idea is to let the motion activity in each block be very low, while the efficiency of the 3-D DCT coding could be increased. Experimental results show that this technique is indeed very efficient. The present approach has substantial improvement over the conventional fixed-length 3-D DCT coding and is also better than that of the Moving Picture Expert Group (MPEG) coding. Yui-Lam Chan, Wan-Chi Siu |
IEEE Trans. Image Process. | 1 |
| 1996 | New adaptive pixel decimation for block motion vector estimationabstractA new adaptive technique based on pixel decimation for the estimation of motion vector is presented. In a traditional approach, a uniform pixel decimation is used. Since part of the pixels in each block do not enter into the matching criterion, this approach limits the accuracy of the motion vector. In this paper, we select the most representative pixels based on image content in each block for the matching criterion. This is due to the fact that high activity in the luminance signal such as edges and texture mainly contributes to the matching criterion. Our approach can compensate the drawback in standard pixel decimation techniques. Computer simulations show that this technique is close to the performance of the exhaustive search with significant computational reduction. Yui-Lam Chan, Wan-Chi Siu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1995 | A new block motion vector estimation using adaptive pixel decimationabstractBlock motion estimation is being widely used in video coding. A new adaptive technique based on pixel decimation for estimating motion vector is presented. In the traditional approach, a uniform pixel decimation is used. Since some pixels in each block do not enter into the matching criterion, this approach might limit the accuracy of the motion vector. We select the most representative pixels based on the image content in each block for the matching criterion. This is due to the fact that high activity in the luminance signal such as edges and texture contributes mainly to the matching criterion. Our approach can compensate the drawback in standard pixel decimation techniques. Computer simulations show that this technique is close to the performance of the exhaustive search method with a significant reduction in computational complexity. Yui-Lam Chan, Wan-Chi Siu |
ICASSP | 1 |
| 1995 | Fast Interframe Transfrom Coding Based on Characteristics of Transform Coefficients and Frame DifferenceabstractThe interframe transform coding has been seldom used in practice because of the considerable computational complexity. To reduce the computational complexity, a fast algorithm is proposed which reduces the number of operations by limiting the calculation of transform coefficients without significant quality degradation. Different modes of transformation are performed according to frame difference. This proposed fast interframe coding, like the MPEG, has an asymmetric property with decoding being much faster than encoding. Computer simulations show that this fast algorithm can significantly reduce the computational burden in interframe transform coding and is even faster than MPEG-like coding. Yui-Lam Chan, Wan-Chi Siu |
ISCAS | 1 |
| 1994 | A New Adaptive Interframe Transform Coding using Directional ClassificationabstractInterframe transform coding is affected not only by the statistics of spatial details within a frame, but also by the variation of the amount of movement and other temporal activities in different regions of the image sequence. Therefore, adaptive techniques have to be used in order to achieve good image quality. In this paper, we propose a new version of the adaptive interframe coding method, namely directional classification, which is based on image sequence statistics. Blocks with different perceptual features such as edges and high motion activity are categorised to different classes. Then, a new adaptive quantization, associated with appropriate scanning and Huffman coding, are employed based on the classification map. Coding tests using computer simulation show this technique is indeed very efficient.> Yui-Lam Chan, Wan-Chi Siu |
ICIP (2) | 1 |