Sik-Ho Tsang

dblp:78/8841 · also Harris Sik-Ho Tsang · DBLP profile ↗
← Back
30ranked-venue papers
10as first author
9since 2021 · last 2026
0000-0002-8578-0696ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 25 · 8 first-author · 8 since 2021Systems, architecture and hardware · 3 · 2 first-authorArtificial intelligence and machine learning · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-Branch Aesthetic and Technical Perspectives With Cross Tri-Fusion Attention for No-Reference Audio-Visual Quality Assessment
Ngai-Wing Kwong, Yui-Lam Chan, Ziyin Huang, Sik-Ho Tsang
IEEE Trans. Circuits Syst. Video Technol.4
2025 Long Short-Term Fusion by Multi-Scale Distillation for Screen Content Video Quality Enhancement
abstract
Different from natural videos, where artifacts distributed evenly, the artifacts of compressed screen content videos mainly occur in the edge areas. Besides, these videos often exhibit abrupt scene switches, resulting in noticeable distortions in video reconstruction. Existing multiple-frame models using a fixed range of neighbor frames face challenges in effectively enhancing frames during scene switches and lack efficiency in reconstructing high-frequency details. To address these limitations, we propose a novel method that effectively handles scene switches and reconstructs high-frequency information. In the feature extraction part, we develop long-term and short-term feature extraction streams, in which the long-term feature extraction stream learns the contextual information, and the short-term feature extraction stream extracts more related information from shorter input to assist the long-term stream to handle fast motion and scene switches. To further enhance the frame quality during scene switches, we incorporate a similarity-based neighbor frame selector before feeding frames into the short-term stream. This selector identifies relevant neighbor frames, aiding in the efficient handling of scene switches. To dynamically fuse the short-term feature and long-term features, the muti-scale feature distillation focuses on adaptively recalibrating channel-wise feature responses to achieve effective feature distillation. In the reconstruction part, a high-frequency reconstruction block is proposed for guiding the model to restore the high-frequency components. Experimental results demonstrate the significant advancements achieved by our proposed Long Short-term Fusion by Multi-Scale Distillation (LSFMD) method in enhancing the quality of compressed screen content videos, surpassing the current state-of-the-art methods.
Ziyin Huang, Yui-Lam Chan, Ngai-Wing Kwong, Sik-Ho Tsang, Kin-Man Lam 0001, Bingo Wing-Kuen Ling
IEEE Trans. Circuits Syst. Video Technol.4
2025 Multi-Frame Spatiotemporal Feature and Hierarchical Learning Approach for No-Reference Screen Content Video Quality Assessment
abstract
The rapid adoption of remote work, online conferencing, and shared-screen collaboration has significantly increased the usage of screen content videos (SCVs), creating a growing need for reliable quality assessment to maintain excellent quality of service. While several full-reference SCV quality assessment (SCVQA) methods have been proposed, their practical application is often limited by the unavailability of reference videos. Existing no-reference SCVQA (NR-SCVQA) methods rely on handcrafted features and focus solely on specific distortions and features, potentially limiting their generalization ability. Moreover, they fail to explore the underlying spatiotemporal information of SCVs, which could hinder their performance. In this work, we propose a novel deep learning-based NR-SCVQA model specifically tailored to capture the comprehensive spatiotemporal features of SCVs to overcome these issues and challenges posed by the SCVQA task. Our approach incorporates a dual-channel spatiotemporal convolutional neural network (DCST-CNN) module to extract both content-aware and edge-aware spatiotemporal quality features, which enables an effective spatiotemporal quality feature representation learning for the downstream SCVQA task. Building upon the DCST-CNN, we further propose a Temporal Pyramid Transformer (TPT) module to fuse spatiotemporal features across multiple temporal scales, enabling the model to capture both short-term and long-term temporal dependencies within an SCV for hierarchical learning. The proposed DCST-CNN and TPT modules work together to provide a robust and accurate NR-SCVQA framework. We conduct experiments on SCVQA databases to validate the effectiveness of our model, which outperforms existing state-of-the-art NR-SCVQA method. The results demonstrate the strength and applicability of our approach in real-world SCVQA tasks.
Ngai-Wing Kwong, Yui-Lam Chan, Sik-Ho Tsang, Ziyin Huang, Kin-Man Lam 0001
IEEE Trans. Multim.3
2024 Frame Similarity-Based Screen Content Video Quality Enhancement via Adaptive Long Short-Term Fusion
abstract
Compressed screen content videos often exhibit artifacts in edge areas and suffer from distortions during scene switches, where content abruptly changes between frames. Existing multi-frame models, which use a fixed range of neighbor frames, struggle with these switches. To address this, we propose a novel method that effectively handles scene switches. Our approach utilizes Long-term Feature Extraction (LFE) to capture contextual information, while the Frame Similarity-based Short-term Feature Extraction (FSFE) focuses on texture information to manage fast motion and scene switches. In FSFE, a Similarity-based Neighbor Frame Selector (SNFS) is designed to choose relevant neighbor frames for the short-term stream, enhancing the quality of scene switch frames. To fuse short-term and long-term features adaptively, we introduce a local-spatial and global-channel attention module, which recalibrates spatial and channel-wise feature responses. Experimental results show that our Frame Similarity-Based via Adaptive Long Short-Term Fusion (FSLST) method significantly improves the quality of compressed videos, outperforming current state-of-the-art methods.
Ziyin Huang, Yui-Lam Chan, Ngai-Wing Kwong, Sik-Ho Tsang, Kin-Man Lam 0001, Bingo Wing-Kuen Ling
VCIP4
2024 Solving the imbalanced dataset problem in surveillance image blur classification
Yikun Pan, Sik-Ho Tsang, Tom Tak-Lam Chan, Yui-Lam Chan, Daniel Pak-Kong Lun
Eng. Appl. Artif. Intell.2
2024 Spatio-temporal feature learning for enhancing video quality based on screen content characteristics
Ziyin Huang, Yui-Lam Chan, Sik-Ho Tsang, Ngai-Wing Kwong, Kin-Man Lam 0001, Bingo Wing-Kuen Ling
J. Vis. Commun. Image Represent.3
2024 Spatiotemporal feature learning for no-reference gaming content video quality assessment
Ngai-Wing Kwong, Yui-Lam Chan, Sik-Ho Tsang, Ziyin Huang, Kin-Man Lam 0001
J. Vis. Commun. Image Represent.3
2023 Optimized Quality Feature Learning for Video Quality Assessment
abstract
Recently, some transfer learning-based methods have been adopted in video quality assessment (VQA) to compensate for the lack of enormous training samples and human annotation labels. But these methods induce a domain gap between source and target domains, resulting in a sub-optimal feature representation that deteriorates the accuracy. This paper proposes the optimized quality feature learning via a multi-channel convolutional neural network (CNN) with the gated recurrent unit (GRU) for no-reference (NR) VQA. First, the multi-channel CNN is pre-trained on the image quality assessment (IQA) domain using non-human annotation labels, which is inspired by self-supervised learning. Then, semi-supervised learning is used to fine-tune CNN and transfer the knowledge from IQA to VQA while considering motion information for optimized quality feature learning. Finally, all frame quality features are extracted as the input of GRU to obtain the final video quality. Experimental results demonstrate that our model achieves better performance than state-of-the-art VQA approaches.
Ngai-Wing Kwong, Yui-Lam Chan, Sik-Ho Tsang, Daniel Pak-Kong Lun
ICASSP3
2021 Efficient Depth Intra Frame Coding in 3D-HEVC by Corner Points
abstract
To improve the coding performance of depth maps, 3D-HEVC includes several new depth intra coding tools at the expense of increased complexity due to a flexible quadtree Coding Unit/Prediction Unit (CU/PU) partitioning structure and a huge number of intra mode candidates. Compared to natural images, depth maps contain large plain regions surrounded by sharp edges at the object boundaries. Our observation finds that the features proposed in the literature either speed up the CU/PU size decision or intra mode decision and they are also difficult to make proper predictions for CUs/PUs with the multi-directional edges in depth maps. In this work, we reveal that the CUs with multi-directional edges are highly correlated with the distribution of corner points (CPs) in the depth map. CP is proposed as a good feature that can guide to split the CUs with multi-directional edges into smaller units until only single directional edge remains. This smaller unit can then be well predicted by the conventional intra mode. Besides, a fast intra mode decision is also proposed for non-CP PUs, which prunes the conventional HEVC intra modes, skips the depth modeling mode decision, and early determines segment-wise depth coding. Furthermore, a two-step adaptive corner point selection technique is designed to make the proposed algorithm adaptive to frame content and quantization parameters, with the capability of providing the flexible tradeoff between the synthesized view quality and complexity. Simulation results show that the proposed algorithm can provide about 66% time reduction of the 3D-HEVC intra encoder without incurring noticeable performance degradation for synthesized views and it also outperforms the previous state-of-the-art algorithms in term of time reduction and ∆ BDBR.
Chang-Hong Fu 0002, Yui-Lam Chan, Hongbin Zhang 0005, Sik-Ho Tsang, Mengting Xu
IEEE Trans. Image Process.4
2020 Low-Complexity Intra Prediction for Screen Content Coding by Convolutional Neural Network
abstract
Screen content coding (SCC) is developed to encode screen content videos, and it is an extension of High Efficiency Video Coding (HEVC). Since screen content videos contain computer-generated content that shows special characteristics, SCC adopts the new Intra Block Copy mode and Palette mode besides the HEVC based Intra mode to improve the coding efficiency. However, the exhaustive mode searching process makes the SCC encoder computational expensive. In this paper, a low-complexity intra prediction algorithm is proposed by the convolutional neural network (CNN). The proposed network skips unnecessary coding units (CUs) and mode candidates by imitating the behavior of the original SCC encoder. The network first decides if a CU size should be checked by analyzing global features, and it decides which mode should be checked by analyzing the local features. Experimental results show that the proposed algorithm achieves 53.44% computational complexity reduction on average with 1.94% Bjentegaard delta bitrate loss under All Intra configuration.
Wei Kuang, Yui-Lam Chan, Sik-Ho Tsang
ISCAS3
2020 360-Degree Intra Coding Mode for Equirectangular Projection Format Videos
abstract
Recent advances in display, networking, and computing technologies have resulted in changing industry focus towards 360-degree/omnidirectional images as witnessed by increased interest in virtual reality (VR) and augmented reality (AR). Numerous curves are generated for 360-degree images due to the lens curvature and projection format. However, in High Efficiency Video Coding (HEVC), the conventional angular intra modes cannot handle them well since only straight lines can be predicted. More enhanced 360-degree image coding is essential for higher efficiency of storage and transmission. Therefore, we propose a new coding mode, called 360-degree intra mode, for predicting the coding units with curves. Experimental results show that our 360-degree intra mode improves the coding efficiency of HEVC by 0.32% on average and up to 0.74% Bjontegaard delta bit rate reduction.
Sik-Ho Tsang, Yui-Lam Chan
ISCAS1
2020 FastSCCNet: Fast Mode Decision in VVC Screen Content Coding via Fully Convolutional Network
abstract
Screen content coding have been supported recently in Versatile Video Coding (VVC) to improve the coding efficiency of screen content videos by adopting new coding modes which are dedicated to screen content video compression. Two new coding modes called Intra Block Copy (IBC) and Palette (PLT) are introduced. However, the flexible quad-tree plus multi-type tree (QTMT) coding structure for coding unit (CU) partitioning in VVC makes the fast algorithm of the SCC particularly challenging. To efficiently reduce the computational complexity of SCC in VVC, we propose a deep learning based fast prediction network, namely FastSCCNet, where a fully convolutional network (FCN) is designed. CUs are classified into natural content block (NCB) and screen content block (SCB). With the use of FCN, only one shot inference is needed to classify the block types of the current CU and all corresponding sub-CUs. After block classification, different subsets of coding modes are assigned according to the block type, to accelerate the encoding process. Compared with the conventional SCC in VVC, our proposed FastSCCNet reduced the encoding time by 29.88% on average, with negligible bitrate increase under all-intra configuration. To the best of our knowledge, it is the first approach to tackle the computational complexity reduction for SCC in VVC.
Sik-Ho Tsang, Ngai-Wing Kwong, Yui-Lam Chan
VCIP1
2020 Overview of current development in depth map coding of 3D video and its future
abstract
3D videos have attracted attention from academia and industry after great success in the film industry. Multiview video plus depth (MVD) is the most popular 3D video format to provide vivid 3D feeling and has been adopted as an international 3D video coding standard, namely 3D extension of high efficiency video coding (HEVC). MVD includes a limited number of textures and depth maps to synthesise virtual views. In MVD, depth samples describe the distance between a camera and an actual object as a grey‐level image. The characteristics of depth maps are quite different from texture images. Consequently, new coding tools are designed for depth maps in 3D‐HEVC to improve the coding efficiency at the expense of high‐computational complexity, which faces great challenges in coding systems. Depth map coding is also an important technique in immersive media to support three degrees of freedom 3DoF+/6DoF applications such as virtual reality/augmented reality. The study starts with an overview of what has been done over the last decade in 3D‐HEVC, especially depth map coding, regarding theories, methodologies, current research and state‐of‐the‐art fast approaches. Following this, a comprehensive comparison of the reviewed techniques is presented, and an outlook on its future trends is provided.
Yui-Lam Chan, Chang-Hong Fu 0002, Hao Chen 0043, Sik-Ho Tsang
IET Signal Process.4
2020 Early termination for fast intra mode decision in depth map coding using DIS-inheritance
Chang-Hong Fu 0002, Hao Chen 0043, Yui-Lam Chan, Sik-Ho Tsang, Xiaohua Zhu 0001
Signal Process. Image Commun.4
2020 Machine Learning-Based Fast Intra Mode Decision for HEVC Screen Content Coding via Decision Trees
abstract
The screen content coding (SCC) extension of high efficiency video coding (HEVC) improves coding gain for screen content videos by introducing two new coding modes, namely, intra block copy (IBC) and palette (PLT) modes. However, the coding gain is achieved at the increased cost of computational complexity. In this paper, we propose a decision tree-based framework for fast intra mode decision by investigating various features in the training sets. To avoid the exhaustive mode searching process, a sequential arrangement of decision trees is proposed to check each mode separately by inserting a classifier before checking a mode. As compared with the previous approaches where both IBC and PLT modes are checked for screen content blocks (SCBs), the proposed coding framework is more flexible which facilitates either the IBC or PLT mode to be checked for SCBs such that computational complexity is further reduced. To enhance the accuracy of decision trees, dynamic features are introduced, which reveal the unique intermediate coding information of a coding unit (CU). Then, if all the modes are decided to be skipped for a CU at the last depth level, at least one possible mode is assigned by a CU-type decision tree. Furthermore, a decision tree constraint technique is developed to reduce the rate-distortion performance loss. Compared with the HEVC-SCC reference software SCM-8.3, the proposed algorithm reduces computational complexity by 47.62% on average with a negligible Bjøntegaard delta bitrate (BDBR) increase of 1.42% under all-intra (AI) configurations, which outperforms all the state-of-the-art algorithms in the literature.
Wei Kuang, Yui-Lam Chan, Sik-Ho Tsang, Wan-Chi Siu
IEEE Trans. Circuits Syst. Video Technol.3
2020 DeepSCC: Deep Learning-Based Fast Prediction Network for Screen Content Coding
abstract
Screen content coding (SCC) is an extension of high efficiency video coding (HEVC), and it is developed to improve the coding efficiency of screen content videos by adopting two new coding modes: Intra Block Copy (IBC) and Palette (PLT). However, the flexible quadtree-based coding tree unit (CTU) partitioning structure and various mode candidates make the fast algorithms of the SCC extremely challenging. To efficiently reduce the computational complexity of SCC, we propose a deep learning-based fast prediction network DeepSCC that contains two parts: DeepSCC-I and DeepSCC-II. Before feeding to DeepSCC, incoming coding units (CUs) are divided into two categories: dynamic CTUs and stationary CTUs. For dynamic CTUs having different content as their collocated CTUs, DeepSCC-I takes raw sample values as the input to make fast predictions. For stationary CTUs having the same content as their collocated CTUs, DeepSCC-II additionally utilizes the optimal mode maps of the stationary CTU to further reduce the computational complexity. Compared with the HEVC-SCC reference software SCM-8.3, the proposed DeepSCC reduces the encoding time by 48.81% on average with a negligible Bjøntegaard delta bitrate increase of 1.18% under all-intra configuration.
Wei Kuang, Yui-Lam Chan, Sik-Ho Tsang, Wan-Chi Siu
IEEE Trans. Circuits Syst. Video Technol.3
2020 Online-Learning-Based Bayesian Decision Rule for Fast Intra Mode and CU Partitioning Algorithm in HEVC Screen Content Coding
abstract
Screen content coding (SCC) is an extension of high efficiency video coding by adopting new coding modes to improve the coding efficiency of SCC at the expense of increased complexity. This paper proposes an online-learning approach for fast mode decision and coding unit (CU) size decision in SCC. To make a fast mode decision, the corner point is first extracted as a unique feature in screen content, which is an essential pre-processing step to guide Bayesian decision modeling. Second, the distinct color number in a CU is derived as another unique feature in screen content to build the precise model using online-learning for skipping unnecessary modes. Third, the correlation of the modes among spatial neighboring CUs is analyzed to further eliminate unnecessary mode candidates. Finally, the Bayesian decision rule using online-learning is applied again to make a fast CU size decision. To ensure the accuracy of the Bayesian decision models, new scene change detection is designed to update the models. Results show that the proposed algorithm achieves 36.69% encoding time reduction with 1.08% Bjøntegaard delta bitrate (BDBR) increment under all intra configuration. By integrating into the existing fast SCC approach, the proposed algorithm reduces 48.83% encoding time with a 1.78% increase in BDBR.
Wei Kuang, Yui-Lam Chan, Sik-Ho Tsang, Wan-Chi Siu
IEEE Trans. Image Process.3
2019 Mode Skipping for HEVC Screen Content Coding via Random Forest
abstract
Screen content coding (SCC) is the extension to high-efficiency video coding (HEVC) for compressing screen content videos. New coding tools, intrablock copy (IBC), and palette (PLT) modes, are introduced to encode screen content (SC) such as texts and graphics. The IBC mode is used for encoding repeating patterns by performing block matching within the same frame, while the PLT mode is designed for SC with few distinct colors by coding the major colors and their corresponding locations using an index map. However, the use of IBC and PLT modes increases the encoder complexity remarkably though coding efficiency can be improved. Therefore, we propose to have a mode skipping approach to reduce the encoder complexity of SCC by making use of SC characteristics, neighbor coding unit (CU) correlations, and intermediate cost information via random forest (RF). Detailed feature analyses and sample preparation are also described. A novel hyperparameter tuning approach with the consideration of coding bitrate and encoding time is proposed for RFs at each CU size to further boost the encoding process. Experimental results show that our proposed approach can obtain 45.06% average encoding time reduction with only a 1.08% increase in Bjøntegaard delta bitrate. Average encoding time can even be reduced to 58.57% by regulating the hyperparameters.
Sik-Ho Tsang, Yui-Lam Chan, Wei Kuang
IEEE Trans. Multim.1
2019 Reduced-Complexity Intra Block Copy (IntraBC) Mode With Early CU Splitting and Pruning for HEVC Screen Content Coding
abstract
A screen content coding (SCC) extension to high efficiency video coding has been developed to incorporate many new coding tools in order to achieve better coding efficiency for videos mixed with camera-captured content and graphics/text/animation. For instance, the Intra Block Copy (IntraBC) mode helps to encode repeating patterns within the same frame while the Palette mode aims at encoding screen content with a few major colors. However, the IntraBC mode brings along high computational complexity due to the exhaustive block matching within the same frame though there are already some constraints and fast approaches applied to the IntraBC mode to reduce its complexity. Thus, we propose a fast intracoding scheme to reduce the complexity of using the IntraBC mode in SCC. Screen content always contains no sensor noise resulting in the characteristics with pixel exactness along both horizontal and vertical directions. These characteristics pave the way for mode skipping and early coding unit (CU) splitting. Besides, early CU pruning and early termination are proposed based on the rate distortion cost to further reduce encoder complexity. Moreover, we also propose reducing the complexity of the IntraBC mode by checking the hash value of each block candidate and the current block during block matching. With our proposed scheme, the encoding time is reduced compared with the SCC while the coding efficiency can still be maintained with a minor increase in the bjontegaard delta bitrate.
Sik-Ho Tsang, Yui-Lam Chan, Wei Kuang, Wan-Chi Siu
IEEE Trans. Multim.1
2018 Fast HEVC to SCC Transcoding Based on Decision Trees
abstract
Screen Content Coding (SCC) is an extension of the High-Efficiency Video Coding (HEVC) for encoding screen content videos. However, there are many legacy screen content videos already encoded by HEVC. To efficiently migrate screen content videos from the existing HEVC to the emerging SCC, a machine learning based fast transcoding algorithm is proposed by using decision trees in this paper. To speed up the transcoding process, the intermediate data from both the HEVC decoder side and the SCC encoder side are jointly analyzed. Then the optimal coding unit (CU) sizes are mapped from HEVC to SCC while the mode candidates are adaptively checked according to the decision tree outcomes in the re-encoding process. Experimental results show that an average of 48.20% re-encoding time reduction is achieved with only 1.47% Bjontegaard delta bitrate loss using All Intra (AI) configuration.
Wei Kuang, Yui-Lam Chan, Sik-Ho Tsang, Wan-Chi Siu
ICME3
2018 Probability-Based Depth Intra-Mode Skipping Strategy and Novel VSO Metric for DMM Decision in 3D-HEVC
abstract
Multiview video plus depth format has been adopted as the emerging 3D video representation recently. It includes a limited number of textures and depth maps to synthesize additional virtual views. Since the quality of depth maps influences the view synthesis process, their sharp edges should be well preserved to avoid mixing foreground with background. To address this issue, 3D-High Efficiency Video Coding (HEVC) introduces new coding tools, a partition-based intra mode [depth modeling mode (DMM)], a residual description technique [segmentwise depth coding (SDC)], and a more complex rate-distortion (RD) evaluation with view synthesis optimization (VSO), to provide more accurate predictions and achieve higher compression rate. However, these new techniques introduce a lot of possible candidates, and each of them requires complicated RD calculation in the process of intra-mode decision. They lead to unacceptable computational burden in a 3D-HEVC encoder. Therefore, in this paper, we raise two efficient techniques for depth intra-mode decision. First, by investigating the statistical characteristics of variance distributions in the two partitions of DMM, a simple but efficient criterion based on the squared Euclidean distance of variances (SEDV) is suggested to evaluate RD costs of the DMM candidates instead of the time-consuming VSO process. Second, a probability-based early depth intra-mode decision is proposed to select only the most promising mode and make the early determination of using SDC based on the low-complexity RD cost in rough mode decision. Experimental results show that the proposed algorithm with these two new techniques provides 33%-48% time reduction with little drop in the coding performance compared with the state-of-the-art algorithms.
Hongbin Zhang 0005, Chang-Hong Fu 0002, Yui-Lam Chan, Sik-Ho Tsang, Wan-Chi Siu
IEEE Trans. Circuits Syst. Video Technol.4
2017 Fast mode decision algorithm for HEVC screen content intra coding
abstract
Screen Content coding (SCC) is one of an extension to High Efficiency Video Coding (HEVC) developed by the Joint Collaborative Team on Video Coding (JCT-VC). It adopts two new coding tools, intra block copy (IBC) and palette (PLT) modes, to improve the compression performance for intra coding. Nevertheless, mode selection causes a substantial increase in encoding complexity. In this paper, a fast mode decision algorithm, which makes use of early mode skip decision based on the Bayesian decision rule using online learning, is proposed. The proposed algorithm is implemented in the SCC reference software SCM-7.0. Experimental results show that the proposed algorithm can achieve 23.2% complexity reduction on average with only 0.58% Bjontegaard delta bitrate loss in All Intra (AI) configurations.
Wei Kuang, Sik-Ho Tsang, Yui-Lam Chan, Wan-Chi Siu
ICIP2
2017 Decoder side merge mode and AMVP in HEVC screen content coding
abstract
Intra Block Copy (IBC) mode in a screen content coding (SCC) extension in High Efficiency Video Coding (HEVC) provides high coding gain by performing motion estimation (ME) and motion compensation (MC) to find the repetitive patterns within the same frame. Merge mode and Advanced Motion Vector Prediction (AMVP), which are originally used for inter mode, are also applied to the IBC mode. However, there are redundant coding bits when they are applied to IBC. Therefore, we propose decoder-side merge mode and AMVP for IBC in SCC so as to remove the redundancy. Experimental shows that the proposed method can achieve up to 0.25% Bjontegaard delta bitrate (BD-rate) reduction compared to the conventional SCC with negligible impact to encoding and decoding complexity.
Sik-Ho Tsang, Wei Kuang, Yui-Lam Chan, Wan-Chi Siu
ICIP1
2016 Quadtree decision for depth intra coding in 3D-HEVC by good feature
abstract
3D-HEVC is a good coding solution for multi-view video plus depth data. It achieves good coding performance of synthesized views. However, depth intra coding brings unbearable complexity, which is the most urgent issue to be solved for the practical applications. Typically, depth maps have a good feature of structure or less texture compared with natural videos. Therefore, in this paper, a fast depth intra coding algorithm is proposed to speed up the quadtree decision by the good feature-corner point (CP). The proposed algorithm can adaptively extract CPs and preallocate the depth level of coding quadtree. The large size of coding units (CUs) can be skipped for blocks, which have higher predicted depth level. On the contrary, the blocks, with lower predicted depth level, do not check the smaller size of CUs. Simulation results show that the proposed algorithm can provide about 41% time reduction while maintaining the BD performance.
Hongbin Zhang 0005, Yui-Lam Chan, Chang-Hong Fu 0002, Sik-Ho Tsang, Wan-Chi Siu
ICASSP4
2015 Fast and efficient intra coding techniques for smooth regions in screen content coding based on boundary prediction samples
abstract
This paper presents fast and efficient intra prediction algorithms for screen content coding (SCC). The proposed algorithms focus on smooth regions frequently appeared in screen content videos, which have the characteristics of noiselessness. All the samples in a noiseless smooth region exhibit exactly the same pixel value. We then propose two intra coding techniques for noiseless smooth regions in SCC based on the smoothness of the boundary samples which are used for intra prediction. Our proposed algorithm can reduce computational complexity by at most 26.7% while keeping nearly the same video quality. Moreover, by removing the redundant coding bits for intra prediction modes, computational complexity can be further reduced to at most 53.3% in terms of encoding time with bitrate reduction up to 1.2%.
Sik-Ho Tsang, Yui-Lam Chan, Wan-Chi Siu
ICASSP1
2015 Efficient depth intra mode decision by reference pixels classification in 3D-HEVC
abstract
The uniform intra prediction increases the intra prediction modes up to 35 and brings better coding efficiency in HEVC. Besides, depth modelling modes (DMMs) are introduced in depth intra coding of 3D-HEVC to preserve sharp edges and avoid ringing artifacts in a synthesized view. Meanwhile, the encoding time of depth intra coding rapidly increases due to a huge number of intra mode candidates. Based on the spatial correlation of a depth map, we find that not all of the intra modes are necessary to be considered in most cases. Hence, a fast content-dependent depth intra mode decision algorithm is raised in this paper by classifying the spatial distribution of the reference pixels. Simulation results show that the proposed adaptive fast algorithm can save 21%-35% time of the depth coding with the insignificant bit rate increase compared with the state-of-the-art algorithm.
Hongbin Zhang 0005, Chang-Hong Fu 0002, Yui-Lam Chan, Sik-Ho Tsang, Wan-Chi Siu
ICIP4
2013 Region-based weighted prediction algorithm for H.264/AVC video coding
abstract
This paper proposes a novel region-based weighted prediction (WP) algorithm to encode scenes with complex brightness variations. It facilitates the use of multiple WP parameter sets in a single reference frame by utilizing the framework of multiple reference frame motion estimation (MRF-ME). With this arrangement, different macroblocks in the current frame can use different WP parameter sets even when they are predicted from the same reference frame. To support this, a region partitioning process is designed to divide the current frame into different regions where each one has some degree of uniformity in its brightness variation. Multiple sets of region-based WP parameters can then be estimated accurately. Consequently, the proposed algorithm can improve prediction in scenes with different degrees of brightness variations in different regions of the same picture. Results show that the region-based algorithm can achieve significant coding gains of scenes with complex brightness variations.
Sik-Ho Tsang, Tsz-Kwan Lee, Yui-Lam Chan, Wan-Chi Siu
ISCAS1
2013 Region-Based Weighted Prediction for Coding Video With Local Brightness Variations
abstract
This paper presents a new region-based scheme for the estimation of weighted prediction (WP) parameter sets for encoders of the H.264/MPEG-4 AVC standard. The proposed scheme is specifically designed for handling local brightness variations (LBVs) in video scenes. It is achieved by making use of multiple WP parameter sets for various regions and assigning them to the same reference frame. An accurate estimation of multiple WP parameter sets is accomplished by: 1) partitioning regions with a simple WP parameter estimator; 2) selecting regions where WP should be applied; and 3) estimating accurate WP parameter sets with a quasioptimal WP parameter estimator. The multiple WP parameter sets of different regions are encoded using the framework of multiple reference frames in the H.264/MPEG-4 AVC standard. With this arrangement, the proposed scheme is compliant with the H.264/MPEG-4 AVC standard. To reduce the implementation cost, a reduction of the memory requirement is realized via look-up tables (LUTs). Experimental results show that the region-based scheme can efficiently handle scenes with global and LBVs and achieve significant coding gain over other WP schemes. Furthermore, our scheme with LUTs can reduce the memory requirement by about 80% while keeping the same coding efficiency as that without LUTs.
Sik-Ho Tsang, Yui-Lam Chan, Wan-Chi Siu
IEEE Trans. Circuits Syst. Video Technol.1
2012 Flash scene video coding using weighted prediction
Sik-Ho Tsang, Yui-Lam Chan, Wan-Chi Siu
J. Vis. Commun. Image Represent.1
2010 H.264 video coding with multiple weighted prediction models
abstract
Weighted prediction is a video coding tool to encode scenes with brightness variations. However, no single WP model works well for all types of brightness variations. In this paper, a novel single reference frame multiple WP models (SRefMWP) scheme is proposed to facilitate the use of multiple WP models in different macroblocks of the current frame even when they are predicted from the same reference. It provides this feature by making a new arrangement of the multiple frame buffers in multiple reference frame motion estimation. Experimental results show that the proposed SRefMWP can improve prediction in scenes with different types of brightness variations, and even benefit to scenes that contain local brightness variation.
Sik-Ho Tsang, Yui-Lam Chan
ICIP1