Ngai-Wing Kwong

dblp:283/8770 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0002-8180-7994ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 7 since 2021
YearPublicationVenuePosition
2026 Multi-Branch Aesthetic and Technical Perspectives With Cross Tri-Fusion Attention for No-Reference Audio-Visual Quality Assessment
Ngai-Wing Kwong, Yui-Lam Chan, Ziyin Huang, Sik-Ho Tsang
IEEE Trans. Circuits Syst. Video Technol.1
2025 Long Short-Term Fusion by Multi-Scale Distillation for Screen Content Video Quality Enhancement
abstract
Different from natural videos, where artifacts distributed evenly, the artifacts of compressed screen content videos mainly occur in the edge areas. Besides, these videos often exhibit abrupt scene switches, resulting in noticeable distortions in video reconstruction. Existing multiple-frame models using a fixed range of neighbor frames face challenges in effectively enhancing frames during scene switches and lack efficiency in reconstructing high-frequency details. To address these limitations, we propose a novel method that effectively handles scene switches and reconstructs high-frequency information. In the feature extraction part, we develop long-term and short-term feature extraction streams, in which the long-term feature extraction stream learns the contextual information, and the short-term feature extraction stream extracts more related information from shorter input to assist the long-term stream to handle fast motion and scene switches. To further enhance the frame quality during scene switches, we incorporate a similarity-based neighbor frame selector before feeding frames into the short-term stream. This selector identifies relevant neighbor frames, aiding in the efficient handling of scene switches. To dynamically fuse the short-term feature and long-term features, the muti-scale feature distillation focuses on adaptively recalibrating channel-wise feature responses to achieve effective feature distillation. In the reconstruction part, a high-frequency reconstruction block is proposed for guiding the model to restore the high-frequency components. Experimental results demonstrate the significant advancements achieved by our proposed Long Short-term Fusion by Multi-Scale Distillation (LSFMD) method in enhancing the quality of compressed screen content videos, surpassing the current state-of-the-art methods.
Ziyin Huang, Yui-Lam Chan, Ngai-Wing Kwong, Sik-Ho Tsang, Kin-Man Lam 0001, Bingo Wing-Kuen Ling
IEEE Trans. Circuits Syst. Video Technol.3
2025 Multi-Frame Spatiotemporal Feature and Hierarchical Learning Approach for No-Reference Screen Content Video Quality Assessment
abstract
The rapid adoption of remote work, online conferencing, and shared-screen collaboration has significantly increased the usage of screen content videos (SCVs), creating a growing need for reliable quality assessment to maintain excellent quality of service. While several full-reference SCV quality assessment (SCVQA) methods have been proposed, their practical application is often limited by the unavailability of reference videos. Existing no-reference SCVQA (NR-SCVQA) methods rely on handcrafted features and focus solely on specific distortions and features, potentially limiting their generalization ability. Moreover, they fail to explore the underlying spatiotemporal information of SCVs, which could hinder their performance. In this work, we propose a novel deep learning-based NR-SCVQA model specifically tailored to capture the comprehensive spatiotemporal features of SCVs to overcome these issues and challenges posed by the SCVQA task. Our approach incorporates a dual-channel spatiotemporal convolutional neural network (DCST-CNN) module to extract both content-aware and edge-aware spatiotemporal quality features, which enables an effective spatiotemporal quality feature representation learning for the downstream SCVQA task. Building upon the DCST-CNN, we further propose a Temporal Pyramid Transformer (TPT) module to fuse spatiotemporal features across multiple temporal scales, enabling the model to capture both short-term and long-term temporal dependencies within an SCV for hierarchical learning. The proposed DCST-CNN and TPT modules work together to provide a robust and accurate NR-SCVQA framework. We conduct experiments on SCVQA databases to validate the effectiveness of our model, which outperforms existing state-of-the-art NR-SCVQA method. The results demonstrate the strength and applicability of our approach in real-world SCVQA tasks.
Ngai-Wing Kwong, Yui-Lam Chan, Sik-Ho Tsang, Ziyin Huang, Kin-Man Lam 0001
IEEE Trans. Multim.1
2024 Frame Similarity-Based Screen Content Video Quality Enhancement via Adaptive Long Short-Term Fusion
abstract
Compressed screen content videos often exhibit artifacts in edge areas and suffer from distortions during scene switches, where content abruptly changes between frames. Existing multi-frame models, which use a fixed range of neighbor frames, struggle with these switches. To address this, we propose a novel method that effectively handles scene switches. Our approach utilizes Long-term Feature Extraction (LFE) to capture contextual information, while the Frame Similarity-based Short-term Feature Extraction (FSFE) focuses on texture information to manage fast motion and scene switches. In FSFE, a Similarity-based Neighbor Frame Selector (SNFS) is designed to choose relevant neighbor frames for the short-term stream, enhancing the quality of scene switch frames. To fuse short-term and long-term features adaptively, we introduce a local-spatial and global-channel attention module, which recalibrates spatial and channel-wise feature responses. Experimental results show that our Frame Similarity-Based via Adaptive Long Short-Term Fusion (FSLST) method significantly improves the quality of compressed videos, outperforming current state-of-the-art methods.
Ziyin Huang, Yui-Lam Chan, Ngai-Wing Kwong, Sik-Ho Tsang, Kin-Man Lam 0001, Bingo Wing-Kuen Ling
VCIP3
2024 Spatio-temporal feature learning for enhancing video quality based on screen content characteristics
Ziyin Huang, Yui-Lam Chan, Sik-Ho Tsang, Ngai-Wing Kwong, Kin-Man Lam 0001, Bingo Wing-Kuen Ling
J. Vis. Commun. Image Represent.4
2024 Spatiotemporal feature learning for no-reference gaming content video quality assessment
Ngai-Wing Kwong, Yui-Lam Chan, Sik-Ho Tsang, Ziyin Huang, Kin-Man Lam 0001
J. Vis. Commun. Image Represent.1
2023 Optimized Quality Feature Learning for Video Quality Assessment
abstract
Recently, some transfer learning-based methods have been adopted in video quality assessment (VQA) to compensate for the lack of enormous training samples and human annotation labels. But these methods induce a domain gap between source and target domains, resulting in a sub-optimal feature representation that deteriorates the accuracy. This paper proposes the optimized quality feature learning via a multi-channel convolutional neural network (CNN) with the gated recurrent unit (GRU) for no-reference (NR) VQA. First, the multi-channel CNN is pre-trained on the image quality assessment (IQA) domain using non-human annotation labels, which is inspired by self-supervised learning. Then, semi-supervised learning is used to fine-tune CNN and transfer the knowledge from IQA to VQA while considering motion information for optimized quality feature learning. Finally, all frame quality features are extracted as the input of GRU to obtain the final video quality. Experimental results demonstrate that our model achieves better performance than state-of-the-art VQA approaches.
Ngai-Wing Kwong, Yui-Lam Chan, Sik-Ho Tsang, Daniel Pak-Kong Lun
ICASSP1
2020 FastSCCNet: Fast Mode Decision in VVC Screen Content Coding via Fully Convolutional Network
abstract
Screen content coding have been supported recently in Versatile Video Coding (VVC) to improve the coding efficiency of screen content videos by adopting new coding modes which are dedicated to screen content video compression. Two new coding modes called Intra Block Copy (IBC) and Palette (PLT) are introduced. However, the flexible quad-tree plus multi-type tree (QTMT) coding structure for coding unit (CU) partitioning in VVC makes the fast algorithm of the SCC particularly challenging. To efficiently reduce the computational complexity of SCC in VVC, we propose a deep learning based fast prediction network, namely FastSCCNet, where a fully convolutional network (FCN) is designed. CUs are classified into natural content block (NCB) and screen content block (SCB). With the use of FCN, only one shot inference is needed to classify the block types of the current CU and all corresponding sub-CUs. After block classification, different subsets of coding modes are assigned according to the block type, to accelerate the encoding process. Compared with the conventional SCC in VVC, our proposed FastSCCNet reduced the encoding time by 29.88% on average, with negligible bitrate increase under all-intra configuration. To the best of our knowledge, it is the first approach to tackle the computational complexity reduction for SCC in VVC.
Sik-Ho Tsang, Ngai-Wing Kwong, Yui-Lam Chan
VCIP2