Ziyin Huang

dblp:235/0755 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-Branch Aesthetic and Technical Perspectives With Cross Tri-Fusion Attention for No-Reference Audio-Visual Quality Assessment
Ngai-Wing Kwong, Yui-Lam Chan, Ziyin Huang, Sik-Ho Tsang
IEEE Trans. Circuits Syst. Video Technol.3
2025 Long Short-Term Fusion by Multi-Scale Distillation for Screen Content Video Quality Enhancement
abstract
Different from natural videos, where artifacts distributed evenly, the artifacts of compressed screen content videos mainly occur in the edge areas. Besides, these videos often exhibit abrupt scene switches, resulting in noticeable distortions in video reconstruction. Existing multiple-frame models using a fixed range of neighbor frames face challenges in effectively enhancing frames during scene switches and lack efficiency in reconstructing high-frequency details. To address these limitations, we propose a novel method that effectively handles scene switches and reconstructs high-frequency information. In the feature extraction part, we develop long-term and short-term feature extraction streams, in which the long-term feature extraction stream learns the contextual information, and the short-term feature extraction stream extracts more related information from shorter input to assist the long-term stream to handle fast motion and scene switches. To further enhance the frame quality during scene switches, we incorporate a similarity-based neighbor frame selector before feeding frames into the short-term stream. This selector identifies relevant neighbor frames, aiding in the efficient handling of scene switches. To dynamically fuse the short-term feature and long-term features, the muti-scale feature distillation focuses on adaptively recalibrating channel-wise feature responses to achieve effective feature distillation. In the reconstruction part, a high-frequency reconstruction block is proposed for guiding the model to restore the high-frequency components. Experimental results demonstrate the significant advancements achieved by our proposed Long Short-term Fusion by Multi-Scale Distillation (LSFMD) method in enhancing the quality of compressed screen content videos, surpassing the current state-of-the-art methods.
Ziyin Huang, Yui-Lam Chan, Ngai-Wing Kwong, Sik-Ho Tsang, Kin-Man Lam 0001, Bingo Wing-Kuen Ling
IEEE Trans. Circuits Syst. Video Technol.1
2025 Multi-Frame Spatiotemporal Feature and Hierarchical Learning Approach for No-Reference Screen Content Video Quality Assessment
abstract
The rapid adoption of remote work, online conferencing, and shared-screen collaboration has significantly increased the usage of screen content videos (SCVs), creating a growing need for reliable quality assessment to maintain excellent quality of service. While several full-reference SCV quality assessment (SCVQA) methods have been proposed, their practical application is often limited by the unavailability of reference videos. Existing no-reference SCVQA (NR-SCVQA) methods rely on handcrafted features and focus solely on specific distortions and features, potentially limiting their generalization ability. Moreover, they fail to explore the underlying spatiotemporal information of SCVs, which could hinder their performance. In this work, we propose a novel deep learning-based NR-SCVQA model specifically tailored to capture the comprehensive spatiotemporal features of SCVs to overcome these issues and challenges posed by the SCVQA task. Our approach incorporates a dual-channel spatiotemporal convolutional neural network (DCST-CNN) module to extract both content-aware and edge-aware spatiotemporal quality features, which enables an effective spatiotemporal quality feature representation learning for the downstream SCVQA task. Building upon the DCST-CNN, we further propose a Temporal Pyramid Transformer (TPT) module to fuse spatiotemporal features across multiple temporal scales, enabling the model to capture both short-term and long-term temporal dependencies within an SCV for hierarchical learning. The proposed DCST-CNN and TPT modules work together to provide a robust and accurate NR-SCVQA framework. We conduct experiments on SCVQA databases to validate the effectiveness of our model, which outperforms existing state-of-the-art NR-SCVQA method. The results demonstrate the strength and applicability of our approach in real-world SCVQA tasks.
Ngai-Wing Kwong, Yui-Lam Chan, Sik-Ho Tsang, Ziyin Huang, Kin-Man Lam 0001
IEEE Trans. Multim.4
2024 Frame Similarity-Based Screen Content Video Quality Enhancement via Adaptive Long Short-Term Fusion
abstract
Compressed screen content videos often exhibit artifacts in edge areas and suffer from distortions during scene switches, where content abruptly changes between frames. Existing multi-frame models, which use a fixed range of neighbor frames, struggle with these switches. To address this, we propose a novel method that effectively handles scene switches. Our approach utilizes Long-term Feature Extraction (LFE) to capture contextual information, while the Frame Similarity-based Short-term Feature Extraction (FSFE) focuses on texture information to manage fast motion and scene switches. In FSFE, a Similarity-based Neighbor Frame Selector (SNFS) is designed to choose relevant neighbor frames for the short-term stream, enhancing the quality of scene switch frames. To fuse short-term and long-term features adaptively, we introduce a local-spatial and global-channel attention module, which recalibrates spatial and channel-wise feature responses. Experimental results show that our Frame Similarity-Based via Adaptive Long Short-Term Fusion (FSLST) method significantly improves the quality of compressed videos, outperforming current state-of-the-art methods.
Ziyin Huang, Yui-Lam Chan, Ngai-Wing Kwong, Sik-Ho Tsang, Kin-Man Lam 0001, Bingo Wing-Kuen Ling
VCIP1
2024 Spatio-temporal feature learning for enhancing video quality based on screen content characteristics
Ziyin Huang, Yui-Lam Chan, Sik-Ho Tsang, Ngai-Wing Kwong, Kin-Man Lam 0001, Bingo Wing-Kuen Ling
J. Vis. Commun. Image Represent.1
2024 Spatiotemporal feature learning for no-reference gaming content video quality assessment
Ngai-Wing Kwong, Yui-Lam Chan, Sik-Ho Tsang, Ziyin Huang, Kin-Man Lam 0001
J. Vis. Commun. Image Represent.4
2023 STADS: Spatial Transcriptomics to Aid Drug-reposition Recommendation
abstract
Drug repurposing is a promising strategy to find new usage for existing drugs outside of their original purpose. We have recently developed a computational framework ASGARD that employs single cell RNA-sequencing (scRNA-seq) to repurpose drugs. However, scRNA-seq lacks the spatial information of the tissue, an important determinant of the heterogeneity within the tissue micro-environment. In this paper, we propose Spatial Transcriptomics to Aid Drug-reposition Recommendation (STADS) framework, a new drug repurposing computational framework that employs Spatial Transcriptomics (ST) datasets with gene expression in situ. STADS reduces the ST gene expression by spatially-aware embedding, integrates multiple healthy and diseased samples, identifies the matched spatial domains between disease and controls, then repurposes drugs by a novel drug score that considers differentially expressed genes between disease and control samples among all matched spatial domains. We apply STADS to a ST dataset of patients with Hepatocellular Carcinoma and demonstrate its superior performance to ASGARD. STADS is a promising personalized drug repurposing prediction method using ST data.
Abdullah Karaaslanli, Ziyin Huang, Yijun Guo, Lana X. Garmire, Haodong Liang
BIBM2
2023 Image super resolution via combination of two dimensional quaternion valued singular spectrum analysis based denoising, empirical mode decomposition based denoising and discrete cosine transform based denoising methods
Yingdan Cheng, Bingo Wing-Kuen Ling, Ziyin Huang, Yui-Lam Chan
Multim. Tools Appl.4
2022 Few Clean Instances Help Denoising Distant Supervision
abstract
Existing distantly supervised relation extractors usually rely on noisy data for both model training and evaluation, which may lead to garbage-in-garbage-out systems. To alleviate the problem, we study whether a small clean dataset could help improve the quality of distantly supervised models. We show that besides getting a more convincing evaluation of models, a small clean dataset also helps us to build more robust denoising models. Specifically, we propose a new criterion for clean instance selection based on influence functions. It collects sample-level evidence for recognizing good instances (which is more informative than loss-level evidence). We also propose a teacher-student mechanism for controlling purity of intermediate results when bootstrapping the clean set. The whole approach is model-agnostic and demonstrates strong performances on both denoising real (NYT) and synthetic noisy datasets.
Yufang Liu, Ziyin Huang, Changzhi Sun, Man Lan, Yuanbin Wu, Xiaofeng Mou
COLING2
2021 Mathematical model for shape description in DCT domain
Ziyin Huang, Bingo Wing-Kuen Ling
Multim. Tools Appl.1
2019 Constrained Heterogeneous Vehicle Path Planning for Large-area Coverage
abstract
There is a strong demand for covering a large area autonomously by multiple UAVs (Unmanned Aerial Vehicles) supported by a ground vehicle. Limited by UAVs' battery life and communication distance, complete coverage of large areas typically involves multiple take-offs and landings to recharge batteries, and the transportation of UAVs between operation areas by a ground vehicle. In this paper, we introduce a novel large-area-coverage planning framework which collectively optimizes the paths for aerial and ground vehicles. Our method first partitions a large area into sub-areas, each of which a given fleet of UAVs can cover without recharging batteries. UAV operation routes, or trails, are then generated for each sub-area. Next, the assignment of trials to different UAVs and the order in which UAVs visit their assigned trails are simultaneously optimized to minimize the total UAV flight distance. Finally, a ground vehicle transportation path which visits all sub-areas is found by solving an asymmetric traveling salesman problem (ATSP). Although finding the globally optimal trail assignment and transition paths can be formulated as a Mixed Integer Quadratic Program (MIQP), the MIQP is intractable even for small problems. We show that the solution time can be reduced to close-to-real-time levels by first finding a feasible solution using a Random Key Genetic Algorithm (RKGA), which is then locally optimized by solving a much smaller MIQP.
Di Deng, Yuhe Fu, Ziyin Huang, Kenji Shimada
IROS4