Zhaolin Wan

dblp:199/8385 · DBLP profile ↗
← Back
18ranked-venue papers
9as first author
13since 2021 · last 2026
0000-0003-0324-7116ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 9 first-author · 11 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Perceptual Quality Assessment of 3D Gaussian Splatting: A Subjective Dataset and Prediction Metric
abstract
With the rapid advancement of 3D visualization, 3D Gaussian Splatting (3DGS) has emerged as a leading technique for real-time, high-fidelity rendering. While prior research has emphasized algorithmic performance and visual fidelity, the perceptual quality of 3DGS-rendered content, especially under varying reconstruction conditions, remains largely underexplored. In practice, factors such as viewpoint sparsity, limited training iterations, point downsampling, noise, and color distortions can significantly degrade visual quality, yet their perceptual impact has not been systematically studied. To bridge this gap, we present 3DGS-QA, the first subjective quality assessment dataset for 3DGS. It comprises 225 degraded reconstructions across 15 object types, enabling a controlled investigation of common distortion factors. Based on this dataset, we introduce a no-reference quality prediction model that directly operates on native 3D Gaussian primitives, without requiring rendered images or ground-truth references. Our model extracts spatial and photometric cues from the Gaussian representation to estimate perceived quality in a structure-aware manner. We further benchmark existing quality assessment methods, spanning both traditional and learning-based approaches. Experimental results show that our method consistently achieves superior performance, highlighting its robustness and effectiveness for 3DGS content evaluation. The dataset and code are made publicly available to facilitate future research in 3DGS quality assessment.
Zhaolin Wan, Yining Diao, Jingqi Xu, Hao Wang 0073, Zhiyang Li 0001, Xiaopeng Fan 0001, Wangmeng Zuo, Debin Zhao
AAAI1
2026 GAPA-3DGS: Dual-Branch Gaussian-Adaptive Perceptual Assessment for 3D Gaussian Splatting
Zhaolin Wan, Jingqi Xu, Zhiyang Li 0001, Wangmeng Zuo, Debin Zhao, Xiaopeng Fan 0001
QoMEX1
2025 CASP: Consistency-aware Audio-induced Saliency Prediction Model for Omnidirectional Video
abstract
Omnidirectional videos (ODVs) present distinct challenges for accurate audio-visual saliency prediction due to their immersive nature, which combines spatial audio with panoramic visuals to enhance the user experience. While auditory cues are crucial for guiding visual attention across the panoramic scene, the interaction between audio and visual stimuli in ODVs remains underexplored. Existing models primarily focus on spatiotemporal visual cues and treat audio signals separately from their spatial and temporal contexts, often leading to misalignments between audio and visual content and undermining temporal consistency across frames. To bridge these gaps, we propose a novel audio-induced saliency prediction model for ODVs that holistically integrates audio and visual inputs through a multi-modal encoder, an audio-visual interaction module, and an audio-visual transformer. Unlike conventional methods that isolate audio cue locations and attributes, our model employs a query-based framework, where learnable audio queries capture comprehensive audio-visual dependencies, thus enhancing saliency prediction by dynamically aligning with audio cues. Besides, we introduce a novel consistency loss to enforce temporal coherence in saliency regions across frames. Extensive experiments demonstrate that our model outperforms state-of-the-art methods in predicting audio-visual salient regions in ODVs, establishing its robustness and superior performance.
Zhaolin Wan, Han Qin, Zhiyang Li 0001, Xiaopeng Fan 0001, Wangmeng Zuo, Debin Zhao
CVPR1
2025 TS-Net: Assembling Task-specific Features from Multiple Feature Levels for Multi-task Learning
abstract
Multi-task learning (MTL) has become an attractive topic that leverages shared knowledge to improve performance and enhance generalization. However, most existing works neglect the varying contribution of multi-level features to sub-task representations. In this paper, we explore the impact of multilevel features on different tasks and propose a novel level-assembling MTL architecture named TS-Net. TS-Net integrates multi-level features into multi-task representations by combining task-specific and task-generic features. We first introduce a Task-Specific Feature Capturing Block (TSFCB) to aggregate task-specific features by dynamically assembling features for input samples and prioritizing more relevant feature levels. In addition, we present a Multi-Task Mixture-of-Experts (MTMoE) module to facilitate cross-task interaction. In MTMoE, task-generic features are captured and integrated with task-specific features through a gating mechanism, allowing TS-Net to effectively share knowledge across tasks. Extensive experiments demonstrate that TS-Net exhibits superior performance across a range of tasks, including detection, segmentation, and image reconstruction.
Zhaolin Wan, Penghong Wang, Xiaopeng Fan 0001
ICASSP2
2025 AMSFormer: A transformer with adaptive multi-scale partitioning and multi-level spectral filtering for time-series forecasting
Yining Diao, Zhaolin Wan, Zhiyang Li 0001
Neurocomputing4
2025 Enhancing No-Reference Audio-Visual Quality Assessment via Joint Cross-Attention Fusion
abstract
As the consumption of multimedia content continues to rise, audio and video have become central to everyday entertainment and social interactions. This growing reliance amplifies the demand for effective and objective audio-visual quality assessment (AVQA) to understand the interaction between audio and visual elements, ultimately enhancing user satisfaction. However, existing state-of-the-art AVQA methods often rely on simplistic machine learning models or fully connected networks for audio-visual signal fusion, which limits their ability to exploit the complementary nature of these modalities. In response to this gap, we propose a novel no-reference AVQA method that utilizes joint cross-attention fusion of audio-visual perception. Our approach begins with a dual-stream feature extraction process that simultaneously captures long-range spatiotemporal visual features and audio features. The fusion model then dynamically adjusts the contributions of features from both modalities, effectively integrating them to provide a more comprehensive perception for quality score prediction. Experimental results on the LIVE-SJTU and UnB-AVC datasets demonstrate that our model outperforms state-of-the-art methods, achieving superior performance in audio-visual quality assessment.
Zhaolin Wan, Xiguang Hao, Xiaopeng Fan 0001, Wangmeng Zuo, Debin Zhao
IEEE Signal Process. Lett.1
2025 Content-Aware Dynamic In-Loop Filter With Adjustable Complexity for VVC Intra Coding
abstract
Recently, neural network-based in-loop filters have been rapidly developed, effectively improving the reconstruction quality and compression efficiency in video coding. Existing deep in-loop filters typically employed networks with fixed structures to process all image blocks. However, under various bitrate conditions, compressed image blocks with different textures exhibit varying degradations, which poses a challenge for high-quality and low-complexity filtering. Additionally, different complexity requirements for coding tools in various scenarios limit the versatility of fixed models. To address these problems, a content-aware dynamic in-loop filter (dubbed DILF) with adjustable complexity is proposed in this paper. Specifically, DILF comprises a policy network and a filtering network. For each reconstructed image block, the policy network dynamically generates a filtering network topology based on pixel information and the quantization parameter (QP), guiding the filtering network to skip redundant layers and conduct content-aware image enhancement, thereby improving the filtering performance. In addition, by introducing a user-defined balancing factor into the policy network, the content-aware filtering network topology can be further adjusted according to user’s requirements, facilitating adjustable complexity with a single model. We integrate DILF into Versatile Video Coding (VVC) to replace the built-in deblocking filter. Extensive experiments demonstrate the efficiency of DILF in processing image blocks with varying degrees of degradation and its flexibility in controlling complexity. When the balancing factor is set to 2e-5, DILF achieves bitrate savings of 8.07%, 17.97%, and 20.93% on average for YUV components over VVC reference software VTM-11.0 under all-intra configuration. Compared to static networks with fixed structures, DILF demonstrates superior performance and lower computational complexity.
Hengyu Man, Hao Wang 0212, Riyu Lu, Zhaolin Wan, Xiaopeng Fan 0001, Debin Zhao
IEEE Trans. Circuits Syst. Video Technol.4
2025 No-Reference Stereoscopic Omnidirectional Image Quality Assessment via a Binocular Viewport Hypergraph Convolutional Network
abstract
Omnidirectional images, offering immersive 360° views, have gained significant attention, but assessing their perceptual quality, especially for stereoscopic content, remains a complex challenge. A major limitation lies in the fact that head-mounted devices restrict the viewer’s experience to a single viewport at a time, necessitating a comprehensive understanding of how multiple viewport images interact and aggregate during the viewing process. Moreover, the depth dimension inherent in stereoscopic content further complicates the 360° visual experience, a factor often oversimplified by existing methods, limiting their ability to accurately differentiate perceptual quality across viewports. To address these challenges, we propose a novel no-reference quality assessment model for stereoscopic omnidirectional images. Our approach integrates binocular vision principles within a viewport hypergraph convolutional network framework. First, guided by the unique viewing patterns of stereoscopic omnidirectional images, our model selects panoramic viewports that align with human visual preferences. Next, we devise an image feature extraction network that simulates the binocular fusion and rivalry mechanisms within the human visual system, leveraging a twin encoder-decoder network and tensor decomposition to capture key features. Finally, to assess overall image quality, we introduce a hypergraph structure module that captures complex positional and content-based interactions among sampled viewports through the Graph Influence Network. Extensive experiments on the NBU-SOID, SOLID, and LIVE 3D VR databases demonstrate the superior accuracy and robustness of our model compared to state-of-the-art methods.
Zhaolin Wan, Zhiyang Li 0001, Xiaopeng Fan 0001, Wangmeng Zuo, Debin Zhao
IEEE Trans. Circuits Syst. Video Technol.1
2024 Dual-stream Perception-driven Blind Quality Assessment for Stereoscopic Omnidirectional Images
abstract
The emergence of virtual reality technology has made stereoscopic omnidirectional images (SOI) easily accessible and prompted the need to evaluate their perceptual quality. At present, most stereoscopic omnidirectional image quality assessment (SOIQA) methods rely on one of the projection formats, i.e., Equirectangular Projection (ERP) or CubeMap Projection (CMP). However, while ERP provides global information and the less distorted CMP complements it by providing local structural guidance, research on leveraging both ERP and CMP in SOIQA remains limited, hindering a comprehensive understanding of both global and local visual cues. Motivated by this gap, our study introduces a novel dual-stream perception-driven network for blind quality assessment of stereoscopic omnidirectional images. By integrating both ERP and CMP, our method effectively captures both global and local information, marking the first attempt to bridge this gap in SOIQA, particularly through deep learning methodologies. We employ an inter-intra feature fusion module, which considers both the inter-complementarity between ERP and CMP and the intra-relationships within CMP images. This module dynamically and complementarily adjusts the contributions of features from both projections and effectively integrates them to achieve a more comprehensive perception. Besides, deformable convolution is employed to extract the local region of interest, simulating the orientation selectivity of the primary visual cortex. Finally, with the features of left and right views of SOI, a stereo cross attention module that simulates the binocular fusion mechanism is proposed to predict the quality score. Extensive experiments are conducted to evaluate our model and the state-of-the-art competitors, demonstrating that our model has achieved the best performance on the databases of LIVE 3D VR, SOLID, and NBU.
Zhaolin Wan, Qiushuang Yang, Zhiyang Li 0001, Xiaopeng Fan 0001, Wangmeng Zuo, Debin Zhao
ACM Multimedia1
2024 Label-Correlation Adaptive Central Similarity Hashing for Multi-label Image Retrieval
Yunpeng Fu, Zhaolin Wan, Jiahao Yao, Zhiyang Li 0001
PRCV (9)2
2023 A Central Similarity Hashing Method via Weighted Partial-Softmax Loss
Mengling Li, Yunpeng Fu, Zhiyang Li 0001, Zhaolin Wan
ICA3PP (7)5
2023 Computing 2D Skeleton via Generalized Electric Potential
Guangzhe Ma, Xiaoshan Wang, Zhiyang Li 0001, Zhaolin Wan
PRCV (10)5
2022 Incorporating Semi-Supervised and Positive-Unlabeled Learning for Boosting Full Reference Image Quality Assessment
abstract
Full-reference (FR) image quality assessment (IQA) evaluates the visual quality of a distorted image by measuring its perceptual difference with pristine-quality reference, and has been widely used in low-level vision tasks. Pairwise labeled data with mean opinion score (MOS) are required in training FR-IQA model, but is time-consuming and cumbersome to collect. In contrast, unlabeled data can be easily collected from an image degradation or restoration process, making it encouraging to exploit unlabeled training data to boost FR-IQA performance. Moreover, due to the distribution inconsistency between labeled and unlabeled data, outliers may occur in unlabeled data, further increasing the training difficulty. In this paper, we suggest to incorporate semi-supervised and positive-unlabeled (PU) learning for exploiting unlabeled data while mitigating the adverse effect of outliers. Particularly, by treating all labeled data as positive samples, PU learning is leveraged to identify negative samples (i.e., outliers) from unlabeled data. Semi-supervised learning (SSL) is further deployed to exploit positive unlabeled data by dynamically generating pseudo-MOS. We adopt a dual-branch network including reference and distortion branches. Furthermore, spatial attention is introduced in the reference branch to concentrate more on the informative regions, and sliced Wasserstein distance is used for robust difference map computation to address the misalignment issues caused by images recovered by GAN models. Extensive experiments show that our method performs favorably against state-of-the-arts on the benchmark datasets PIPAL, KADID-10k, TID2013, LIVE and CSIQ. The source code and model are available at https://github.com/happycaoyue/JSPL.
Yue Cao 0009, Zhaolin Wan, Dongwei Ren, Zifei Yan, Wangmeng Zuo
CVPR2
2020 Reduced Reference Stereoscopic Image Quality Assessment Using Sparse Representation and Natural Scene Statistics
abstract
An ideal quality assessment model should simulate the properties of the visual brain to be consistent with human evaluation. The visual brain appears to have both evolved to seek an efficient, decorrelated representation of image information and to “match” the statistics of the natural image. On one hand, the theoretical studies suggest that sparse representation resembles the strategy in the primary visual cortex of brain for representing natural images. On the other hand, the natural scene statistics have driven the evolution of human visual system and have also inspired the understanding and simulating of visual perception. Inspired by these observations, in this paper, we propose a novel reduced-reference stereoscopic image quality assessment metric using sparse representation and natural scene statistics to simulate the visual perception of the brain. Specifically, the distribution statistics of the classified visual primitives extracted by sparse representation are used to measure the visual information, which is closely related to the hierarchical progressive process of human visual perception. Particularly, the mutual information of classified primitives between two view images is derived as a binocular cue to simulate the binocular fusion process. The maximum mechanism that is applied to select the visual information is a pooling mechanism with which complex cells use the maximal stimuli from a group of simple cells during the transfer process in the primary visual cortex. The natural scene statistics of locally normalized luminance coefficients are used to evaluate the natural losses due to the presence of distortions. The differences of the visual information and the natural scene statistics between the original and distorted images are used to compute the quality score by a prediction function which is trained using support vector regression. Experimental results show that the proposed metric outperforms the state-of-the-art stereoscopic image quality assessment metrics on LIVE 3D IQA database and NBU-MDSID Phase-II database, and delivers competitive performance on Waterloo IVC 3D database.
Zhaolin Wan, Ke Gu 0001, Debin Zhao
IEEE Trans. Multim.1
2017 Convolutional Neural Networks Based Intra Prediction for HEVC
abstract
Summary form only given. Traditional intra prediction methods for HEVC rely on using the nearest reference lines for predicting a block, which ignore much richer context between the current block and its neighboring blocks and therefore cause inaccurate prediction especially when weak spatial correlation exists between the current block and the reference lines. To overcome this problem, in this paper, an intra-prediction convolutional neural network (IPCNN) is proposed for intra prediction, which exploits the rich context of the current block and therefore is capable of improving the accuracy of predicting the current block. Meanwhile, the reconstruction of the three nearest blocks can also be refined. To the best of our knowledge, this is the first paper that directly applies CNNs to intra prediction for HEVC. Experimental results validate the effectiveness of applying CNNs to intra prediction and the proposed method can achieve 0.70% bitrate reduction compared to HEVC reference software HM-14.0.
Wenxue Cui, Tao Zhang 0013, Shengping Zhang, Feng Jiang 0001, Wangmeng Zuo, Zhaolin Wan, Debin Zhao
DCC6
2017 Reduced Reference Image Quality Assessment Based on Entropy of Classified Primitives
abstract
The human visual perception is a layered progressive process that brain assimilates visual information gradually, from primary information, structural information to detailed information. Recently, the visual primitives (atoms in the dictionary) extracted by sparse representation have been shown to be highly related to the layered progressive process of human visual perception. In this paper, the visual primitives are first classified into three categories: DCprimary, sketch and texture in terms of their inherent properties regarding tothe perceptual information. Then, we propose a novel reduced reference (RR) image quality assessment (IQA) metric using perceptual information represented by entropy of classified primitives (EoCP). Specifically, EoCP is a measurement of the distribution statistics of the visual primitives, which can represent the visual information. The differences of EoCPs between the reference image and its distorted version are calculated as features to characterize perceptual loss. The extracted features (only three scalars) are used to compute the quality score by a prediction function which is trained using support vector regression(SVR). Experimental results on LIVE, CSIQ and TID2013 image databases demonstrate that the proposed metric achieves high consistency with the human perception and show competitive performance with state-of-the-art IQA metrics.
Zhaolin Wan, Yutao Liu 0002, Debin Zhao
DCC1
2017 Global quality of assessment and optimization for the backward-compatible stereoscopic display system
abstract
The backward-compatible stereoscopic display is a technology that a stereoscopic view is perceived with 3D glasses while a 2D version of the 3D image is concurrently available for naked-eye viewers on the same physical display medium. This unique functionality is achieved by an information display technology Temporal Psychovisual Modulation (TPVM), an interesting interplay between high refresh rate optoelectronic display, signal processing and psychophysics. However, the current performance of the system is not satisfactory, and it is a trade-off keeping simultaneously the best performance of the 3D view and the 2D view. The global quality of 3D scene and 2D scene plays a great important impact on user's quality-of-experience. In this paper, we are the first to put forward the concept of global quality, including two components: the quality of 3D view and 2D view. Then we have constructed the display system prototype and first proposed quality assessment criteria to evaluate the quality of both 3D scene and 2D scene, towards as guidance to improve their performances. We conduct subjective experiments to figure out when the system runs the best under the criteria. Experimental results demonstrate that we can effectively calculate the optimal quality of this system using the quality assessment criteria.
Yuanchun Chen, Guangtao Zhai, Jiantao Zhou 0001, Zhaolin Wan
ICIP4
2017 Reduced reference stereoscopic image quality assessment based on entropy of classified primitives
abstract
Stereoscopic vision is a complex system which receives and integrates perceptual information from both monocular and binocular cues. In this paper, a novel reduced-reference stereoscopic image quality assessment scheme is proposed, based on the visual perceptual information measured by entropy of classified primitives (EoCP) and mutual information of classified primitives (MIoCP), named as DCprimary, sketch and texture primitives respectively, which is in accordance with the hierarchical progressive process of human visual perception. Specifically, EoCP of each-view image are calculated as monocular cue, and MIoCP between two-view images is derived as binocular cue. The Maximum (MAX) mechanism is applied to determine the perceptual information. The perceptual information differences between the original and distorted images are used to predict the stereoscopic image quality by support vector regression (SVR). Experimental results on LIVE phase II asymmetric database validate the proposed metric achieves significantly higher consistency with subjective ratings and outperforms state-of-the-art stereoscopic image quality assessment methods.
Zhaolin Wan, Yutao Liu 0002, Debin Zhao
ICME1