Yu Zhou 0009

dblp:36/2728-9 · DBLP profile ↗
← Back
25ranked-venue papers
11as first author
12since 2021 · last 2026
0000-0002-8375-0784ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 10 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Temporal Consistency and Variation-Guided Spatio-Temporal Aggregation for Few-Shot Action Recognition
abstract
Few-shot Action Recognition (FSAR) aims to recognize novel actions from only a few labeled examples, posing challenges due to limited supervision and complex temporal dynamics. Existing methods often adopt a unified motion modeling strategy for both short- and long-term dynamics, overlooking the need to adapt motion pattern extraction to the specific temporal properties inherent to different timescales. This forces models to hedge against multi-scale relevance through exhaustive searches over temporal tuples, followed by heavy spatio-temporal fusion, which substantially increases parameters and computation and ultimately limits efficiency. To this end, we propose the efficient Temporal Consistency and Variation-Guided Spatio-Temporal Aggregation Network (TCV-STA), which comprises four key components: the Temporal Consistency Module (TCM), the Temporal Variation Module (TVM), the Spatio-Temporal Aggregation attention (STA), and the Shifted Window Temporal Attention (SWTA). The TCM captures stable motion patterns to suppress short-term perturbations and enhance temporal consistency for robust motion representation, while the TVM models dynamic motion patterns to highlight long-term variations that improve inter-class discriminability and facilitate intra-class alignment. Built upon these complementary motion cues, the STA selectively aggregates spatial and temporal representations under the guidance of the learned stable and dynamic motion patterns, avoiding global dense fusion. Finally, to address the limited receptive field and discontinuous modeling caused by frame grouping in TCM and TVM, we adapt a SWTA to capture longer-range temporal dependencies and ensure smooth transitions across subaction segments for few-shot action recognition. Experiments demonstrate that TCV-STA achieves competitive accuracy across four widely-used FSAR benchmarks while reducing parameters by up to 27.9% and computational cost by 21.3%, striking a favorable balance between accuracy and efficiency for deployment in resource-constrained scenarios.
Kaiwen Dong, Quanyi Li, Yanjing Sun, Xiao Yun, Yu Zhou 0009, Kévin Riou, Xiaofeng Hou, Patrick Le Callet
IEEE Trans. Circuits Syst. Video Technol.5
2025 A Robust Quality Evaluator for Panoramic Videos
abstract
Most of the existing methods to evaluate the quality of panoramic content mainly focus on studying the quality evaluation of static panoramic images, rather than the more widely used dynamic panoramic videos. Also, the few panoramic video quality metrics that are available have obvious weaknesses in terms of robustness. To this end, we propose a robust quality evaluator for panoramic videos (RQE-PV). First, a viewport prediction module is proposed by developing a multi-step fixation screening mechanism to simulate the characteristics of limited visual range and excavate the process of human eye movement. Further, both the temporal and spatial features are explored and fused for quality prediction. The superior robustness of the proposed method has been demonstrated on two public panoramic video databases.
Jiabao Feng, Yu Zhou 0009, Lijuan Tang, Ruirui Chen 0001, Yanjing Sun, Jicun Ding
ICASSP3
2025 Time Switching Protocol based Wireless-Powered OAM Communications for Green IoT
abstract
In recent years, Internet of Things (IoT) develops rapidly all over the world, and the massive IoT equipment (IE) is envisioned to consume huge amount of energy, thus becoming the leading energy guzzler in future 6G communications. As the green energy technology, the time switching (TS) protocol can support the perpetual energy supply for the IE. Furthermore, orbital angular momentum (OAM) can provide the IE with new dimensional orthogonal resource. Therefore, this paper proposes the TS protocol based wireless-powered OAM communications for green IoT, which can transfer power and multiple independent information simultaneously with different OAM modes. For the energy-limited IE, we first design the framework of the TS protocol based wireless-powered OAM communications. Then, the capacity maximization problem is formulated with the quality-of-service requirement of the IE. Finally, we prove that the formulated capacity maximization problem is convex, and derive the closed-form expression of the optimal TS factor. Simulation results are presented to demonstrate the performance of the proposed TS protocol based wireless-powered OAM communications.
Ruirui Chen 0001, Jiale Zheng, Yu Zhou 0009
VTC2025-Fall5
2025 Multi-Task Guided No-Reference Omnidirectional Image Quality Assessment With Feature Interaction
abstract
Omnidirectional image quality assessment (OIQA) has become an increasingly vital problem in recent years. Most previous no-reference OIQA methods only extract local features from the distorted viewports, or extract global features from the entire distorted image, lacking the interaction and fusion between local and global features. Moreover, the lack of reference information also limits their performance. Thus, we propose a no-reference OIQA model which consists of three novel modules, including a bidirectional pseudo-reference module, a Mamba-based global feature extraction module, and a multi-scale local-global feature aggregation module. Specifically, by considering the image distortion degradation process, a bidirectional pseudo-reference module capturing the error maps on viewports is first constructed to refine the multi-scale local visual features, which can supply rich quality degradation reference information without the reference image. To well complement the local features, the VMamba module is adopted to extract the representative multi-scale global visual features. Inspired by human hierarchical visual perception characteristics, a novel multi-scale aggregation module is built to strengthen the feature interaction and effective fusion which can extract deep semantic information. Finally, motivated by the multi-task managing mechanism of human brain, a multi-task learning module is introduced to assist the main quality assessment task by digging the hidden information in compression type and distortion degree. Extensive experimental results demonstrate that our proposed method achieves the state-of-the-art performance on the no-reference OIQA task compared to other models.
Yun Liu 0009, Huiyu Duan, Yu Zhou 0009, Daoxin Fan, Guangtao Zhai
IEEE Trans. Circuits Syst. Video Technol.4
2024 Quality Assessment for Stitched Panoramic Images via Patch Registration and Bidimensional Feature Aggregation
abstract
Quality assessment for stitched panoramic images (SPIQA) is of great significance for the stitching algorithm optimization. By contrast, this task is much more challenging and arduous than traditional IQA task due to the high resolution of stitched panoramic images and the particularity and complexity of stitching distortions. For this task, we propose an effective method based on patch registration and bidimensional feature aggregation (PRBFA). First, inspired by the attention mechanism of the human visual system and the limited range of human vision, a soft patch segmentation and selection method is presented to determine the key patches in panoramic images to participate in the following patch matching and feature alignment stages, achieving patch registration between the panoramic image and the corresponding constituent images. Further, to fully simulate the human visual perception process from local viewport to panorama, the feature exploration is successively performed from local to global, which is also adaptive to the complexity of the distortions in stitched panoramic images. For performance testification, extensive experiments are conducted on the publicly released SPIQA database, the results of which prove the performance superiority of the PRBFA method.
Yu Zhou 0009, Weikang Gong, Yanjing Sun, Leida Li, Ke Gu 0001, Jinjian Wu
IEEE Trans. Multim.1
2023 Bi-level deep mutual learning assisted multi-task network for occluded person re-identification
abstract
Abstract An occluded person re‐identification (ReID) approach is presented by constructing a Bi‐level deep Mutual learning assisted Multi‐task network (BMM), where the holistic and occluded person ReID tasks are treated as two related but not identical tasks. This is inspired by the human perception characteristic that there exist both similarities and differences when human views a holistic image and the occluded one. Specifically, a multi‐task network with two branches is designed, where the convolutional neural network based feature representation part shares the weights by two tasks for commonality extraction, while the following output layers have respective weights for difference representation. Furthermore, as the non‐occluded regions convey discriminative information, a bi‐level mutual learning strategy is proposed and applied mutually on two branches to obtain more effective information from the non‐occluded regions in the occluded images for better identity recognition. This is achieved by both feature‐level and output‐level mutual loss functions. Extensive experiments prove the advantages of the BMM for person ReID.
Yi Wang 0105, Liangbo Wang, Yu Zhou 0009
IET Image Process.3
2023 Pyramid Feature Aggregation for Hierarchical Quality Prediction of Stitched Panoramic Images
abstract
Panoramic image quality assessment (PIQA) is crucial to the successful application of technologies that can provide immersive visual experience. Stitching distortions are one of the main types of distortions that result in panoramic image degradation. However, most existing PIQA methods are general-purpose ones, which ignore the special characteristics of the stitching distortions caused by imperfect stitching algorithms. This results in unsatisfactory performance. To this end, we propose an effective stitched PIQA method, which consists of an imaginary reference generation (IRG) module and a hierarchical quality prediction (HQP) module. Among them, the IRG module is proposed to mimic the capability of the human visual system in imagining the raw version in the face of a degraded image. For the IRG module learning, we construct a large-scale database. The HQP module is presented to adapt to the particularity and complexity of stitching distortions, which is achieved by the pyramid feature aggregation. Extensive experiments and comparisons have been performed on the stitched PIQA database and the experimental results demonstrate the superiority of the proposed method in evaluating the quality of stitched panoramic images.
Yu Zhou 0009, Weikang Gong, Yanjing Sun, Leida Li, Jinjian Wu, Xinbo Gao 0001
IEEE Trans. Multim.1
2022 Occluded person re-identification based on differential attention siamese network
Liangbo Wang, Yu Zhou 0009, Yanjing Sun, Song Li 0001
Appl. Intell.2
2022 Knowledge self-distillation for visible-infrared cross-modality person re-identification
Yu Zhou 0009, Yanjing Sun, Kaiwen Dong, Song Li 0001
Appl. Intell.1
2022 Omnidirectional Image Quality Assessment by Distortion Discrimination Assisted Multi-Stream Network
abstract
Omnidirectional image (OI) quality assessment is crucial to facilitate the development of virtual reality (VR) related technology. In this work, a distortion discrimination assisted multi-stream network is proposed for OI quality assessment. The multi-stream architecture is constructed by generating the viewport images received by the retina at one point to simulate the characteristics of humans perceiving VR contents. Additionally, the strategy of generating several viewport image sets from one OI is proposed for data augmentation. Furthermore, the facts that the human brain has the ability for both quality assessment and distortion type distinguishment, and the process of human brain handling two tasks exists information interaction inspire us to employ an auxiliary distortion discrimination task to facilitate the quality assessment task learning. Extensive experiments conducted on two public OI databases demonstrate the superiority of the proposed method to both traditional 2D quality metrics and existing metrics specific for OIs. Moreover, utilizing the assistant task is proven to be more effective than the single task learning for OI quality evaluation. Better generalization performance is also verified to be another valuable trait of the proposed method.
Yu Zhou 0009, Yanjing Sun, Leida Li, Ke Gu 0001, Yuming Fang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2021 A confidence prior for image dehazing
Feiniu Yuan, Yu Zhou 0009, Xue Xia 0005, Xueming Qian
Pattern Recognit.2
2021 Quality Index for View Synthesis by Measuring Instance Degradation and Global Appearance
abstract
Virtual view synthesis plays a vital role in the application of multi-view and free-viewpoint videos. Depth-image-based rendering (DIBR) is the most commonly used approach in view synthesis, and many DIBR algorithms have been proposed. However, how to evaluate the quality of DIBR-synthesized images and benchmark the DIBR algorithms are still very challenging, which may hinder the further development of the view synthesis technique. Hence, an effective quality metric for evaluating the distortions in view synthesis is urgently needed. With this motivation, this paper presents a quality index for view synthesis by simultaneously measuring local Instance DEgradation and global Appearance (IDEA). Due to the imperfection of rendering algorithms, local geometric distortions are easily introduced around instance contours, causing instance degradation, which is the dominant distortion in synthesized views. In this work, image instances are first detected and local instance degradation is measured based on discrete orthogonal moments. Meantime, we propose to measure the global appearance of synthesized images based on the superpixel representation. By integrating both local and global aspects of the distortions, a more accurate quality model is built for view synthesis. Extensive experiments and comparisons have demonstrated the superiority of the proposed method in evaluating the quality of DIBR-synthesized images and benchmarking the performance of view synthesis algorithms.
Leida Li, Yu Zhou 0009, Jinjian Wu, Fu Li 0002, Guangming Shi
IEEE Trans. Multim.2
2020 Image dehazing based on a transmission fusion strategy by automatic image matting
Feiniu Yuan, Yu Zhou 0009, Xue Xia 0005, Jinting Shi, Yuming Fang 0001, Xueming Qian
Comput. Vis. Image Underst.2
2020 No-reference quality assessment for live broadcasting videos in temporal and spatial domains
abstract
Nowadays, live broadcasting video has become increasingly popular and high‐quality live broadcasting video is highly needed. In practice, live broadcasting videos usually undergo several processing stages, which inevitably introduce multiple distortions, e. g. frame freezing and intensity mutation, causing the degraded quality of experience. However, little work has been done to the quality evaluation of live broadcasting videos, which may hinder the further development of more advanced live broadcasting video delivery systems. Motivated by this, this study presents a no‐reference quality evaluation model for live broadcasting videos (LBVQA) in temporal and spatial domains. In the temporal domain, statistic features are extracted to measure the frame freezing and intensity mutation, and the entropy‐based feature is extracted to describe the global jitter. In the spatial domain, blurring is measured based on phase coherence, and abnormal exposure ratio is calculated based on an adaptive threshold. Finally, all features are fed into a backpropagation neural network to train the quality prediction model. Experimental results on the Live Broadcasting Video Database demonstrate the advantages of the proposed metric over the state‐of‐the‐art image and video quality metrics.
Yipo Huang, Leida Li, Yu Zhou 0009, Bo Hu 0008
IET Image Process.3
2020 Blind Realistic Blur Assessment Based on Discrepancy Learning
abstract
Blur is one of the most common distortions that degrade natural images. This stimulates the blossom of sharpness assessment metrics. Existing sharpness metrics possess good performance for evaluating simulated blur, but are limited for the more common realistic blur that are introduced during image capture and processing in real life. To this end, we propose an effective Realistic Blur Assessment method (RBA) based on discrepancy learning. First, motivated by the fact that the distortion-free reference images are usually unavailable in practice, but the Human Visual System (HVS) can still accurately perceive image sharpness by quantifying the perceptual discrepancy between the distorted image and the hallucinated reference image in mind, we propose to train a discrepancy generation model to automatically generate the discrepancy map from the distorted image analogous to the HVS. This is achieved by using a deep neural network with rich training images. With the discrepancy map, two sharpness-aware features, i.e. sparse representation based entropy of primitive and content-guided variation of power, are then extracted to severally quantify spatial visual information amount and spectral power. Finally, the two features are integrated to produce the overall sharpness score. Extensive experiments demonstrate the superiority of the proposed method over the state-of-the-arts.
Leida Li, Yu Zhou 0009, Ke Gu 0001, Yuzhe Yang 0001, Yuming Fang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2019 No-reference quality assessment for contrast-distorted images based on multifaceted statistical representation of structure
Yu Zhou 0009, Leida Li, Hancheng Zhu, Hantao Liu, Shiqi Wang 0001, Yao Zhao 0001
J. Vis. Commun. Image Represent.1
2019 Quality assessment for view synthesis using low-level and mid-level structural representation
Yu Zhou 0009, Leida Li, Suiyi Ling, Patrick Le Callet
Signal Process. Image Commun.1
2019 No-Reference Quality Assessment for View Synthesis Using DoG-Based Edge Statistics and Texture Naturalness
abstract
View synthesis is a key technique in free-viewpoint video, which renders virtual views based on texture and depth images. The distortions in synthesized views come from two stages, i.e., the stage of the acquisition and processing of texture and depth images, and the rendering stage using depth-image-based-rendering (DIBR) algorithms. The existing view synthesis quality metrics are designed for the distortions caused by a single stage, which cannot accurately evaluate the quality of the entire view synthesis process. With the considerations that the distortions introduced by two stages both cause edge degradation and texture unnaturalness, and the Difference-of-Gaussian (DoG) representation is powerful in capturing image edge and texture characteristics by simulating the center-surrounding receptive fields of retinal ganglion cells of human eyes, this paper presents a no-reference quality index for Synthesized views using DoG-based Edge statistics and Texture naturalness (SET). To mimic the multi-scale property of the Human Visual System (HVS), DoG images are first calculated at multiple scales. Then the orientation selective statistics features and the texture naturalness features are calculated on the DoG images and the coarsest scale image, producing two groups of quality-aware features. Finally, the quality model is learnt from these features using the random forest regression model. Experimental results on two view synthesis image databases demonstrate that the proposed metric is advantageous over the relevant state-of-the-arts in dealing with the distortions in the whole view synthesis process.
Yu Zhou 0009, Leida Li, Shiqi Wang 0001, Jinjian Wu, Yuming Fang 0001, Xinbo Gao 0001
IEEE Trans. Image Process.1
2018 No-reference quality assessment of DIBR-synthesized videos by measuring temporal flickering
Yu Zhou 0009, Leida Li, Shiqi Wang 0001, Jinjian Wu, Yun Zhang 0002
J. Vis. Commun. Image Represent.1
2018 Reduced-reference quality assessment of DIBR-synthesized images based on multi-scale edge intensity similarity
Yu Zhou 0009, Leida Li, Ke Gu 0001, Lijuan Tang
Multim. Tools Appl.1
2018 Quality Assessment of DIBR-Synthesized Images by Measuring Local Geometric Distortions and Global Sharpness
abstract
Depth-image-based rendering (DIBR) is a fundamental technique in free viewpoint video, which is widely adopted to synthesize virtual viewpoints. The warping and rendering operations in DIBR generally introduce geometric distortions and sharpness change. The state-of-the-art quality indices are limited in dealing with such images since they are sensitive to geometric changes. In this paper, a new quality model for DIBR-synthesized view images is presented by measuring LOcal Geometric distortions in disoccluded regions and global Sharpness (LOGS). A disoccluded region detection method is first proposed using SIFT-flow-based warping. Then, the sizes and distortion strength of local disoccluded regions are combined to generate a score. Furthermore, a reblurring-based strategy is proposed to quantify the global sharpness. Finally, the overall quality score is calculated by pooling the scores of local disoccluded regions and global sharpness. Experiments on four public DIBR-synthesized image/video databases show the superiority of the proposed metric over the state-of-the-art quality models. The proposed method is further adopted for boosting the performances of existing quality metrics and benchmarking DIBR algorithms, both achieving very promising results.
Leida Li, Yu Zhou 0009, Ke Gu 0001, Weisi Lin, Shiqi Wang 0001
IEEE Trans. Multim.2
2018 Blind Quality Index for Multiply Distorted Images Using Biorder Structure Degradation and Nonlocal Statistics
abstract
In the past decade, extensive image quality metrics have been proposed. The majority of them are tailored for the images that contain a specific type of distortion. However, in practice, the images are usually degraded by different types of distortions simultaneously. This poses great challenges to the existing quality metrics. Motivated by this, this paper proposes a no-reference quality index for the multiply distorted images using the biorder structure degradation and the nonlocal statistics. The design philosophy is inspired by the fact that the human visual system (HVS) is highly sensitive to the degradations of both the spatial contrast and the spatial distribution, which are prone to be changed by the joint effects of the multiple distortions. Specifically, the multiresolution representation of the image is first built by downsampling to simulate the hierarchical property of the HVS. Then, the structure degradation is calculated to measure the spatial contrast. Considering the fact that the human visual cortex has the separate mechanisms to perceive the first- and second-order structures, dubbed biorder structures, the degradations of biorder structures are calculated to account for the spatial contrast, producing the first group of the quality-aware features. Furthermore, the nonlocal self-similarity statistics is calculated to measure the spatial distribution, producing the second group of features. Finally, all the features are fed into the random forest regression model to learn the quality model for the multiply distorted images. Extensive experimental results conducted on the three public databases demonstrate the superiority of the proposed metric to the state-of-the-art metrics. Moreover, the proposed metric is also advantageous over the existing metrics in terms of the generalization ability.
Yu Zhou 0009, Leida Li, Jinjian Wu, Ke Gu 0001, Weisheng Dong, Guangming Shi
IEEE Trans. Multim.1
2016 Quality assessment of 3D synthesized images via disoccluded region discovery
abstract
Depth-Image-Based-Rendering (DIBR) is fundamental in free-viewpoint 3D video, which has been widely used to generate synthesized views from multi-view images. The majority of DIBR algorithms cause disoccluded regions, which are the areas invisible in original views but emerge in synthesized views. The quality of synthesized images is mainly contaminated by distortions in these disoccluded regions. Unfortunately, traditional image quality metrics are not effective for these synthesized images because they are sensitive to geometric distortions. To solve the problem, this paper proposes an objective quality evaluation method for 3D Synthesized images via Disoccluded Region Discovery (SDRD). A self-adaptive scale transform model is first adopted to preprocess the images on account of the impacts of view distance. Then disoccluded regions are detected by comparing the absolute difference between the preprocessed synthesized image and the warped image of preprocessed reference image. Furthermore, the disoccluded regions are weighted by a weighting function proposed to account for the varying sensitivities of human eyes to the size of disoccluded regions. Experiments conducted on IRCCyN/IVC DIBR image database demonstrate that the proposed SDRD method remarkably outperforms traditional 2D and existing DIBR-related quality metrics.
Yu Zhou 0009, Leida Li, Ke Gu 0001, Yuming Fang 0001, Weisi Lin
ICIP1
2016 No-reference quality assessment of deblocked images
Leida Li, Yu Zhou 0009, Weisi Lin, Jinjian Wu, Xinfeng Zhang 0001, Beijing Chen
Neurocomputing2
2015 GridSAR: Grid strength and regularity for robust evaluation of blocking artifacts in JPEG images
Leida Li, Yu Zhou 0009, Jinjian Wu, Weisi Lin, Haoliang Li
J. Vis. Commun. Image Represent.2