VLDB 2026 Research / reviewers in the wild / expert
Yifan Zuo 0001
dblp:116/7225-1
· DBLP profile ↗
49ranked-venue papers
18as first author
34since 2021 · last 2026
0000-0003-4980-7211ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 40 · 15 first-author · 28 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 11 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PanFoMa: A Lightweight Foundation Model and Benchmark for Pan-CancerabstractSingle-cell RNA sequencing (scRNA-seq) is essential for decoding tumor heterogeneity. However, pan-cancer research still faces two key challenges: learning discriminative and efficient single-cell representations, and establishing a comprehensive evaluation benchmark. In this paper, we introduce \algoname, a lightweight hybrid neural network that combines the strengths of Transformers and state-space models to achieve a balance between performance and efficiency. \algoname consists of a front-end local-context encoder with shared self-attention layers to capture complex, order-independent gene interactions; and a back-end global sequential feature decoder that efficiently integrates global context using a linear-time state-space model. This modular design preserves the expressive power of Transformers while leveraging the scalability of Mamba to enable transcriptome modeling, effectively capturing both local and global regulatory signals. To enable robust evaluation, we also construct a large-scale pan-cancer single-cell benchmark, \algoname Bench, containing over 3.5 million high-quality cells across 33 cancer subtypes, curated through a rigorous preprocessing pipeline. Experimental results show that \algoname outperforms state-of-the-art models on our pan-cancer benchmark (+4.0\%) and across multiple public tasks, including cell type annotation (+7.4\%), batch integration (+4.0\%) and multi-omics integration (+3.1\%). Xiaoshui Huang, Tianlin Zhu, Yifan Zuo 0001, Xue Xia 0005, Zonghan Wu, Jiebin Yan, Dingli Hua, Zongyi Xu, Yuming Fang 0001, Jian Zhang 0002 |
AAAI | 3 |
| 2026 | Inter-modality feature prediction through multimodal fusion for 3D shape defect detection
Mujtaba Asad, Waqar Azeem, Hafiz Tayyab Mustafa, Yuming Fang 0001, Jie Yang 0002, Yifan Zuo 0001, Wei Liu 0044 |
Neural Networks | 6 |
| 2026 | CMF-IoU: Multi-Stage Cross-Modal Fusion 3D Object Detection With IoU Joint PredictionabstractMulti-modal methods based on camera and LiDAR sensors have garnered significant attention in the field of 3D detection. However, many prevalent works focus on single or partial stage fusion, leading to insufficient feature extraction and suboptimal performance. In this paper, we introduce a multi-stage cross-modal fusion 3D detection framework, termed CMF-IOU, to effectively address the challenge of aligning 3D spatial and 2D semantic information. Specifically, we first project the pixel information into 3D space via a depth completion network to get the pseudo points, which unifies the representation of the LiDAR and camera information. Then, a bilateral cross-view enhancement 3D backbone is designed to encode LiDAR points and pseudo points. The first sparse-to-distant (S2D) branch utilizes an encoder-decoder structure to reinforce the representation of sparse LiDAR points. The second residual view consistency (ResVC) branch is proposed to mitigate the influence of inaccurate pseudo points via both the 3D and 2D convolution processes. Subsequently, we introduce an iterative voxel-point aware fine grained pooling module, which captures the spatial information from LiDAR points and textural information from pseudo points in the proposal refinement stage. To achieve more precise refinement during iteration, an intersection over union (IoU) joint prediction branch integrated with a novel proposals generation technique is designed to preserve the bounding boxes with both high IoU and classification scores. Extensive experiments show the superior performance of our method on the KITTI, nuScenes and Waymo datasets. The code is available at https://github.com/pami-zwning/CMF-IOU. Zhiwei Ning, Zhaojiang Liu, Xuanang Gao, Yifan Zuo 0001, Jie Yang 0002, Yuming Fang 0001, Wei Liu 0044 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | PSReg: Prior-guided Sparse Mixture of Experts for Point Cloud RegistrationabstractThe discriminative feature is crucial for point cloud registration. Recent methods improve the feature discriminative by distinguishing between non-overlapping and overlapping region points. However, they still face challenges in distinguishing the ambiguous structures in the overlapping regions. Therefore, the ambiguous features they extracted resulted in a significant number of outlier matches from overlapping regions. To solve this problem, we propose a prior-guided SMoE-based registration method to improve the feature distinctiveness by dispatching the potential correspondences to the same experts. Specifically, we propose a prior-guided SMoE module by fusing prior overlap and potential correspondence embeddings for routing, assigning tokens to the most suitable experts for processing. In addition, we propose a registration framework by a specific combination of Transformer layer and prior-guided SMoE module. The proposed method not only pays attention to the importance of locating the overlapping areas of point clouds, but also commits to finding more accurate correspondences in overlapping areas. Our extensive experiments demonstrate the effectiveness of our method, achieving state-of-the-art registration recall (95.7%/79.3%) on the 3DMatch/3DLoMatch benchmark. Moreover, we also test the performance on ModelNet40 and demonstrate excellent performance. Xiaoshui Huang, Zhou Huang 0006, Yifan Zuo 0001, Yongshun Gong, Chengdong Zhang, Deyang Liu, Yuming Fang 0001 |
AAAI | 3 |
| 2025 | PointGAC: Geometric-Aware Codebook for Masked Point Cloud ModelingabstractMost masked point cloud modeling (MPM) methods follow a regression paradigm to reconstruct the coordinate or feature of masked regions. However, they tend to over-constrain the model to learn the details of the masked region, resulting in failure to capture generalized features. To address this limitation, we propose \textbf{\textit{PointGAC}}, a novel clustering-based MPM method that aims to align the feature distribution of masked regions. Specially, it features an online codebook-guided teacher-student framework. Firstly, it presents a geometry-aware partitioning strategy to extract initial patches. Then, the teacher model updates a codebook via online k-means based on features extracted from the complete patches. This procedure facilitates codebook vectors to become cluster centers. Afterward, we assigns the unmasked features to their corresponding cluster centers, and the student model aligns the assignment for the reconstructed masked features. This strategy focuses on identifying the cluster centers to which the masked features belong, enabling the model to learn more generalized feature representations. Benefiting from a proposed codebook maintenance mechanism, codebook vectors are actively updated, which further increases the efficiency of semantic feature learning. Experiments validate the effectiveness of the proposed method on various downstream tasks. Code is available at https://github.com/LAB123-tech/PointGAC Abiao Li, Chenlei Lv, Yuming Fang 0001, Yifan Zuo 0001, Jian Zhang 0002, Guofeng Mei |
ICCV | 4 |
| 2025 | Dynamic 3D Gaussian Reconstruction with Specular Reflectionabstract3D Gaussian Splatting (3DGS) has shown remarkable potential in novel view synthesis. However, it still encounters significant challenges in reconstructing dynamic scenes, particularly when dealing with reflective surfaces. To address this issue, we propose a novel 3DGS-based method for dynamic scene reconstruction with explicit reflection modeling. Our approach integrates deferred shading with a dual-environment map that combines static and dynamic components, enabling effective modeling of specular reflections. This allows our method to capture both steady and temporally varying lighting, resulting in more realistic renderings. We evaluate the proposed method on the benchmark dynamic reflection dataset, NERF-DS, and compare it with state-of-the-art approaches. Experimental results show that our method achieves superior or comparable performance in terms of PSNR, SSIM, and LPIPS metrics compared to competing approaches. Mingyang Zhao 0001, Yuanzhi Xu, Yifan Zuo 0001, Xiaoshui Huang, Yuming Fang 0001 |
ICIP | 3 |
| 2025 | Max360IQ: Blind omnidirectional image quality assessment with multi-axis attention
Jiebin Yan, Ziwen Tan, Yuming Fang 0001, Jiale Rao, Yifan Zuo 0001 |
Pattern Recognit. | 5 |
| 2025 | Learning Stage-wise Fusion Transformer for light field saliency detection
Wenhui Jiang 0001, Qi Shu, Hongwei Cheng, Yuming Fang 0001, Yifan Zuo 0001 |
Pattern Recognit. Lett. | 5 |
| 2025 | Perceptual Transform Fusion of Infrared and Visible ImagesabstractInfrared and visible image fusion aims to generate fused images with rich textures and clear target representations. Existing methods generally assume high-quality input images, thus overlooking issues such as reduced contrast and loss of details in visible images under low-light conditions. The naive enhance-then-fuse strategy cannot perform fuse-oriented image enhancement, which always reaches a sub-optimal result. To address this challenge, we propose a perceptual transform fusion of infrared and visible images, which simultaneously optimizes low-light enhancement and image fusion. Specifically, to improve computational efficiency and optimize key feature representations while suppressing noise interactions caused by lighting variations, we introduce a lightweight adaptive sparse Transformer block (ASTBlock). This model adaptively integrates sparse and dense attention mechanisms to enhance feature representations and employs a feed-forward network to eliminate redundant information, thereby ensuring the quality of image fusion. Subsequently, to retain significant details while reducing the impact of noise introduced by low-light enhancement, we incorporate discrete wavelet transform (DWT) for feature decomposition and fusion, further enhancing the representation capability and feature preservation of fused images. Meanwhile, to tackle the issues of insufficient contrast and hidden details in low-light conditions, we design an illumination perception module and an illumination consistency loss to improve the contrast and clarity of fused images. Experimental results on multiple public benchmark datasets for quality assessment and downstream tasks, e.g., pedestrian detection, demonstrate that our method significantly outperforms the state-of-the-art (SOTA) methods. The code is available at https://github.com/hinmouc/PIVFusion. Dingli Hua, Qingmao Chen, Zhiliang Wu, Yifan Zuo 0001, Wenying Wen, Yuming Fang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Learning Guided Implicit Depth Function With Scale-Aware Feature FusionabstractRecently, the single image super-resolution based on implicit image function is a hot topic, which learns a universal model for arbitrary upsampling scales. By contrast, color-guided depth map super-resolution is less explored based on implicit function learning. The related research faces three questions. First, is it also necessary and applicable to fuse the depth feature and the color feature in the encoder with continuous upsampling scales? Second, is the scale information in the encoder as important as that in the decoder? Third, how to efficiently and effectively model the affinity of location distance and content similarity within cross domains in the decoder? This paper proposes a transformer-based network to answer the above questions, which includes a depth super-resolution branch and a guidance extraction branch. Specifically, in the encoder, the effective implicit cross transformer is designed to fuse the guidance from the color feature with continuous coordinate mapping. In addition, the unrelated guidance is filtered out by correlation evaluation in the high-dimension feature space. Unlike the scale only introduced in the decoder, this paper additionally embeds the scale into the position encoding and the feed-forward network in the encoder to learn the scale-aware feature representation. In the decoder, the high-resolution depth feature is reconstructed by using the internal prior and the external guidance. The internal prior is implemented by implicit self-attention in the depth super-resolution branch, and the external guidance is exploited via implicit cross-attention between both branches. Finally, the above decoded features are complementary to generate the high-resolution depth map. The sufficient experiments on the synthetic and real datasets for in-distribution and out-of-distribution upsampling scales validate the improved performance. The code and the models are public via https://github.com/NaNRan13/GIDF. Yifan Zuo 0001, Yuming Fang 0001, Jiebin Yan, Wenhui Jiang 0001, Yuxin Peng 0001, Yan Huang 0023 |
IEEE Trans. Image Process. | 1 |
| 2025 | Subjective and Objective Quality Assessment of Non-Uniformly Distorted Omnidirectional ImagesabstractOmnidirectional image quality assessment (OIQA) has been one of the hot topics in IQA with the continuous development of VR techniques, and achieved much success in the past few years. However, most studies devote themselves to the uniform distortion issue, i.e., all regions of an omnidirectional image are perturbed by the “same amount” of noise, while ignoring the non-uniform distortion issue, i.e., partial regions undergo “different amount” of perturbation with the other regions in the same omnidirectional image. Additionally, nearly all OIQA models are verified on the platforms containing a limited number of samples, which largely increases the over-fitting risk and therefore impedes the development of OIQA. To alleviate these issues, we elaborately explore this topic from both subjective and objective perspectives. Specifically, we construct a large OIQA database containing 10,320 non-uniformly distorted omnidirectional images, each of which is generated by considering quality impairments on one or two camera len(s). Then we meticulously conduct psychophysical experiments and delve into the influence of both holistic and individual factors (i.e., distortion range and viewing condition) on omnidirectional image quality. Furthermore, we propose a perception-guided OIQA model for non-uniform distortion by adaptively simulating users' viewing behavior. Experimental results demonstrate that the proposed model outperforms state-of-the-art methods. Jiebin Yan, Jiale Rao, Xuelin Liu, Yuming Fang 0001, Yifan Zuo 0001, Weide Liu |
IEEE Trans. Multim. | 5 |
| 2025 | Computational Analysis of Degradation Modeling in Blind Panoramic Image Quality AssessmentabstractBlind panoramic image quality assessment (BPIQA) has recently brought a new challenge to the visual quality community, due to the complex interaction between immersive content and human behavior. Although many efforts have been made to advance BPIQA from both conducting psychophysical experiments and designing performance-driven objective algorithms, limited content and few samples in those closed sets inevitably would result in shaky conclusions, thereby hindering the development of BPIQA; we refer to it as the easy-database issue. In this article, we present a sufficient computational analysis of degradation modeling in BPIQA to thoroughly explore the easy-database issue , where we carefully design three types of experiments via investigating the gap between BPIQA and blind image quality assessment (BIQA), the necessity of specific design in BPIQA models, and the generalization ability of BPIQA models. From extensive experiments, we find that easy databases narrow the gap between the performance of BPIQA and BIQA models, which is unconducive to the development of BPIQA. And the easy databases make the BPIQA models be closed to saturation; therefore, the effectiveness of the associated specific designs cannot be well verified. Besides, the BPIQA models trained on our recently proposed databases with complicated degradation show better generalization ability. Thus, we believe that much more efforts are highly desired to put into BPIQA from both subjective viewpoint and objective viewpoint. Jiebin Yan, Ziwen Tan, Jiale Rao, Yifan Zuo 0001, Yuming Fang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2024 | Frozen CLIP Transformer Is an Efficient Point Cloud EncoderabstractThe pretrain-finetune paradigm has achieved great success in NLP and 2D image fields because of the high-quality representation ability and transferability of their pretrained models. However, pretraining such a strong model is difficult in the 3D point cloud field due to the limited amount of point cloud sequences. This paper introduces Efficient Point Cloud Learning (EPCL), an effective and efficient point cloud learner for directly training high-quality point cloud models with a frozen CLIP transformer. Our EPCL connects the 2D and 3D modalities by semantically aligning the image features and point cloud features without paired 2D-3D data. Specifically, the input point cloud is divided into a series of local patches, which are converted to token embeddings by the designed point cloud tokenizer. These token embeddings are concatenated with a task token and fed into the frozen CLIP transformer to learn point cloud representation. The intuition is that the proposed point cloud tokenizer projects the input point cloud into a unified token space that is similar to the 2D images. Comprehensive experiments on 3D detection, semantic segmentation, classification and few-shot learning demonstrate that the CLIP transformer can serve as an efficient point cloud encoder and our method achieves promising performance on both indoor and outdoor benchmarks. In particular, performance gains brought by our EPCL are 19.7 AP50 on ScanNet V2 detection, 4.4 mIoU on S3DIS segmentation and 1.2 mIoU on SemanticKITTI segmentation compared to contemporary pretrained models. Code is available at \url{https://github.com/XiaoshuiHuang/EPCL}. Xiaoshui Huang, Zhou Huang 0006, Sheng Li 0020, Wentao Qu, Tong He 0001, Yuenan Hou, Yifan Zuo 0001, Wanli Ouyang |
AAAI | 7 |
| 2024 | GSTran: Joint Geometric and Semantic Coherence for Point Cloud Segmentation
Abiao Li, Chenlei Lv, Guofeng Mei, Yifan Zuo 0001, Jian Zhang 0002, Yuming Fang 0001 |
ICPR (18) | 4 |
| 2024 | CD-iNet: Deep Invertible Network for Perceptual Image Color Difference Measurement
Zhihua Wang 0002, Keshuo Xu, Keyan Ding, Qiuping Jiang, Yifan Zuo 0001, Zhangkai Ni, Yuming Fang 0001 |
Int. J. Comput. Vis. | 5 |
| 2024 | CFNet: Conditional filter learning with dynamic noise estimation for real image denoising
Yifan Zuo 0001, Wenhao Yao, Yifeng Zeng, Yuming Fang 0001, Yan Huang 0023, Wenhui Jiang 0001 |
Knowl. Based Syst. | 1 |
| 2024 | Learning content-aware feature fusion for guided depth map super-resolution
Yifan Zuo 0001, Xiaoshui Huang, Xue Xia 0005, Yuming Fang 0001 |
Signal Process. Image Commun. | 1 |
| 2024 | A2 GSTran: Depth Map Super-Resolution via Asymmetric Attention With Guidance SelectionabstractCurrently, Convolutional Neural Network (CNN) has dominated guided depth map super-resolution (SR). However, the inefficient receptive field growing and input-independent convolution limit the generalization of CNN. Motivated by vision transformer, this paper proposes an efficient transformer-based backbone A2GSTran for guided depth map SR, which resolves the above intrinsic defect of CNN. In addition, state-of-the-art (SOTA) models only refine depth features with the guidance which is implicitly selected without supervision. So, there is no explicit guarantee to mitigate the artifacts of texture copying and edge blurring. Accordingly, the proposed A2GSTran simultaneously solves two sub-problems,i.e., guided monocular depth estimation and guided depth SR, in separate branches. Specifically, the explicit supervision upon monocular depth estimation lifts the efficiency of guidance selection. The feature fusion between branches is designed via bi-directional cross attention. Moreover, since guidance domain is defined in high resolution (HR), we propose asymmetric cross attention to maintain the guidance information via pixel unshuffle instead of pooling which has unequal channel number to depth features. Based on the supervisions to depth reconstruction and guidance selection, the final depth features are refined by fusing the output features of the corresponding branches via channel attention to generate the HR depth map. Sufficient experimental results on synthetic and real datasets for multiple scales validate our contributions compared with SOTA models. The code and models are public via https://github.com/alex-cate/Depth_Map_Super-resolution_via_Asymmetric_Attention_with_Guidance_Selection Yifan Zuo 0001, Yifeng Zeng, Yuming Fang 0001, Xiaoshui Huang, Jiebin Yan |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Image Super-Resolution via Efficient Transformer Embedding Frequency Decomposition With RestartabstractRecently, transformer-based backbones show superior performance over the convolutional counterparts in computer vision. Due to quadratic complexity with respect to the token number in global attention, local attention is always adopted in low-level image processing with linear complexity. However, the limited receptive field is harmful to the performance. In this paper, motivated by Octave convolution, we propose a transformer-based single image super-resolution (SISR) model, which explicitly embeds dynamic frequency decomposition into the standard local transformer. All the frequency components are continuously updated and re-assigned via intra-scale attention and inter-scale interaction, respectively. Specifically, the attention in low resolution is enough for low-frequency features, which not only increases the receptive field, but also decreases the complexity. Compared with the standard local transformer, the proposed FDRTran layer simultaneously decreases FLOPs and parameters. By contrast, Octave convolution only decreases FLOPs of the standard convolution, but keeps the parameter number unchanged. In addition, the restart mechanism is proposed for every a few frequency updates, which first fuses the low and high frequency, then decomposes the features again. In this way, the features can be decomposed in multiple viewpoints by learnable parameters, which avoids the risk of early saturation for frequency representation. Furthermore, based on the FDRTran layer with restart mechanism, the proposed FDRNet is the first transformer backbone for SISR which discusses the Octave design. Sufficient experiments show our model reaches state-of-the-art performance on 6 synthetic and real datasets. The code and the models are available at https://github.com/catnip1029/FDRNet. Yifan Zuo 0001, Wenhao Yao, Yuming Fang 0001, Wei Liu 0044, Yuxin Peng 0001 |
IEEE Trans. Image Process. | 1 |
| 2024 | Visual Security Index Combining CNN and Filter for Perceptually Encrypted Light Field ImagesabstractVisual security index (VSI) represents a quantitative index for the visual security evaluation of perceptually encrypted images. Recently, the research on visual security of encrypted light field (LF) images faces two challenges. One is that the existing perceptually encrypted image databases are often too small, which is easy to cause overfitting in convolutional neural network (CNN). The other is that existing VSI models did not take a full account the intrinsic characteristics of the LF images and highly relied on handcrafted feature extraction. In this article, we construct a new database of perceptually encrypted LF images, called the PE-SLF, which is 2.6 times as big as the existing largest perceptual encrypted image database. Moreover, a novel visual security index (VSI) model is proposed by taking into full consideration the intrinsic spatial-angular characteristics of the LF images and the outstanding capabilities of CNN in feature extraction. First, we exploit CNN to detect the texture and structure features of encrypted sub-aperture images in the spatial domain. Second, we apply the Gabor filter to detect the Gabor feature over the epi-polar plane images in angular domain. Last, the spatial and angular similarity measurements are subsequently calculated for jointly yielding the final visual security score. Experimental results on the constructed PE-SLF demonstrate that the proposed VSI model is closer to the perception of HVS in visual security evaluation of encrypted LF images compared to other classical and state-of-the-art models. Wenying Wen, Minghui Huang, Yushu Zhang 0001, Yuming Fang 0001, Yifan Zuo 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | Laptran: Transformer Embedding Graph Laplacian for Point Cloud Part SegmentationabstractSince the feature representations of the points located at the junction regions of various parts are ambiguous, it is still challenging to exploit the fine-grained semantic features of point clouds on part segmentation tasks. To resolve the issue, we design a modified transformer module, named Laplacian transformer, to investigate the local differences between each point and its corresponding neighbors based on graph Laplacian theory. This module constructs a more accurate local geometric representation of the point cloud. It concentrates on the points located at the junction areas of various parts while boosting the recognition effect of these points. Encapsulated with the Laplacian module, we propose a Unet-like transformer framework to perform part segmentation for point clouds. Experimental results demonstrate that the proposed framework achieves more accurate results on public benchmark datasets. Abiao Li, Chenlei Lv, Yuming Fang 0001, Yifan Zuo 0001 |
ICIP | 4 |
| 2023 | Low complexity inter coding scheme for Versatile Video Coding (VVC)
Xiwu Shang, Xiaoli Zhao 0003, Yifan Zuo 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2023 | Geometry-assisted multi-representation view reconstruction network for Light Field image angular super-resolution
Deyang Liu, Zaidong Tong, Yan Huang 0023, Yifan Zuo 0001, Yuming Fang 0001 |
Knowl. Based Syst. | 5 |
| 2023 | Fast CU size decision algorithm for VVC intra coding
Xiwu Shang, Xiaoli Zhao 0003, Hua Han 0002, Yifan Zuo 0001 |
Multim. Tools Appl. | 5 |
| 2023 | Gradient-Guided Single Image Super-Resolution Based on Joint Trilateral Feature FilteringabstractThe state of the arts (SOTAs) of single image super-resolution always exploit guidance from gradient prior. The fusion of gradient guidance is implemented by channel-wise concatenation followed by a convolutional layer. However, the kernels sharing in spatial positions cannot adaptively tune the effect of gradient guidance for all feature positions. To resolve this problem, a novel network module is proposed to simulate the traditional Joint Trilateral Filter (JTF) by extending the definition domain from pixels to features. Moreover, to improve the efficiency and flexibility, the functions of JTF kernel generation for image features and gradient features are explicitly learned instead of individual kernel weights, e.g., the exponential functions in the traditional JTF. Based on the proposed JTF modules, this paper follows the gradient-guided framework which simultaneously infers high-resolution (HR) image features and HR gradient features within two parallel branches, respectively. Specifically, by treating image features and gradient features as cross guidance to each other, the proposed JTF modules adaptively adjust the fusion patterns for local features via a bi-directional way. By doing so, the quality of image features and gradient features is alternatively enhanced. Compared with SOTAs, the proposed JTF-SISR shows improvement which is evaluated for multiple upsampling scales and degradation modes on 5 synthetic datasets, i.e., Set5, Set14, B100, Urban100 and Manga109, and 1 real dataset, i.e., RealSRSet. The code is public inhttps://github.com/a239xjc/JTF-SISR. Yifan Zuo 0001, Yuming Fang 0001, Deyang Liu, Wenying Wen |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Multi-Stream Dense View Reconstruction Network for Light Field Image CompressionabstractRecently, many view synthesis-based methods are proposed for high-efficiency light field (LF) image compression. However, most existing methods fail to recover more texture details on occlusion regions, which reduces the compression efficiency. In this paper, we propose a multi-stream dense view reconstruction network to further improve LF image compression performance. In our method, only sparsely-sampled LF views are transmitted and the rest of the views are reconstructed at the decoder side. During the reconstruction process, we firstly constitute a multi-disparity geometry (MDG) structure based on the decoded sparse LF views, which can reflect abundant disparity characteristics. Subsequently, a multi-stream view reconstruction network (MSVRNet) is put forward to reconstruct a high-quality dense LF image, which consists of a multi-scale feature fusion sub-network, a fusion reconstruction sub-network, and a detail refinement sub-network. The multi-scale feature fusion sub-network can implicitly lean abundant multiscale geometric structure features from the constituted MDG structure. The fusion reconstruction sub-network and the detail refinement sub-network are respectively utilized to fuse the learned multiscale geometric features and restore more texture details, especially for occlusion regions. Moreover, 3D convolutional operations are adopted in the whole reconstruction process, which allow information propagation among the learned multiscale geometric features. Comprehensive experimental results demonstrate the effectiveness of the proposed method. The perceptual quality of reconstructed views and application on depth estimation also demonstrate that the proposed method can keep structural consistency of the reconstructed LF image and recover more texture details. Deyang Liu, Yan Huang 0023, Yuming Fang 0001, Yifan Zuo 0001, Ping An 0001 |
IEEE Trans. Multim. | 4 |
| 2022 | Evaluating the Robustness of Depth Image Super-Resolution ModelsabstractDepth image super-resolution (DISR) is one of the hot topics in computer vision. Although great progress has been made in this research topic, the robustness of DISR models is not sufficiently investigated, which is of great importance in the real applications. Accordingly, in this paper, we make an initial attempt to investigate the robustness of DISR models. Specifically, we test their generalization ability when the input depth image suffers from visual quality degradation. To facilitate this study, we construct a large-scale depth image dataset in which the reference depth images are perturbed to generate the degraded depth images automatically. Then, we test six top-performing DISR models on the constructed dataset and then compare their strengths and weaknesses. By conducting comprehensive experiments, we find that depth image super-resolution models perform poorly on Gaussian noise, and that the higher the level, the lower the quality of the predicted depth map. Furthermore, some DISR models only outperform at lower magnifications (such as 2x and 4x). Dengxiang Wang, Jiebin Yan, Xuelin Liu, Yifan Zuo 0001 |
MMSP | 4 |
| 2022 | Color-Sensitivity-Based Rate-Distortion Optimization for H.265/HEVCabstractRate-Distortion Optimization (RDO) is an important step in video coding to achieve the best quality under a certain compression ratio constraint. The traditional RDO assigns equal importance to different color components. However, Human Visual System (HVS) has different sensitivities to different components. In this paper, the color-sensitivity-based combined PSNR (CSPSNR) is utilized as the distortion measurement in the process of RDO, where the characteristics of the color sensitivities of HVS are taken into account. Firstly, the distortion weights of luma and chroma components are derived from the criterion of maximizing CSPSNR. Then Lagrange multiplier and quantization parameter (QP) are adjusted according to the variation of distortion weights among different components. Finally, the CSPSNR-based RDO (CSRDO) adaptively calculates the RD costs of luma and chroma components under different sampling rates to improve the coding efficiency of the whole sequence. Experimental results in H.265/HEVC demonstrate that the proposed method can achieve 3.11% and 3.58% BD-RATE gain for AI and RA configurations in terms of CSPSNR on average. Xiwu Shang, Jie Liang 0001, Xiaoli Zhao 0003, Hua Han 0002, Yifan Zuo 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2022 | MIG-Net: Multi-Scale Network Alternatively Guided by Intensity and Gradient Features for Depth Map Super-ResolutionabstractThe studies of previous decades have shown that the quality of depth maps can be significantly lifted by introducing the guidance from intensity images describing the same scenes. With the rising of deep convolutional neural network, the performance of guided depth map super-resolution is further improved. The variants always consider deep structure, optimized gradient flow and feature reusing. Nevertheless, it is difficult to obtain sufficient and appropriate guidance from intensity features without any prior. In fact, features in the gradient domain, e.g., edges, present strong correlations between the intensity image and the corresponding depth map. Therefore, the guidance in the gradient domain can be more efficiently explored. In this paper, the depth features are iteratively upsampled by 2×. In each upsampling stage, the low-quality depth features and the corresponding gradient features are iteratively refined by the guidance from the intensity features via two parallel streams. Then, to make full use of depth features in the image and gradient domains, the depth features and gradient features are alternatively complemented with each other. Compared with state-of-the-art counterparts, the sufficient experimental results show improvements according to the objective and subjective assessments. The code is available athttps://github.com/Yifan-Zuo/MIG-net-gradient_guided_depth_enhancement. Yifan Zuo 0001, Yuming Fang 0001, Xiaoshui Huang, Xiwu Shang, Qiang Wu 0001 |
IEEE Trans. Multim. | 1 |
| 2021 | Asymmetrically distorted 3D video quality assessment: From the motion variation to perceived quality
Yuming Fang 0001, Xiangjie Sui, Jiebin Yan, Yifan Zuo 0001, Jiheng Wang, Zhaoqian Li |
Signal Process. | 4 |
| 2021 | Visual attention prediction for Autism Spectrum Disorder with hierarchical semantic fusion
Yuming Fang 0001, Yifan Zuo 0001, Wenhui Jiang 0001, Hanqin Huang, Jiebin Yan |
Signal Process. Image Commun. | 3 |
| 2021 | Objective quality assessment of synthesized images by local variation measurement
Xiangjie Sui, Mengna Ding, Jiebin Yan, Yuming Fang 0001, Yifan Zuo 0001, Zuowen Tan |
Signal Process. Image Commun. | 5 |
| 2021 | Blind Quality Assessment for Tone-Mapped Images by Analysis of Gradient and Chromatic StatisticsabstractA tone-mapped image (TMI) obtained from the corresponding high dynamic range (HDR) image induces artifacts and distortion, which might result in the loss of structure information and impaired color. By analyzing the visual characteristics of TMIs, this work proposes a robust blind visual quality evaluation method for TMIs by using gradient and chromatic statistics (VQGC). First, motivated by the perceptual mechanism that the human visual system (HVS) is sensitive to image structure variation, we employ the gradient features to measure structure degradation in TMIs. To predict structure distortion accurately, we compute the gradient magnitude and orientation to measure image structure variation, and the relative gradient magnitude and orientation are also computed to capture microstructure change. Second, the color invariance descriptors are utilized to capture the visual degradation of colorfulness by local binary pattern (LBP) on four chromatic feature maps. Finally, the gradient and chromatic features are combined together as the final quality-aware feature vector, which is applied to assess the perceptual quality of TMIs by support vector regression (SVR). Comparison experiments show that the performance of the proposed method is better than other existing blind quality assessment methods on public databases. Yuming Fang 0001, Jiebin Yan, Rengang Du, Yifan Zuo 0001, Wenying Wen, Yan Zeng 0001, Leida Li |
IEEE Trans. Multim. | 4 |
| 2021 | Frequency-Dependent Depth Map Enhancement via Iterative Depth-Guided Affine Transformation and Intensity-Guided RefinementabstractRecently, deep convolutional neural network sho-ws significant improvement for intensity-guided depth map enhancement. The most networks focus on either increasing depth or easing features propagation via residual learning and dense connection. However, it has not been explicitly considered yet to mitigate the artifacts caused by the differences of the distributions between the depth map and the corresponding color image, e.g., edge misalignment. In this paper, a novel depth-guided affine transformation is used to filter out the unrelated intensity features, which is further used to refine the depth features. Since the quality of initial depth features is low, the depth-guided intensity features filtering and the intensity-guided depth features refinement are iteratively performed, which progressively promotes effects of such tasks. To make full use of the iterations, all the refined depth features are dense connected followed by a 1 × 1 convolution layer. In addition, to improve the performance in the case of large upsampling factors (e.g., 16×), the depth features are enhanced from coarse to fine. In each frequency-dependent refinement of the depth features, the above iterative subnetwork as well as the residual learning are introduced. The proposed method is tested for the noise-free and noisy cases which compares against 16 state-of-the-art methods. Our experimental results show the improved performances based on the qualitative and quantitative evaluations. Yifan Zuo 0001, Yuming Fang 0001, Ping An 0001, Xiwu Shang, Junnan Yang |
IEEE Trans. Multim. | 1 |
| 2020 | Blind quality assessment for tone-mapped images based on local and global features
Xuelin Liu, Yuming Fang 0001, Rengang Du, Yifan Zuo 0001, Wenying Wen |
Inf. Sci. | 4 |
| 2020 | Perceptual Quality Assessment for Screen Content Images by Spatial ContinuityabstractIn this paper, we propose an effective blind quality assessment method for screen content images (SCIs), called perceptual quality measure by spatial continuity (PQSC). With the center-surround mechanism in the human visual system (HVS), the proposed method extracts the statistical features on chromatic and textural variations in SCIs to measure the visual distortion. First, by considering the chromatic continuity between spatially adjacent pixels, photo-metric invariant chromatic descriptors are extracted as zero-order and first-order features. Second, motivated by the perceptual mechanism that the HVS is sensitive to image texture variation, we employ local ternary pattern operator to effectively depict the spatial continuity of texture. With these extracted chromatic and textural features, we further adopt histogram to compute the statistical chromatic and textural features. Support vector regression (SVR) is used to train the quality prediction model from visual features to human ratings. Experimental results on three public benchmark databases demonstrate that the performance of our method is superior to the current blind image quality assessment methods, even better than some full reference image quality assessment counterparts. Yuming Fang 0001, Rengang Du, Yifan Zuo 0001, Wenying Wen, Leida Li |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Depth Map Enhancement by Revisiting Multi-Scale Intensity Guidance Within Coarse-to-Fine StagesabstractBeing different from the most methods of guided depth map enhancement based on deep convolutional neural network which focus on increasing the depth of networks, this paper is to improve the effectiveness of intensity guidance when the network goes deep. Overall, the proposed network upsamples the low-resolution depth maps from coarse to fine. Within each refinement stage of certain-scale depth features, the current-scale and all coarse-scales of the guidance features are revisited by dense connection. Therefore, the multi-scale guidance is efficiently maintained as the propagation of features. Furthermore, the proposed network maintains the intensity features in the high-resolution domain from which the multi-scale guidance is directly extracted. This design further improves the quality of intensity guidance. In addition, the shallow depth features upsampled via transposed convolution layer are directly transferred to the final depth features for reconstruction, which is called global residual learning in feature domain. Similarly, the global residual learning in pixel domain learns the difference between the depth ground truth and the coarsely upsampled depth map. Also, the local residual learning is to maintain the low frequency within each refinement stage and progressively recover the high frequency. The proposed method is tested for noise-free and noisy cases which compares against 16 state-of-the-art methods. Our experimental results show the improved performances based on the qualitative and quantitative evaluations. Yifan Zuo 0001, Yuming Fang 0001, Yong Yang 0001, Xiwu Shang, Qiang Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Multi-Scale Frequency Reconstruction for Guided Depth Map Super-Resolution via Deep Residual NetworkabstractThe depth maps obtained by the consumer-level sensors are always noisy in the low-resolution (LR) domain. Existing methods for the guided depth super-resolution, which are based on the pre-defined local and global models, perform well in general cases (e.g., joint bilateral filter and Markov random field). However, such model-based methods may fail to describe the potential relationship between RGB-D image pairs. To solve this problem, this paper proposes a data-driven approach based on the deep convolutional neural network with global and local residual learning. It progressively upsamples the LR depth map guided by the high-resolution intensity image in multiple scales. A global residual learning is adopted to learn the difference between the ground truth and the coarsely upsampled depth map, and the local residual learning is introduced in each scale-dependent reconstruction sub-network. This scheme can restore the depth structure from coarse to fine via multi-scale frequency synthesis. In addition, batch normalization layers are used to improve the performance of depth map denoising. Our method is evaluated in noise-free and noisy cases. A comprehensive comparison against 17 state-of-the-art methods is carried out. The experimental results show that the proposed method has faster convergence speed as well as improved performances based on the qualitative and quantitative evaluations. Yifan Zuo 0001, Qiang Wu 0001, Yuming Fang 0001, Ping An 0001, Liqin Huang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | No Reference Quality Assessment for 3D Synthesized Views by Local Structure Variation and Global Naturalness ChangeabstractDepth image based rendering (DIBR) has been widely used to generate different virtual viewpoints of the same scene from the new perspective. However, DIBR tends to introduce annoying artifacts including blurring, discontinuity, blocking, and stretching, etc.. Thus, to improve DIBR performance, it is important to accurately measure the visual quality of synthesized views. In this paper, we propose a novel and effective no reference (NR) quality assessment method for 3D synthesized views by local variation and global change (LVGC). More specifically, we firstly compute the Gaussian derivatives for the input image to extract structure and chromatic features. Then, we use the local binary pattern (LBP) operator to encode the structure and chromatic feature maps, which are used to calculate quality-aware features to measure the local structural and chromatic distortion. Besides, we extract luminance features by global change to evaluate the naturalness of 3D synthesized views. With these extracted features, we utilize random forest regression (RFR) to train the quality prediction model from visual features to human ratings. Experimental results on three public benchmark databases demonstrate the effectiveness of our method on estimating visual quality of 3D synthesized views. Jiebin Yan, Yuming Fang 0001, Rengang Du, Yan Zeng 0001, Yifan Zuo 0001 |
IEEE Trans. Image Process. | 5 |
| 2019 | Blind Quality Assessment for DIBR-Synthesized Images Based on Chromatic and Disoccluded Information
Mengna Ding, Yuming Fang 0001, Yifan Zuo 0001, Zuowen Tan |
PRCV (2) | 3 |
| 2019 | Residual dense network for intensity-guided depth map enhancement
Yifan Zuo 0001, Yuming Fang 0001, Yong Yang 0001, Xiwu Shang |
Inf. Sci. | 1 |
| 2019 | Pyramid-Structured Depth MAP Super-Resolution Based on Deep Dense-Residual NetworkabstractAlthough deep convolutional neural networks (DCNN) show significant improvement for single depth map (SD) super-resolution (SR) over the traditional counterparts, most SDSR DCNNs do not reuse the hierarchical features for depth map SR resulting in blurred high-resolution (HR) depth maps. They always stack convolutional layers to make network deeper and wider. In addition, most SDSR networks generate HR depth maps at a single level, which is not suitable for large up-sampling factors. To solve these problems, we present pyramid-structured depth map super-resolution based on deep dense-residual network. Specially, our networks are made up of dense residual blocks that use densely connected layers and residual learning to model the mapping between high-frequency residuals and low-resolution (LR) depth map. Furthermore, based on the pyramid structure, our network can progressively generate depth maps of various levels by taking advantages of features from different levels. The proposed network adopts a deep supervision scheme to reduce the difficulty of model training and further improve the performance. The proposed method is evaluated on Middlebury datasets which shows improved performance compared with 6 state-of-the-art methods. Liqin Huang, Jianjia Zhang, Yifan Zuo 0001, Qiang Wu 0001 |
IEEE Signal Process. Lett. | 3 |
| 2018 | Explicit Edge Inconsistency Evaluation Model for Color-Guided Depth Map EnhancementabstractColor-guided depth enhancement is used to refine depth maps according to the assumption that the depth edges and the color edges at the corresponding locations are consistent. In methods on such low-level vision tasks, the Markov random field (MRF), including its variants, is one of the major approaches that have dominated this area for several years. However, the assumption above is not always true. To tackle the problem, the state-of-the-art solutions are to adjust the weighting coefficient inside the smoothness term of the MRF model. These methods lack an explicit evaluation model to quantitatively measure the inconsistency between the depth edge map and the color edge map, so they cannot adaptively control the efforts of the guidance from the color image for depth enhancement, leading to various defects such as texture-copy artifacts and blurring depth edges. In this paper, we propose a quantitative measurement on such inconsistency and explicitly embed it into the smoothness term. The proposed method demonstrates promising experimental results compared with the benchmark and state-of-the-art methods on the Middlebury ToF-Mark, and NYU data sets. Yifan Zuo 0001, Qiang Wu 0001, Jian Zhang 0002, Ping An 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Minimum Spanning Forest With Embedded Edge Inconsistency Measurement Model for Guided Depth Map EnhancementabstractGuided depth map enhancement based on Markov Random Field (MRF) normally assumes edge consistency between the color image and the corresponding depth map. Under this assumption, the low-quality depth edges can be refined according to the guidance from the high-quality color image. However, such consistency is not always true, which leads to texture-copying artifacts and blurring depth edges. In addition, the previous MRF-based models always calculate the guidance affinities in the regularization term via a non-structural scheme which ignores the local structure on the depth map. In this paper, a novel MRF-based method is proposed. It computes these affinities via the distance between pixels in a space consisting of the Minimum Spanning Trees (Forest) to better preserve depth edges. Furthermore, inside each Minimum Spanning Tree, the weights of edges are computed based on explicit edge inconsistency measurement model, which significantly mitigates texture-copying artifacts. To further tolerate the effects caused by noise and better preserve depth edges, a bandwidth adaption scheme is proposed. Our method is evaluated for depth map super-resolution and depth map completion problems on synthetic and real datasets including Middlebury, ToF-Mark and NYU. A comprehensive comparison against 16 state-of-the-art methods is carried out. Both qualitative and quantitative evaluation present the improved performances. Yifan Zuo 0001, Qiang Wu 0001, Jian Zhang 0002, Ping An 0001 |
IEEE Trans. Image Process. | 1 |
| 2017 | Minimum spanning forest with embedded edge inconsistency measurement for color-guided depth map upsamplingabstractColor-guided depth map up-sampling, such as Markov-Random-Field-based (MRF-based) methods, is a popular depth map enhancement solution, which normally assumes edge consistency between color image and corresponding depth map. It calculates the coefficients of smoothness term in MRF according to such assumption. However, such consistency is not always true which leads to texture-copying artifacts and blurring depth edges. In this paper, we propose a novel coefficient computing scheme for smoothness term in MRF which is based on the distance between pixels in the Minimum Spanning Trees (Forest) to better preserve depth edges. The explicit edge inconsistency measurement is embedded into weights of edges in Minimum Spanning Trees, which significantly mitigates texture-copying artifacts. The proposed method is evaluated on Middlebury datasets and ToF-Mark datasets which demonstrates improved results compared with state-of-the-art methods. Yifan Zuo 0001, Qiang Wu 0001, Jian Zhang 0002, Ping An 0001 |
ICME | 1 |
| 2016 | Explicit measurement on depth-color inconsistency for depth completionabstractColor-guided depth completion is to refine depth map through structure light sensing by filling missing depth structure and de-nosing. It is based on the assumption that depth discontinuity and color edge at the corresponding location are consistent. Among all proposed methods, MRF-based method including its variants is one of major approaches. However, the assumption above is not always true, which causes texture-copy and depth discontinuity blurring artifacts. The state-of-the-art solutions usually are to modify the weighting inside smoothness term of MRF model. Because there is no any method explicitly considering the inconsistency occurring between depth discontinuity and the corresponding color edge, they cannot adaptively control the effect of guidance from color image when completing depth map. In this paper, we propose quantitative measurement on such inconsistency and explicitly embed it into weighting value of smoothness term. The proposed method is evaluated on NYU Kinect datasets and demonstrates promising results. Yifan Zuo 0001, Qiang Wu 0001, Ping An 0001, Jian Zhang 0002 |
ICIP | 1 |
| 2016 | Explicit modeling on depth-color inconsistency for color-guided depth up-samplingabstractColor-guided depth up-sampling is to enhance the resolution of depth map according to the assumption that the depth discontinuity and color image edge at the corresponding location are consistent. Through all methods reported, MRF including its variants is one of major approaches, which has dominated in this area for several years. However, the assumption above is not always true. Solution usually is to adjust the weighting inside smoothness term in MRF model. But there is no any method explicitly considering the inconsistency occurring between depth discontinuity and the corresponding color edge. In this paper, we propose quantitative measurement on such inconsistency and explicitly embed it into weighting value of smoothness term. Such solution has not been reported in the literature. The improved depth up-sampling based on the proposed method is evaluated on Middlebury datasets and ToFMark datasets and demonstrate promising results. Yifan Zuo 0001, Qiang Wu 0001, Jian Zhang 0002, Ping An 0001 |
ICME | 1 |
| 2015 | Depth upsampling method via Markov random fields without edge-misaligned artifactsabstractRecently, the widely use of time-of-flight sensors captures depth information for dynamic scenes in real time, which promotes the developing of many 3D image or video processing applications. However, such depth maps are noisy and have low resolutions. In this paper, we propose an edge-based depth map super-resolution method via solving a labeling optimization problem in MRF. The inputs are low quality depth map and the according high-resolution color image. The proposed method not only avoids the texture-copy artifacts, but also preserves the edges of depth which do not exist in the color image. We compare our algorithm with the state of the art on the benchmark dataset. The experimental results prove the validity and robustness of our approach. Yifan Zuo 0001, Ping An 0001, Zhaoyang Zhang 0002 |
ICIP | 1 |
| 2012 | Fast Segment-Based Algorithm for Multi-view Depth Map Generation
Yifan Zuo 0001, Ping An 0001, Qiuwen Zhang, Zhaoyang Zhang 0002 |
ICIC (2) | 1 |