VLDB 2026 Research / reviewers in the wild / expert
You Yang 0002
dblp:09/6306-2
· DBLP profile ↗
78ranked-venue papers
11as first author
44since 2021 · last 2026
0000-0002-5695-1046ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 57 · 8 first-author · 29 since 2021Artificial intelligence and machine learning · 17 · 1 first-author · 14 since 2021Databases, data management, data science and information retrieval · 10 · 3 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Spatially adaptive representation of facial meshes for face video super-resolution
Shangchen Cai, Qiong Liu 0001, You Yang 0002 |
Expert Syst. Appl. | 5 |
| 2026 | Density-aware few-parametric networks for robust few-shot point cloud semantic segmentation
Yudong Liang, Pei An, Qiong Liu 0001, You Yang 0002 |
Neurocomputing | 4 |
| 2026 | FM-MFIF: A motion-aware multi-focus image fusion network based on multi-scale focus migration
Zhilong Li, Zhulun Yang, You Yang 0002, Qiong Liu 0001 |
Image Vis. Comput. | 3 |
| 2026 | Space-Time Correlation Adaptive-Layered Optimization for Microimage Motion Search in Plenoptic video coding
Jingyang Luo, Qiong Liu 0001, Xiatian Xie, Pengpeng Han, You Yang 0002 |
J. Vis. Commun. Image Represent. | 6 |
| 2026 | FSF-Net: Enhance 4D occupancy forecasting with coarse BEV scene flow for autonomous driving
Erxin Guo, Pei An, You Yang 0002, Qiong Liu 0001, Anan Liu |
Pattern Recognit. | 3 |
| 2026 | DFS-Net: A Dense Focal Stack Image Generation Network From Misaligned Multi-Focus ImagesabstractDense focal stack images inherently encode depth cues and are crucial for various 3D vision applications. However, existing generation methods are susceptible to misalignment and introduce a domain gap between synthetic and real-world data due to off-axis aberrations. To address these challenges, we introduce DFS-Net, an aberration-aware dense focal stack image generation network. DFS-Net consists of two core modules: all-in-focus image synthesis and aberration-aware point spread function (PSF) generation. The all-in-focus image synthesis is achieved through a densely connected fusion network based on multi-scale focus migration and focus property detection. This fusion network can effectively fuse misaligned multi-focus images into an all-in-focus image. The aberration-aware PSF generation is realized through a multi-layer perceptron (MLP) network. Supervised by ray-tracing-based PSFs, the MLP network can generate spatially varying PSFs for arbitrary spatial positions and focus distances. By selecting a set of focus distances, the generated PSF maps are locally convolved with the all-in-focus image to produce an aberration-aware dense focal stack. We conduct extensive comparative experiments on all-in-focus image fusion and focal stack generation against state-of-the-art methods. The experimental results demonstrate that DFS-Net can synthesize all-in-focus images with high subjective and objective quality, as well as generate dense focal stacks that closely approximate ray-tracing results. In addition, we conduct comparative experiments on the depth-from-focus and salient object detection tasks using the generated focal stacks. The experimental results demonstrate that our DFS-Net can significantly enhance the performance of existing depth-from-focus and salient object detection models. The code and dataset will be publicly available at https://github.com/North-Li/DFS-Net. Zhilong Li, Pei An, You Yang 0002, Qiong Liu 0001, Dan Song 0006, Anan Liu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | PNProRL: Self-Supervised Neural Relighting via Photometric Perception and Progressive OptimizationabstractPortrait relighting shows great potential in photography, film, and AR by simulating diverse lighting effects. Existing state-of-the-art methods often rely on expensive paired OLAT or synthetic data, which limits scalability. Moreover, accurately modeling the interaction between physics-guided rendering, neural rendering, and real-world remains challenging. To address these issues, we propose a novel multi-stage self-supervised relighting framework. It progressively refines intrinsic scene properties via a simple-to-complex training strategy, removing the need for expensive paired data while adapting to various lighting conditions. One core design introduces a novel pre-training method approach using diverse shading-based masking for self-reconstruction, which improves the model's perception of complex lighting variations. Furthermore, we introduce two perceptual modules that leverage the linear superposition of light to narrow the gap between physics-guided and neural rendering, and better align relit results with real-world observations. Extensive experiments demonstrate that our unified framework achieves new state-of-the-art performance in portrait relighting, surpassing recent methods in photorealism, synthesis quality, and identity preservation. It provides a practical paradigm for high-fidelity relighting under diverse lighting. Chenhao Guo, Zhulun Yang, You Yang 0002, Qiong Liu 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | MinCD-PnP: Learning 2D-3D Correspondences with Approximate Blind PnPabstractImage-to-point-cloud (I2P) registration is a fundamental problem in computer vision, focusing on establishing 2D-3D correspondences between an image and a point cloud. The differential perspective-n-point (PnP) has been widely used to supervise I2P registration networks by enforcing the projective constraints on 2D-3D correspondences. However, differential PnP is highly sensitive to noise and outliers in the predicted correspondences. This issue hinders the effectiveness of correspondence learning. Inspired by the robustness of blind PnP against noise and outliers in correspondences, we propose an approximated blind PnP based correspondence learning approach. To mitigate the high computational cost of blind PnP, we simplify blind PnP to an amenable task of minimizing Chamfer distance between learned 2D and 3D keypoints, called MinCD-PnP. To effectively solve MinCD-PnP, we design a lightweight multi-task learning module, named as MinCD-Net, which can be easily integrated into the existing I2P registration architectures. Extensive experiments on 7-Scenes, RGBD-V2, ScanNet, and self-collected datasets demonstrate that MinCD-Net outperforms state-of-the-art methods and achieves a higher inlier ratio (IR) and registration recall (RR) in both cross-scene and cross-dataset settings. Pei An, Jiaqi Yang 0002, Muyao Peng, You Yang 0002, Qiong Liu 0001, Liangliang Nan |
ICCV | 4 |
| 2025 | Top-I2P: Explore Open-Domain Image-to-Point Cloud Registration Using Topology RelationshipabstractImage-to-point cloud (I2P) registration is a fundamental task in computer vision, which aims to align pixels in 2D images with corresponding points in 3D point clouds. While deep learning based methods dominate this field, they often fail to generalize to the open domain. In this paper, we address open-domain I2P registration from the topology relationships perspective. Firstly, we find that topology relationships reflect sparse connections between pixels and points, which shows the significant potential in enhancing cross-modality feature interaction in the open domain. Building on this insight, we develop an I2P registration framework using topology relationships. After that, to construct and leverage the topology relationships between the heterogeneous 2D and 3D spaces, we design a registration network, Top-I2P, with correction-based topology reasoning and fast topology feature interaction modules. Extensive experiments on 7-Scenes, RGBD-V2, ScanNet, and self-collected I2P datasets demonstrate that Top-I2P achieves superior registration performance in open-domain scenarios. Pei An, Jiaqi Yang 0002, Muyao Peng, You Yang 0002, Qiong Liu 0001, Jie Ma 0003, Liangliang Nan |
IJCAI | 4 |
| 2025 | Enhance Image-to-Point-Cloud Registration with Beltrami Flow
Pei An, You Yang 0002, Jiaqi Yang 0002, Muyao Peng, Qiong Liu 0001, Liangliang Nan |
Int. J. Comput. Vis. | 2 |
| 2025 | Adaptive CLIP for open-domain 3D model retrieval
Dan Song 0006, Zekai Qiang, Chumeng Zhang, Lanjun Wang, Qiong Liu 0001, You Yang 0002, Anan Liu |
Inf. Process. Manag. | 6 |
| 2025 | Corner selection and dual network blender for efficient view synthesis in outdoor scenes
Mohannad A. M. Al-Ja'afari, Firas Abedi, You Yang 0002, Qiong Liu 0001 |
Pattern Recognit. | 3 |
| 2025 | Unsupervised learning non-uniform face enhancement under physics-guided model of illumination decoupling
Zhongyuan Wang 0001, Qiong Liu 0001, You Yang 0002, Zhenyu Shu |
Pattern Recognit. | 5 |
| 2025 | Real-time small object detection using adaptive weighted fusion of efficient positional features
Qiong Liu 0001, You Yang 0002 |
Pattern Recognit. | 4 |
| 2025 | Progressive Contrastive Label Optimization for Source-Free Universal 3D Model RetrievalabstractUnsupervised Cross-Domain 3D Model Retrieval (UCD3DMR) has emerged as an effective tool for managing 3D model data recently. However, existing UCD3DMR algorithms typically demand accessibility to source data and cross-domain label consistency, limiting their deployment in real-world industrial scenarios. Therefore, we relax the two demanding constraints and explore to address a newly challenging task, source-free universal 3D model retrieval (SFU3DMR). However, the inaccessibility to source data results in significant label noise in target pseudo-labels, while cross-domain label inconsistency introduces interference from target-private models, presenting tremendous challenges to model transfer. To address these challenges, we propose a novel SFU3DMR algorithm, Progressive Contrastive Label Optimization (PCLO). Specifically, we introduce the Neighbor-based Soft Label Optimization (NSLO) strategy, which refines target pseudo-labels based on the pseudo-label confidence of their nearest neighbors. Additionally, we design the Adaptive Hybrid Label Optimization (AHLO) strategy, which conducts positive label optimization to maximize label semantics for target-common models and executes negative label optimization to minimize label noise for target-private models. Experimental results confirm that the combined NSLO and AHLO strategies effectively refine the target pseudo-labels, and our PCLO achieves state-of-the-art performance for SFU3DMR on two well-established cross-domain benchmarks (MI3DOR and NTU/PSB). Jiayu Li 0004, Yuting Su 0001, Dan Song 0006, Wenhui Li 0001, You Yang 0002, Anan Liu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | MV-CLIP: Multi-View CLIP for Zero-Shot 3D Shape RecognitionabstractLarge-scale pre-trained models have demonstrated impressive performance in vision and language tasks within open-world scenarios. Due to the lack of comparable pre-trained models for 3D shapes, recent methods utilize language-image pre-training to realize zero-shot 3D shape recognition. However, due to the modality gap, pretrained language-image models are not confident enough in the generalization to 3D shape recognition. Consequently, this paper aims to improve the confidence with view selection and hierarchical prompts. Building on the well-established CLIP model, we introduce view selection in the vision side that minimizes entropy to identify the most informative views for 3D shape. On the textual side, hierarchical prompts combined of hand-crafted and GPT-generated prompts are proposed to refine predictions. The first layer prompts several classification candidates with traditional class-level descriptions, while the second layer refines the prediction based on function-level descriptions or further distinctions between the candidates. Extensive experiments demonstrate the effectiveness of the proposed modules for zero-shot 3D shape recognition. Remarkably, without the need for additional training, our proposed method achieves impressive zero-shot 3D classification accuracies of 84.44%, 91.51%, and 66.17% on ModelNet40, ModelNet10, and ShapeNet Core55, respectively. Furthermore, we will make the code publicly available to facilitate reproducibility and further research in this area. Dan Song 0006, Xinwei Fu, Weizhi Nie, Wenhui Li 0001, Lanjun Wang, You Yang 0002, Anan Liu |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | Geometry-Aware Self-Supervised Indoor 360$^{\circ }$ Depth Estimation via Asymmetric Dual-Domain Collaborative LearningabstractBeing able to estimate monocular depth for spherical panoramas is of fundamental importance in 3D scene perception. However, spherical distortion severely limits the effectiveness of vanilla convolutions. To push the envelope of accuracy, recent approaches attempt to utilize Tangent projection (TP) to estimate the depth of$360 ^{\circ }$images. Yet, these methods still suffer from discrepancies and inconsistencies among patch-wise tangent images, as well as the lack of accurate ground truth depth maps under a supervised fashion. In this paper, we propose a geometry-aware self-supervised$360 ^{\circ }$image depth estimation methodology that explores the complementary advantages of TP and Equirectangular projection (ERP) by an asymmetric dual-domain collaborative learning strategy. Especially, we first develop a lightweight asymmetric dual-domain depth estimation network, which enables to aggregate depth-related features from a single TP domain, and then produce depth distributions of the TP and ERP domains via collaborative learning. This effectively mitigates stitching artifacts and preserves fine details in depth inference without overspending model parameters. In addition, a frequent-spatial feature concentration module is devised to simultaneously capture non-local Fourier features and local spatial features, such that facilitating the efficient exploration of monocular depth cues. Moreover, we introduce a geometric structural alignment module to further improve geometric structural consistency among tangent images. Extensive experiments illustrate that our designed approach outperforms existing self-supervised$360 ^{\circ }$depth estimation methods on three publicly available benchmark datasets. Xu Wang 0006, Ziyan He, Qiudan Zhang, You Yang 0002, Tiesong Zhao, Jianmin Jiang |
IEEE Trans. Multim. | 4 |
| 2025 | End-to-End Deep Video Compression Based on Hierarchical Temporal Context LearningabstractEmerging learning-based video compression suffers from error propagation in long group of pictures (GOP), yielding limited coding performance. To address this problem, a novel end-to-end Deep Video Compression method based on Hierarchical Temporal Context Learning (DVCH) is proposed in this paper. DVCH aims to fully exploit temporal contexts and suppress error propagation for better coding performance. It first divides video frames into several hierarchies with different compression qualities. The frames in lower hierarchies have high compression quality, and serve as reference frames. To mine high-quality reference information, we propose a Hierarchical Temporal Context Learning (HTCL) network as the fundamental module of our DVCH. The informative temporal context features from hierarchical prediction structure can be extracted by the network. Motion vectors (MVs) between the to-be-coded frame and its reference frames are estimated by the MV Learning module and used to align the extracted contexts. The contexts are fed into Context Coding module to generate the prediction of the decoded frame. Moreover, a multi-stage training strategy is developed to solve the imbalanced training challenge. Experimental results demonstrate that the proposed DVCH exceeds x264 and other end-to-end video compression methods, regardless of objective, subjective, error propagation suppression, GOP sizes, and sequence length evaluations. As much as 49.27% bitrate savings and 2.52 dB PSNR gains can be achieved in large GOP. Kejun Wu, You Yang 0002, Qiong Liu 0001, Xiao-Ping Zhang 0002 |
IEEE Trans. Multim. | 3 |
| 2025 | A Serial Perspective on Photometric Stereo of Filtering and Serializing Spatial InformationabstractIn this paper, we introduce a novel method of Filtering and Serializing Spatial Information to tackle uncalibrated photometric stereo tasks, termed FSSI-PS. Photometric stereo aims to recover surface normals from images with varying lighting and is crucial for tasks like 3D reconstruction and defect detection. Current methods in complex surface reconstruction are costly and inaccurate due to redundant feature representations from GCN or Transformer modules, caused by the weak global information extraction capability of GCNs or the large computational cost of Transformers. Furthermore, the trainset's lack of richness in texture complexity makes reconstruction more difficult. We address these issues by optimizing feature maps and dataset richness through serializing and filtering. First, we use Mamba-RNN to optimize feature representation by directly fusing feature maps, which reduces redundancy and uses minimal computational resources. Specifically, we treat input spatial information as a sequence and serialize it by sorting. Furthermore, we introduce the Mean Angular Variation metric to assess reconstruction difficulty by measuring texture complexity. It classifies PS-Sculpture and PS-Blobby into three categories: Difficult, Normal, and Simple. We use this to construct DNS-S+B, a photometric stereo training set with rich complexity levels. Our method is compared with state-of-the-art methods on the DiLiGenT and LUCES benchmarks to highlight effectiveness. Minzhe Xu, You Yang 0002, Yinqiang Zheng, Qiong Liu 0001 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2024 | Low-Rank Completion Based Normal Guided Lidar Point Cloud Up-SamplingabstractCommercial inexpensive LiDAR sensor generally suffers low vertical resolutions, whose point cloud is sparse and may not be able to satisfy future metaverse applications. LiDAR point cloud up-sampling is a task to increase the vertical resolution while preserving the structural details. Scene representation is the central pillar of point cloud up-sampling. However, the sparsity of point cloud hinders the extraction of scene representation. In this paper, we find that low-rank representation can describe the primary scene structure approximately, and convert up-sampling as low-rank tensor completion problem. To decrease problem complexity, we leverage range view projection to convert the problem as low-rank depth completion, and present a low-rank normal guided up-sampling approach. It uses normal as guidance to smooth range depth. Extensive experiments show that our method outperforms current methods. In 2× up-sampling task, it achieves as low as 41cm of mean absolute error (MAE), which is 282% and 32% smaller than interpolation and traditional matrix completion methods, respectively. Hence, we believe the proposed method benefits to the field of metaverse. Pei An, You Yang 0002, Jie Ma 0003 |
ICASSP | 3 |
| 2024 | Multi-View Multi-Focus Image Fusion: A Novel Benchmark Dataset and MethodabstractMulti-focus image fusion fuses multiple images focused on different depths to generate a clear image covering the whole scene. However, the existing multi-focus image fusion methods do not consider the movement of the camera or objects in actual shooting. To address this, we propose an end-to-end deep learning network to generate the all-in-focus image from multi-view multi-focus images. Specifically, our method first warps the multi-view multi-focus images to a unified camera view by homography transformation matrices, and measures the defocus degree of co-located image patches through a focus information evaluation mechanism. Finally, our fusion network applies an adaptive fusion scheme to fuse the detected image patches into a clear image. For testing our fusion network, a multi-view multi-focus image benchmark dataset (MVMFI) is constructed with more than 1000 image sequences. Experiments results demonstrate that our method outperforms the state-of-the-art methods both qualitatively and quantitatively. MVMFI dataset is available at https://github.com/North-Li/MVMFI. Zhilong Li, Kejun Wu, Junhao Liu 0001, Qiong Liu 0001, You Yang 0002 |
ICIP | 5 |
| 2024 | Deep video compression based on Long-range Temporal Context Learning
Kejun Wu, You Yang 0002, Qiong Liu 0001 |
Comput. Vis. Image Underst. | 3 |
| 2024 | OL-Reg: Registration of Image and Sparse LiDAR Point Cloud With Object-Level Dense CorrespondencesabstractImage and point cloud registration (2D-3D registration) is an essential prerequisite for multi-modal feature fusion. However, due to the significant feature difference of point cloud and image, it is challenging to establish 2D-3D correspondences. Targeting for the background of autonomous driving, we propose 2D-3D registration method with object-level correspondence (OL-Reg) in this paper. Object-level correspondence consists of object bounding box and object contour in 2D image and 3D space. The first step is to match 2D-3D objects. Due to sensor pose and field of view (FoV) difference, object shape and occlusion is different in image and point cloud, causing the difficulty of object matching. To solve this issue, we represent object as 3D bounding box, and design 2D-3D object matching with 3D box projection (Box-Proj) constraint. It aligns object 3D bounding box in image and point cloud. After that, the next step is to build 2D-3D correspondence from the matched objects. To extract correspondence from object with irregular shape, we notice the distance constraint of object surface and rays back-projected from object contour, and present projection based iterative closest point (Proj-ICP). Towards the stability of Proj-ICP, object-level regularization term is designed. Experiment is conducted in KITTI object and odometry dataset. With the pre-trained 3D object detector, results suggest that OL-Reg has the better performance than current approaches in tasks of re-localization and extrinsic calibration. And source code will be released soon1. Pei An, Xuzhong Hu, Junfeng Ding, Jun Zhang 0062, Jie Ma 0003, You Yang 0002, Qiong Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Survey of Extrinsic Calibration on LiDAR-Camera System for Intelligent Vehicle: Challenges, Approaches, and TrendsabstractA system with light detection and ranging (LiDAR) and camera (named as LiDAR-camera system) plays the essential role in intelligent vehicle (IV), for it provides 3D spatial and 2D texture features for 3D scene understanding. To leverage LiDAR point cloud and image, extrinsic calibration is a crucial technique, for it can align 2D pixel and 3D point in the pixel-level accuracy. With the rapid development of IV, calibration demand is shifted from offline to online, from the specific scenes to the open scenes. It brings new challenge to the calibration task. Although numbers of approaches have been proposed in the last decade, there lacks an in-depth summary about this topic. Thus, we conduct a survey of extrinsic calibration. Theoretically, the key of calibration is to build correspondence from LiDAR point cloud and optical image. From the viewpoint of correspondence, we attempt to divide the mainstream approaches into explicit and implicit correspondence based methods. After that, we summarize both the strength and weakness of the current works, provide the methods comparison, and list the open-source implementations. Finally, we analyze the tendency of calibration approach, discuss the remained problems in this field. We believe that this survey benefits to the community of autonomous driving. Pei An, Junfeng Ding, Siwen Quan, Jiaqi Yang 0002, You Yang 0002, Qiong Liu 0001, Jie Ma 0003 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2024 | SP-Det: Leveraging Saliency Prediction for Voxel-Based 3D Object Detection in Sparse Point CloudabstractVoxel is one of the common structural representation of 3D point cloud. Due to the sparsity of point cloud generated by light detection and ranging (LiDAR), there is the extreme imbalance in the foreground and background voxels. It decreases the accuracy of 3D object detection, has the negative effect on intelligent driving safety. To overcome this problem, we present a saliency prediction based 3D object detector SP-Det in this article. Although foreground voxels have the sufficient feature of object, it is difficult to localize the foreground region from voxel space with the larger background region. We design an auxiliary learning task, saliency prediction (SP). It benefits 3D detector in identifying the foreground region. SP task uses label diffusion to alleviate the label imbalance. It reduces the learning difficulty of saliency in voxel and bird's eye view (BEV) spaces. After that, to strengthen feature interaction from the sparse foreground region, we design saliency fusion (SF) module to fuse the learning result in SP task. It utilizes voxel and BEV saliency maps as progressive attention to resist the redundant feature from background region. To aggregate more foreground feature inside 3D and BEV region of interest (RoI), we design hybrid grid maps based RoI pooling (Hybrid-RoI pooling). Experiments are conducted in STF dataset. The adverse weather enlarges the sparsity of LiDAR point cloud, increasing the difficulty of object detection. SP-Det identifies and leverages foreground region, and achieves the performance better than the current methods. Hence, we believe that SP-Det benefits to LiDAR based 3D scene understanding in the adverse weather. Pei An, Yucong Duan, Yuliang Huang, Jie Ma 0003, Yanfei Chen, Liheng Wang, You Yang 0002, Qiong Liu 0001 |
IEEE Trans. Multim. | 7 |
| 2024 | ESC-Net: Alleviating Triple Sparsity on 3D LiDAR Point Clouds for Extreme Sparse Scene Completionabstract3D scene completion (SC) has made progress in the last three years. From the application of mobile robot system, SC should support the downstream task (i.e. mapping or perception), instead of only predicting the completed scenes. However, as the low-cost few-beam LiDAR is widely applied in mobile robot, gap between SC and downstream tasks is large. To generate the high quality completion result, the bottleneck lies in the triple sparsity of input, ground truth (GT) occupancy, and GT foreground. To deal with the triple sparsity, we present an extreme sparse scene completion network (ESC-Net). At first, input sparsity hides most of the spatial information of the scene. A feature completion (FC) decoder is designed to mine the spatial feature using feature-level completion. Then, GT occupancy sparsity hinders representation learning of the real scene with continuous surfaces. A multi-view multi-task attention (MMA) loss is presented to recover the high-quality object boundaries via correcting occupancy and semantic labels of regions from 3D and bird's eye view (BEV) spaces. After that, GT foreground sparsity is the imbalance of foreground and background GT labels. It causes the inaccuracy of local 3D object completion. A combination network (ESC-Net-D) is presented to recover 3D structural details of both foreground and background. Experiment is conducted on KITTI and SemanticPOSS datasets. It shows that ESC-Net has the performance higher than current methods not only on completion task, but also on the downstream tasks (i.e. 3D registration, 3D object detection). Hence, we believe that ESC-Net benefits to the community of mobile robot. Source code is released soon. Pei An, Siwen Quan, Junfeng Ding, Jie Ma 0003, You Yang 0002, Qiong Liu 0001 |
IEEE Trans. Multim. | 6 |
| 2024 | Distortion-Aware Self-Supervised Indoor 360$^{\circ }$ Depth Estimation via Hybrid Projection Fusion and Structural RegularitiesabstractOwing to the rapid development of emerging 360$^{\circ }$panoramic imaging techniques, indoor 360$^{\circ }$depth estimation has aroused extensive attention in the community. Due to the lack of available ground truth depth data, it is extremely urgent to model indoor 360$^{\circ }$depth estimation in self-supervised mode. However, self-supervised 360$^{\circ }$depth estimation suffers from two major limitations. One is the distortion and network training problems caused by Equirectangular projection (ERP), and the other is that texture-less regions are quite difficult to back-propagate in self-supervised mode. Hence, to address the above issues, we introduce spherical view synthesis for learning self-supervised 360$^{\circ }$depth estimation. Specifically, to alleviate the ERP-related problems, we first propose a dual-branch distortion-aware network to produce the coarse depth map, including a distortion-aware module and a hybrid projection fusion module. Subsequently, the coarse depth map is utilized for spherical view synthesis, in which a spherically weighted loss function for view reconstruction and depth smoothing is investigated to optimize the projection distribution problem of 360$^{\circ }$images. In addition, two structural regularities of indoor 360$^{\circ }$scenes are devised as two additional supervisory signals to efficiently optimize our self-supervised 360$^{\circ }$depth estimation model, containing the principal-direction normal constraint and the co-planar depth constraint. The principal-direction normal constraint is designed to align the normal of the 360$^{\circ }$image with the direction of the vanishing points. Meanwhile, we employ the co-planar depth constraint to fit the estimated depth of each pixel through its 3D plane. Finally, a depth map is obtained for the 360$^{\circ }$image. Experimental results illustrate that our proposed method achieves superior performance than the current advanced depth estimation methods on four publicly available datasets. Xu Wang 0006, Weifeng Kong, Qiudan Zhang, You Yang 0002, Tiesong Zhao, Jianmin Jiang |
IEEE Trans. Multim. | 4 |
| 2024 | Hierarchical Independent Coding Scheme for Varifocal Multiview Images Based on Angular-Focal Joint PredictionabstractVarifocal multiview (VFMV) images are dense views that focus on variable focal planes. Thus, VFMV images are highly redundant in the angular, spatial and focal dimensions. In this article, the redundancies of VFMV images are analyzed and represented by full parallaxes and focal inconsistency. To exploit these distinctive redundancies, we propose a hierarchical independent coding scheme based on angular-focal joint prediction. The scheme is constructed by hierarchical independent prediction structure (HIPS) and angular-focal joint prediction (AFJP). The HIPS separates all views into several independent subdivisions and assigns different hierarchies inside each subdivision, which enhances random access capability and scalability. The AFJP conducts motion estimation and focal approximation simultaneously to predict parallaxes and focal inconsistency. Therefore, the redundancies in the angular and focal dimensions can be exploited by the proposed coding scheme. We construct a VFMV dataset with 10 test sequences for different acquisition methods. The experimental results on these test sequences demonstrate that the proposed scheme outperforms all comparison schemes in objective quality, subjective quality and random access capability. Specifically, the proposed coding scheme achieves up to 2.661 dB PSNR gains and 52.817% bitrate savings compared with the HEVC random access benchmark scheme. Kejun Wu, You Yang 0002, Qiong Liu 0001, Gangyi Jiang, Xiao-Ping Zhang 0002 |
IEEE Trans. Multim. | 2 |
| 2024 | WaRENet: A Novel Urban Waterlogging Risk Evaluation NetworkabstractIn this article, we propose a novel urban waterlogging risk evaluation network (WaRENet) to evaluate the risk of waterlogging. The WaRENet distinguishes whether an urban image involves waterlogging by classification module, and estimates the waterlogging risk levels by multi-class reference objects detection module (MCROD). First, in the waterlogging scene classification, ResNet combined with Se-block is used to identify the waterlogging scene, and lightweight gradient-weighted class activation mapping (Grad-CAM) is also integrated to roughly locate overall waterlogging areas with low computational burden. Second, in the MCROD module, we detect reference objects, e.g., cars and persons in waterlogging scenes. The positional relationship between water depths and reference objects serves as risk indicators for accurately evaluating waterlogging risk. Specifically, we incorporate switchable atrous convolution (SAC) into YOLOv5 to solve occlusions and varying scales problems in complex waterlogging scenes. Moreover, we construct a large-scale urban waterlogging dataset called UrbanWaterloggingRiskDataset (UWRDataset) with 6,351 images for waterlogging scene classification and 3,217 images for reference objects detection. Experimental results on the dataset show that our WaRENet outperforms all comparison methods. The waterlogging scene classification module achieves accuracy of 95.99%. The MCROD module obtains mAP of 54.9%, while maintaining a high processing speed of 70.04 fps. Xiaoya Yu, Kejun Wu, You Yang 0002, Qiong Liu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | Extending Depth of Field by Varifocal Multi-View Computational Imaging for MetaverseabstractThe flexible field of view (FoV) and large depth of field (DoF) are the main bricks that build the strong immersive experience in Metaverse. However, due to the nature of optics, the captured multi-view images are generally with flexible FoV but limited DoF. To extend the DoF of captured data, in this paper, we propose an all-in-focus image fusion scheme by varifocal multi-view computational imaging. Varifocal multi-view images are a series of multi-view images in different DoF, where different scene contents and blur degrees are the main features among views. Due to the complex inter-view features, the existing extending DoF methods on varifocal multi-view images yield severe ghosting problem. To alleviate the ghosting problem, a patch-based DenseNet image fusion network is designed and embedded in the proposed scheme. The patch-based image fusion network enables to mitigate the ghosting problem in the fused image. Experiments on varifocal multi-view images of different scenes demonstrate that the proposed all-in-focus image fusion scheme can synthesize all-in-focus results with higher visual quality and accuracy. The proposed all-in-focus image fusion scheme is expected to benefit metaverse, photo-realistic novel view synthesis, interactive and immersive experience. Zhilong Li, Kejun Wu, Gangyi Jiang, You Yang 0002 |
MMSP | 4 |
| 2023 | A High Dynamic Range Imaging Method for Short Exposure Multiview Images
You Yang 0002, Kejun Wu, Atif Mehmood, Zahid Hussain Qaisar, Zhonglong Zheng |
Pattern Recognit. | 2 |
| 2023 | RS-Aug: Improve 3D Object Detection on LiDAR With Realistic Simulator Based Data AugmentationabstractLight detection and ranging (LiDAR) is an essential sensor for three dimensional (3D) object detection via generating 3D point cloud of the surroundings, and it has been widely used in the various visual applications, especially autonomous driving. However, limited numbers of labeled LiDAR datasets brutally restrain the development of 3D object detector, and this situation breeds an urgent demand on data augmentation in this field. By far, most of the traditional methods reuse the labeled samples, while those unlabeled are hastily untaken. Motivated by this, we propose aRealisticSimulator based data augmentation (RS-Aug). It aims to construct augmented real scenes to enrich the diversity of training dataset. To train 3D object detector in a supervised learning way, the first step of RS-Aug is auto-annotation. Time-continuous LiDAR frames are used to construct the dense scene, which is beneficial to annotation and the subsequent rendering augmentation. However, 3D points with incorrect semantic labels are naturally gathered during multi-view reconstruction, causing the negative effect on auto-annotation. We propose an algorithm of cluster guided$k$-nearest neighbor (c-$k$NN). It emphasizes on de-nosing semantic labels of clustered points using distance and intensity constraints. Then, the next step of RS-Aug is rendering augmentation on the real scene. To enhance the rendering quality using collision and distance constraints with the less computation complexity, we propose a scheme of heuristic search (HS) based object insertion. It estimates the proper position of the inserted object from 2D bird’s eye view (BEV). Experiments demonstrate the de-noising accuracy of c-$k$NN, rendering quality of HS based object insertion, and improvement of RS-Aug on object detection. Pei An, Junxiong Liang, Jie Ma 0003, Yanfei Chen, Liheng Wang, You Yang 0002, Qiong Liu 0001 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2023 | Focal Stack Image Compression Based on Basis-Quadtree RepresentationabstractIn this paper, we propose an efficient compression scheme for focal stack images (FoSIs) based on a new basis-quadtree representation. In the new basis-quadtree representation, FoSIs are initially reorganized as co-located block groups in the depth dimension. In each group, selective basis blocks and adaptive quadtree partition are optimized to predict the focused or defocused co-located blocks by intra-group approximation. By solving a joint optimization problem, FoSIs can be efficiently represented by the optimal basis blocks, corresponding quadtree partition and approximation parameters, which will be compressed separately. Then, these basis blocks are stitched into several new frames (basis frames) according to their original locations and partition modes. Basis frames are compressed by our designed encoder, where the intra-group approximation is embedded into the high efficiency video coding (HEVC) encoder. Thus, the redundancies of basis blocks can be further eliminated. Finally, the approximation parameters are refined to suppress the amplified errors caused by introduced compression blur after basis frame coding. The refined parameters are compressed losslessly and multiplexed with the bitstream of the basis frames to ensure the reconstruction quality of FoSIs. Experiments on 12 test sequences demonstrate that the proposed scheme can obtain higher coding performance than the state-of-the-art comparison schemes. Specifically, the proposed scheme achieves up to 5.23 dB PSNR gains and 71.59% bitrate savings over the HEVC baseline scheme on sequences I03 and I05, respectively. Kejun Wu, You Yang 0002, Qiong Liu 0001, Xiao-Ping Zhang 0002 |
IEEE Trans. Multim. | 2 |
| 2022 | Iterative enhancement scheme of synthesized color and depth images for immersive video systemabstractImmersive video allows viewers to freely switch the viewpoints. The intensity of realistic experience greatly relies on the quality of synthesized depth maps. However, there exist distorted regions due to inaccurate depth estimation or compression. Yongquan Su, Qiong Liu 0001, Kejun Wu, Gangyi Jiang, You Yang 0002 |
DCC | 5 |
| 2022 | Self-supervised Indoor 360-Degree Depth Estimation via Structural Regularization
Weifeng Kong, Qiudan Zhang, You Yang 0002, Tiesong Zhao, Wenhui Wu 0001, Xu Wang 0006 |
PRICAI (3) | 3 |
| 2022 | Gaussian-Wiener Representation and Hierarchical Coding Scheme for Focal Stack ImagesabstractFocal stack images (FoSIs) are a set of 2D images that captured one scene with serial focal depths. The redundancies of FoSIs mainly come from gradual focused depths changes rather than motion of objects. Conventional coding schemes cannot fully exploit such redundancies, leading to coding inefficiency. In this paper, we propose a new Gaussian-Wiener representation to model the gradual focused depths changes among FoSIs. In the representation, image degradation-restoration relations are utilized to describe the focus-defocus changing characteristics of FoSIs. Based on this representation, we propose a new hierarchical coding scheme for fully exploiting the inter-frame redundancies of FoSIs. In the scheme, a Gaussian-Wiener representation based inter prediction (GWR-IP) is presented by embedding Gaussian convolution and Wiener deconvolution into normal video encoder. Block-wise focus-defocus changing of FoSIs can be predicted in bi-directional manner by solving optimization problem. For higher coding efficiency, a Gaussian-Wiener representation based hierarchical prediction structure (GWR-HPS) is also designed and applied in the coding scheme. The proposed coding scheme is performed on 10 test sequences, including 5 synthetic scenes and 5 realistic scenes. Experimental results show that proposed coding scheme can obtain 2.640 dB PSNR gains and 51.830% bitrate savings on average of all test sequences in Low Delay P configuration, 2.123 dB PSNR gains and 43.975% bitrate savings in Low Delay B configuration, and 1.044 dB PSNR gains and 26.078% bitrate savings in Random Access configuration. Particularly, it achieves up to 65.544% bit rate savings and 3.901 dB PSNR increments for test sequence I09 in Low Delay P configuration. Furthermore, ablation test demonstrates that Gaussian representation contributes more on coding performance than Wiener representation and GWR-HPS. Kejun Wu, You Yang 0002, Qiong Liu 0001, Xiao-Ping Zhang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Deep Learning Based Just Noticeable Difference and Perceptual Quality Prediction Models for Compressed VideoabstractHuman visual system has a limitation of sensitivity in detecting small distortion in an image/video and the minimum perceptual threshold is so called Just Noticeable Difference (JND). JND modelling is challenging since it highly depends on visual contents and perceptual factors are not fully understood. In this paper, we propose deep learning based JND and perceptual quality prediction models, which are able to predict the Satisfied User Ratio (SUR) and Video Wise JND (VWJND) of compressed videos with different resolutions and coding parameters. Firstly, the SUR prediction is modeled as a regression problem that fits deep learning tools. Then, Video Wise Spatial SUR method (VW-SSUR) is proposed to predict the SUR value for compressed video, which mainly considers the spatial distortion. Thirdly, we further propose Video Wise Spatial-Temporal SUR (VW-STSUR) method to improve the SUR prediction accuracy by considering the spatial and temporal information. Two fusion schemes that fuse the spatial and temporal information in quality score level and in feature level, respectively, are investigated. Finally, key factors including key frame and patch selections, cross resolution prediction and complexity are analyzed. Experimental results demonstrate the proposed VW-SSUR method outperforms in both SUR and VWJND prediction as compared with the state-of-the-art schemes. Moreover, the proposed VW-STSUR further improves the accuracy as compared with the VW-SSUR and the conventional JND models, where the mean SUR prediction error is 0.049, and mean VWJND prediction error is 1.69 in quantization parameter and 0.84 dB in peak signal-to-noise ratio. Yun Zhang 0002, Huanhua Liu, You Yang 0002, Xiaoping Fan, Sam Kwong, C.-C. Jay Kuo |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | SA-GNN: Stereo Attention and Graph Neural Network for Stereo Image Super-Resolution
Qiong Liu 0001, You Yang 0002 |
ICIG (3) | 3 |
| 2021 | Dedark+Detection: A Hybrid Scheme for Object Detection under Low-light SurveillanceabstractObject detection under low-light surveillance is a crucial problem that less efforts have been made on it. In this paper, we proposed a hybrid method that jointly use enhancement and object detection for the above challenge, namely Dedark+Detection. In this method, the low-light surveillance video is processed by the proposed de-dark method, and the video can thus be converted to appearance under normal lighting condition. This enhancement bring more benefits to the subsequent stage of object detection. After that, an object detection network is trained on the enhanced dataset for practical applications under low-light surveillance. Experiments are performed on 18 low-light surveillance video test sequences, and superior performance can be found when comparing to state-of-the-arts. Xiaolei Luo, Sen Xiang, Yingfeng Wang, Qiong Liu 0001, You Yang 0002, Kejun Wu |
MMAsia | 5 |
| 2021 | Divide and conquer: Ill-light image enhancement via hybrid deep network
You Yang 0002, Qiong Liu 0001, Zahid Hussain Qaisar |
Expert Syst. Appl. | 2 |
| 2021 | A ghostfree contrast enhancement method for multiview images without depth information
You Yang 0002, Qiong Liu 0001, Zahid Hussain Qaisar |
J. Vis. Commun. Image Represent. | 2 |
| 2021 | Cubemap-Based Perception-Driven Blind Quality Assessment for 360-degree Imagesabstractimage can be represented with different formats, such as the equirectangular projection (ERP) image, viewport images or spherical image, for its different processing procedures and applications. Accordingly, the 360-degree image quality assessment (360-IQA) can be performed on these different formats. However, the performance of 360-IQA with the ERP image is not equivalent with those with the viewport images or spherical image due to the over-sampling and the resulted obvious geometric distortion of ERP image. This imbalance problem brings challenge to ERP image based applications, such as 360-degree image/video compression and assessment. In this paper, we propose a new blind 360-IQA framework to handle this imbalance problem. In the proposed framework, cubemap projection (CMP) with six inter-related faces is used to realize the omnidirectional viewing of 360-degree image. A multi-distortions visual attention quality dataset for 360-degree images is firstly established as the benchmark to analyze the performance of objective 360-IQA methods. Then, the perception-driven blind 360-IQA framework is proposed based on six cubemap faces of CMP for 360-degree image, in which human attention behavior is taken into account to improve the effectiveness of the proposed framework. The cubemap quality feature subset of CMP image is first obtained, and additionally, attention feature matrices and subsets are also calculated to describe the human visual behavior. Experimental results show that the proposed framework achieves superior performances compared with state-of-the-art IQA methods, and the cross dataset validation also verifies the effectiveness of the proposed framework. In addition, the proposed framework can also be combined with new quality feature extraction method to further improve the performance of 360-IQA. All of these demonstrate that the proposed framework is effective in 360-IQA and has a good potential for future applications. Hao Jiang 0014, Gangyi Jiang, Mei Yu 0001, Yun Zhang 0002, You Yang 0002, Zongju Peng |
IEEE Trans. Image Process. | 5 |
| 2021 | I Understand You: Blind 3D Human Attention Inference From the Perspective of Third-PersonabstractInferring object-wise human attention in 3D space from the third-person perspective (e.g., a camera) is crucial to many visual tasks and applications, including human-robot collaboration, unmanned vehicle driving, etc. Challenges arise from classical human attention when human eyes are not visible to cameras, gaze point is outside the field of vision, or the gazed object is occluded by others in the 3D space. In this case, blind 3D human attention inference brings a new paradigm to the community. In this paper, we address these challenges by proposing a scene-behavior associated mechanism, in which both 3D scene and temporal behavior of human are adopted to infer object-wise human attention and its transition. Specifically, point cloud is reconstructed and used for the spatial representation of 3D scene, which is beneficial to handle the blind problem from the perspective of a camera. Based on this, in order to address the blind human attention inference without eye information, we propose a Sequential Skeleton Based Attention Network (S2BAN) for behavior-based attention modeling. As is embedded in the scene-behavior associated mechanism, the proposed S2BAN is built under the temporal architecture of Long-Short-Term-Memory (LSTM). Our network employs human skeleton as behavior representation, and maps it to the attention direction frame by frame, which makes attention inference a temporal-correlated issue. With the help of S2BAN, 3D gaze spot and further the attended objects can be obtained frame by frame via intersection and segmentation on the previously reconstructed point cloud. Finally, we conduct experiments from various aspects to verify the object-wise attention localization accuracy, the angular error of attention direction calculation, as well as the subjective results. The experimental results show that the proposed outperforms other competitors. You Yang 0002, Qiong Liu 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Make Full Use of Priors: Cross-View Optimized Filter for Multi-View Depth EnhancementabstractMulti-view video plus depth (MVD) is the promising and widely adopted data representation for future 3D visual applications and interactive media. However, compression distortions on depth videos impede the development of such applications, and filters are crucially needed for the quality enhancement at the terminal side. Cross-view priors can intuitively be involved in filter design, but these priors are also distorted in compression and thus the contribution of them can hardly be considered in previous research. In this article, we propose a cross-view optimized filter for depth map quality enhancement by making full use of inner- and cross-view priors. We dedicate to evaluate the contributions of distorted cross-view priors in filtering the current view of depth, and then both inner- and cross-view priors can be involved in the filter design. Thus, distortions of cross-view priors are not barriers again as before. For the purpose of that, mutual information guided cross-view consistency is designed to evaluate the contributions of cross-view priors from compression distortions of MVD. After that, under the framework of global optimization, both inner- and cross-view priors are modeled and taken to minimize the designed energy function where both data accuracy and spatial smoothness are modeled. The experimental results show that the proposed model outperforms state-of-the-art methods, where 3.289 dB and 0.0407 average gains on peak signal-to-noise ratio and structural similarity metrics can be obtained, respectively. For the subjective evaluations, object details and structure information are recovered in the compressed depth video. We also verify our method via several practical applications, including virtual view synthesis for smooth interaction and point cloud for 3D modeling for accuracy evaluation. In these verifications, the ringing and malposition artifacts on object contours are properly handled for interactive video, and discontinuous object surfaces are restored for 3D modeling. All of these results suggest that compression distortions in MVD can be properly filtered by the proposed model, which provides a promising solution for future bandwidth constrained 3D and interactive visual applications. Qiong Liu 0001, You Yang 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2020 | Gaussian Guided Inter Prediction for Focal Stack Images CompressionabstractFocal stack is an intermediate data representation obtained by projecting 4D light field (LF) in z-dimension. This kind of representation is fundamental for future interactive and immersive visual applications. However, focal stack images are a series of samples focused at varying depths of static scenes, which yields considerable redundancy among them. In this paper, we propose a Gaussian guided inter prediction model to eliminate the visual redundancy. In our work, the effect of the varying focus distance on plenoptic imaging system is characterized by the point spread function (PSF), and we propose a simplified Gaussian-like PSF to fit this characteristic according to the features of focal stack images. After that, Gaussian guided motion estimation and motion compensation are both implemented in this model. Experimental results show that our model can achieve smaller residual distribution. There are 10.33% bit rate saving and 0.397 dB PSNR gain on average in three configurations compared with HEVC anchor. Particularly, it brings about up to 16.60% bit rate saving with 0.649 dB PSNR increment in Low Delay P configuration. Kejun Wu, Qiong Liu 0001, Yaguang Yin, You Yang 0002 |
DCC | 4 |
| 2020 | Similarity Graph Convolutional Construction Network for Interactive Action Recognition
Qiong Liu 0001, You Yang 0002 |
MMM (2) | 3 |
| 2020 | Depth map artefacts reduction: a reviewabstractDepth maps are crucial for many visual applications, where they represent the positioning information of the objects in a three‐dimensional scene. Usually, depth maps can be acquired via various devices, including Time of Flight, Kinect or light field camera, in practical applications. However, a brutal truth is that both intrinsic and extrinsic artefacts can be found in these depth maps which limits the prosperity of three‐dimensional visual applications. In this study, the authors survey the depth map artefacts reduction methods proposed in the literature, from mono‐ to multi‐view, via spatial to temporal dimension, in local to global manner, with signal processing to learning‐based methods. They also compare the state‐of‐the‐arts via different metrics to show their potentials in future visual applications. Mostafa Mahmoud Ibrahim, Qiong Liu 0001, Ehsan Adeli-Mosabbeb, You Yang 0002 |
IET Image Process. | 6 |
| 2020 | Adaptive colour-guided non-local means algorithm for compound noise reduction of depth mapsabstractDepth maps are used to describe object positioning information in three‐dimensional (3D) space, and they are crucial for RGB‐D data representation, which is useful for numerous interactive visual applications. In practice, depth maps are often contaminated by compound noise, including intrinsic noise and missing regions owing to active illumination shadows. As existing noise models cannot describe the above‐mentioned compound noise effectively, the subsequent filter design is a challenging task. In this study, an adaptive colour‐guided non‐local mean (NLM) filter is proposed to address such compound noise. First, the authors classify the depth map into hole and non‐hole pixels. Then, the proposed filter is designed on the basis of the NLM framework, where the colour image is used as a guide prior for hole‐artifact removal. Finally, the authors use a shock filter to effectively address the non‐regularisation of the restored depth map edges and remove the remaining noise. Experiments show that the proposed filter qualitatively and quantitatively outperforms existing colour‐guided and unguided filters. Moreover, the authors verify the superiority of the proposed filter through virtual view synthesis and 3D scene reconstruction applications. Mostafa Mahmoud Ibrahim, Qiong Liu 0001, You Yang 0002 |
IET Image Process. | 3 |
| 2020 | Novel calibration method for camera array in spherical arrangement
Pei An, Qiong Liu 0001, Firas Abedi, You Yang 0002 |
Signal Process. Image Commun. | 4 |
| 2020 | MV-GNN: Multi-View Graph Neural Network for Compression Artifacts ReductionabstractInevitable compression artifacts in multi-view video (MVV) can clearly degrade the quality of experience in many interaction-oriented 3D visual applications. Under the framework of asymmetric coding, low-quality images can be enhanced with high-quality images from the neighboring viewpoints considering the similarity among different views. However, compression artifacts and warping error cause different cross-view quality gaps for various sequences, and thus the contribution of cross-view priors can hardly be located and considered in previous works. In this paper, we propose a multi-view graph neural network (MV-GNN) to reduce compression artifacts in multi-view compressed images. We dedicate to design a fusion mechanism which can exploit contributions from neighboring viewpoints and meanwhile suppress the misleading information. In our method, a GNN-based fusion mechanism is designed to fuse the cross-view information under the aggregation and update mechanism of GNN. Experiments show that 1.672 dB and 0.0242 average gains on PSNR and SSIM metrics can be obtained, respectively. For the subjective evaluations, blocking effect in the compressed images are clearly suppressed and the damaged object boundary are better recovered. The experimental results demonstrate that our MV-GNN outperforms the state-of-the-art methods. Qiong Liu 0001, You Yang 0002 |
IEEE Trans. Image Process. | 3 |
| 2019 | Multi-view Multi-modality Priors Residual Network of Depth Video Enhancement for Bandwidth Limited Asymmetric Coding FrameworkabstractAsymmetric coding methodology for multi-view video plus depth is a promising technique for future three-dimensional and multi-view driven visual applications for its superior coding performance in bandwidth limited conditions. Since the depth video suffers from asymmetric distortions corresponding to viewpoint, it's a challenge in smooth and quality consistent content based interaction. To solve this challenge, we propose a residual learning framework to enhance the quality of compression distorted multi-view depth video. In this work, we exploit the correlation between viewpoints to restore the target viewpoint depth maps by using multi-modality priors, which are depth maps from adjacent viewpoints with better quality and color frames in the same viewpoint. A residual network is designed to fully exploit the contribution from these priors. Experimental results show the superiority of our framework in the quality improvement on both decoded depth video and synthesized virtual viewpoint images. Qiong Liu 0001, You Yang 0002 |
DCC | 3 |
| 2019 | A Global Co-Saliency Guided Bit Allocation for Light Field Image CompressionabstractLight field is the most prospective technology for interactive and immersive visual applications. and light field image is an intermediate data format that demands a large amount of storage space and higher transmission bandwidth. Therefore, compression of light field images is highly desired for further applications. In this paper, we propose a co-saliency guided bit allocation scheme with constraints of consistency among sub-aperture images. Firstly, saliency is jointly detected on color and depth images of sub-aperture by improving our previous model. The obtained pixel-wise co-saliency map is converted into block-wise via K-means clustering. In this way, the saliency weight of each coding tree unit (CTU) can be calculated. Then, target bits of each CTU are initially determined by the weight of each block. The allocation is adjusted dynamically under the guidance of co-saliency map and the image texture complexity. The experimental results show that BD-PSNR of 0.384 dB can be achieved for the salient region at the cost of less than 0.107 dB decrease for the whole image compared to HTM anchor. Moreover, subjective quality of proposed scheme outperforms the anchor for the salient region, and there is no noticeable distortion for non-salient region. Kejun Wu, Zongbang Liao, Qiong Liu 0001, Yaguang Yin, You Yang 0002 |
DCC | 5 |
| 2019 | Multi-view high dynamic range reconstruction via gain estimationabstractMulti-view high dynamic range reconstruction is a challenging problem, especially if the multi-view low dynamic range images are obtained from cameras arranged sparsely with limited shared view of vision among them. In this paper, we address the above challenge in addition to the back-lighting problem. We first enclose the geometry characteristic of the scene to rectify the outlier feature points. Consequently, an exposure gain is calculated according to those rectified features. After that, we extend the dynamic range for the multi-view low dynamic range images based on the estimated gain, then, generate a final high dynamic range image per view. Experimental results demonstrate superior performance for the proposed method over state-of-the-art methods in both objective and subject comparisons. These results suggest that our method is suitable to improve the visual quality of multi-view low dynamic range images captured in low back-lighting conditions via commercial cameras sparsely located among each other. Firas Abedi, Qiong Liu 0001, You Yang 0002 |
VCIP | 3 |
| 2019 | Cross-View Multi-Lateral Filter for Compressed Multi-View Depth VideoabstractMulti-view depth is crucial for describing positioning information in 3D space for virtual reality, free viewpoint video, and other interaction- and remote-oriented applications. However, in cases of lossy compression for bandwidth limited remote applications, the quality of multi-view depth video suffers from quantization errors, leading to the generation of obvious artifacts in consequent virtual view rendering during interactions. Considerable efforts must be made to properly address these artifacts. In this paper, we propose a cross-view multi-lateral filtering scheme to improve the quality of compressed depth maps/videos within the framework of asymmetric multi-view video with depth compression. Through this scheme, a distorted depth map is enhanced via non-local candidates selected from current and neighboring viewpoints of different time-slots. Specifically, these candidates are clustered into a macro super pixel denoting the physical and semantic cross-relationships of the cross-view, spatial and temporal priors. The experimental results show that gains from static depth maps and dynamic depth videos can be obtained from PSNR and SSIM metrics, respectively. In subjective evaluations, even object contours are recovered from a compressed depth video. We also verify our method via several practical applications. For these verifications, artifacts on object contours are properly managed for the development of interactive video and discontinuous object surfaces are restored for 3D modeling. Our results suggest that the proposed filter outperforms state-of-the-art filters and is suitable for use in multi-view color plus depth-based interaction- and remote-oriented applications. You Yang 0002, Qiong Liu 0001, Zhen Liu 0012 |
IEEE Trans. Image Process. | 1 |
| 2019 | A Two-Stage Clustering Based 3D Visual Saliency Model for Dynamic ScenariosabstractThree-dimensional (3D) visual saliency is fundamental for vision-guided applications such as human-computer interaction in virtual reality, image quality assessment, object tracking, and event retrieval. Classical models for 3D visual saliency can draw an appropriate saliency map when the quality of the required depth maps or auxiliary cues is high enough. However, the depth map is usually impaired with artifacts (such as holes or noise) from faults in stereo matching or multipaths in range sensors. In these cases, challenges arise in those 3D visual saliency models because the core preliminary processes, such as the detection of low-level visual features, may fail. To solve this problem, we proposed a two-stage clustering-based 3D visual saliency model for human visual fixation prediction in dynamic scenarios. In this model, a two-stage clustering scheme is designed to handle the negative influence of impaired depth videos. With the help of this scheme, representative cues are selected for saliency modeling. After that, multimodal saliency maps are obtained from depth, color, and 3D motion cues. Finally, a cross-Bayesian model is designed for the pooling of multimodal saliency maps. The experimental results demonstrate that the proposed 3D saliency model based on two-stage clustering outperforms other state-of-the-art models on a variety of metrics. Furthermore, the consistency and robustness of our model are also verified. You Yang 0002, Pian Li, Qiong Liu 0001 |
IEEE Trans. Multim. | 1 |
| 2018 | Graph-Based Saliency Fusion with Superpixel-Level Belief Propagation for 3D Fixation PredictionabstractIn recent years, many 3D visual attention models (VAMs) have been proposed with diverse fusion methods, of which the main challenge lies in the inconsistence, or even conflicts of different saliency maps. To address the challenge, we propose a graph-based fusion method with superpixel-level belief propagation for 3D fixation prediction on stereoscopic video, which models the aggregation as a global optimization issue. After extracting multi-modality saliency maps, the fusion step is based on the graph constructed at superpixel level, and we design for the graph an energy function considering multi-modality constraints, which is minimized using the belief propagation algorithm. The experimental results on two databases demonstrate that the proposed model achieves competitive performance. Qiong Liu 0001, You Yang 0002 |
ICIP | 4 |
| 2017 | Illumination Attributes Coding for Virtual Reality Broadcasting SystemabstractIn this paper, we propose a method of illumination attribute coding method for virtual reality broadcasting system. As for the virtual reality content, it is captured with local illumination variations. Our method is motivated by the Phong illumination model, and illumination attribute is extracted from images and then an illumination reference is synthesized with higher correlation to the current image. You Yang 0002, Qiong Liu 0001 |
DCC | 1 |
| 2017 | A robust 3D visual saliency computation model for human fixation prediction of stereoscopic videosabstract3D saliency have been gaining an increasing amount of attention because of the emergence of 3D contents and applications. Most of existing works on 3D saliency are based on the assumption of fine quality of depth maps. However, depth maps from stereo matching or range sensors are usually with holes and artifacts, which severely drop the performance of those 3D saliency models. In this paper, we propose a robust 3D saliency computation model for human fixation prediction. First, a cluster-contrast saliency prediction model is proposed for depth maps. The prediction is obtained with the centroid of the largest clusters of each depth super-pixel, and thus those bad effects originate from holes and other artifacts in depth are then eliminated. The cluster-contrast strategy is exploited both to depth texture and motions in 3D video. Finally, a Bayesian integration model is proposed for the multi-modality fusion between depth saliency and color saliency. The experimental results demonstrate that our saliency model has better performance of accuracy than other state-of-arts models. Qiong Liu 0001, You Yang 0002, Pian Li |
VCIP | 2 |
| 2016 | Cluster-based cross-view filtering for compressed multi-view depth mapsabstractIn the field of multi-view video coding, multi-view plus depth video is an important data format, but it always suffers from quantization errors, which result in obvious artifacts in consequent virtual view rendering. In this paper, we propose a cluster-based cross-view filtering (CBF) scheme for the enhancement of compressed depth maps. In this scheme, reconstructed depth information are mapped from cross-view, and this information is benefit to the proposed filter. Then in filtering one viewpoint depth map with candidate information that are selected from non-locally current and neighboring viewpoints. Specifically, in our scheme, candidates are clustered in 3D super-pixel wise rather than block wise due to cross-relationship among pixels in depth maps. The experimental results show that 2.0074 dB average gain can be obtained by our scheme, which suggests that the scheme outperforms than state-of-the-art and classical filters in filtering the reconstructed depth maps. Zhen Liu 0002, Qiong Liu 0001, You Yang 0002, Yuchi Liu, Gangyi Jiang, Mei Yu 0001 |
VCIP | 3 |
| 2016 | User models of subjective image quality assessment on virtual viewpoint in free-viewpoint video system
You Yang 0002, Xu Wang 0006, Qiong Liu 0001, Mingliang Xu 0001 |
Multim. Tools Appl. | 1 |
| 2015 | Multi-camera interference cancellation of time-of-flight (TOF) camerasabstractIn the applications based on depth, multiple TOF cameras are often required to capture the same scene. But if multiple cameras operate simultaneously on the same frequency, they interfere with each other. The multi-camera interference causes a lot of errors in depth measurement. The depth quality is severely reduced, which limits the application of TOF cameras and needs to be resolved. In this paper, a multi-camera interference model is presented. The interference signal for multiple frames is proved to be an ergodic and wide-sense stationary stochastic process. The least square estimation of noninterference signal is proposed to remove the multi-camera interference. The results of experiments prove the approach can recover the depth and amplitude information of multiple TOF cameras from the severe interference. Lianhua Li, Sen Xiang, You Yang 0002, Li Yu 0003 |
ICIP | 3 |
| 2015 | A database of reflected irradiance field with depth for image based relightingabstractImage based relighting is an important application for light field, which represents the reflectance properties of the captured object/scene. Many systems were proposed without depth information of the captured object/scene. In this paper, we propose a lighting system named as Light Cube, and the system can capture color and depth information synchronously. This is important to consequent relighting process for both image based or model based methods. Utilizing the proposed lighting system, we calibrate the illumination and reflectance performance, and capture a database for reflected irradiance field. You Yang 0002, Qiong Liu 0001 |
ICIP | 1 |
| 2015 | Depth map reconstruction and rectification through coding parameters for mobile 3D video system
You Yang 0002, Huiping Deng, Li Yu 0003 |
Neurocomputing | 1 |
| 2015 | A bundled-optimization model of multiview dense depth map synthesis for dynamic scene reconstruction
You Yang 0002, Xu Wang 0006, Qiong Liu 0001, Li Yu 0003 |
Inf. Sci. | 1 |
| 2015 | Dense depth image synthesis via energy minimization for three-dimensional video
You Yang 0002, Qiong Liu 0001, Hao Liu 0019, Li Yu 0003, Fanglin Wang |
Signal Process. | 1 |
| 2015 | Interfered depth map recovery with texture guidance for multiple structured light depth cameras
Sen Xiang, Li Yu 0003, You Yang 0002, Qiong Liu 0001, Jialiang Zhou |
Signal Process. Image Commun. | 3 |
| 2015 | Depth Error Elimination for RGB-D CamerasabstractThe rapid spreading of RGB-D cameras has led to wide applications of 3D videos in both academia and industry, such as 3D entertainment and 3D visual understanding. Under these circumstances, extensive research efforts have been dedicated to RGB-D camera--oriented topics. In these topics, quality promotion of depth videos with the temporal characteristic is emerging and important. Due to the limited exposure time of RGB-D cameras, object movement can easily lead to motion blurs in intensive images, which can further result in obvious artifacts (holes or fake boundaries) in the corresponding depth frames. With regard to this problem, we propose a depth error elimination method based on time series analysis to remove the artifacts in depth images. In this method, we first locate the regions with erroneous depths in intensive images by using motion blur detection based on a time series analysis model. This is based on the fact that the depth image is calculated by intensive color images that are captured synchronously by RGB-D cameras. Then, the artifacts, such as holes or fake boundaries, are fixed by a depth error elimination method. To evaluate the performance of the proposed method, we conducted experiments on 250 images. Experimental results demonstrate that the proposed method can locate the error regions correctly and eliminate these artifacts effectively. The quality of depth video can be improved significantly by using the proposed method. Yue Gao 0002, You Yang 0002, Yi Zhen, Qionghai Dai |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2014 | A multi-dimensional image quality prediction model for user-generated images in social networks
You Yang 0002, Xu Wang 0006, Jialie Shen 0001, Li Yu 0003 |
Inf. Sci. | 1 |
| 2013 | Stereotime: a wireless 2D and 3D switchable video communication systemabstractMobile 3D video communication, especially with 2D and 3D compatible, is a new paradigm for both video communication and 3D video processing. Current techniques face challenges in mobile devices when bundled constraints such as computation resource and compatibility should be considered. In this work, we present a wireless 2D and 3D switchable video communication to handle the previous challenges, and name it as Stereotime. The methods of Zig-Zag fast object segmentation, depth cues detection and merging, and texture-adaptive view generation are used for 3D scene reconstruction. We show the functionalities and compatibilities on 3D mobile devices in WiFi network environment. You Yang 0002, Qiong Liu 0001, Yue Gao 0002, Binbin Xiong, Li Yu 0003, Huan-Bo Luan, Rongrong Ji, Qi Tian 0001 |
ACM Multimedia | 1 |
| 2013 | Quality Assessment on User Generated Image for Mobile Search Application
Qiong Liu 0001, You Yang 0002, Xu Wang 0006, Liujuan Cao |
MMM (2) | 2 |
| 2013 | Texture-adaptive hole-filling algorithm in raster-order for three-dimensional video applications
Qiong Liu 0001, You Yang 0002, Yue Gao 0002, Richang Hong |
Neurocomputing | 2 |
| 2013 | A Bayesian framework for dense depth estimation based on spatial-temporal correlation
Qiong Liu 0001, You Yang 0002, Yue Gao 0002, Rongrong Ji, Li Yu 0003 |
Neurocomputing | 2 |
| 2012 | Cross-View Down/Up-Sampling Method for Multiview Depth Video CodingabstractIn this letter, we propose a cross-view down/up-sampling (CDU) method for the framework of reduced resolution multiview depth video coding, which exploits cross-view information to assist the up-sampling at the decoder. In the down-sampling procedure of CDU, the odd-even interlaced extraction is employed to preserve more confident information of the original depth video with reduced resolution. In the decoder, the cross-view information is exploited for up-sampling the reconstructed depth video. An iterative interpolation process is proposed to eliminate the effect of compression distortion on this up-sampling. Experimental results demonstrate the gains of up to 3.88 dB for the proposed algorithm and better quality of synthesized views. Qiong Liu 0001, You Yang 0002, Rongrong Ji, Yue Gao 0002, Li Yu 0003 |
IEEE Signal Process. Lett. | 2 |
| 2010 | Representative views re-ranking for 3D model retrieval with multi-bipartite graph reinforcement modelabstractIn this paper, we propose a multi-bipartite graph reinforcement model for representative views re-ranking in 3D model retrieval. Given the views of one query 3D model, all query views are grouped into clusters to generate representative views and corresponding original weights. In the retrieval procedure, labeled positive retrieval results are employed to refine the query information. Each group of views from positive retrieval results and the group of representative query views are employed to construct a bipartite graph, and a multi-bipartite graph reinforcement algorithm is performed on these bipartite graphs to re-rank all views. Then the weights of all representative query views are updated. Experimental results on two 3D model databases are provided to justify the effectiveness of the proposed method. Yue Gao 0002, You Yang 0002, Qionghai Dai, Naiyao Zhang |
ACM Multimedia | 2 |
| 2010 | 3D object retrieval with bag-of-region-wordsabstractView-based method becomes an essential approach to 3D object retrieval in recent years. In the view-based 3D object retrieval framework, each object is described by a set of views and representative features are extracted from these views to match the objects in database. In this paper, we propose a novel 3D multi-view representation method, Bag-of-Region-Words (BoRW). It first gridly selects points in each view and extracts local SIFT features. Each local feature is encoded into a visual word with a trained visual vocabulary. Then each view is split into several regions, and each region is represented by a bag-of-visual-words feature vector. All the obtained regions are further grouped into clusters based on the bag-of-visual-words feature, and one feature is selected from each cluster with corresponding weight. In this way, each object is described by a set of BoRW. The Earth Movers Distance is employed to estimate the distance between two BoRW feature vectors. Experimental results show that the proposed method can achieve better retrieval performance than existing methods. Yue Gao 0002, You Yang 0002, Qionghai Dai, Naiyao Zhang |
ACM Multimedia | 2 |
| 2010 | Depth perceptual region-of-interest based multiview video coding
Yun Zhang 0002, Gangyi Jiang, Mei Yu 0001, You Yang 0002, Zongju Peng, Ken Chen 0003 |
J. Vis. Commun. Image Represent. | 4 |
| 2006 | Fast Multi-view Disparity Estimation for Multi-view Video Systems
Gangyi Jiang, Mei Yu 0001, Feng Shao 0001, You Yang 0002 |
ACIVS | 4 |
| 2006 | Parallel Process of Hyper-Space-Based Multiview Video CompressionabstractMultiview video coding (MVC) is a key technology in free-viewpoint television. MVC based on traditional existing codec system has been studied widely, but all of them need powerful computational capacity in processing. Parallel process of MVC can facilitate the efficient implementation of encoder and decoder and has been required as a function by MPEG. In this paper, a parallelization methodology for MVC based on hyper-space theory is presented and tested on the local area multi-computer - message passing interface (LAM-MPI) parallel platform and modified H.264 codec. Experimental results show that the proposed method can speed up processing of multiview video compression and obtain high rate-distortion results. You Yang 0002, Gangyi Jiang, Mei Yu 0001, Dingju Zhu |
ICIP | 1 |