VLDB 2026 Research / reviewers in the wild / expert
Qiong Liu 0001
dblp:20/718-1
· DBLP profile ↗
78ranked-venue papers
10as first author
39since 2021 · last 2026
0000-0002-2407-806XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 53 · 5 first-author · 25 since 2021Artificial intelligence and machine learning · 18 · 3 first-author · 13 since 2021Databases, data management, data science and information retrieval · 9 · 1 first-author · 3 since 2021Computer networks · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Spatially adaptive representation of facial meshes for face video super-resolution
Shangchen Cai, Qiong Liu 0001, You Yang 0002 |
Expert Syst. Appl. | 4 |
| 2026 | Density-aware few-parametric networks for robust few-shot point cloud semantic segmentation
Yudong Liang, Pei An, Qiong Liu 0001, You Yang 0002 |
Neurocomputing | 3 |
| 2026 | FM-MFIF: A motion-aware multi-focus image fusion network based on multi-scale focus migration
Zhilong Li, Zhulun Yang, You Yang 0002, Qiong Liu 0001 |
Image Vis. Comput. | 4 |
| 2026 | PatchNeRF: Patch-based Neural Radiance Fields for real time view synthesis in wide-scale scenes
Xiaoguang Jiang, Qiong Liu 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2026 | Space-Time Correlation Adaptive-Layered Optimization for Microimage Motion Search in Plenoptic video coding
Jingyang Luo, Qiong Liu 0001, Xiatian Xie, Pengpeng Han, You Yang 0002 |
J. Vis. Commun. Image Represent. | 2 |
| 2026 | FSF-Net: Enhance 4D occupancy forecasting with coarse BEV scene flow for autonomous driving
Erxin Guo, Pei An, You Yang 0002, Qiong Liu 0001, Anan Liu |
Pattern Recognit. | 4 |
| 2026 | DFS-Net: A Dense Focal Stack Image Generation Network From Misaligned Multi-Focus ImagesabstractDense focal stack images inherently encode depth cues and are crucial for various 3D vision applications. However, existing generation methods are susceptible to misalignment and introduce a domain gap between synthetic and real-world data due to off-axis aberrations. To address these challenges, we introduce DFS-Net, an aberration-aware dense focal stack image generation network. DFS-Net consists of two core modules: all-in-focus image synthesis and aberration-aware point spread function (PSF) generation. The all-in-focus image synthesis is achieved through a densely connected fusion network based on multi-scale focus migration and focus property detection. This fusion network can effectively fuse misaligned multi-focus images into an all-in-focus image. The aberration-aware PSF generation is realized through a multi-layer perceptron (MLP) network. Supervised by ray-tracing-based PSFs, the MLP network can generate spatially varying PSFs for arbitrary spatial positions and focus distances. By selecting a set of focus distances, the generated PSF maps are locally convolved with the all-in-focus image to produce an aberration-aware dense focal stack. We conduct extensive comparative experiments on all-in-focus image fusion and focal stack generation against state-of-the-art methods. The experimental results demonstrate that DFS-Net can synthesize all-in-focus images with high subjective and objective quality, as well as generate dense focal stacks that closely approximate ray-tracing results. In addition, we conduct comparative experiments on the depth-from-focus and salient object detection tasks using the generated focal stacks. The experimental results demonstrate that our DFS-Net can significantly enhance the performance of existing depth-from-focus and salient object detection models. The code and dataset will be publicly available at https://github.com/North-Li/DFS-Net. Zhilong Li, Pei An, You Yang 0002, Qiong Liu 0001, Dan Song 0006, Anan Liu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | PNProRL: Self-Supervised Neural Relighting via Photometric Perception and Progressive OptimizationabstractPortrait relighting shows great potential in photography, film, and AR by simulating diverse lighting effects. Existing state-of-the-art methods often rely on expensive paired OLAT or synthetic data, which limits scalability. Moreover, accurately modeling the interaction between physics-guided rendering, neural rendering, and real-world remains challenging. To address these issues, we propose a novel multi-stage self-supervised relighting framework. It progressively refines intrinsic scene properties via a simple-to-complex training strategy, removing the need for expensive paired data while adapting to various lighting conditions. One core design introduces a novel pre-training method approach using diverse shading-based masking for self-reconstruction, which improves the model's perception of complex lighting variations. Furthermore, we introduce two perceptual modules that leverage the linear superposition of light to narrow the gap between physics-guided and neural rendering, and better align relit results with real-world observations. Extensive experiments demonstrate that our unified framework achieves new state-of-the-art performance in portrait relighting, surpassing recent methods in photorealism, synthesis quality, and identity preservation. It provides a practical paradigm for high-fidelity relighting under diverse lighting. Chenhao Guo, Zhulun Yang, You Yang 0002, Qiong Liu 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2025 | MinCD-PnP: Learning 2D-3D Correspondences with Approximate Blind PnPabstractImage-to-point-cloud (I2P) registration is a fundamental problem in computer vision, focusing on establishing 2D-3D correspondences between an image and a point cloud. The differential perspective-n-point (PnP) has been widely used to supervise I2P registration networks by enforcing the projective constraints on 2D-3D correspondences. However, differential PnP is highly sensitive to noise and outliers in the predicted correspondences. This issue hinders the effectiveness of correspondence learning. Inspired by the robustness of blind PnP against noise and outliers in correspondences, we propose an approximated blind PnP based correspondence learning approach. To mitigate the high computational cost of blind PnP, we simplify blind PnP to an amenable task of minimizing Chamfer distance between learned 2D and 3D keypoints, called MinCD-PnP. To effectively solve MinCD-PnP, we design a lightweight multi-task learning module, named as MinCD-Net, which can be easily integrated into the existing I2P registration architectures. Extensive experiments on 7-Scenes, RGBD-V2, ScanNet, and self-collected datasets demonstrate that MinCD-Net outperforms state-of-the-art methods and achieves a higher inlier ratio (IR) and registration recall (RR) in both cross-scene and cross-dataset settings. Pei An, Jiaqi Yang 0002, Muyao Peng, You Yang 0002, Qiong Liu 0001, Liangliang Nan |
ICCV | 5 |
| 2025 | CWC-DNERF: Compact Dynamic Neural Radiance Field VIA Discrete Wavelet Transform And Learnable CodebooksabstractNeural radiance fields have significantly advanced dynamic scene reconstruction and novel view synthesis. However, relying on multiple implicit multi-layer perceptrons for reconstructing dynamic scenes is computationally expensive. Recent methods have alleviated this challenge by introducing explicit data structures, such as voxel grids and feature planes, but these significantly increase storage demands and complicate network transmission. We propose Cwc-DNeRF, a compact dynamic NeRF representation that leverages discrete wavelet transform (DWT) and learnable codebooks to achieve superior storage efficiency while maintaining competitive rendering quality compared to K-Planes. In Stage I, DWT and trainable masks are employed to optimize parameter efficiency, resulting in sparse space planes. In Stage II, learnable codebooks are introduced for the space-time planes to merge redundant spatio-temporal features further reducing the storage demand. Additionally, a data compression pipeline is applied to compress both sparse space plane parameters and codebooks. Experimental results on D-NeRF and DyNeRF datasets show that our method achieves state-of-the-art rendering quality within a 10MB storage budget while retaining the benefits of explicit feature planes. Yaojian Xu, Qiudan Zhang, Longhao Zou, Qiong Liu 0001, Xu Wang 0006 |
ICIP | 5 |
| 2025 | Top-I2P: Explore Open-Domain Image-to-Point Cloud Registration Using Topology RelationshipabstractImage-to-point cloud (I2P) registration is a fundamental task in computer vision, which aims to align pixels in 2D images with corresponding points in 3D point clouds. While deep learning based methods dominate this field, they often fail to generalize to the open domain. In this paper, we address open-domain I2P registration from the topology relationships perspective. Firstly, we find that topology relationships reflect sparse connections between pixels and points, which shows the significant potential in enhancing cross-modality feature interaction in the open domain. Building on this insight, we develop an I2P registration framework using topology relationships. After that, to construct and leverage the topology relationships between the heterogeneous 2D and 3D spaces, we design a registration network, Top-I2P, with correction-based topology reasoning and fast topology feature interaction modules. Extensive experiments on 7-Scenes, RGBD-V2, ScanNet, and self-collected I2P datasets demonstrate that Top-I2P achieves superior registration performance in open-domain scenarios. Pei An, Jiaqi Yang 0002, Muyao Peng, You Yang 0002, Qiong Liu 0001, Jie Ma 0003, Liangliang Nan |
IJCAI | 5 |
| 2025 | Enhance Image-to-Point-Cloud Registration with Beltrami Flow
Pei An, You Yang 0002, Jiaqi Yang 0002, Muyao Peng, Qiong Liu 0001, Liangliang Nan |
Int. J. Comput. Vis. | 5 |
| 2025 | Adaptive CLIP for open-domain 3D model retrieval
Dan Song 0006, Zekai Qiang, Chumeng Zhang, Lanjun Wang, Qiong Liu 0001, You Yang 0002, Anan Liu |
Inf. Process. Manag. | 5 |
| 2025 | Corner selection and dual network blender for efficient view synthesis in outdoor scenes
Mohannad A. M. Al-Ja'afari, Firas Abedi, You Yang 0002, Qiong Liu 0001 |
Pattern Recognit. | 4 |
| 2025 | Unsupervised learning non-uniform face enhancement under physics-guided model of illumination decoupling
Zhongyuan Wang 0001, Qiong Liu 0001, You Yang 0002, Zhenyu Shu |
Pattern Recognit. | 4 |
| 2025 | Real-time small object detection using adaptive weighted fusion of efficient positional features
Qiong Liu 0001, You Yang 0002 |
Pattern Recognit. | 3 |
| 2025 | End-to-End Deep Video Compression Based on Hierarchical Temporal Context LearningabstractEmerging learning-based video compression suffers from error propagation in long group of pictures (GOP), yielding limited coding performance. To address this problem, a novel end-to-end Deep Video Compression method based on Hierarchical Temporal Context Learning (DVCH) is proposed in this paper. DVCH aims to fully exploit temporal contexts and suppress error propagation for better coding performance. It first divides video frames into several hierarchies with different compression qualities. The frames in lower hierarchies have high compression quality, and serve as reference frames. To mine high-quality reference information, we propose a Hierarchical Temporal Context Learning (HTCL) network as the fundamental module of our DVCH. The informative temporal context features from hierarchical prediction structure can be extracted by the network. Motion vectors (MVs) between the to-be-coded frame and its reference frames are estimated by the MV Learning module and used to align the extracted contexts. The contexts are fed into Context Coding module to generate the prediction of the decoded frame. Moreover, a multi-stage training strategy is developed to solve the imbalanced training challenge. Experimental results demonstrate that the proposed DVCH exceeds x264 and other end-to-end video compression methods, regardless of objective, subjective, error propagation suppression, GOP sizes, and sequence length evaluations. As much as 49.27% bitrate savings and 2.52 dB PSNR gains can be achieved in large GOP. Kejun Wu, You Yang 0002, Qiong Liu 0001, Xiao-Ping Zhang 0002 |
IEEE Trans. Multim. | 4 |
| 2025 | A Serial Perspective on Photometric Stereo of Filtering and Serializing Spatial InformationabstractIn this paper, we introduce a novel method of Filtering and Serializing Spatial Information to tackle uncalibrated photometric stereo tasks, termed FSSI-PS. Photometric stereo aims to recover surface normals from images with varying lighting and is crucial for tasks like 3D reconstruction and defect detection. Current methods in complex surface reconstruction are costly and inaccurate due to redundant feature representations from GCN or Transformer modules, caused by the weak global information extraction capability of GCNs or the large computational cost of Transformers. Furthermore, the trainset's lack of richness in texture complexity makes reconstruction more difficult. We address these issues by optimizing feature maps and dataset richness through serializing and filtering. First, we use Mamba-RNN to optimize feature representation by directly fusing feature maps, which reduces redundancy and uses minimal computational resources. Specifically, we treat input spatial information as a sequence and serialize it by sorting. Furthermore, we introduce the Mean Angular Variation metric to assess reconstruction difficulty by measuring texture complexity. It classifies PS-Sculpture and PS-Blobby into three categories: Difficult, Normal, and Simple. We use this to construct DNS-S+B, a photometric stereo training set with rich complexity levels. Our method is compared with state-of-the-art methods on the DiLiGenT and LUCES benchmarks to highlight effectiveness. Minzhe Xu, You Yang 0002, Yinqiang Zheng, Qiong Liu 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | Multi-View Multi-Focus Image Fusion: A Novel Benchmark Dataset and MethodabstractMulti-focus image fusion fuses multiple images focused on different depths to generate a clear image covering the whole scene. However, the existing multi-focus image fusion methods do not consider the movement of the camera or objects in actual shooting. To address this, we propose an end-to-end deep learning network to generate the all-in-focus image from multi-view multi-focus images. Specifically, our method first warps the multi-view multi-focus images to a unified camera view by homography transformation matrices, and measures the defocus degree of co-located image patches through a focus information evaluation mechanism. Finally, our fusion network applies an adaptive fusion scheme to fuse the detected image patches into a clear image. For testing our fusion network, a multi-view multi-focus image benchmark dataset (MVMFI) is constructed with more than 1000 image sequences. Experiments results demonstrate that our method outperforms the state-of-the-art methods both qualitatively and quantitatively. MVMFI dataset is available at https://github.com/North-Li/MVMFI. Zhilong Li, Kejun Wu, Junhao Liu 0001, Qiong Liu 0001, You Yang 0002 |
ICIP | 4 |
| 2024 | Deep video compression based on Long-range Temporal Context Learning
Kejun Wu, You Yang 0002, Qiong Liu 0001 |
Comput. Vis. Image Underst. | 4 |
| 2024 | OL-Reg: Registration of Image and Sparse LiDAR Point Cloud With Object-Level Dense CorrespondencesabstractImage and point cloud registration (2D-3D registration) is an essential prerequisite for multi-modal feature fusion. However, due to the significant feature difference of point cloud and image, it is challenging to establish 2D-3D correspondences. Targeting for the background of autonomous driving, we propose 2D-3D registration method with object-level correspondence (OL-Reg) in this paper. Object-level correspondence consists of object bounding box and object contour in 2D image and 3D space. The first step is to match 2D-3D objects. Due to sensor pose and field of view (FoV) difference, object shape and occlusion is different in image and point cloud, causing the difficulty of object matching. To solve this issue, we represent object as 3D bounding box, and design 2D-3D object matching with 3D box projection (Box-Proj) constraint. It aligns object 3D bounding box in image and point cloud. After that, the next step is to build 2D-3D correspondence from the matched objects. To extract correspondence from object with irregular shape, we notice the distance constraint of object surface and rays back-projected from object contour, and present projection based iterative closest point (Proj-ICP). Towards the stability of Proj-ICP, object-level regularization term is designed. Experiment is conducted in KITTI object and odometry dataset. With the pre-trained 3D object detector, results suggest that OL-Reg has the better performance than current approaches in tasks of re-localization and extrinsic calibration. And source code will be released soon1. Pei An, Xuzhong Hu, Junfeng Ding, Jun Zhang 0062, Jie Ma 0003, You Yang 0002, Qiong Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | Survey of Extrinsic Calibration on LiDAR-Camera System for Intelligent Vehicle: Challenges, Approaches, and TrendsabstractA system with light detection and ranging (LiDAR) and camera (named as LiDAR-camera system) plays the essential role in intelligent vehicle (IV), for it provides 3D spatial and 2D texture features for 3D scene understanding. To leverage LiDAR point cloud and image, extrinsic calibration is a crucial technique, for it can align 2D pixel and 3D point in the pixel-level accuracy. With the rapid development of IV, calibration demand is shifted from offline to online, from the specific scenes to the open scenes. It brings new challenge to the calibration task. Although numbers of approaches have been proposed in the last decade, there lacks an in-depth summary about this topic. Thus, we conduct a survey of extrinsic calibration. Theoretically, the key of calibration is to build correspondence from LiDAR point cloud and optical image. From the viewpoint of correspondence, we attempt to divide the mainstream approaches into explicit and implicit correspondence based methods. After that, we summarize both the strength and weakness of the current works, provide the methods comparison, and list the open-source implementations. Finally, we analyze the tendency of calibration approach, discuss the remained problems in this field. We believe that this survey benefits to the community of autonomous driving. Pei An, Junfeng Ding, Siwen Quan, Jiaqi Yang 0002, You Yang 0002, Qiong Liu 0001, Jie Ma 0003 |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2024 | SP-Det: Leveraging Saliency Prediction for Voxel-Based 3D Object Detection in Sparse Point CloudabstractVoxel is one of the common structural representation of 3D point cloud. Due to the sparsity of point cloud generated by light detection and ranging (LiDAR), there is the extreme imbalance in the foreground and background voxels. It decreases the accuracy of 3D object detection, has the negative effect on intelligent driving safety. To overcome this problem, we present a saliency prediction based 3D object detector SP-Det in this article. Although foreground voxels have the sufficient feature of object, it is difficult to localize the foreground region from voxel space with the larger background region. We design an auxiliary learning task, saliency prediction (SP). It benefits 3D detector in identifying the foreground region. SP task uses label diffusion to alleviate the label imbalance. It reduces the learning difficulty of saliency in voxel and bird's eye view (BEV) spaces. After that, to strengthen feature interaction from the sparse foreground region, we design saliency fusion (SF) module to fuse the learning result in SP task. It utilizes voxel and BEV saliency maps as progressive attention to resist the redundant feature from background region. To aggregate more foreground feature inside 3D and BEV region of interest (RoI), we design hybrid grid maps based RoI pooling (Hybrid-RoI pooling). Experiments are conducted in STF dataset. The adverse weather enlarges the sparsity of LiDAR point cloud, increasing the difficulty of object detection. SP-Det identifies and leverages foreground region, and achieves the performance better than the current methods. Hence, we believe that SP-Det benefits to LiDAR based 3D scene understanding in the adverse weather. Pei An, Yucong Duan, Yuliang Huang, Jie Ma 0003, Yanfei Chen, Liheng Wang, You Yang 0002, Qiong Liu 0001 |
IEEE Trans. Multim. | 8 |
| 2024 | ESC-Net: Alleviating Triple Sparsity on 3D LiDAR Point Clouds for Extreme Sparse Scene Completionabstract3D scene completion (SC) has made progress in the last three years. From the application of mobile robot system, SC should support the downstream task (i.e. mapping or perception), instead of only predicting the completed scenes. However, as the low-cost few-beam LiDAR is widely applied in mobile robot, gap between SC and downstream tasks is large. To generate the high quality completion result, the bottleneck lies in the triple sparsity of input, ground truth (GT) occupancy, and GT foreground. To deal with the triple sparsity, we present an extreme sparse scene completion network (ESC-Net). At first, input sparsity hides most of the spatial information of the scene. A feature completion (FC) decoder is designed to mine the spatial feature using feature-level completion. Then, GT occupancy sparsity hinders representation learning of the real scene with continuous surfaces. A multi-view multi-task attention (MMA) loss is presented to recover the high-quality object boundaries via correcting occupancy and semantic labels of regions from 3D and bird's eye view (BEV) spaces. After that, GT foreground sparsity is the imbalance of foreground and background GT labels. It causes the inaccuracy of local 3D object completion. A combination network (ESC-Net-D) is presented to recover 3D structural details of both foreground and background. Experiment is conducted on KITTI and SemanticPOSS datasets. It shows that ESC-Net has the performance higher than current methods not only on completion task, but also on the downstream tasks (i.e. 3D registration, 3D object detection). Hence, we believe that ESC-Net benefits to the community of mobile robot. Source code is released soon. Pei An, Siwen Quan, Junfeng Ding, Jie Ma 0003, You Yang 0002, Qiong Liu 0001 |
IEEE Trans. Multim. | 7 |
| 2024 | Hierarchical Independent Coding Scheme for Varifocal Multiview Images Based on Angular-Focal Joint PredictionabstractVarifocal multiview (VFMV) images are dense views that focus on variable focal planes. Thus, VFMV images are highly redundant in the angular, spatial and focal dimensions. In this article, the redundancies of VFMV images are analyzed and represented by full parallaxes and focal inconsistency. To exploit these distinctive redundancies, we propose a hierarchical independent coding scheme based on angular-focal joint prediction. The scheme is constructed by hierarchical independent prediction structure (HIPS) and angular-focal joint prediction (AFJP). The HIPS separates all views into several independent subdivisions and assigns different hierarchies inside each subdivision, which enhances random access capability and scalability. The AFJP conducts motion estimation and focal approximation simultaneously to predict parallaxes and focal inconsistency. Therefore, the redundancies in the angular and focal dimensions can be exploited by the proposed coding scheme. We construct a VFMV dataset with 10 test sequences for different acquisition methods. The experimental results on these test sequences demonstrate that the proposed scheme outperforms all comparison schemes in objective quality, subjective quality and random access capability. Specifically, the proposed coding scheme achieves up to 2.661 dB PSNR gains and 52.817% bitrate savings compared with the HEVC random access benchmark scheme. Kejun Wu, You Yang 0002, Qiong Liu 0001, Gangyi Jiang, Xiao-Ping Zhang 0002 |
IEEE Trans. Multim. | 3 |
| 2024 | WaRENet: A Novel Urban Waterlogging Risk Evaluation NetworkabstractIn this article, we propose a novel urban waterlogging risk evaluation network (WaRENet) to evaluate the risk of waterlogging. The WaRENet distinguishes whether an urban image involves waterlogging by classification module, and estimates the waterlogging risk levels by multi-class reference objects detection module (MCROD). First, in the waterlogging scene classification, ResNet combined with Se-block is used to identify the waterlogging scene, and lightweight gradient-weighted class activation mapping (Grad-CAM) is also integrated to roughly locate overall waterlogging areas with low computational burden. Second, in the MCROD module, we detect reference objects, e.g., cars and persons in waterlogging scenes. The positional relationship between water depths and reference objects serves as risk indicators for accurately evaluating waterlogging risk. Specifically, we incorporate switchable atrous convolution (SAC) into YOLOv5 to solve occlusions and varying scales problems in complex waterlogging scenes. Moreover, we construct a large-scale urban waterlogging dataset called UrbanWaterloggingRiskDataset (UWRDataset) with 6,351 images for waterlogging scene classification and 3,217 images for reference objects detection. Experimental results on the dataset show that our WaRENet outperforms all comparison methods. The waterlogging scene classification module achieves accuracy of 95.99%. The MCROD module obtains mAP of 54.9%, while maintaining a high processing speed of 70.04 fps. Xiaoya Yu, Kejun Wu, You Yang 0002, Qiong Liu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Multi-scale Non-local Bidirectional Fusion for Video Super-Resolution
Qinglin Zhou, Qiong Liu 0001, Zongju Peng |
ICIG (5) | 2 |
| 2023 | RS-Aug: Improve 3D Object Detection on LiDAR With Realistic Simulator Based Data AugmentationabstractLight detection and ranging (LiDAR) is an essential sensor for three dimensional (3D) object detection via generating 3D point cloud of the surroundings, and it has been widely used in the various visual applications, especially autonomous driving. However, limited numbers of labeled LiDAR datasets brutally restrain the development of 3D object detector, and this situation breeds an urgent demand on data augmentation in this field. By far, most of the traditional methods reuse the labeled samples, while those unlabeled are hastily untaken. Motivated by this, we propose aRealisticSimulator based data augmentation (RS-Aug). It aims to construct augmented real scenes to enrich the diversity of training dataset. To train 3D object detector in a supervised learning way, the first step of RS-Aug is auto-annotation. Time-continuous LiDAR frames are used to construct the dense scene, which is beneficial to annotation and the subsequent rendering augmentation. However, 3D points with incorrect semantic labels are naturally gathered during multi-view reconstruction, causing the negative effect on auto-annotation. We propose an algorithm of cluster guided$k$-nearest neighbor (c-$k$NN). It emphasizes on de-nosing semantic labels of clustered points using distance and intensity constraints. Then, the next step of RS-Aug is rendering augmentation on the real scene. To enhance the rendering quality using collision and distance constraints with the less computation complexity, we propose a scheme of heuristic search (HS) based object insertion. It estimates the proper position of the inserted object from 2D bird’s eye view (BEV). Experiments demonstrate the de-noising accuracy of c-$k$NN, rendering quality of HS based object insertion, and improvement of RS-Aug on object detection. Pei An, Junxiong Liang, Jie Ma 0003, Yanfei Chen, Liheng Wang, You Yang 0002, Qiong Liu 0001 |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2023 | Focal Stack Image Compression Based on Basis-Quadtree RepresentationabstractIn this paper, we propose an efficient compression scheme for focal stack images (FoSIs) based on a new basis-quadtree representation. In the new basis-quadtree representation, FoSIs are initially reorganized as co-located block groups in the depth dimension. In each group, selective basis blocks and adaptive quadtree partition are optimized to predict the focused or defocused co-located blocks by intra-group approximation. By solving a joint optimization problem, FoSIs can be efficiently represented by the optimal basis blocks, corresponding quadtree partition and approximation parameters, which will be compressed separately. Then, these basis blocks are stitched into several new frames (basis frames) according to their original locations and partition modes. Basis frames are compressed by our designed encoder, where the intra-group approximation is embedded into the high efficiency video coding (HEVC) encoder. Thus, the redundancies of basis blocks can be further eliminated. Finally, the approximation parameters are refined to suppress the amplified errors caused by introduced compression blur after basis frame coding. The refined parameters are compressed losslessly and multiplexed with the bitstream of the basis frames to ensure the reconstruction quality of FoSIs. Experiments on 12 test sequences demonstrate that the proposed scheme can obtain higher coding performance than the state-of-the-art comparison schemes. Specifically, the proposed scheme achieves up to 5.23 dB PSNR gains and 71.59% bitrate savings over the HEVC baseline scheme on sequences I03 and I05, respectively. Kejun Wu, You Yang 0002, Qiong Liu 0001, Xiao-Ping Zhang 0002 |
IEEE Trans. Multim. | 3 |
| 2022 | Iterative enhancement scheme of synthesized color and depth images for immersive video systemabstractImmersive video allows viewers to freely switch the viewpoints. The intensity of realistic experience greatly relies on the quality of synthesized depth maps. However, there exist distorted regions due to inaccurate depth estimation or compression. Yongquan Su, Qiong Liu 0001, Kejun Wu, Gangyi Jiang, You Yang 0002 |
DCC | 2 |
| 2022 | Gaussian-Wiener Representation and Hierarchical Coding Scheme for Focal Stack ImagesabstractFocal stack images (FoSIs) are a set of 2D images that captured one scene with serial focal depths. The redundancies of FoSIs mainly come from gradual focused depths changes rather than motion of objects. Conventional coding schemes cannot fully exploit such redundancies, leading to coding inefficiency. In this paper, we propose a new Gaussian-Wiener representation to model the gradual focused depths changes among FoSIs. In the representation, image degradation-restoration relations are utilized to describe the focus-defocus changing characteristics of FoSIs. Based on this representation, we propose a new hierarchical coding scheme for fully exploiting the inter-frame redundancies of FoSIs. In the scheme, a Gaussian-Wiener representation based inter prediction (GWR-IP) is presented by embedding Gaussian convolution and Wiener deconvolution into normal video encoder. Block-wise focus-defocus changing of FoSIs can be predicted in bi-directional manner by solving optimization problem. For higher coding efficiency, a Gaussian-Wiener representation based hierarchical prediction structure (GWR-HPS) is also designed and applied in the coding scheme. The proposed coding scheme is performed on 10 test sequences, including 5 synthetic scenes and 5 realistic scenes. Experimental results show that proposed coding scheme can obtain 2.640 dB PSNR gains and 51.830% bitrate savings on average of all test sequences in Low Delay P configuration, 2.123 dB PSNR gains and 43.975% bitrate savings in Low Delay B configuration, and 1.044 dB PSNR gains and 26.078% bitrate savings in Random Access configuration. Particularly, it achieves up to 65.544% bit rate savings and 3.901 dB PSNR increments for test sequence I09 in Low Delay P configuration. Furthermore, ablation test demonstrates that Gaussian representation contributes more on coding performance than Wiener representation and GWR-HPS. Kejun Wu, You Yang 0002, Qiong Liu 0001, Xiao-Ping Zhang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Decoupled Dynamic Filter NetworksabstractConvolution is one of the basic building blocks of CNN architectures. Despite its common use, standard convolution has two main shortcomings: Content-agnostic and Computation-heavy. Dynamic filters are content-adaptive, while further increasing the computational overhead. Depth-wise convolution is a lightweight variant, but it usually leads to a drop in CNN performance or requires a larger number of channels. In this work, we propose the Decoupled Dynamic Filter (DDF) that can simultaneously tackle both of these shortcomings. Inspired by recent advances in attention, DDF decouples a depth-wise dynamic filter into spatial and channel dynamic filters. This decomposition considerably reduces the number of parameters and limits computational costs to the same level as depth-wise convolution. Meanwhile, we observe a significant boost in performance when replacing standard convolution with DDF in classification networks. ResNet50 / 101 get improved by 1.9% and 1.3% on the top-1 accuracy, while their computational costs are reduced by nearly half. Experiments on the detection and joint upsampling networks also demonstrate the superior performance of the DDF upsampling variant (DDF-Up) in comparison with standard convolution and specialized content-adaptive layers. The project page with code is available1. Jingkai Zhou, Varun Jampani, Zhixiong Pi, Qiong Liu 0001, Ming-Hsuan Yang 0001 |
CVPR | 4 |
| 2021 | Multi-Objective Network Congestion Control via Constrained Reinforcement LearningabstractTraditional congestion control algorithms rely on various model-based methods to improve the end-to-end (E2E) performance of packet transmission. The resulting decisions quickly become less effective amid the dynamics of network conditions. In order to perform congestion control adaptively, reinforcement learning (RL) can be adopted to continuously learn the optimal strategy from the network environment. Oftentimes, the reward of such a learning problem is a weighted sum of multiple E2E performance metrics, such as throughput, delay, and fairness. Unfortunately, those weights can be only manually tuned based on extensive experiments. To address this issue, in this paper, we design a constrained RL algorithm for congestion control named CRL-CC to adaptively tune those weights, with the objective of effectively improving the overall E2E packet transmission performance. In particular, the multi-objective optimization problem is firstly formulated as a constrained optimization problem. Then, the Lagrangian relaxation method is leveraged to transform the constrained optimization problem into a single-objective optimization problem, which is solved by designing a multi-objective reward function with Lagrangian multipliers. Extensive experiments based on OpenAI-Gym show that the proposed CRL-CC algorithm can achieve higher overall performance in various network conditions. In particular, the CRL-CC algorithm outperforms the benchmark algorithm on Pantheon by 21.7%, 27.4%, and 5.3% in throughput, delay, and fairness, respectively. Qiong Liu 0001, Peng Yang 0004, Feng Lyu 0001, Ning Zhang 0007, Li Yu 0003 |
GLOBECOM | 1 |
| 2021 | SA-GNN: Stereo Attention and Graph Neural Network for Stereo Image Super-Resolution
Qiong Liu 0001, You Yang 0002 |
ICIG (3) | 2 |
| 2021 | Dedark+Detection: A Hybrid Scheme for Object Detection under Low-light SurveillanceabstractObject detection under low-light surveillance is a crucial problem that less efforts have been made on it. In this paper, we proposed a hybrid method that jointly use enhancement and object detection for the above challenge, namely Dedark+Detection. In this method, the low-light surveillance video is processed by the proposed de-dark method, and the video can thus be converted to appearance under normal lighting condition. This enhancement bring more benefits to the subsequent stage of object detection. After that, an object detection network is trained on the enhanced dataset for practical applications under low-light surveillance. Experiments are performed on 18 low-light surveillance video test sequences, and superior performance can be found when comparing to state-of-the-arts. Xiaolei Luo, Sen Xiang, Yingfeng Wang, Qiong Liu 0001, You Yang 0002, Kejun Wu |
MMAsia | 4 |
| 2021 | Divide and conquer: Ill-light image enhancement via hybrid deep network
You Yang 0002, Qiong Liu 0001, Zahid Hussain Qaisar |
Expert Syst. Appl. | 3 |
| 2021 | A ghostfree contrast enhancement method for multiview images without depth information
You Yang 0002, Qiong Liu 0001, Zahid Hussain Qaisar |
J. Vis. Commun. Image Represent. | 3 |
| 2021 | I Understand You: Blind 3D Human Attention Inference From the Perspective of Third-PersonabstractInferring object-wise human attention in 3D space from the third-person perspective (e.g., a camera) is crucial to many visual tasks and applications, including human-robot collaboration, unmanned vehicle driving, etc. Challenges arise from classical human attention when human eyes are not visible to cameras, gaze point is outside the field of vision, or the gazed object is occluded by others in the 3D space. In this case, blind 3D human attention inference brings a new paradigm to the community. In this paper, we address these challenges by proposing a scene-behavior associated mechanism, in which both 3D scene and temporal behavior of human are adopted to infer object-wise human attention and its transition. Specifically, point cloud is reconstructed and used for the spatial representation of 3D scene, which is beneficial to handle the blind problem from the perspective of a camera. Based on this, in order to address the blind human attention inference without eye information, we propose a Sequential Skeleton Based Attention Network (S2BAN) for behavior-based attention modeling. As is embedded in the scene-behavior associated mechanism, the proposed S2BAN is built under the temporal architecture of Long-Short-Term-Memory (LSTM). Our network employs human skeleton as behavior representation, and maps it to the attention direction frame by frame, which makes attention inference a temporal-correlated issue. With the help of S2BAN, 3D gaze spot and further the attended objects can be obtained frame by frame via intersection and segmentation on the previously reconstructed point cloud. Finally, we conduct experiments from various aspects to verify the object-wise attention localization accuracy, the angular error of attention direction calculation, as well as the subjective results. The experimental results show that the proposed outperforms other competitors. You Yang 0002, Qiong Liu 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | Make Full Use of Priors: Cross-View Optimized Filter for Multi-View Depth EnhancementabstractMulti-view video plus depth (MVD) is the promising and widely adopted data representation for future 3D visual applications and interactive media. However, compression distortions on depth videos impede the development of such applications, and filters are crucially needed for the quality enhancement at the terminal side. Cross-view priors can intuitively be involved in filter design, but these priors are also distorted in compression and thus the contribution of them can hardly be considered in previous research. In this article, we propose a cross-view optimized filter for depth map quality enhancement by making full use of inner- and cross-view priors. We dedicate to evaluate the contributions of distorted cross-view priors in filtering the current view of depth, and then both inner- and cross-view priors can be involved in the filter design. Thus, distortions of cross-view priors are not barriers again as before. For the purpose of that, mutual information guided cross-view consistency is designed to evaluate the contributions of cross-view priors from compression distortions of MVD. After that, under the framework of global optimization, both inner- and cross-view priors are modeled and taken to minimize the designed energy function where both data accuracy and spatial smoothness are modeled. The experimental results show that the proposed model outperforms state-of-the-art methods, where 3.289 dB and 0.0407 average gains on peak signal-to-noise ratio and structural similarity metrics can be obtained, respectively. For the subjective evaluations, object details and structure information are recovered in the compressed depth video. We also verify our method via several practical applications, including virtual view synthesis for smooth interaction and point cloud for 3D modeling for accuracy evaluation. In these verifications, the ringing and malposition artifacts on object contours are properly handled for interactive video, and discontinuous object surfaces are restored for 3D modeling. All of these results suggest that compression distortions in MVD can be properly filtered by the proposed model, which provides a promising solution for future bandwidth constrained 3D and interactive visual applications. Qiong Liu 0001, You Yang 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2020 | Gaussian Guided Inter Prediction for Focal Stack Images CompressionabstractFocal stack is an intermediate data representation obtained by projecting 4D light field (LF) in z-dimension. This kind of representation is fundamental for future interactive and immersive visual applications. However, focal stack images are a series of samples focused at varying depths of static scenes, which yields considerable redundancy among them. In this paper, we propose a Gaussian guided inter prediction model to eliminate the visual redundancy. In our work, the effect of the varying focus distance on plenoptic imaging system is characterized by the point spread function (PSF), and we propose a simplified Gaussian-like PSF to fit this characteristic according to the features of focal stack images. After that, Gaussian guided motion estimation and motion compensation are both implemented in this model. Experimental results show that our model can achieve smaller residual distribution. There are 10.33% bit rate saving and 0.397 dB PSNR gain on average in three configurations compared with HEVC anchor. Particularly, it brings about up to 16.60% bit rate saving with 0.649 dB PSNR increment in Low Delay P configuration. Kejun Wu, Qiong Liu 0001, Yaguang Yin, You Yang 0002 |
DCC | 2 |
| 2020 | Similarity Graph Convolutional Construction Network for Interactive Action Recognition
Qiong Liu 0001, You Yang 0002 |
MMM (2) | 2 |
| 2020 | Deformed Phase Prediction Using SVM for Structured Light Depth Generation
Sen Xiang, Qiong Liu 0001, Huiping Deng, Li Yu 0003 |
MMM (2) | 2 |
| 2020 | Depth map artefacts reduction: a reviewabstractDepth maps are crucial for many visual applications, where they represent the positioning information of the objects in a three‐dimensional scene. Usually, depth maps can be acquired via various devices, including Time of Flight, Kinect or light field camera, in practical applications. However, a brutal truth is that both intrinsic and extrinsic artefacts can be found in these depth maps which limits the prosperity of three‐dimensional visual applications. In this study, the authors survey the depth map artefacts reduction methods proposed in the literature, from mono‐ to multi‐view, via spatial to temporal dimension, in local to global manner, with signal processing to learning‐based methods. They also compare the state‐of‐the‐arts via different metrics to show their potentials in future visual applications. Mostafa Mahmoud Ibrahim, Qiong Liu 0001, Ehsan Adeli-Mosabbeb, You Yang 0002 |
IET Image Process. | 2 |
| 2020 | Adaptive colour-guided non-local means algorithm for compound noise reduction of depth mapsabstractDepth maps are used to describe object positioning information in three‐dimensional (3D) space, and they are crucial for RGB‐D data representation, which is useful for numerous interactive visual applications. In practice, depth maps are often contaminated by compound noise, including intrinsic noise and missing regions owing to active illumination shadows. As existing noise models cannot describe the above‐mentioned compound noise effectively, the subsequent filter design is a challenging task. In this study, an adaptive colour‐guided non‐local mean (NLM) filter is proposed to address such compound noise. First, the authors classify the depth map into hole and non‐hole pixels. Then, the proposed filter is designed on the basis of the NLM framework, where the colour image is used as a guide prior for hole‐artifact removal. Finally, the authors use a shock filter to effectively address the non‐regularisation of the restored depth map edges and remove the remaining noise. Experiments show that the proposed filter qualitatively and quantitatively outperforms existing colour‐guided and unguided filters. Moreover, the authors verify the superiority of the proposed filter through virtual view synthesis and 3D scene reconstruction applications. Mostafa Mahmoud Ibrahim, Qiong Liu 0001, You Yang 0002 |
IET Image Process. | 2 |
| 2020 | Novel calibration method for camera array in spherical arrangement
Pei An, Qiong Liu 0001, Firas Abedi, You Yang 0002 |
Signal Process. Image Commun. | 2 |
| 2020 | MV-GNN: Multi-View Graph Neural Network for Compression Artifacts ReductionabstractInevitable compression artifacts in multi-view video (MVV) can clearly degrade the quality of experience in many interaction-oriented 3D visual applications. Under the framework of asymmetric coding, low-quality images can be enhanced with high-quality images from the neighboring viewpoints considering the similarity among different views. However, compression artifacts and warping error cause different cross-view quality gaps for various sequences, and thus the contribution of cross-view priors can hardly be located and considered in previous works. In this paper, we propose a multi-view graph neural network (MV-GNN) to reduce compression artifacts in multi-view compressed images. We dedicate to design a fusion mechanism which can exploit contributions from neighboring viewpoints and meanwhile suppress the misleading information. In our method, a GNN-based fusion mechanism is designed to fuse the cross-view information under the aggregation and update mechanism of GNN. Experiments show that 1.672 dB and 0.0242 average gains on PSNR and SSIM metrics can be obtained, respectively. For the subjective evaluations, blocking effect in the compressed images are clearly suppressed and the damaged object boundary are better recovered. The experimental results demonstrate that our MV-GNN outperforms the state-of-the-art methods. Qiong Liu 0001, You Yang 0002 |
IEEE Trans. Image Process. | 2 |
| 2019 | Multi-view Multi-modality Priors Residual Network of Depth Video Enhancement for Bandwidth Limited Asymmetric Coding FrameworkabstractAsymmetric coding methodology for multi-view video plus depth is a promising technique for future three-dimensional and multi-view driven visual applications for its superior coding performance in bandwidth limited conditions. Since the depth video suffers from asymmetric distortions corresponding to viewpoint, it's a challenge in smooth and quality consistent content based interaction. To solve this challenge, we propose a residual learning framework to enhance the quality of compression distorted multi-view depth video. In this work, we exploit the correlation between viewpoints to restore the target viewpoint depth maps by using multi-modality priors, which are depth maps from adjacent viewpoints with better quality and color frames in the same viewpoint. A residual network is designed to fully exploit the contribution from these priors. Experimental results show the superiority of our framework in the quality improvement on both decoded depth video and synthesized virtual viewpoint images. Qiong Liu 0001, You Yang 0002 |
DCC | 2 |
| 2019 | A Global Co-Saliency Guided Bit Allocation for Light Field Image CompressionabstractLight field is the most prospective technology for interactive and immersive visual applications. and light field image is an intermediate data format that demands a large amount of storage space and higher transmission bandwidth. Therefore, compression of light field images is highly desired for further applications. In this paper, we propose a co-saliency guided bit allocation scheme with constraints of consistency among sub-aperture images. Firstly, saliency is jointly detected on color and depth images of sub-aperture by improving our previous model. The obtained pixel-wise co-saliency map is converted into block-wise via K-means clustering. In this way, the saliency weight of each coding tree unit (CTU) can be calculated. Then, target bits of each CTU are initially determined by the weight of each block. The allocation is adjusted dynamically under the guidance of co-saliency map and the image texture complexity. The experimental results show that BD-PSNR of 0.384 dB can be achieved for the salient region at the cost of less than 0.107 dB decrease for the whole image compared to HTM anchor. Moreover, subjective quality of proposed scheme outperforms the anchor for the salient region, and there is no noticeable distortion for non-salient region. Kejun Wu, Zongbang Liao, Qiong Liu 0001, Yaguang Yin, You Yang 0002 |
DCC | 3 |
| 2019 | Optimized Color-guided Filter for Depth Image DenoisingabstractColor Guided Depth image denoising often suffers from the texture coping from the color image as well as the blurry effect at the depth discontinuities. Motivated by this, we propose an optimized color-guided filter for depth image denoising from different types of noises. This is a new framework that helps to mitigate the texture coping and enhance the depth discontinuities, especially in heavy noises. This framework consists of two parts namely depth driven color flattening model and patch synthesis-based Markov random field model. The first part which is a prepare step for the second part is used to mitigate the texture coping problem that faces all color guided methods. This first model consists of a modified joint bilateral filter which is used to mitigate the noise from the noisy depth image and an iterative guided bilateral filter that is proposed to flatten the colors in the color image for mitigating the texture coping problem. Based on the first part, Markov random field with an optimization technique is used for mitigating the blurry effect. Experiments indicate that our method outperforms counterpart filters with guided and non-guided manners in terms of a variety of evaluation metrics. Mostafa Mahmoud Ibrahim, Qiong Liu 0001 |
ICASSP | 2 |
| 2019 | Multi-view high dynamic range reconstruction via gain estimationabstractMulti-view high dynamic range reconstruction is a challenging problem, especially if the multi-view low dynamic range images are obtained from cameras arranged sparsely with limited shared view of vision among them. In this paper, we address the above challenge in addition to the back-lighting problem. We first enclose the geometry characteristic of the scene to rectify the outlier feature points. Consequently, an exposure gain is calculated according to those rectified features. After that, we extend the dynamic range for the multi-view low dynamic range images based on the estimated gain, then, generate a final high dynamic range image per view. Experimental results demonstrate superior performance for the proposed method over state-of-the-art methods in both objective and subject comparisons. These results suggest that our method is suitable to improve the visual quality of multi-view low dynamic range images captured in low back-lighting conditions via commercial cameras sparsely located among each other. Firas Abedi, Qiong Liu 0001, You Yang 0002 |
VCIP | 2 |
| 2019 | Adaptive-BBR: Fine-Grained Congestion Control with Improved Fairness and Low LatencyabstractTraditional loss-based congestion control protocols interpret packet loss as network congestion. Recently, Google proposed BBR, which is a congestion-based congestion control protocol. It employs delivery rate as the knob for congestion control, which achieves higher throughput and lower latency. Interestingly, BBR is found to have a preference for longer round-trip time (RTT) flows, which enjoy higher bandwidth ratio compared to flows with shorter RTT. To address this fairness issue, we proposed Adaptive-BBR, which creatively uses adaptive pacing gain to adjust the sending rate. The objective is that, via the proposed fine-grained adaptive mechanism, flows with different RTTs share similar portion of bottleneck bandwidth. Simulation results show that Adaptive-BBR can improve fairness by at least 47.8 %, and reduce average queuing delay by up to 93.3%, compared with that of BBR. Peng Yang 0004, Chaozhun Wen, Qiong Liu 0001, Jingjing Luo, Li Yu 0003 |
WCNC | 4 |
| 2019 | Cross-View Multi-Lateral Filter for Compressed Multi-View Depth VideoabstractMulti-view depth is crucial for describing positioning information in 3D space for virtual reality, free viewpoint video, and other interaction- and remote-oriented applications. However, in cases of lossy compression for bandwidth limited remote applications, the quality of multi-view depth video suffers from quantization errors, leading to the generation of obvious artifacts in consequent virtual view rendering during interactions. Considerable efforts must be made to properly address these artifacts. In this paper, we propose a cross-view multi-lateral filtering scheme to improve the quality of compressed depth maps/videos within the framework of asymmetric multi-view video with depth compression. Through this scheme, a distorted depth map is enhanced via non-local candidates selected from current and neighboring viewpoints of different time-slots. Specifically, these candidates are clustered into a macro super pixel denoting the physical and semantic cross-relationships of the cross-view, spatial and temporal priors. The experimental results show that gains from static depth maps and dynamic depth videos can be obtained from PSNR and SSIM metrics, respectively. In subjective evaluations, even object contours are recovered from a compressed depth video. We also verify our method via several practical applications. For these verifications, artifacts on object contours are properly managed for the development of interactive video and discontinuous object surfaces are restored for 3D modeling. Our results suggest that the proposed filter outperforms state-of-the-art filters and is suitable for use in multi-view color plus depth-based interaction- and remote-oriented applications. You Yang 0002, Qiong Liu 0001, Zhen Liu 0012 |
IEEE Trans. Image Process. | 2 |
| 2019 | A Two-Stage Clustering Based 3D Visual Saliency Model for Dynamic ScenariosabstractThree-dimensional (3D) visual saliency is fundamental for vision-guided applications such as human-computer interaction in virtual reality, image quality assessment, object tracking, and event retrieval. Classical models for 3D visual saliency can draw an appropriate saliency map when the quality of the required depth maps or auxiliary cues is high enough. However, the depth map is usually impaired with artifacts (such as holes or noise) from faults in stereo matching or multipaths in range sensors. In these cases, challenges arise in those 3D visual saliency models because the core preliminary processes, such as the detection of low-level visual features, may fail. To solve this problem, we proposed a two-stage clustering-based 3D visual saliency model for human visual fixation prediction in dynamic scenarios. In this model, a two-stage clustering scheme is designed to handle the negative influence of impaired depth videos. With the help of this scheme, representative cues are selected for saliency modeling. After that, multimodal saliency maps are obtained from depth, color, and 3D motion cues. Finally, a cross-Bayesian model is designed for the pooling of multimodal saliency maps. The experimental results demonstrate that the proposed 3D saliency model based on two-stage clustering outperforms other state-of-the-art models on a variety of metrics. Furthermore, the consistency and robustness of our model are also verified. You Yang 0002, Pian Li, Qiong Liu 0001 |
IEEE Trans. Multim. | 4 |
| 2018 | Graph-Based Saliency Fusion with Superpixel-Level Belief Propagation for 3D Fixation PredictionabstractIn recent years, many 3D visual attention models (VAMs) have been proposed with diverse fusion methods, of which the main challenge lies in the inconsistence, or even conflicts of different saliency maps. To address the challenge, we propose a graph-based fusion method with superpixel-level belief propagation for 3D fixation prediction on stereoscopic video, which models the aggregation as a global optimization issue. After extracting multi-modality saliency maps, the fusion step is based on the graph constructed at superpixel level, and we design for the graph an energy function considering multi-modality constraints, which is minimized using the belief propagation algorithm. The experimental results on two databases demonstrate that the proposed model achieves competitive performance. Qiong Liu 0001, You Yang 0002 |
ICIP | 2 |
| 2017 | Illumination Attributes Coding for Virtual Reality Broadcasting SystemabstractIn this paper, we propose a method of illumination attribute coding method for virtual reality broadcasting system. As for the virtual reality content, it is captured with local illumination variations. Our method is motivated by the Phong illumination model, and illumination attribute is extracted from images and then an illumination reference is synthesized with higher correlation to the current image. You Yang 0002, Qiong Liu 0001 |
DCC | 2 |
| 2017 | A robust 3D visual saliency computation model for human fixation prediction of stereoscopic videosabstract3D saliency have been gaining an increasing amount of attention because of the emergence of 3D contents and applications. Most of existing works on 3D saliency are based on the assumption of fine quality of depth maps. However, depth maps from stereo matching or range sensors are usually with holes and artifacts, which severely drop the performance of those 3D saliency models. In this paper, we propose a robust 3D saliency computation model for human fixation prediction. First, a cluster-contrast saliency prediction model is proposed for depth maps. The prediction is obtained with the centroid of the largest clusters of each depth super-pixel, and thus those bad effects originate from holes and other artifacts in depth are then eliminated. The cluster-contrast strategy is exploited both to depth texture and motions in 3D video. Finally, a Bayesian integration model is proposed for the multi-modality fusion between depth saliency and color saliency. The experimental results demonstrate that our saliency model has better performance of accuracy than other state-of-arts models. Qiong Liu 0001, You Yang 0002, Pian Li |
VCIP | 1 |
| 2017 | Special issue on dynamic depth field data driven learning, recognition and computation
Qiong Liu 0001, Yi Zhen |
Neurocomputing | 1 |
| 2016 | Cluster-based cross-view filtering for compressed multi-view depth mapsabstractIn the field of multi-view video coding, multi-view plus depth video is an important data format, but it always suffers from quantization errors, which result in obvious artifacts in consequent virtual view rendering. In this paper, we propose a cluster-based cross-view filtering (CBF) scheme for the enhancement of compressed depth maps. In this scheme, reconstructed depth information are mapped from cross-view, and this information is benefit to the proposed filter. Then in filtering one viewpoint depth map with candidate information that are selected from non-locally current and neighboring viewpoints. Specifically, in our scheme, candidates are clustered in 3D super-pixel wise rather than block wise due to cross-relationship among pixels in depth maps. The experimental results show that 2.0074 dB average gain can be obtained by our scheme, which suggests that the scheme outperforms than state-of-the-art and classical filters in filtering the reconstructed depth maps. Zhen Liu 0002, Qiong Liu 0001, You Yang 0002, Yuchi Liu, Gangyi Jiang, Mei Yu 0001 |
VCIP | 2 |
| 2016 | Reorganized DCT-based image representation for reduced reference stereoscopic image quality assessment
Lin Ma 0002, Xu Wang 0006, Qiong Liu 0001, King Ngi Ngan |
Neurocomputing | 3 |
| 2016 | User models of subjective image quality assessment on virtual viewpoint in free-viewpoint video system
You Yang 0002, Xu Wang 0006, Qiong Liu 0001, Mingliang Xu 0001 |
Multim. Tools Appl. | 3 |
| 2015 | A database of reflected irradiance field with depth for image based relightingabstractImage based relighting is an important application for light field, which represents the reflectance properties of the captured object/scene. Many systems were proposed without depth information of the captured object/scene. In this paper, we propose a lighting system named as Light Cube, and the system can capture color and depth information synchronously. This is important to consequent relighting process for both image based or model based methods. Utilizing the proposed lighting system, we calibrate the illumination and reflectance performance, and capture a database for reflected irradiance field. You Yang 0002, Qiong Liu 0001 |
ICIP | 2 |
| 2015 | Natural image statistics based 3D reduced reference image quality assessment in contourlet domain
Xu Wang 0006, Qiong Liu 0001, Ran Wang 0001, Zhuo Chen 0006 |
Neurocomputing | 2 |
| 2015 | A bundled-optimization model of multiview dense depth map synthesis for dynamic scene reconstruction
You Yang 0002, Xu Wang 0006, Qiong Liu 0001, Li Yu 0003 |
Inf. Sci. | 3 |
| 2015 | Dense depth image synthesis via energy minimization for three-dimensional video
You Yang 0002, Qiong Liu 0001, Hao Liu 0019, Li Yu 0003, Fanglin Wang |
Signal Process. | 2 |
| 2015 | Interfered depth map recovery with texture guidance for multiple structured light depth cameras
Sen Xiang, Li Yu 0003, You Yang 0002, Qiong Liu 0001, Jialiang Zhou |
Signal Process. Image Commun. | 4 |
| 2014 | Gradient-domain-based enhancement of multi-view depth video
Qiong Liu 0001, Zhengjun Zha, Yang Yang 0002 |
Inf. Sci. | 1 |
| 2014 | Spectral-Spatial Constraint Hyperspectral Image ClassificationabstractHyperspectral image classification has attracted extensive research efforts in the recent decade. The main difficulty lies in the few labeled samples versus the high dimensional features. To this end, it is a fundamental step to explore the relationship among different pixels in hyperspectral image classification, toward jointly handing both the lack of label and high dimensionality problems. In the hyperspectral images, the classification task can be benefited from the spatial layout information. In this paper, we propose a hyperspectral image classification method to address both the pixel spectral and spatial constraints, in which the relationship among pixels is formulated in a hypergraph structure. In the constructed hypergraph, each vertex denotes a pixel in the hyperspectral image. And the hyperedges are constructed from both the distance between pixels in the feature space and the spatial locations of pixels. More specifically, a feature-based hyperedge is generated by using distance among pixels, where each pixel is connected with its K nearest neighbors in the feature space. Second, a spatial-based hyperedge is generated to model the layout among pixels by linking where each pixel is linked with its spatial local neighbors. Both the learning on the combinational hypergraph is conducted by jointly investigating the image feature and the spatial layout of pixels to seek their joint optimal partitions. Experiments on four data sets are performed to evaluate the effectiveness and and efficiency of the proposed method. Comparisons to the state-of-the-art methods demonstrate the superiority of the proposed method in the hyperspectral image classification. Rongrong Ji, Yue Gao 0002, Richang Hong, Qiong Liu 0001, Dacheng Tao, Xuelong Li 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2013 | A gradient-based approach for interference cancelation in systems with multiple Kinect camerasabstractMicrosoft Kinect cameras provide a fast and convenient way to acquire depth information. However, the interference problem of multiple Kinect cameras dramatically degrades the depth quality. In this paper, we study the interference problem and propose an interference cancelation approach based on the statistical properties of depth maps. The gradient values are investigated and propagated from the interference-free region to the interfered region. The gradient values are obtained based on the statistic gradient features of the depth maps and the optimal solution of depth values is derived with a least error criterion. Experiment results demonstrate that our proposed method can eliminate interference efficiently, and lead to better qualities of depth maps and rendered virtual views. Sen Xiang, Li Yu 0003, Qiong Liu 0001, Zixiang Xiong |
ISCAS | 3 |
| 2013 | Stereotime: a wireless 2D and 3D switchable video communication systemabstractMobile 3D video communication, especially with 2D and 3D compatible, is a new paradigm for both video communication and 3D video processing. Current techniques face challenges in mobile devices when bundled constraints such as computation resource and compatibility should be considered. In this work, we present a wireless 2D and 3D switchable video communication to handle the previous challenges, and name it as Stereotime. The methods of Zig-Zag fast object segmentation, depth cues detection and merging, and texture-adaptive view generation are used for 3D scene reconstruction. We show the functionalities and compatibilities on 3D mobile devices in WiFi network environment. You Yang 0002, Qiong Liu 0001, Yue Gao 0002, Binbin Xiong, Li Yu 0003, Huan-Bo Luan, Rongrong Ji, Qi Tian 0001 |
ACM Multimedia | 2 |
| 2013 | Geographical Retagging
Liujuan Cao, Yue Gao 0002, Qiong Liu 0001, Rongrong Ji |
MMM (2) | 3 |
| 2013 | Quality Assessment on User Generated Image for Mobile Search Application
Qiong Liu 0001, You Yang 0002, Xu Wang 0006, Liujuan Cao |
MMM (2) | 1 |
| 2013 | Texture-adaptive hole-filling algorithm in raster-order for three-dimensional video applications
Qiong Liu 0001, You Yang 0002, Yue Gao 0002, Richang Hong |
Neurocomputing | 1 |
| 2013 | A Bayesian framework for dense depth estimation based on spatial-temporal correlation
Qiong Liu 0001, You Yang 0002, Yue Gao 0002, Rongrong Ji, Li Yu 0003 |
Neurocomputing | 1 |
| 2012 | Geometric mapping assisted multi-view depth video codingabstractMulti-view plus depth (MVD), as a video representation supporting view synthesis based on depth video, has attracted more and more attention for the free view video (FVV) application. It is a challenge to efficiently compress the multi-view depth data in MVD format. In this paper, we explore the geometric relationships in 3D space and propose a geometric mapping assisted (GMA) multi-view depth video coding algorithm. The proposed GMA utilizes the mapped depth image as a reference candidate during prediction. Furthermore, the inpainting method is employed to fill in the holes in mapped depth images. Experimental results demonstrate the gains of up to 2.45 dB for the depth coding, as well as better quality of synthesized views. Qiong Liu 0001, Yongbing Zhang 0002, Xiangyang Ji, Qionghai Dai |
ICASSP | 1 |
| 2012 | View-based 3D object retrieval by bipartite graph matchingabstractBipartite graph matching has been investigated in multiple view matching for 3D object retrieval. However, existing methods employ one-to-one vertex matching scheme while more than two views may share close semantic meanings in practice. In this work, we propose a bipartite graph matching method to measure the distance between two objects based on multiple views. In the proposed method, representative views are first selected by using view clustering for each object, and the corresponding weights are given based on the cluster results. A bipartite graph is constructed by using the two groups of representative views from two compared objects. To calculate the similarity between two objects, the bipartite graph is first partitioned to several subsets, and the views in the same sub-set are with high possibility to be with similar semantic meanings. The distances between two objects within individual subsets are then assembled through the graph to obtain the final similarity. Experimental results and comparison with the state-of-the-art methods demonstrate the effectiveness of the proposed algorithm. Yue Wen, Yue Gao 0002, Richang Hong, Huan-Bo Luan, Qiong Liu 0001, Jialie Shen 0001, Rongrong Ji |
ACM Multimedia | 5 |
| 2012 | Cross-View Down/Up-Sampling Method for Multiview Depth Video CodingabstractIn this letter, we propose a cross-view down/up-sampling (CDU) method for the framework of reduced resolution multiview depth video coding, which exploits cross-view information to assist the up-sampling at the decoder. In the down-sampling procedure of CDU, the odd-even interlaced extraction is employed to preserve more confident information of the original depth video with reduced resolution. In the decoder, the cross-view information is exploited for up-sampling the reconstructed depth video. An iterative interpolation process is proposed to eliminate the effect of compression distortion on this up-sampling. Experimental results demonstrate the gains of up to 3.88 dB for the proposed algorithm and better quality of synthesized views. Qiong Liu 0001, You Yang 0002, Rongrong Ji, Yue Gao 0002, Li Yu 0003 |
IEEE Signal Process. Lett. | 1 |
| 2011 | Cooperative cross-layer transmission for scalable video to resource-constrained receiverabstractFor resource-constrained wireless scenarios, this paper proposes a cooperative cross-layer scheme to improve the performance of scalable video transmission. At the application layer, the extraction parameters of scalable video coding are first optimized to maximize the visual quality and fulfill the resource constraints of decoding capacity and bandwidth. At the link layer, a priority-based cooperative scheduling strategy is further proposed to transmit the video packets of extracted substream. In this strategy, the relay link may substitute for the unreliable direct link if it has a relatively smaller packet error rate (PER). The packet priorities are determined by both the link PERs and the video layers properties. Simulations demonstrate the effectiveness of our proposed scheme. Hongjiang Xiao, Qiong Liu 0001, Qionghai Dai |
MUM | 2 |
| 2007 | A Novel Intra/Inter Mode Decision Algorithm for H.264/AVC Based on Spatio-temporal Correlation
Qiong Liu 0001, Shengfeng Ye, Ruimin Hu, Zhen Han 0002 |
MMM (1) | 1 |