Qiang Zhao 0005

dblp:00/2166-5 · DBLP profile ↗
← Back
35ranked-venue papers
8as first author
20since 2021 · last 2027
0000-0002-5867-3828ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 26 · 5 first-author · 13 since 2021Artificial intelligence and machine learning · 17 · 3 first-author · 12 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2027 MusicIllustrator: A multimodal music-to-image generation approach based on emotion analysis
Wenting Zhao 0003, Qiang Zhao 0005, Yan Zhao 0012
Expert Syst. Appl.3
2026 A Comprehensive Survey on the Research and Development of RGB-T Salient Object Detection
abstract
Salient object detection (SOD) aims to mimic human visual perception by identifying the most eye-catching objects within a scene, and has remained a popular research topic for many years. The introduction of thermal (T) images offers additional information for challenging scenarios such as those with low light and complex backgrounds, and thus enhance performance when combined with RGB images. In this paper, we have, to the best of our ability, conducted the first comprehensive survey of dual-modality RGB-T SOD. We summarize and categorize published RGB-T SOD models, emphasizing their characteristics and features. Important components of these models are classified and elaborated, such as feature extraction, modality fusion, and loss function design. Following this, we analyze existing RGB-T SOD datasets and evaluation metrics. We evaluate a selection of representative SOD models using unified protocols and statistical analysis. We also present comparative experiments on the impact of the choice of loss function and the usage of datasets on performance. Finally, we consider several key issues and potential solutions in RGB-T SOD research, revealing promising directions for future efforts. We hope this survey will offer an effective way to understand the current state of the technology and, more importantly, stimulate discussion within the community.
Hongfa Wen, Qiang Zhao 0005, Junbo Ma, Zongpeng Li, Shuai Wang 0003, Chenggang Yan 0001
Comput. Vis. Media3
2026 Feature distribution learning based on variance transfer and center shift for long-tailed classification
Chenxi Hong, Qiang Zhao 0005, Tao Tan 0002, Chenggang Yan 0001
Neurocomputing2
2026 AND-GS: Adaptive supervision of normal and depth in Gaussian splatting for accurate and efficient surface reconstruction
Xiang Le, Qiang Zhao 0005, Haofan Ren, Zhongtian Zheng, Tingyu Wang 0002, Jiyong Zhang 0001, Chenggang Yan 0001
Neurocomputing2
2026 Empirical Study on Fusion Strategy in RGB-T Salient Object Detection
abstract
In the research field of RGB-Thermal saliency object detection (RGB-T SOD), the effective exploitation of the complementary characteristics of the two modalities represents a major challenge for enhancing detection performance. Current fusion methodologies can be roughly classified into early fusion and middle fusion strategies, with prevalent techniques primarily encompassing concatenation, summation, and multiplication of the two modalities. To in depth assess the efficacy of these fusion strategies, we took an empirical investigation on them. Our findings demonstrate that the concatenation of middle features constitutes a more advantageous fusion strategy, yielding superior performance and demonstrating enhanced stability. Furthermore, observing the unique properties of thermal (T) images, we introduced gamma correction as a novel data augmentation methodology to RGB-T SOD. We subsequently evaluated the responses across varying correction parameter ranges, revealing that while the response to this data augmentation technique differs across various models, data augmentation is found to be effective in general. Building upon these findings, we proposed the Gamma Correction Network (GaCNet). Specifically, we also integrated image pyramid mechanism in a lightweight manner, which facilitates a more effective recovery of fine-grained image details. Significant improvement was achieved on commonly used RGB-T testing datasets, especially in VT821 dataset, manifesting the effectiveness of our method.
Shuai Wang 0003, Qiang Zhao 0005, Junbo Ma, Xichun Sheng, Yaoqi Sun, Hongfa Wen, Chenggang Yan 0001
IEEE Trans. Circuits Syst. Video Technol.3
2026 Anatomy-Aware MR-Imaging-Only Radiotherapy
abstract
The synthesis of computed tomography images can supplement electron density information and eliminate MR-CT image registration errors. Consequently, an increasing number of MR-to-CT image translation approaches are being proposed for MR-only radiotherapy planning. However, due to substantial anatomical differences between various regions, traditional approaches often require each model to undergo independent development and use. In this paper, we propose a unified model driven by prompts that dynamically adapt to the different anatomical regions and generates CT images with high structural consistency. Specifically, it utilizes a region-specific attention mechanism, including a region-aware vector and a dynamic gating factor, to achieve MRI-to-CT image translation for multiple anatomical regions. Qualitative and quantitative results on three datasets of anatomical parts demonstrate that our models generate clearer and more anatomically detailed CT images than other state-of-the-art translation models. The results of the dosimetric analysis also indicate that our proposed model generates images with dose distributions more closely aligned to those of the real CT images. Thus, the proposed model demonstrates promising potential for enabling MR-only radiotherapy across multiple anatomical regions. we have released the source code for our RSAM model. The repository is accessible to the public at: https://github.com/yhyumi123/RSAM.
Hao Yang 0026, Yue Sun 0001, Chi Kin Lam, Qiang Zhao 0005, Xiangyu Xiong, Kunyan Cai, Behdad Dashtbozorg, Chenggang Yan 0001, Tao Tan 0002
IEEE Trans. Image Process.6
2026 Enhance Panoramic Object Detection Using Planar Image Datasets
abstract
Panoramic images have been used in various applications because of their ability to provide comprehensive spatial information. However, the high cost of obtaining panoramic images and the complexity of annotation pose serious obstacles to enhancing the performance of panoramic tasks by restricting the size and quality of datasets. The use of annotated planar images to synthesize panoramic images proves to be an effective approach to narrow this gap. Prior synthesis methods introduced new distortions that lead to inconsistent object shapes before and after synthesis. Concurrently, since the angle of view of the planar image is much smaller than that of the panoramic image, there is an issue of missing spatial information in the synthesized panoramic image. To address these challenges, we introduce a novel approach for converting planar images into panoramic ones with reduced deformation. For annotating targets in synthetic images, we develop a new algorithm based on determining the minimum spherical area to calculate spherical bounding boxes that closely adhere to object boundaries, rather than relying on an estimated target center point which results in inevitable calculation errors as in previous studies. Subsequently, we propose a new method for effectively filling any blank areas in the synthetic panoramic images to compensate for the loss of precision caused by absence of spatial information. Additionally, with a focus on the distortion characteristics of panoramic images, an innovative data augmentation strategy is devised to further enhance the model's ability to recognize objects in different positions. These methods are conducive to a more effective utilization of the rich planar image datasets for panoramic object detection tasks. In the experiments, we generated two synthetic panoramic datasets based on COCO for training models. Experimental results demonstrate that training with these synthetic datasets significantly improves prediction accuracy, far surpasses the state-of-the-art methods, and helps to unleash the full potential of panoramic object detection models. Source code is available athttps://github.com/longlong-yu/official-panorama-coco.git.
Longlong Yu 0001, Qiang Zhao 0005, Xinyuan Liu 0003, Tingyu Wang 0002, Chenggang Yan 0001
IEEE Trans. Multim.2
2025 Region-aware Difference Distilling with Attribute-guided Contrastive Regularization for Change Captioning
abstract
Change captioning aims to describe the differences between two similar images using natural language, significantly aiding in understanding and monitoring changes. This challenging task requires a fine-grained understanding of subtle changes while resisting disturbances like viewpoint shifts and illumination variations. Existing methods often rely solely on global difference features and lack comprehensive alignment of linguistic and visual information, leading to overlooking fine-grained details and generating semantic hallucinated sentences. To address these limitations, we propose the region-aware difference distilling (RDD) network with attribute-guided contrastive regularization (ACR). The RDD uses global difference features to progressively distill regional difference features using learnable vectors, allowing for more precise identification of changed regions. The ACR enhances comprehensive alignment between linguistic and visual information by formulating Nouns-to-Objects (N2O) and Verbs-to-Actions (V2A) alignment losses to regularize the regional difference features. Promising results on three datasets demonstrate that our method outperforms the state-of-the-art change captioning methods.
Liang Li 0003, Qiang Zhao 0005, Hongkui Wang, Chenggang Yan 0001
AAAI4
2025 Lightweight three-stream encoder-decoder network for multi-modal salient object detection
Junzhe Lu 0002, Tingyu Wang 0002, Bin Wan, Qiang Zhao 0005, Shuai Wang 0003, Yaoqi Sun, Yang Zhou 0052, Chenggang Yan 0001
J. Vis. Commun. Image Represent.4
2025 P2FCN: Environment-Independent UAV-View Geo-Localization via Pixel-to-Feature Co-Enhancement
abstract
This paper investigates the challenges of UAV-view geo-localization under extreme environmental changes, where significant cross-domain style differences can lead to degradation of model performance. Existing methods primarily focus on mitigating domain shift issues caused by environmental factors but generally overlook the direct interference of environmental noise. We argue that mitigating environmental noise is equally critical for extracting discriminative cross-view features and introduce a pixel-to-feature co-enhancement network (P2FCN). P2FCN comprises a style-noise dual suppression module (SNDS) and a part-based multi-dimensional feature learning strategy (PMDFL). Specifically, the SNDS module mitigates the stylistic discrepancies between cross-environment images through pixel-level dynamic adjustment, while reducing the introduction of environmental noise to a certain degree. PMDFL improves the representation generalization by purifying discriminative information within partitioned channel-and spatial-wise feature subspaces. Extensive experiments on two widely used benchmarks,i.e., University-1652 and SUES-200, demonstrate that the proposed method achieves state-of-the-art performance in multiple environments compared to existing methods.
Qiang Zhao 0005, Tingyu Wang 0002, Rongfeng Lu, Chenggang Yan 0001
IEEE Trans. Geosci. Remote. Sens.1
2025 Loose-tight cluster regularization for unsupervised person re-identification
Yixiu Liu, Long Zhan, Pengju Si, Shaowei Jiang, Qiang Zhao 0005, Chenggang Yan 0001
Vis. Comput.6
2024 Aerial-view geo-localization based on multi-layer local pattern cross-attention network
Haoran Li 0025, Tingyu Wang 0002, Qiang Zhao 0005, Shaowei Jiang, Chenggang Yan 0001, Bolun Zheng
Appl. Intell.4
2023 Gaussian Label Distribution Learning for Spherical Image Object Detection
abstract
Spherical image object detection emerges in many applications from virtual reality to robotics and automatic driving, while many existing detectors use$l_{n}$-norms loss for regression of spherical bounding boxes. There are two intrinsic flaws for$l_{n}$-norms loss, i.e., independent optimization of parameters and inconsistency between metric (dominated by IoU) and loss. These problems are common in planar image detection but more significant in spherical image detection. Solution for these problems has been extensively discussed in planar image detection by using IoU loss and related variants. However, these solutions cannot be migrated to spherical image object detection due to the undifferentiable of the Spherical IoU (SphIoU). In this paper, we design a simple but effective regression loss based on Gaussian Label Distribution Learning (GLDL) for spherical image object detection. Besides, we observe that the scale of the object in a spherical image varies greatly. The huge differences among objects from different categories make the sample selection strategy based on SphIoU challenging. Therefore, we propose GLDL-ATSS as a better training sample selection strategy for objects of the spherical image, which can alleviate the drawback of IoU threshold-based strategy of scale-sample imbalance. Extensive results on various two datasets with different baseline detectors show the effectiveness of our approach.
Xinyuan Liu 0003, Qiang Zhao 0005, Yike Ma, Chenggang Yan 0001
CVPR3
2023 Sph2Pob: Boosting Object Detection on Spherical Images with Planar Oriented Boxes Methods
abstract
Object detection on panoramic/spherical images has been developed rapidly in the past few years, where IoU-calculator is a fundamental part of various detector components, i.e. Label Assignment, Loss and NMS. Due to the low efficiency and non-differentiability of spherical Unbiased IoU, spherical approximate IoU methods have been proposed recently. We find that the key of these approximate methods is to map spherical boxes to planar boxes. However, there exists two problems in these methods: (1) they do not eliminate the influence of panoramic image distortion; (2) they break the original pose between bounding boxes. They lead to the low accuracy of these methods. Taking the two problems into account, we propose a new sphere-plane boxes transform, called Sph2Pob. Based on the Sph2Pob, we propose (1) an differentiable IoU, Sph2Pob-IoU, for spherical boxes with low time-cost and high accuracy and (2) an agent Loss, Sph2Pob-Loss, for spherical detection with high flexibility and expansibility. Extensive experiments verify the effectiveness and generality of our approaches, and Sph2Pob-IoU and Sph2Pob-Loss together boost the performance of spherical detectors. The source code is available at https://github.com/AntXinyuan/sph2pob.
Xinyuan Liu 0003, Bin Chen 0021, Qiang Zhao 0005, Yike Ma, Chenggang Yan 0001
IJCAI4
2022 Unbiased IoU for Spherical Image Object Detection
abstract
As one of the fundamental components of object detection, intersection-over-union (IoU) calculations between two bounding boxes play an important role in samples selection, NMS operation and evaluation of object detection algorithms. This procedure is well-defined and solved for planar images, while it is challenging for spherical ones. Some existing methods utilize planar bounding boxes to represent spherical objects. However, they are biased due to the distortions of spherical objects. Others use spherical rectangles as unbiased representations, but they adopt excessive approximate algorithms when computing the IoU. In this paper, we propose an unbiased IoU as a novel evaluation criterion for spherical image object detection, which is based on the unbiased representations and utilize unbiased analytical method for IoU calculation. This is the first time that the absolutely accurate IoU calculation is applied to the evaluation criterion, thus object detection algorithms can be correctly evaluated for spherical images. With the unbiased representation and calculation, we also present Spherical CenterNet, an anchor free object detection algorithm for spherical images. The experiments show that our unbiased IoU gives accurate results and the proposed Spherical CenterNet achieves better performance on one real-world and two synthetic spherical object detection datasets than existing methods.
Bin Chen 0021, Yike Ma, Bailan Feng, Chenggang Yan 0001, Qiang Zhao 0005
AAAI9
2022 PANDORA: A Panoramic Detection Dataset for Object with Orientation
Qiang Zhao 0005, Yike Ma, Bailan Feng, Chenggang Yan 0001
ECCV (8)2
2022 Global Boundary Refinement for Semantic Segmentation via Optimal Transport
Shuaibin Zhang, Hao Liu 0068, Yike Ma, Qiang Zhao 0005
PRICAI (3)5
2021 Bipartite Matching for Crowd Counting with Point Supervision
abstract
For crowd counting task, it has been demonstrated that imposing Gaussians to point annotations hurts generalization performance. Several methods attempt to utilize point annotations as supervision directly. And they have made significant improvement compared with density-map based methods. However, these point based methods ignore the inevitable annotation noises and still suffer from low robustness to noisy annotations. To address the problem, we propose a bipartite matching based method for crowd counting with only point supervision (BM-Count). In BM-Count, we select a subset of most similar pixels from the predicted density map to match annotated pixels via bipartite matching. Then loss functions can be defined based on the matching pairs to alleviate the bad effect caused by those annotated dots with incorrect positions. Under the noisy annotations, our method reduces MAE and RMSE by 9% and 11.2% respectively. Moreover, we propose a novel ranking distribution learning framework to address the imbalanced distribution problem of head counts, which encodes the head counts as classification distribution in the ranking domain and refines the estimated count map in the continuous domain. Extensive experiments on four datasets show that our method achieves state-of-the-art performance and performs better crowd localization.
Hao Liu 0068, Qiang Zhao 0005, Yike Ma
IJCAI2
2021 Dense Scale Network for Crowd Counting
abstract
Crowd counting has been widely studied by computer vision community in recent years. Due to the large scale variation, it remains to be a challenging task. Previous methods adopt either multi-column CNN or single-column CNN with multiple branches to deal with this problem. However, restricted by the number of columns or branches, these methods can only capture a few different scales and have limited capability. In this paper, we propose a simple but effective network called DSNet for crowd counting, which can be easily trained in an end-to-end fashion. The key component of our network is the dense dilated convolution block, in which each dilation layer is densely connected with the others to preserve information from continuously varied scales. The dilation rates in dilation layers are carefully selected to prevent the block from gridding artifacts. To further enlarge the range of scales covered by the network, we cascade three blocks and link them with dense residual connections. We also introduce a novel multi-scale density level consistency loss for performance improvement. To evaluate our method, we compare it with state-of-the-art algorithms on five crowd counting datasets (ShanghaiTech, UCF-QNRF, UCF_CC_50, UCSD and WorldExpo'10). Experimental results demonstrate that DSNet can achieve the best overall performance and make significant improvements.
Hao Liu 0068, Yike Ma, Xi Zhang 0008, Qiang Zhao 0005
ICMR5
2021 Image stitching via deep homography estimation
Qiang Zhao 0005, Yike Ma, Chunfeng Yao, Bailan Feng
Neurocomputing1
2020 Dilated Convolutional Neural Networks for Panoramic Image Saliency Prediction
abstract
Saliency prediction is an important way to understand human's behavior and has a wide range of applications. Although lots of algorithms have been designed to predict saliency for planar images, there are few works for 360° images. In this paper, we propose an encoder-decoder network for panoramic image saliency prediction. Dilated convolutional layers are deployed in the network, which can extract more representative features and improve the accuracy of saliency prediction. To deal with the image distortions in 360° images, our network takes cube map format as input and processes six faces of cube map simultaneously. Respecting the saliency distribution of ground truth, we also propose a new data augmentation method to train the network, which is validated to be helpful for performance improvement. Extensive experiments show that our method gives new state-of-the-art results on 360° image saliency prediction.
Youqiang Zhang, Yike Ma, Qiang Zhao 0005
ICASSP5
2020 Light Field Reconstruction Using Dynamically Generated Filters
Xiuxiu Jing, Yike Ma, Qiang Zhao 0005, Ke Lyu
MMM (1)3
2020 Change detection with absolute difference of multiscale deep features
Rui Huang 0006, Qiang Zhao 0005, Yaobin Zou
Neurocomputing3
2020 Panoramic Light Field From Hand-Held Video and Its Sampling for Real-Time Rendering
abstract
By providing angular and spatial information of light rays, light field images are widely used in many applications. To capture large field of view light field, existing approaches either stitch small field of view light fields, whose apertures are also small, or leverage specialized equipments, which are not accessible to ordinary users. In this paper, we present a method to extract fully 360° field of view panoramic light fields from densely sampled hand-held videos. Unlike previous works, our method handles the large exposure variation and motion blur problems, which are common in the panoramic capture and hand-held videos. To provide real-time rendering performance, our method applies light field sampling to the extracted panoramic light field. Based on the unstructured representation, our method can sample the light field without explicit parameterization. We formulate the sampling problem into a set multicover problem and solve it globally using integer linear programming. Compared with the greedy based method, our approach can use much less video frames to represent the panoramic light field, but still can get better rendering results.
Qiang Zhao 0005, Jing Lv, Yike Ma, Yongdong Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2019 Near-infrared Image Guided Neural Networks for Color Image Denoising
abstract
Noisy color image and guided near-infrared (NIR) image can be jointly employed to eliminate noise and enhance details. Existing methods mostly rely on explicit designed filters and hand-crafted objective function optimization. These methods usually introduce erroneous structures from guidance signal. Besides, they are time-consuming and not suitable for real time applications. In this paper, we come up with a learning based method. The noisy color image and NIR image are fused, then fed into a fully convolutional neural network. The network learns a directly map from degraded image to restored sharp image. Our architecture can effectively eliminate image noise and transfer detail structure from guided image. Our trained network accepts any resolution of input image and runs in constant time. We evaluate the presented approach on both synthetic and real images. Results show that our approach outperforms the state-of-art methods.
Yike Ma, Junbo Guo, Qiang Zhao 0005, Yongdong Zhang 0001
ICASSP5
2019 Freely Explore the Scene with 360°Field of View
abstract
By providing 360° field of view, spherical panoramas are widely used in virtual reality (VR) systems and street view services. However, due to bandwidth or storage limitations, existing systems only provide sparsely captured panoramas and have limited interaction modes. Although there are methods that can synthesize novel views based on captured panoramas, the generated novel views all lie on the lines connecting existing views. Therefore these methods do not support free viewpoint navigation. In this paper, we propose a new panoramic image based rendering method. Our method takes pre-captured images as input and can synthesize panoramas at novel views that are far from input camera positions. Thus, it supports to freely explore the scene with 360° field of view.
Yike Ma, Juan Cao 0001, Qiang Zhao 0005, Yongdong Zhang 0001
VR5
2018 Spherical Superpixels: Benchmark and Evaluation
Xiaorui Xu, Qiang Zhao 0005, Wei Feng 0005
ACCV (6)3
2018 Wide Range Depth Estimation from Binocular Light Field Camera
Yike Ma, Guoqing Jin, Qiang Zhao 0005
BMVC5
2018 Panoramic Light Field Video Acquisition
abstract
Due to the limited field of view (FOV), the light field cameras can only capture a portion of the scene. Although there are methods designed to increase the FOV of light field images, these approaches stitch multiple sequentially captured light field images into a panoramic one and are restricted to static scenes. In this paper, we present a panoramic light field video acquisition system for dynamic scenes. Our system consists of 48 light field cameras and can capture panoramic light field videos at 6.3 fps. Two challenging problems, i.e. large volume data transmission, storage and camera synchronization, have been solved in our system. By using the light field stitching algorithm, the generated panoramic light field videos have 9×9 angular resolution, 6283×1200 spatial resolution and 360°×70° FOV, enabling us to do video refocusing, video focus tracking and other features not supported by light field static images.
Jing Lv, Qiang Zhao 0005, Yike Ma, Yongdong Zhang 0001
ICME3
2018 Distortion-aware CNNs for Spherical Images
abstract
Convolutional neural networks are widely used in computer vision applications. Although they have achieved great success, these networks can not be applied to 360 spherical images directly due to varying distortion effect. In this paper, we present distortion-aware convolutional network for spherical images. For each pixel, our network samples a non-regular grid based on its distortion level, and convolves the sampled grid using square kernels shared by all pixels. The network successively approximates large image patches from different tangent planes of viewing sphere with small local sampling grids, thus improves the computational efficiency. Our method also deals with the boundary problem, which is an inherent issue for spherical images. To evaluate our method, we apply our network in spherical image classification problems based on transformed MNIST and CIFAR-10 datasets. Compared with the baseline method, our method can get much better performance. We also analyze the variants of our network.
Qiang Zhao 0005, Yike Ma, Guoqing Jin, Yongdong Zhang 0001
IJCAI1
2018 Variable aperture panoramic imaging
Jing Lv, Qiang Zhao 0005, Yike Ma, Yongdong Zhang 0001
Multim. Tools Appl.2
2018 Spherical Superpixel Segmentation
abstract
These days, superpixel algorithms are widely used in computer vision and multimedia applications. However, existing algorithms are designed for planar images, which are less suited to deal with wide angle images. In this paper, we present a superpixel segmentation method for 360° spherical images. Unlike previous methods, our approach explicitly considers the geometry for spherical images and makes clustering to spherical image pixels. It starts with the seeds defined by Hammersley points sampled on the sphere, then iterates between assignment step and update step, which are both based on the distance metric respecting spherical geometry. We evaluate our method on the transformed Berkeley segmentation dataset and panorama segmentation dataset collected by ourselves. Experimental results show that our method can gain better performance in terms of adherence to image boundaries and superpixel structural regularity. Furthermore, superpixels generated by our method can reserve the coherence across image boundaries and all have closed contours.
Qiang Zhao 0005, Yike Ma, Jiawan Zhang, Yongdong Zhang 0001
IEEE Trans. Multim.1
2016 Spherical superpixel segmentation
abstract
In this paper, we present a superpixel generation method for spherical images, which cover 360° field-of-view. Unlike previous works that directly use existing superpixel algorithms on unrolled spherical images, our approach explicitly considers the geometry for spherical images and uses sphere as the underlying representation. For quantitative evaluation, we make a spherical image segmentation database by transforming Berkeley segmentation dataset to the spherical domain. Experimental results show that our method can get better performance in terms of adherence to image boundaries and spherical size variance. What's more, superpixels generated by our method all have closed contours.
Qiang Zhao 0005, Jiawan Zhang
ICME1
2015 SPHORB: A Fast and Robust Binary Feature on the Sphere
Qiang Zhao 0005, Wei Feng 0005, Jiawan Zhang
Int. J. Comput. Vis.1
2013 Cube2Video: Navigate Between Cubic Panoramas in Real-Time
abstract
Online virtual navigation systems enable users to hop from one 360° panorama to another, which belong to a sparse point-to-point collection, resulting in a less pleasant viewing experience. In this paper, we present a novel method, namely Cube2Video, to support navigating between cubic panoramas in a video-viewing mode. Our method circumvents the intrinsic challenge of cubic panoramas, i.e., the discontinuities between cube faces, in an efficient way. The proposed method extends the matching-triangulation-interpolation procedure with special considerations of the spherical domain. A triangle-to-triangle homography-based warping is developed to achieve physically plausible and visually pleasant interpolation results. The temporal smoothness of the synthesized video sequence is improved by means of a compensation transformation. As experimental results demonstrate, our method can synthesize pleasant video sequences in real time, thus mimicking walking or driving navigation.
Qiang Zhao 0005, Wei Feng 0005, Jiawan Zhang, Tien-Tsin Wong
IEEE Trans. Multim.1