Pan Gao 0001

dblp:87/5856-1 · DBLP profile ↗
← Back
53ranked-venue papers
17as first author
39since 2021 · last 2026
0000-0002-4492-5430ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 47 · 17 first-author · 33 since 2021Artificial intelligence and machine learning · 11 · 11 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 CloudMamba: Grouped Selective State Spaces for Point Cloud Analysis
abstract
Due to the long-range modeling ability and linear complexity property, Mamba has attracted considerable attention in point cloud analysis. Despite some interesting progress, related work still suffers from imperfect point cloud serialization, insufficient high-level geometric perception, and overfitting of the selective state space model (S6) at the core of Mamba. To this end, we resort to an SSM-based point cloud network termed CloudMamba to address the above challenges. Specifically, we propose sequence expanding and sequence merging, where the former serializes points along each axis separately and the latter serves to fuse the corresponding higher-order features causally inferred from different sequences, enabling unordered point sets to adapt more stably to the causal nature of Mamba without parameters. Meanwhile, we design chainedMamba that chains the forward and backward processes in the parallel bidirectional Mamba, capturing high-level geometric information during scanning. In addition, we propose a grouped selective state space model (GS6) via parameter sharing on S6, alleviating the overfitting problem caused by the computational mode in S6. Experiments on various point cloud tasks validate CloudMamba's ability to achieve state-of-the-art results with significantly less complexity.
Kanglin Qu, Pan Gao 0001, Qun Dai, Zhanzhi Ye, Rui Ye 0003, Yuanhao Sun
AAAI2
2026 Simba: Towards High-Fidelity and Geometrically-Consistent Point Cloud Completion via Transformation Diffusion
abstract
Point cloud completion is a fundamental task in 3D vision. A persistent challenge in this field is simultaneously preserving fine-grained details present in the input while ensuring the global structural integrity of the completed shape. While recent works leveraging local symmetry transformations via direct regression have significantly improved the preservation of geometric structure details, these methods suffer from two major limitations: (1) These regression-based methods are prone to overfitting which tend to memorize instant-specific transformations instead of learning a generalizable geometric prior. (2) Their reliance on point-wise transformation regression lead to high sensitivity to input noise, severely degrading their robustness and generalization. To address these challenges, we introduce Simba, a novel framework that reformulates point-wise transformation regression as a distribution learning problem. Our approach integrates symmetry priors with the powerful generative capabilities of diffusion models, avoiding instance-specific memorization while capturing robust geometric structures. Additionally, we introduce a hierarchical Mamba-based architecture to achieve high-fidelity upsampling. Extensive experiments across the PCN, ShapeNet, and KITTI benchmarks validate our method's state-of-the-art (SOTA) performance.
Lirui Zhang, Zhengkai Zhao, Zhi Zuo, Pan Gao 0001, Jie Qin 0004
AAAI4
2026 DP-Retinex: Dual-Prior Guided Low-Light Image Enhancement With YUV-Domain Reflectance-Illumination Decomposition
abstract
Accurate estimation of reflectance and illumination maps in the Retinex framework remains a significant challenge for low-light image enhancement due to inherent decomposition ambiguity. To address this, we propose DualPrior-Retinex, a novel framework that, inspired by Retinex theory, leverages the practical advantages of the YUV color space for robust enhancement. Our framework introduces a dual-prior architecture that effectively decouples the restoration process. It combines a diffusion-based global prior, responsible for ensuring low-frequency content consistency, with a YUV-based local prior designed to preserve high-frequency structural details. These complementary components are integrated by our Hierarchical Prior Fusion Module (HPFM), which balances perceptual quality with pixel-level fidelity in complex low-light scenarios. Extensive evaluations on multiple benchmarks demonstrate that our method achieves state-of-the-art performance across diverse metrics and visual qualities. Codes will be released at https://github.com/I2-Multimedia-Lab/DP-Retinex.
Zhengkai Zhao, Longmi Gao, Pan Gao 0001, Jie Qin 0004
IEEE Trans. Circuits Syst. Video Technol.3
2026 Fast Wrong-Way Cycling Detection in CCTV Videos: Sparse Sampling Is All You Need
abstract
Effective monitoring of unusual transportation behaviors, such as wrong-way cycling (i.e., riding a bicycle or e-bike against designated traffic flow), is crucial for optimizing law enforcement deployment and traffic planning. However, accurately recording all wrong-way cycling events is both unnecessary and infeasible in resource-constrained environments, as it requires high-resolution cameras for evidence collection and event detection. To address this challenge, we propose WWC-Predictor, a novel method for efficiently estimating the wrong-way cycling ratio, defined as the proportion of wrong-way cycling events relative to the total number of cycling movements over a given time period. The core innovation of our method lies in accurately detecting wrong-way cycling events in sparsely sampled frames using a light-weight detector, then estimating the overall ratio using an autoregressive moving average model. To evaluate the effectiveness of our method, we construct a benchmark dataset consisting of 35 minutes of video sequences with minute-level annotations. Our method achieves an average error rate of a mere 1.475% while consuming only 19.12% GPU time required by conventional tracking methods, validating its effectiveness in estimating the wrong-way cycling ratio. Our source code is publicly available at:https://github.com/VICA-Lab-HKUST-GZ/WWC-Predictor
Jing Xu 0029, Wentao Shi 0003, Lijuan Zhang 0003, Weikai Yang, Pan Gao 0001, Jie Qin 0004
IEEE Trans. Intell. Transp. Syst.6
2026 Distance-Attention Augmented Reinforcement Learning: A Robust Approach for 3D Cooperative UAV Navigation in Dense Urban Environments
abstract
Autonomous navigation is one of the key techniques of the extensive applications of unmanned aerial vehicles (UAVs) in various fields, such as urban traffic management, disaster response, and intelligent logistics. However, 3D cooperative navigation of UAVs in dense urban environments continues to face significant challenges, including extensive state spaces, trajectory smoothness, collaborative operations, and complex motion control. To address these challenges, we propose a distance-attention augmented reinforcement learning (DA2RL) algorithm to enable safe and smooth cooperative navigation of UAVs. In DA2RL, a distance-attention-based actor network is developed to refine the observation space and capture sequential information during training. A historical feature flow-based critic network is then developed for more accurate action decision evaluation. Additionally, several non-sparse reward functions are designed to further accelerate the training process. Finally, numerical experiments and comparison results demonstrate that DA2RL achieves superior navigation performance and generalization capability compared to recent benchmark works.
Lijuan Zhang 0003, Shihong Zhao, Pan Gao 0001
IEEE Trans. Mob. Comput.6
2025 Spatial-Spectral Aware Learning with Deformable Affinity for Weakly Supervised Semantic Segmentation
abstract
Weakly supervised semantic segmentation (WSSS) leverages image-level labels for semantic segmentation, reducing the reliance on pixel-level annotations. While recent studies have explored CLIP for this task, challenges remain in capturing sufficient local information and achieving complete object representations. To address these limitations, we propose a spatial-spectral awareness learning strategy that leverages spectral features to extract intricate structural information, compensating for the lack of pixel-level supervision. By effectively integrating spectral and spatial tokens, our method enhances feature representation and preserves fine-grained details. Furthermore, we introduce deformable offsets to incorporate positional constraints, refine the affinity matrix, reduce noise interference, and accurately delineate object boundaries. Our framework significantly improves CLIP’s performance for WSSS. Experiments on PASCAL VOC 2012 and MS COCO 2014 show that our single-stage WSSS approach outperforms state-of-the-art, and even surpasses some multi-stage methods in segmentation results. Code is available at: https://github.com/I2-Multimedia-Lab/CLIP-SSD
Yuzhen Zhou, Pan Gao 0001, Li Yu 0004
ICME2
2025 HydraMamba: Multi-Head State Space Model for Global Point Cloud Learning
abstract
The attention mechanism has become a dominant operator in point cloud learning, but its quadratic complexity leads to limited inter-point interactions, hindering long-range dependency modeling between objects. Due to excellent long-range modeling capability with linear complexity, the selective state space model (S6), as the core of Mamba, has been exploited in point cloud learning for long-range dependency interactions over the entire point cloud. Despite some significant progress, related works still suffer from imperfect point cloud serialization and lack of locality learning. To this end, we explore a state space model-based point cloud network termed HydraMamba to address the above challenges. Specifically, we design a shuffle serialization strategy, making unordered point sets better adapted to the causal nature of S6. Meanwhile, to overcome the deficiency of existing techniques in locality learning, we propose a ConvBiS6 layer, which is capable of capturing local geometries and global context dependencies synergistically. Besides, we propose MHS6 by extending the multi-head design to S6, further enhancing its modeling capability. HydraMamba achieves state-of-the-art results on various tasks at both object-level and scene-level. The code is available at https://github.com/Point-Cloud-Learning/HydraMamba.
Kanglin Qu, Pan Gao 0001, Qun Dai, Yuanhao Sun
ACM Multimedia2
2025 AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual Grounding
abstract
Weakly supervised visual grounding (VG) aims to locate objects in images based on text descriptions. Despite significant progress, existing methods lack strong cross-modal reasoning to distinguish subtle semantic differences in text expressions due to category-based and attribute-based ambiguity. To address these challenges, we introduce AlignCAT, a novel query-based semantic matching framework for weakly supervised VG. To enhance visual-linguistic alignment, we propose a coarse-grained alignment module that utilizes category information and global context, effectively mitigating interference from category-inconsistent objects. Subsequently, a fine-grained alignment module leverages descriptive information and captures word-level text features to achieve attribute consistency. By exploiting linguistic cues to their fullest extent, our proposed AlignCAT progressively filters out misaligned visual queries and enhances contrastive learning efficiency. Extensive experiments on three VG benchmarks, namely RefCOCO, RefCOCO+, and RefCOCOg, verify the superiority of AlignCAT against existing weakly supervised methods on two VG tasks. Our code is available at: https://github.com/I2-Multimedia-Lab/AlignCAT.
Chenyi Zhuang, Wutao Liu, Pan Gao 0001, Nicu Sebe
ACM Multimedia4
2025 HyperDiff: Hypergraph Guided Diffusion Model for 3D Human Pose Estimation
abstract
Monocular 3D human pose estimation (HPE) often encounters challenges such as depth ambiguity and occlusion during the 2D-to-3D lifting process. Additionally, traditional methods may overlook multi-scale skeleton features when utilizing skeleton structure information, which can negatively impact the accuracy of pose estimation. To address these challenges, this paper introduces a novel 3D pose estimation method, HyperDiff, which integrates diffusion models with HyperGCN. The diffusion model effectively captures data uncertainty, alleviating depth ambiguity and occlusion. Meanwhile, HyperGCN, serving as a denoiser, employs multi-granularity structures to accurately model high-order correlations between joints. This improves the model’s denoising capability especially for complex poses. Experimental results demonstrate that HyperDiff achieves state-of-the-art performance on the Human3.6M and MPI-INF-3DHP datasets and can flexibly adapt to varying computational resources to balance performance and efficiency. Code is released at https://github.com/IHENL/HyperDiff
Yuhua Huang, Pan Gao 0001
MMSP3
2025 ACGFormer: Attribute Classification Guided Transformer for Camouflaged Object Detection
Wutao Liu, Yao Yuan, Pan Gao 0001, Zheng Lin 0005, Jie Qin 0004
PRCV (16)3
2025 Attention-Based Augmented Reinforcement Learning Algorithm for Multi-UAV Navigation
abstract
Multi-UAV autonomous navigation serves as a fundamental basis for a wide array of applications across various fields. However, achieving cooperative navigation in complex and unknown three-dimensional environments remains a significant challenge. In this paper, a novel attention-based augmented reinforcement learning algorithm is proposed for multi-UAV navigation (A2RL-NAV) in 3D environments. Specifically, we incorporate an LSTM network along with an attention mechanism within the actor network. Additionally, we design some goal-oriented reward functions to address the issue of sparse rewards. Finally, we validate the effectiveness and robustness of our approach by comparing the proposed algorithm with several baseline algorithms through simulation experiments.
Pan Gao 0001, Lijuan Zhang 0003
VTC2025-Spring5
2025 Attribute reduction using self-information uncertainty measures in optimistic neighborhood extreme-granulation rough set
Kanglin Qu, Pan Gao 0001, Qun Dai, Yuanhao Sun, Xu Hua
Inf. Sci.2
2025 Bridging Domain Gap of Point Cloud Representations via Self-Supervised Geometric Augmentation
abstract
Recent progress of semantic point clouds analysis is largely driven by synthetic data (e.g., the ModelNet and the ShapeNet), which are typically complete, well-aligned and noisy-free. Therefore, representations of those ideal synthetic point clouds have limited variations in the geometric perspective and can gain good performance on a number of 3D vision tasks such as point cloud classification. In the context of unsupervised domain adaptation (UDA), representation learning designed for synthetic point clouds can hardly capture domain invariant geometric patterns from incomplete and noisy point clouds. To address such a problem, we introduce a novel scheme for induced geometric invariance of point cloud representations across domains, via regularizing representation learning with two self-supervised geometric augmentation tasks. On one hand, a novel pretext task of predicting translation distances of augmented samples is proposed to alleviate centroid shift of point clouds due to occlusion and noises. On the other hand, we pioneer an integration of the self-supervised relational learning on geometrically-augmented point clouds in a cascade manner, utilizing the intrinsic relationship of augmented variants and other samples as extra constraints of cross-domain geometric features. Experiments on the PointDA-10 dataset demonstrate the effectiveness of the proposed method, achieving the state-of-the-art performance.
Li Yu 0004, Hongchao Zhong, Longkun Zou, Ke Chen 0004, Pan Gao 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 Uncertainty-Guided Refinement for Fine-Grained Salient Object Detection
abstract
Recently, salient object detection (SOD) methods have achieved impressive performance. However, salient regions predicted by existing methods usually contain unsaturated regions and shadows, which limits the model for reliable fine-grained predictions. To address this, we introduce the uncertainty guidance learning approach to SOD, intended to enhance the model's perception of uncertain regions. Specifically, we design a novel Uncertainty Guided Refinement Attention Network (UGRAN), which incorporates three important components, i.e., the Multilevel Interaction Attention (MIA) module, the Scale Spatial-Consistent Attention (SSCA) module, and the Uncertainty Refinement Attention (URA) module. Unlike conventional methods dedicated to enhancing features, the proposed MIA facilitates the interaction and perception of multilevel features, leveraging the complementary characteristics among multilevel features. Then, through the proposed SSCA, the salient information across diverse scales within the aggregated features can be integrated more comprehensively and integrally. In the subsequent steps, we utilize the uncertainty map generated from the saliency prediction map to enhance the model's perception capability of uncertain regions, generating a highly-saturated fine-grained saliency prediction map. Additionally, we devise an adaptive dynamic partition (ADP) mechanism to minimize the computational overhead of the URA module and improve the utilization of uncertainty guidance. Experiments on seven benchmark datasets demonstrate the superiority of the proposed UGRAN over the state-of-the-art methodologies. Codes will be released at https://github.com/I2-Multimedia-Lab/UGRAN.
Yao Yuan, Pan Gao 0001, Qun Dai, Jie Qin 0004, Wei Xiang 0001
IEEE Trans. Image Process.2
2025 RLGrid: Reinforcement Learning Controlled Grid Deformation for Coarse-to-Fine Point Cloud Completion
abstract
Many point cloud completion methods typically rely on two steps: coarse generation and 2D Grid deformed fine output. However, in the fine generation, the expansion range (2D Grid Scale) required by each point cloud sample may be vastly different. For example, if the expansion range for a vessel shape is applied to a table shape, the final output may be blurry or sparse. To this end, we propose the RLGrid, Reinforcement Learning Controlled Grid Deformation. In detail, we firstly obtain two point cloud skeletons by two branches. One is to use an autoencoder, and the other is to convert the randomly generated normal distribution to coarse point cloud by GAN. We choose the one with smaller chamfer distance between coarse output and incomplete input as the input of the second stage. Then, a Reinforcement Learning (RL) agent is designed to select the appropriate expansion range based on the feature of each point cloud, and generate a 2D Grid. Finally, all the features are concatenated and sent into a Multilayer Perceptron to obtain the detailed complete point cloud. Experimental results show that RLGrid achieves state-of-the-art performance on various datasets. To the best of our knowledge, RL is not widely used in point cloud completion task due to lack of custom environment, and the proposed RLGrid provides an insight on how to formulate 2D Grid deformation as a sequential decision making problem. Further, it can also be plug-and-play on any 2D Grid features.
Pan Gao 0001, Xiaoyang Tan, Wei Xiang 0001
IEEE Trans. Multim.2
2024 Transformer-Based No-Reference Image Quality Assessment via Supervised Contrastive Learning
abstract
Image Quality Assessment (IQA) has long been a research hotspot in the field of image processing, especially No-Reference Image Quality Assessment (NR-IQA). Due to the powerful feature extraction ability, existing Convolution Neural Network (CNN) and Transformers based NR-IQA methods have achieved considerable progress. However, they still exhibit limited capability when facing unknown authentic distortion datasets. To further improve NR-IQA performance, in this paper, a novel supervised contrastive learning (SCL) and Transformer-based NR-IQA model SaTQA is proposed. We first train a model on a large-scale synthetic dataset by SCL (no image subjective score is required) to extract degradation features of images with various distortion types and levels. To further extract distortion information from images, we propose a backbone network incorporating the Multi-Stream Block (MSB) by combining the CNN inductive bias and Transformer long-term dependence modeling capability. Finally, we propose the Patch Attention Block (PAB) to obtain the final distorted image quality score by fusing the degradation features learned from contrastive learning with the perceptual distortion information extracted by the backbone network. Experimental results on six standard IQA datasets show that SaTQA outperforms the state-of-the-art methods for both synthetic and authentic datasets. Code is available at https://github.com/I2-Multimedia-Lab/SaTQA.
Jinsong Shi 0002, Pan Gao 0001, Jie Qin 0004
AAAI2
2024 CDFormer: When Degradation Prediction Embraces Diffusion Model for Blind Image Super-Resolution
abstract
Existing Blind image Super-Resolution (BSR) methods focus on estimating either kernel or degradation infor-mation, but have long overlooked the essential content details. In this paper, we propose a novel BSR approach, Content-aware Degradation-driven Transformer (CDFormer), to capture both degradation and content rep-resentations. However, low-resolution images cannot pro-vide enough content details, and thus we introduce a diffusion-based module CD Former dif f to first learn Con-tent Degradation Prior (CDP) in both low- and high-resolution images, and then approximate the real distribution given only low-resolution information. Moreover, we apply an adaptive SR network CDFormersR that effectively utilizes CDP to refine features. Compared to previous diffusion-based SR methods, we treat the diffusion model as an estimator that can overcome the limitations of expensive sampling time and excessive diversity. Experiments show that CDFormer can outperform existing methods, establishing a new state-of-the-art performance on various bench-marks under blind settings. Codes and models will be avail-able at https://github.com/I2-Multimedia-Lab/CDFormer.
Qingguo Liu, Chenyi Zhuang, Pan Gao 0001, Jie Qin 0004
CVPR3
2024 Puff-Net: Efficient Style Transfer with Pure Content and Style Feature Fusion Network
abstract
Style transfer aims to render an image with the artistic features of a style image, while maintaining the origi-nal structure. Various methods have been put forward for this task, but some challenges still exist. For instance, it is difficult for CNN-based methods to handle global information and long-range dependencies between input images, for which transformer-based methods have been proposed. Although transformers can better model the relationship between content and style images, they require high-cost hard-ware and time-consuming inference. To address these is-sues, we design a novel transformer model that includes only the encoder, thus significantly reducing the computational cost. In addition, we also find that existing style transfer methods may lead to images under-stylied or missing content. In order to achieve better stylization, we de-sign a content feature extractor and a style feature extrac-tor, based on which pure content and style images can be fed to the transformer. Finally, we propose a novel network termed Puff-Net, i.e., pure content and style feature fusion network. Through qualitative and quantitative experiments, we demonstrate the advantages of our model compared to state-of-the-art ones in the literature. The code is available at https://github.com/ZszYmy9/Puff-Net.
Sizhe Zheng, Pan Gao 0001, Peng Zhou 0010, Jie Qin 0004
CVPR2
2024 DSMix: Distortion-Induced Sensitivity Map Based Pre-training for No-Reference Image Quality Assessment
Jinsong Shi 0002, Pan Gao 0001, Xiaojiang Peng, Jie Qin 0004
ECCV (70)2
2024 Pointsoup: High-Performance and Extremely Low-Decoding-Latency Learned Geometry Codec for Large-Scale Point Cloud Scenes
Kang You, Li Yu 0004, Pan Gao 0001, Dandan Ding
IJCAI4
2024 Unified Unsupervised Salient Object Detection via Knowledge Transfer
Yao Yuan, Wutao Liu, Pan Gao 0001, Qun Dai, Jie Qin 0004
IJCAI3
2024 DiffuseST: Unleashing the Capability of the Diffusion Model for Style Transfer
Chenyi Zhuang, Pan Gao 0001
MMAsia3
2024 An Efficient Reinforcement Learning-Based Cooperative Navigation Algorithm for Multiple UAVs in Complex Environments
abstract
Autonomous navigation of multiple unmanned aerial vehicles (UAVs) serves as the foundation for their widespread applications in various fields. However, multi-UAV cooperative navigation is a challenging problem because it is difficult for UAVs to plan efficient paths while avoiding collisions with both neighboring UAVs and other obstacles. In this work, an efficient reinforcement learning (RL)-based cooperative navigation algorithm (RL-CN) is proposed to make optimal decisions in dense and dynamic environments. In RL-CN, the multi-UAV cooperative navigation is formulated as a Markov decision process and an enhanced RL method is proposed for continuous control of multiple UAVs. To solve the sparse reward problem, a group of reward functions are designed in reward shaping. Next, a staged-tuning with two subactor networks strategy is developed to accelerate the training process and improve the navigation performance. Then, an enhanced prioritized experience replay strategy is designed by considering both Temporal Difference (TD) error and connectivity trend. Thus, high-value transitions are more frequently selected to further enhance the training process. Finally, we conduct comprehensive simulation experiments and provide comparative results to substantiate the effectiveness and robustness of RL-CN algorithm.
Lijuan Zhang 0003, Weiguo Yi, Jiabin Peng, Pan Gao 0001
IEEE Trans. Ind. Informatics5
2024 Degradation-Aware Self-Attention Based Transformer for Blind Image Super-Resolution
abstract
Compared to CNN-based methods, Transformer-based methods achieve impressive image restoration outcomes due to their ability to model remote dependencies. However, how to apply Transformer-based methods to the field of blind super-resolution (SR) and further make an SR network adaptive to degradation information is still an open problem. In this paper, we propose a new degradation-aware self-attention-based Transformer model, where we incorporate contrastive learning into the Transformer network for learning the degradation representations of input images with unknown noise. In particular, we integrate both CNN and Transformer components into the SR network, where we first use the CNN modulated by the degradation information to extract local features, and then employ the degradation-aware Transformer to extract global semantic features. We apply our proposed model to several popular large-scale benchmark datasets for testing, and achieve the state-of-the-art performance compared to existing methods. In particular, our method yields a PSNR of 32.43 dB on the Urban100 dataset at ×2 scale, 0.94 dB higher than DASR, and 26.62 dB on the Urban100 dataset at ×4 scale, 0.26 dB improvement over KDSR, setting a new benchmark in this area. The source code is available at:https://github.com/I2-Multimedia-Lab/DSAT/tree/main.
Qingguo Liu, Pan Gao 0001, Kang Han, Ningzhong Liu, Wei Xiang 0001
IEEE Trans. Multim.2
2024 Blind Image Quality Assessment via Transformer Predicted Error Map and Perceptual Quality Token
abstract
Image quality assessment is a fundamental problem in the field of image processing, and due to the lack of reference images in most practical scenarios, no-reference image quality assessment (NR-IQA), has gained increasing attention recently. With the development of deep learning technology, many deep neural network-based NR-IQA methods have been developed, which try to learn the image quality based on the understanding of database information. Currently, Transformer has achieved remarkable progress in various vision tasks. Since the characteristics of the attention mechanism in Transformer fit the global perceptual impact of artifacts perceived by a human, Transformer is thus well suited for image quality assessment tasks. In this paper, we propose a Transformer based NR-IQA model using a predicted objective error map and perceptual quality token. Specifically, we firstly generate the predicted error map by pre-training one model consisting of a Transformer encoder and decoder, in which the objective difference between the distorted and the reference images is used as supervision. Then, we freeze the parameters of the pre-trained model and design another branch using the vision Transformer to extract the perceptual quality token for feature fusion with the predicted error map. Finally, the fused features are regressed to the final image quality score. Extensive experiments have shown that our proposed method outperforms the state-of-the-art methods in both authentic and synthetic image datasets. Moreover, the attentional map extracted by the perceptual quality token also does conform to the characteristics of the human visual system.
Jinsong Shi 0002, Pan Gao 0001, Aljoscha Smolic
IEEE Trans. Multim.2
2023 ProxyFormer: Proxy Alignment Assisted Point Cloud Completion with Missing Part Sensitive Transformer
abstract
Problems such as equipment defects or limited view-points will lead the captured point clouds to be incomplete. Therefore, recovering the complete point clouds from the partial ones plays an vital role in many practical tasks, and one of the keys lies in the prediction of the missing part. In this paper, we propose a novel point cloud completion approach namely ProxyFormer that divides point clouds into existing (input) and missing (to be predicted) parts and each part communicates information through its proxies. Specifically, we fuse information into point proxy via feature and position extractor, and generate features for missing point proxies from the features of existing point proxies. Then, in order to better perceive the position of missing points, we design a missing part sensitive transformer, which converts random normal distribution into reasonable position information, and uses proxy alignment to refine the missing proxies. It makes the predicted point proxies more sensitive to the features and positions of the missing part, and thus makes these proxies more suitable for subsequent coarse-to-fine processes. Experimental results show that our method outperforms state-of-the-art completion networks on several benchmark datasets and has the fastest inference speed.
Pan Gao 0001, Xiaoyang Tan, Mingqiang Wei
CVPR2
2023 Video Frame Interpolation with Flow Transformer
abstract
Video frame interpolation has been actively studied with the development of convolutional neural networks. However, due to the intrinsic limitations of kernel weight sharing in convolution, the interpolated frame generated by it may lose details. In contrast, the attention mechanism in Transformer can better distinguish the contribution of each pixel, and it can also capture long-range pixel dependencies, which provides great potential for video interpolation. Nevertheless, the original Transformer is commonly used for 2D images; how to develop a Transformer-based framework with consideration of temporal self-attention for video frame interpolation remains an open issue. In this paper, we propose Video Frame Interpolation Flow Transformer to incorporate motion dynamics from optical flows into the self-attention mechanism. Specifically, we design a Flow Transformer Block that calculates the temporal self-attention in a matched local area with the guidance of flow, making our framework suitable for interpolating frames with large motion while maintaining reasonably low complexity. In addition, we construct a multi-scale architecture to account for multi-scale motion, further improving the overall performance. Extensive experiments on three benchmarks demonstrate that the proposed method can generate interpolated frames with better visual quality than state-of-the-art methods.
Pan Gao 0001, Haoyue Tian, Jie Qin 0004
ACM Multimedia1
2023 StylePrompter: All Styles Need Is Attention
abstract
GAN inversion aims at inverting given images into corresponding latent codes for Generative Adversarial Networks (GANs), especially StyleGAN where exists a disentangled latent space that allows attribute-based image manipulation. As most inversion methods build upon Convolutional Neural Networks (CNNs), we transfer a hierarchical vision Transformer backbone innovatively to predict W+ latent codes at token level. We further apply a Style-driven Multi-scale Adaptive Refinement Transformer (SMART) in ℱ space to refine the intermediate style features of the generator. By treating style features as queries to retrieve lost identity information from the encoder's feature maps, SMART can not only produce high-quality inverted images but also surprisingly adapt to editing tasks. We then prove that StylePrompter lies in a more disentangled W+ and show the controllability of SMART. Finally, quantitative and qualitative experiments demonstrate that Style Prompter can achieve desirable performance in balancing reconstruction quality and editability, and is "smart" enough to fit into most edits, outperforming other ℱ -involved inversion methods. Our code is available at: https://github.com/I2-Multimedia-Lab/StylePrompter.
Chenyi Zhuang, Pan Gao 0001, Aljoscha Smolic
ACM Multimedia2
2023 Rate-Distortion Modeling for Bit Rate Constrained Point Cloud Compression
abstract
As being one of the main representation formats of 3D real world and well-suited for virtual reality and augmented reality applications, point clouds have gained a lot of popularity. In order to reduce the huge amount of data, a considerable amount of research on point cloud compression has been done. However, given a target bit rate, how to properly choose the color and geometry quantization parameters for compressing point clouds is still an open issue. In this paper, we propose a rate-distortion model based quantization parameter selection scheme for bit rate constrained point cloud compression. Firstly, to overcome the measurement uncertainty in evaluating the distortion of the point clouds, we propose a unified model to combine the geometry distortion and color distortion. In this model, we take into account the correlation between geometry and color variables of point clouds and derive a dimensionless quantity to represent the overall quality degradation. Then, we derive the relationships of overall distortion and bit rate with the quantization parameters. Finally, we formulate the bit rate constrained point cloud compression as a constrained minimization problem using the derived polynomial models and deduce the solution via an iterative numerical method. Experimental results show that the proposed algorithm can achieve optimal decoded point cloud quality at various target bit rates, and substantially outperform the video-rate-distortion model based point cloud compression scheme.
Pan Gao 0001, Shengzhou Luo, Manoranjan Paul
IEEE Trans. Circuits Syst. Video Technol.1
2022 Dilated Convolutional Neural Network-Based Deep Reference Picture Generation for Video Compression
abstract
Motion estimation and motion compensation are indispensable parts of inter prediction in video coding. Since the motion vector of objects is mostly in fractional pixel units, original reference pictures may not accurately provide a suitable reference for motion compensation. In this paper, we propose a deep reference picture generator which can create a picture that is more relevant to the cur-rent encoding frame, thereby further reducing temporal redundancy and improving video compression efficiency. Inspired by the recent progress of Convolutional Neural Network(CNN), this paper pro-poses to use a dilated CNN to build the generator. Moreover, we insert the generated deep picture into Versatile Video Coding(VVC) as a reference picture and perform a comprehensive set of experiments to evaluate the effectiveness of our network on the latest VVC Test Model–VTM. The experimental results demonstrate that our pro-posed method achieves on average 9.7% bit saving compared with VVC under low-delay P configuration.
Haoyue Tian, Pan Gao 0001, Manoranjan Paul
ICASSP2
2022 Video Frame Interpolation Based on Deformable Kernel Region
abstract
Video frame interpolation task has recently become more and more prevalent in the computer vision field. At present, a number of researches based on deep learning have achieved great success. Most of them are either based on optical flow information, or interpolation kernel, or a combination of these two methods. However, these methods have ignored that there are grid restrictions on the position of kernel region during synthesizing each target pixel. These limitations result in that they cannot well adapt to the irregularity of object shape and uncertainty of motion, which may lead to irrelevant reference pixels used for interpolation. In order to solve this problem, we revisit the deformable convolution for video interpolation, which can break the fixed grid restrictions on the kernel region, making the distribution of reference points more suitable for the shape of the object, and thus warp a more accurate interpolation frame. Experiments are conducted on four datasets to demonstrate the superior performance of the proposed model in comparison to the state-of-the-art alternatives.
Haoyue Tian, Pan Gao 0001, Xiaojiang Peng
IJCAI2
2022 SSformer: A Lightweight Transformer for Semantic Segmentation
abstract
It is well believed that Transformer performs better in semantic segmentation compared to convolutional neural networks. Nevertheless, the original Vision Transformer [2] may lack of inductive biases of local neighborhoods and possess a high time complexity. Recently, Swin Transformer [3] sets a new record in various vision tasks by using hierarchical architecture and shifted windows while being more efficient. However, as Swin Transformer is specifically designed for image classification, it may achieve suboptimal performance on dense prediction-based segmentation task. Further, simply combing Swin Transformer with existing methods would lead to the boost of model size and parameters for the final segmentation model. In this paper, we rethink the Swin Transformer for semantic segmentation, and design a lightweight yet effective transformer model, called SSformer. In this model, considering the inherent hierarchical design of Swin Transformer, we propose a decoder to aggregate information from different layers, thus obtaining both local and global attentions. Experimental results show the proposed SSformer yields comparable mIoU performance with state-of-the-art models, while maintaining a smaller model size and lower compute. Source code and pretrained models are available at: https://github.com/shiwt03/SSformer.
Wentao Shi 0003, Jing Xu 0029, Pan Gao 0001
MMSP3
2022 Block size selection in rate-constrained geometry based point cloud compression
Pan Gao 0001, Mingqiang Wei
Multim. Tools Appl.1
2022 Learning-based high-efficiency compression framework for light field videos
Wei Xiang 0001, Eric Wang 0001, Qiang Peng, Pan Gao 0001, Xiao Wu 0001
Multim. Tools Appl.5
2022 Quality Assessment for Omnidirectional Video: A Spatio-Temporal Distortion Modeling Approach
abstract
Omnidirectional video, also known as 360-degree video, has become increasingly popular nowadays due to its ability to provide immersive and interactive visual experiences. However, the ultra high resolution and the spherical observation space brought by the large spherical viewing range make omnidirectional video distinctly different from traditional 2D video. To date, the video quality assessment (VQA) for omnidirectional video is still an open issue. The existing VQA metrics for omnidirectional video only consider the spatial characteristics of distortions, but the temporal change of spatial distortions can also considerably influence human visual perception. In this paper, we propose a spatiotemporal modeling approach to evaluate the quality of the omnidirectional video. Firstly, we construct a spatioral quality assessment unit to evaluate the average distortion in temporal dimension at the eye fixation level, based upon which the smoothed distortion value is recursively calculated and consolidated by the characteristics of temporal variations. Then, we give a detailed solution of how to to integrate the three existing spatial VQA metrics into our approach. Besides, the cross-format omnidirectional video distortion measurement is also investigated. Finally, the spatiotemporal distortion of the whole video sequence is obtained by pooling. Based on the modeling approach, a full reference objective quality assessment metric for omnidirectional video is derived, namely OV-PSNR. The experimental results show that our proposed OV-PSNR greatly improves the prediction performance of the existing VQA metrics for omnidirectional video.
Pan Gao 0001, Aljoscha Smolic
IEEE Trans. Multim.1
2021 Coding and Quality Evaluation of Affordable 6DoF Video Content
Pan Gao 0001, Manoranjan Paul
ICIG (3)1
2021 Patch-Based Deep Autoencoder for Point Cloud Geometry Compression
abstract
The ever-increasing 3D application makes the point cloud compression unprecedentedly important and needed. In this paper, we propose a patch-based compression process using deep learning, focusing on the lossy point cloud geometry compression. Unlike existing point cloud compression networks, which apply feature extraction and reconstruction on the entire point cloud, we divide the point cloud into patches and compress each patch independently. In the decoding process, we finally assemble the decompressed patches into a complete point cloud. In addition, we train our network by a patch-to-patch criterion, i.e., use the local reconstruction loss for optimization, to approximate the global reconstruction optimality. Our method outperforms the state-of-the-art in terms of rate-distortion performance, especially at low bitrates. Moreover, the compression process we proposed can guarantee to generate the same number of points as the input. The network model of this method can be easily applied to other point cloud reconstruction problems, such as upsampling.
Kang You, Pan Gao 0001
MMAsia2
2021 Disocclusion filling for depth-based view synthesis with adaptive utilization of temporal correlations
Pan Gao 0001, Manoranjan Paul
J. Vis. Commun. Image Represent.1
2021 Cross-scene foreground segmentation with supervised and unsupervised model communication
Dong Liang 0008, Bin Kang, Pan Gao 0001, Xiaoyang Tan, Shun'ichi Kaneko
Pattern Recognit.4
2020 Textured Mesh vs Coloured Point Cloud: A Subjective Study for Volumetric Video Compression
abstract
Volumetric video (VV) pipelines reached a high level of maturity, creating interest to use such content in interactive visualisation scenarios. VV allows real world content to be captured and represented as 3D models, which can be viewed from any chosen viewpoint and direction. Thus, VV is ideal to be used in augmented reality (AR) or virtual reality (VR) applications. Both textured polygonal meshes and point clouds are popular methods to represent VV. Even though the signal and image processing community slightly favours the point cloud due to its simpler data structure and faster acquisition, textured polygonal meshes might have other benefits such as better visual quality and easier integration with computer graphics pipelines. To better understand the difference between them, in this study, we compare these two different representation formats for a VV compression scenario utilising state-of-the-art compression techniques. For this purpose, we build a database and collect user opinion scores for subjective quality assessment of the compressed VV. The results show that meshes provide the best quality at high bitrates, while point clouds perform better for low bitrate cases. The created VV quality database will be made available online to support further scientific studies on VV quality assessment.
Emin Zerman, Cagri Ozcinar, Pan Gao 0001, Aljoscha Smolic
QoMEX3
2020 Rate-Distortion Optimal Joint Texture and Depth Map Coding for 3-D Video Streaming
abstract
For high compression efficiency, 3-D video coding usually employs a multimode methodology to exploit the dependencies between multiple views as well as between texture and depth. However, different coding modes will posses differentiating error propagation behaviour when the compressed 3-D video bit stream is transmitted over packet-switched networks, and thus lead to different amount of visual distortions. Further, the texture and depth distortions are combined in a highly complex fashion to produce the overall view synthesis distortion. To minimize the expected view synthesis distortion, this paper proposes an efficient rate-distortion optimized algorithm for joint selection of texture and depth modes. Firstly, a statistical model is developed to estimate the overall view synthesis distortion, in which the channel distortions caused by error propagation under different coding modes are analyzed. Then, joint optimization of texture and depth modes is derived within an operational rate-distortion framework using the Lagrange multiplier method. The adjacent block dependency caused by warping operation is explicitly considered in optimization, for which we develop a dynamic programming method to find the optimal solution. Finally, we extend the Lagrange minimization method to the more general variable-block-size prediction case, where the optimal quadtree tree structure and the combined coding modes are jointly determined using a multi-level dual trellis. Experimental results are presented for a wide range of packet loss rates to illustrate the effectiveness of the proposed algorithm.
Pan Gao 0001, Manoranjan Paul
IEEE Trans. Multim.1
2019 Quality Assessment for Omnidirectional Video with Consideration of Temporal Distortion Variations
abstract
Omnidirectional video, also known as 360-degree video, offers an immersive visual experience by providing viewers with an ability to look in all directions within a scene. The quality assessment for omnidirectional video is still a quite difficult task compared to 2D video. As the temporal changes of spatial distortions can considerably influence human visual perception, this paper proposes a full reference objective video quality assessment metric by considering both the spatial characteristics of omnidirectional video and the temporal variation of distortions across frames. Firstly, we construct a spatio-temporal quality assessment unit to evaluate the average distortion in temporal dimension at eye fixation level. The smoothed distortion value is then consolidated by the characteristics of temporal variations. Afterwards, a global quality score of the whole video sequence is produced by pooling. Finally, our experimental results show that our proposed VQA method improves the prediction performance of existing VQA methods for omnidirectional video.
Pan Gao 0001
VCIP2
2019 An Improved Gaussian Mixture Model Based Hole-filling Algorithm Exploiting Depth Information
abstract
Virtual views generation is of great significance in free viewpoint video (FVV) as it can avoid the need to transmit a large volume of video data. An important issue in generating virtual views is how to fill the holes caused by occlusion. Using the Gaussian mixture model (GMM) to generate the background reference image is a commonly used hole-filling method. However, GMM usually has poor performance for sequences with reciprocal motion. In this paper, we propose an improved GMM-based method. To avoid the foreground pixels misclassified as the background pixels, we use depth information to adjust the learning rate in GMM. Foreground pixel is given a smaller learning rate than the background. Further, a refined foreground depth correlation (FDC) algorithm is proposed, which generates the background frame by tracking the change of the foreground depth in the temporal direction. In contrast to existing algorithms, we use a sliding window to obtain multiple background reference frames. These reference frames are then fused together to generate a more accurate background frame. Finally, we adaptively choose the background pixel from the GMM and FDC for hole filling. The experimental results show that subjective gain can be achieved, and significant objective gain can be observed in reciprocal motion sequences.
Pan Gao 0001
VCIP2
2019 Texture-Distortion-Constrained Joint Source-Channel Coding of Multi-View Video Plus Depth-Based 3D Video
abstract
A novel joint source and channel coding scheme tailored to 3D video is proposed in this paper to minimize the end-to-end view synthesis distortion within a given total bit rate for both texture and depth as well as a maximum tolerable distortion constraint for texture. First, we formulate a joint texture and depth coding mode selection strategy for error-resilient source coding of multi-view video plus depth-based 3D video through using the Lagrange multiplier method. Then, by considering the effect of residual errors after channel coding, we evolve to a more general formulation that jointly optimizes error-resilient source coding and channel coding in an integrated manner for unequal error protection between texture and depth, for which a theoretic solution using a proposed dual-trellis is derived. Finally, we extend the general formulation by including the texture distortion constraint. We show how to optimize the view synthesis quality while simultaneously catering to the texture quality constraint. Experimental results demonstrate the proposed algorithm has much better performance than existing related work.
Pan Gao 0001, Wei Xiang 0001, Dong Liang 0008
IEEE Trans. Circuits Syst. Video Technol.1
2019 Occlusion-Aware Depth Map Coding Optimization Using Allowable Depth Map Distortions
abstract
In depth map coding, rate-distortion optimization for those pixels that will cause occlusion in view synthesis is a rather challenging task, since the synthesis distortion estimation is complicated by the warping competition and the occlusion order can be easily changed by the adopted optimization strategy. In this paper, an efficient depth map coding approach using allowable depth map distortions is proposed for occlusion-inducing pixels. First, we derive the range of allowable depth level change for both the zero disparity error case and non-zero disparity error case with theoretic and geometrical proofs. Then, we formulate the problem of optimally selecting the depth distortion within allowable depth distortion range with the objective to minimize the overall synthesis distortion involved in the occlusion. The unicity and occlusion order invariance properties of allowable depth distortion range is demonstrated. Finally, we propose a dynamic programming based algorithm to locate the optimal depth distortion for each pixel. Simulation results illustrate the performance improvement of the proposed algorithm over the other state-of-the-art depth map coding optimization schemes.
Pan Gao 0001, Aljoscha Smolic
IEEE Trans. Image Process.1
2018 Optimization of Occlusion-Inducing Depth Pixels in 3-D Video Coding
abstract
The optimization of occlusion-inducing depth pixels in depth map coding has received little attention in the literature, since their associated texture pixels are occluded in the synthesized view and their effect on the synthesized view is considered negligible. However, the occlusion-inducing depth pixels still need to consume the bits to be transmitted, and will induce geometry distortion that inherently exists in the synthesized view. In this paper, we propose an efficient depth map coding scheme specifically for the occlusion-inducing depth pixels by using allowable depth distortions. Firstly, we formulate a problem of minimizing the overall geometry distortion in the occlusion subject to the bit rate constraint, for which the depth distortion is properly adjusted within the set of allowable depth distortions that introduce the same disparity error as the initial depth distortion. Then, we propose a dynamic programming solution to find the optimal depth distortion vector for the occlusion. The proposed algorithm can improve the coding efficiency without alteration of the occlusion order. Simulation results confirm the performance improvement compared to other existing algorithms.
Pan Gao 0001, Cagri Ozcinar, Aljoscha Smolic
ICIP1
2017 Joint texture and depth map coding for error-resilient 3-D video transmission
abstract
This paper addresses the problem of error-resilient source coding for 3-D video transmission over packet-loss networks. The proposed approach jointly optimizes the texture coding mode and the depth coding mode for each macroblock in the reference views. Firstly, a distortion model is developed to capture the effect of the texture distortion and depth distortion on the synthesized view. Then, joint optimization of texture and depth coding modes is derived based upon an operational rate-distortion framework using Lagrange multiplier method. In particular, a dual trellis-based algorithm is introduced in order to overcome the macroblock interdependencies of texture and depth map in the optimization procedure. Simulation results demonstrate that significant and consistent gains can be achieved over currently used techniques.
Pan Gao 0001, Wei Xiang 0001, D. M. Motiur Rahaman, Manoranjan Paul
ICIP1
2017 Analysis of Packet-Loss-Induced Distortion in View Synthesis Prediction-Based 3D Video Coding
abstract
View synthesis prediction (VSP) is a crucial coding tool for improving compression efficiency in the next generation 3D video systems. However, VSP is susceptible to catastrophic error propagation when multi-view video plus depth (MVD) data are transmitted over lossy networks. This paper aims at accurately modeling the transmission errors propagated in the inter-view direction caused by VSP. Toward this end, we first study how channel errors gradually propagate along the VSP-based inter-view prediction path. Then, a new recursive model is formulated to estimate the expected end-to-end distortion caused by those channel losses. For the proposed model, the compound impact of the transmission distortions of both the texture video and depth map on the quality of the synthetic reference view is mathematically analyzed. Especially, the expected view synthesis distortion due to depth errors is characterized in the frequency domain using a new approach, which combines the energy densities of the reconstructed texture image and the channel errors. The proposed model also explicitly considers the disparity rounding operation invoked for the sub-pixel precision rendering of the synthesized reference view. Experimental results are presented to demonstrate that the proposed analytic model is capable of effectively modeling the channel-induced distortion for MVD-based 3D video transmission.
Pan Gao 0001, Qiang Peng, Wei Xiang 0001
IEEE Trans. Image Process.1
2015 Transmission distortion modeling for view synthesis prediction based 3-D video streaming
abstract
View synthesis prediction (VSP) is an important tool for improving the coding efficiency in the next generation three-dimensional (3-D) video systems. However, VSP will result in a new type of inter-view error propagation when the multi-view video plus depth (MVD) data are transmitted over the lossy networks. In this paper, this new type of error propagation is characterized and modeled. Firstly, a new analytic model is formulated to estimate the expected transmission distortion caused by error propagation from the synthesized reference view. Then, the compound impact of the transmission distortions of both the texture video and the depth map on the quality of the synthetic reference view is mathematically analysed. Our extensive simulation results demonstrate that the proposed transmission distortion model is very accurate.
Pan Gao 0001, Wei Xiang 0001, Lijuan Zhang 0003
ICASSP1
2015 Modeling of packet-loss-induced distortion in 3-D synthesized views
abstract
This paper analyzes how transmission errors in the texture and depth map jointly affect the synthesized virtual view in 3-D video coding. In particular, we propose a framework that decouples the effects attributed to transmission errors in texture and depth to facilitate theoretical analysis. The synthesis distortion due to depth map errors is characterized in the frequency domain using a new approach that combines the energy density of the reconstructed texture and channel errors. Experimental results show that our analytical model can accurately estimate the rendering view quality.
Pan Gao 0001, Wei Xiang 0001
VCIP1
2015 Error-resilient multi-view video coding using Wyner-Ziv techniques
Pan Gao 0001, Qiang Peng, Wei Xiang 0001
Multim. Tools Appl.1
2015 Disparity Vector Correction for View Synthesis Prediction-Based 3-D Video Transmission
abstract
View synthesis prediction (VSP) is an important tool for enhancing the coding efficiency in the next-generation three- dimensional (3-D) video systems. However, VSP will lead to prediction position errors when the depth maps are corrupted by packet losses during transmission. In order to mitigate the prediction position errors, a novel disparity vector correction algorithm is proposed in this paper. Firstly, we investigate the relationship between the rendering position errors and the depth errors according to the VSP procedure. The depth map errors due to packet losses are then recursively estimated at the decoder without the use of the error-free reconstructed frames. Finally, based on the estimation of the reconstructed depth errors, the received disparity vectors can be corrected to find the matching synthesized pixels as those used at the encoder, and thereby the view synthesis-based inter-view error propagation can be effectively stopped. Experimental results show that the proposed methods with the estimated and actual depth errors can provide significant improvements in terms of both objective and subjective evaluations.
Pan Gao 0001, Wei Xiang 0001
IEEE Trans. Multim.1
2014 Rate-Distortion Optimized Mode Switching for Error-Resilient Multi-View Video Plus Depth Based 3-D Video Coding
abstract
In this paper, a rate-distortion optimized coding mode switching scheme is proposed to improve error resilience for multi-view video plus depth (MVD) based 3-D video transmission over lossy networks. First, we derive a new end-to-end distortion model for MVD-based 3-D video transmission. As compared with the previous MVD-based video distortion models in which distortion is measured by only investigating the expected texture video errors and depth errors on the synthesized virtual view, the proposed scheme characterizes both the end-to-end distortions in the rendered virtual view and the coded texture video due to packet losses. Moreover, inter-view error propagation for the texture video and depth map is also considered. Based on the proposed distortion model, an optimal mode decision algorithm is then performed in the texture video and depth map coding process. Experimental results show that the proposed method provides significant improvements in terms of both objective and subjective evaluations.
Pan Gao 0001, Wei Xiang 0001
IEEE Trans. Multim.1