Longguang Wang

dblp:202/1700 · DBLP profile ↗
← Back
80ranked-venue papers
12as first author
73since 2021 · last 2026
0000-0003-0429-0263ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 55 · 8 first-author · 48 since 2021Artificial intelligence and machine learning · 46 · 11 first-author · 43 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 6 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Not all regions are equal: Spatially adaptive representation learning for efficient visual object tracking
abstract
The sparsity of information in natural images poses great challenges for trackers to strike a balance between accuracy and efficiency. Existing methods commonly process all regions in the template and search area equally without considering their difference. As a result, considerable redundant computation is involved and limited inference efficiency is achieved. To remedy this, in this paper, we argue that not all regions are equal during the representation learning for visual object tracking. Specifically, we develop a sparse mask Transformer (SMTransformer) that is able to achieve spatially adaptive representation learning. Particularly, a deformable patch embedding module is constructed to adapt the receptive field to focus on the object in the template. In addition, sparse mask module is developed to dynamically identify regions with low object existence probabilities, thereby reducing the search region progressively for higher computational efficiency. With these two modules, our SMTransformer can significantly reduce the redundant computation for superior efficiency while maintaining high performance. Extensive experiments are conducted on a wide range of benchmark datasets and the results demonstrate the state-of-the-art performance of the proposed SMTransformer against previous methods in terms of both accuracy and efficiency.
Hongke Xu, Zhanwen Liu, Longguang Wang
Neural Networks6
2026 Deep Lookup Network
abstract
Convolutional neural networks are constructed with massive operations with different types and are highly computationally intensive. Among these operations, multiplication operation is higher in computational complexity and usually requires more energy consumption with longer inference time than other operations, which hinders the deployment of convolutional neural networks on mobile devices. In many resource-limited edge devices, complicated operations can be calculated via lookup tables to reduce computational cost. Motivated by this, in this paper, we introduce a generic and efficient lookup operation which can be used as a basic operation for the construction of neural networks. Instead of calculating the multiplication of weights and activation values, simple yet efficient lookup operations are adopted to compute their responses. To enable end-to-end optimization of the lookup operation, we construct the lookup tables in a differentiable manner and propose several training strategies to promote their convergence. By replacing computationally expensive multiplication operations with our lookup operations, we develop lookup networks for the image classification, image super-resolution, and point cloud classification tasks. It is demonstrated that our lookup networks can benefit from the lookup operations to achieve higher efficiency in terms of energy consumption and inference speed while maintaining competitive performance to vanilla convolutional networks. Extensive experiments show that our lookup networks produce state-of-the-art performance on different tasks (both classification and regression tasks) and different data types (both images and point clouds).
Yulan Guo, Longguang Wang, Wendong Mao, Yingqian Wang 0002, Li Liu 0002, Wei An 0003
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 Simulating the Real World: A Unified Survey of Multimodal Generative Models
abstract
Understanding and replicating the real world is a critical challenge in Artificial General Intelligence (AGI) research. To achieve this, many existing approaches, such as world models, aim to capture the fundamental principles governing the physical world, enabling more accurate simulations and meaningful interactions. However, current methods often treat different modalities, including 2D (images), videos, 3D, and 4D representations, as independent domains, overlooking their interdependencies. Additionally, these methods typically focus on isolated dimensions of reality without systematically integrating their connections. In this survey, we present a unified survey for multimodal generative models that investigate the progression of data dimensionality in real-world simulation. Specifically, this survey starts from 2D generation (appearance), then moves to video (appearance+dynamics) and 3D generation (appearance+ geometry), and finally culminates in 4D generation that integrate all dimensions. To the best of our knowledge, this is the first attempt to systematically unify the study of 2D, video, 3D, and 4D generation within a single framework. To guide future research, we provide a comprehensive review of datasets, evaluation metrics, and future directions to foster insights for newcomers. This survey serves as a bridge to advance the study of multimodal generative models and real-world simulation within a unified framework.
Longguang Wang, Yuwei Guo 0002, Yukai Shi, Anyi Rao, Zeyu Wang 0003, Hui Xiong 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 Probing Deep Into Temporal Profile Makes the Infrared Small Target Detector Much Better
abstract
Infrared small target (IRST) detection is challenging in simultaneously achieving precise, robust, and efficient performance due to extremely dim targets and strong interference. Current learning-based methods attempt to leverage "more" information from both the spatial and the short-term temporal domains, but suffer from unreliable performance under complex conditions while incurring computational redundancy. In this paper, we explore the "more essential" information from a more crucial domain for the detection. Through theoretical analysis, we reveal that the global temporal saliency and correlation information in the temporal profile demonstrate significant superiority in distinguishing target signals from other signals. To investigate whether such superiority is preferentially leveraged by well-trained networks, we built the first prediction attribution tool in this field and verified the importance of the temporal profile information. Inspired by the above conclusions, we remodel the IRST detection task as a one-dimensional signal anomaly detection task, and propose an efficient deep temporal probe network (DeepPro) that only performs calculations in the time dimension for IRST detection. We conducted extensive experiments to fully validate the effectiveness of our method. The experimental results are exciting, as our DeepPro outperforms existing state-of-the-art IRST detection methods on widely-used benchmarks with extremely high efficiency, and achieves a significant improvement on dim targets and in complex scenarios. We provide a new modeling domain, a new insight, a new method, and a new performance, which can promote the development of IRST detection.
Ruojing Li, Wei An 0003, Yingqian Wang 0002, Xinyi Ying, Yimian Dai, Longguang Wang, Yulan Guo, Li Liu 0002
IEEE Trans. Pattern Anal. Mach. Intell.6
2026 Diving Into Epipolar Transformers for Light Field Super-Resolution and Disparity Estimation
abstract
Light field (LF) cameras capture the light rays of a 3D scene from multiple views simultaneously, and thus provide a more immersive experience of the real world as compared to traditional cameras. Although significant progress has been made in various LF image processing tasks, it remains challenging to effectively model the non-local spatial-angular correlations inherent in LF images, particularly when dealing with complex disparity variations. In this paper, we focus on orthogonal epipolar geometry of LF images and propose a generic Epipolar Transformer mechanism that incorporates geometrically meaningful correlations along the epipolar lines. Our Epipolar Transformer mechanism enjoys the following benefits: learning effective and diverse LF feature representations, delivering satisfactory results without redundant architectural designs, and enabling flexible extension to various LF-related tasks with simple adaptations. For LF spatial and angular super-resolution, our methods not only achieve state-of-the-art performance on benchmark datasets, but also demonstrate superior and robust performance on large disparity variations. For disparity estimation, we explore the use of geometry information encoded in our Epipolar Transformer to directly regress the disparity results, effectively avoiding the limitation of a fixed maximum disparity.
Zhengyu Liang, Yingqian Wang 0002, Longguang Wang, Jun-Gang Yang, Yulan Guo, Li Liu 0002, Shilin Zhou 0001, Wei An 0003
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 Triple Spectral Fusion for Sensor-Based Human Activity Recognition
abstract
The field of sensor-based human activity recognition (HAR) mainly uses posture, motion and context data of Inertial Measurement Units (IMUs) to identify daily activities. Despite the advancements in learning-based methods, it is challenging to perform information fusion from the temporal perspective due to the complexities in fusing heterogeneous sensor data and establishing long-term context correlations. This paper proposes a novel triple spectral fusion framework tailored for HAR. First, we develop an adaptive complementary filtering technique for noise suppression and organize each IMU's sensors into posture and motion modality nodes. Given that IMU nodes form a dynamic heterogeneous graph, we then apply adaptive filtering within the graph Fourier domain to merge both homogeneous and heterogeneous node information. Furthermore, an adaptive wavelet frequency selection approach is implemented to suppress context redundancy and shorten the length of features. This approach enhances both timestamp-based graph aggregation and the correlation of long-term contexts. Our framework uses adaptive filtering in the Fourier, graph Fourier, and wavelet domains, enabling effective multi-sensor fusion and context correlation. Extensive experiments on ten benchmark datasets demonstrate the superior performance of our framework.
Ye Zhang 0037, Longguang Wang, Qing Gao 0002, Chaocan Xiang, Mohammed Bennamoun, Yulan Guo
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 Structured Grouping Collaborative Decorrelated Regularization for Model Pruning in Infrared Small Target Detection
Yonghao Li, Jun Chen 0005, Boyang Li 0007, Yulan Guo, Longguang Wang, Siyi Deng
Pattern Recognit.5
2026 P3D: Plug-and-play prompt-driven framework for RGB-thermal semantic segmentation
abstract
• A plug-and-play prompt-driven framework for RGB-thermal image semantic segmentation. • LoRA-based fine-tuning strategy for SAM series model integration. • A model-agnostic encoder to generate statistical distributed prompts for training. The semantic segmentation of RGB-thermal images is critical for applications with low-light conditions. Existing works primarily focus on feature fusion strategies and model design to enhance performance. While Visual Foundation Models (VFMs) have been introduced in previous studies to improve generalization and segmentation accuracy, they suffer from poor compatibility with other models thus requiring full model retraining. Additionally, the domain gap and modality gap between VFM pre-training datasets and RGB-thermal semantic segmentation datasets pose significant challenges to VFM adaptation for downstream tasks. To address these issues, in this paper a plug-and-play prompt driven framework P 3 D is proposed. Unlike existing VFM-based methods that require complete retraining for each specific architecture, P 3 D is designed with a model-agnostic training strategy that enables one-time training and seamless integration with various existing methods without requiring retraining. First, a dual-branch LoRA (Low-Rank Adaptation) fine-tuned (DBLF) image encoder for the RGB and thermal image branches is proposed to narrow the domain gap and modality gap when incorporating SAM series models into our task. Second, a unified prompt generation and representation (UPGR) encoder is proposed. It generates diverse prompts using semantic labels during the training stage, ensuring the generated prompts are model-agnostic and compatible with existing methods. Finally, a cross-modality spatial-channel attention (CM-SCA) decoder is developed to fuse the embeddings from two-modality images and prompts for the final prediction. Extensive experiments are conducted on three popular benchmarks. Results demonstrate that P 3 D not only improves the performance of existing models but also outperforms current state-of-the-art (SOTA) methods leveraging < 1% trainable parameters. More importantly, by simply plugging P 3 D into existing methods, we consistently achieve significant performance improvements without retraining these base models, demonstrating the practical value of our plug-and-play design.
Yongqi Sun, Chenguang Dai, Hanyun Wang, Longguang Wang, Wenke Li, Anzhu Yu
Pattern Recognit.4
2026 RelightFlow: An Inversion-Free Video Relighting Model via Dual-Trajectory Diffusion Editing
abstract
Video relighting is a fundamental task with wide-ranging applications in contemporary visual computing, including film production, immersive virtual reality, augmented reality, and interactive digital worlds. Its objective is to generate temporally stable lighting effects while preserving the structural integrity, visual appearance, and intrinsic physical properties of objects in the source video. Recent studies combined image relighting models with video diffusion models, achieving notable progress in training-free video relighting. However, these methods rely on mapping noisy latents back to the pixel space during the relighting process, leading to degraded fidelity, consistency, and stability. In this work, we propose RelightFlow, a training-free video relighting framework built upon a flow-matching-based video DiT model, which requires neither inversion nor additional training, obtaining more feasible and robust relighting effects. Specifically, we first design a detail-preserving relighting module coupled with an interpolation trajectory, which injects stable lighting cues in the early generation stages while maintaining fine spatial details. Next, we develop a temporally consistent relighting module to form a flow-editing trajectory, leveraging the velocity fields predicted by the video DiT model to enhance temporal coherence. Finally, we introduce a dynamic fusion strategy that adaptively integrates these two trajectories to balance relighting intensity and temporal stability. Extensive experiments demonstrate that RelightFlow achieves high-quality video relighting with stable relighting intensity, superior fidelity, and temporal consistency. The code are provided in https://github.com/Yukun66/RelightFlow.
Qi Zhang 0029, Longguang Wang, Yulan Guo
IEEE Trans. Circuits Syst. Video Technol.4
2026 Revisiting Subspace Disentangling for Light Field Spatial Super-Resolution
Yingqian Wang 0002, Xueying Wang 0001, Zhengyu Liang, Longguang Wang, Lvli Tian, Jun-Gang Yang
IEEE Trans. Circuits Syst. Video Technol.5
2026 GeoStyler: A Generalizable Geometry-Aware Diffusion-Based Approach for Direct 3D Gaussian Style Transfer
abstract
Direct 3D scene stylization from sparse views remains a significant challenge, as existing optimization-based methods are prohibitively slow and require dense inputs to prevent geometric corruption. While recent direct methods accelerate this process, their rigid decoupling of a static geometry from appearance often leads to visual artifacts, where stylistic textures conflict with and distort the underlying scene structure. To address these limitations, we introduce GeoStyler, a direct framework that generates high-fidelity, multi-view consistent stylized 3D scenes in seconds. Our approach reformulates the conventional pipeline by first leveraging a diffusion model to generate a set of geometrically consistent stylized 2D images. The core of this stage is a novel hybrid query formulation for the self-attention mechanism. Specifically, cross-view geometric information is directly embedded into the query to enforce 3D consistency, while style information is independently injected via the key and value to preserve scene structure. This process is further stabilized by a geometrically-aware latent initialization that provides a coherent starting point for the denoising process. Subsequently, a decoupled reconstruction network lifts these 2D stylized images to 3D Gaussians. A geometry branch predicts a robust 3D scaffold from the original content images, while a parallel style branch predicts the final appearance from our generated stylized images, ensuring structural integrity is not compromised. Extensive experiments on large-scale benchmarks, including RealEstate10K and ACID, demonstrate that GeoStyler significantly outperforms prior arts in stylization quality and multi-view consistency, achieving state-of-the-art performance with a dramatic speedup. Our project page: https://huhuhuxiao.github.io/Geo-Styler/.
Qibin Hu, Ye Zhang 0037, Jisheng Dang, Minglin Chen, Longguang Wang, Yulan Guo
IEEE Trans. Image Process.5
2026 POSITION: Open World 3D Scene CAD Recomposition
abstract
3D scene CAD recomposition aims to reconstruct a given scene by retrieving and assembling CAD models from a database, so as to accurately simulate the geometric properties and spatial arrangement of the original environment. Recent methods learn this task through training on limited scan-to-CAD annotation data, which hinders their generalization to diverse real-world scenes. In this paper, we propose POSITION, an open-world 3D scene CAD recomposition method to construct the 3D scene with CADs retrieved from an open-set database. POSITION is designed following a divide-and-conquer strategy. Firstly, we extract open-world multi-modal object representations from a captured 3D scene. Secondly, on top of the representations, we propose a coarse-to-fine retrieval method to retrieve CADs that are visually, geometrically and semantically match real objects. Thirdly, we present a physically plausible pose alignment method to adjust retrieved CAD models to maintain consistent geometry and layout with the observation. By decomposing the problem into well-defined subtasks, our approach achieves generalization across various scene types and scalable CAD databases without retraining or fine-tuning. Our approach demonstrates superior CAD recomposition performance on both the Scan2CAD and diverse real-world 3D scene datasets. Our project page: https://yangrongkun.github.io/position/.
Rongkun Yang, Hongda Liu 0001, Sheng Ao, Longguang Wang, Shunbo Zhou, Yulan Guo
IEEE Trans. Image Process.6
2025 AIQViT: Architecture-Informed Post-Training Quantization for Vision Transformers
abstract
Post-training quantization (PTQ) has emerged as a promising solution for reducing the storage and computational cost of vision transformers (ViTs). Recent advances primarily target at crafting quantizers to deal with peculiar activations characterized by ViTs. However, most existing methods underestimate the information loss incurred by weight quantization, resulting in significant performance deterioration, particularly in low-bit cases. Furthermore, a common practice in quantizing post-Softmax activations of ViTs is to employ logarithmic transformations, which unfortunately prioritize less informative values around zero. This approach introduces additional redundancies, ultimately leading to suboptimal quantization efficacy. To handle these, this paper proposes an innovative PTQ method tailored for ViTs, termed AIQViT (Architecture-Informed Post-training Quantization for ViTs). First, we design an architecture-informed low-rank compensation mechanism, wherein learnable low-rank weights are introduced to compensate for the degradation caused by weight quantization. Second, we design a dynamic focusing quantizer to accommodate the unbalanced distribution of post-Softmax activations, which dynamically selects the most valuable interval for higher quantization resolution. Extensive experiments on five vision tasks, including image classification, object detection, instance segmentation, point cloud classification, and point cloud part segmentation, demonstrate the superiority of AIQViT over state-of-the-art PTQ methods.
Runqing Jiang, Ye Zhang 0037, Longguang Wang, Pengpeng Yu, Yulan Guo
AAAI3
2025 SaMam: Style-aware State Space Model for Arbitrary Image Style Transfer
abstract
Global effective receptive field plays a crucial role for image style transfer (ST) to obtain high-quality stylized results. However, existing ST backbones (e.g., CNNs and Transformers) suffer huge computational complexity to achieve global receptive fields. Recently, State Space Model (SSM), especially the improved variant Mamba, has shown great potential for long-range dependency modeling with linear complexity, which offers an approach to resolve the above dilemma. In this paper, we develop a Mamba-based style transfer framework, termed SaMam. Specifically, a mamba encoder is designed to efficiently extract content and style information. In addition, a style-aware mamba decoder is developed to flexibly adapt to various styles. Moreover, to address the problems of local pixel forgetting, channel redundancy and spatial discontinuity of existing SSMs, we introduce local enhancement and zigzag scan mechanisms. Qualitative and quantitative results demonstrate that our SaMam outperforms state-of-the-art methods in terms of both accuracy and efficiency.
Hongda Liu 0001, Longguang Wang, Ye Zhang 0037, Ziru Yu, Yulan Guo
CVPR2
2025 VideoDirector: Precise Video Editing via Text-to-Video Models
abstract
Despite the typical inversion-then-editing paradigm using text-to-image (T2I) models has demonstrated promising results, directly extending it to text-to-video (T2V) models still suffers severe artifacts such as color flickering and content distortion. Consequently, current video editing methods primarily rely on T2I models, which inherently lack temporal-coherence generative ability, often resulting in inferior editing results. In this paper, we attribute the failure of the typical editing paradigm to: 1) Tightly Spatial-temporal Coupling. The vanilla pivotal-based inversion strategy struggles to disentangle spatial-temporal information in the video diffusion model; 2) Complicated Spatial-temporal Layout. The vanilla cross-attention control is deficient in preserving the unedited content. To address these limitations, we propose a spatial-temporal decoupled guidance (STDG) and multi-frame null-text optimization strategy to provide pivotal temporal cues for more precise pivotal inversion. Furthermore, we introduce a self-attention control strategy to maintain higher fidelity for precise partial content editing. Experimental results demonstrate that our method (termed VideoDirector) effectively harnesses the powerful temporal generation capabilities of T2V models, producing edited videos with state-of-the-art performance in accuracy, motion smoothness, realism, and fidelity to unedited content.
Longguang Wang, Qibin Hu, Kai Xu 0004, Yulan Guo
CVPR2
2025 DropoutGS: Dropping Out Gaussians for Better Sparse-view Rendering
abstract
Although 3D Gaussian Splatting (3DGS) has demonstrated promising results in novel view synthesis, its performance degrades dramatically with sparse inputs and generates undesirable artifacts. As the number of training views decreases, the novel view synthesis task degrades to a highly under-determined problem such that existing methods suffer from the notorious overfitting issue. Interestingly, we observe that models with fewer Gaussian primitives exhibit less overfitting under spare inputs. Inspired by this observation, we propose a Random Dropout Regularization (RDR) to exploit the advantages of low-complexity models to alleviate overfitting. In addition, to remedy the lack of high-frequency details for these models, an Edge-guided Splitting Strategy (ESS) is developed. With these two techniques, our method (termed DropoutGS) provides a simple yet effective plug-in approach to improve the generalization performance of existing 3DGS methods. Extensive experiments show that our DropoutGS produces state-of-the-art performance under sparse views on benchmark datasets including Blender, LLFF, and DTU. The project page is at: https://xuyx55.github.io/DropoutGS/.
Yexing Xu, Longguang Wang, Minglin Chen, Sheng Ao, Li Li 0100, Yulan Guo
CVPR2
2025 Learning Robust Stereo Matching in the Wild with Selective Mixture-of-Experts
abstract
Recently, learning-based stereo matching networks have advanced significantly. However, they often lack robustness and struggle to achieve impressive cross-domain performance due to domain shifts and imbalanced disparity distributions among diverse datasets. Leveraging Vision Foundation Models (VFMs) can intuitively enhance the model's robustness, but integrating such a model into stereo matching cost-effectively to fully realize their robustness remains a key challenge. To address this, we propose SMoEStereo, a novel framework that adapts VFMs for stereo matching through a tailored, scene-specific fusion of Low-Rank Adaptation (LoRA) and Mixture-of-Experts (MoE) modules. SMoEStereo introduces MoE-LoRA with adaptive ranks and MoE-Adapter with adaptive kernel sizes. The former dynamically selects optimal experts within MoE to adapt varying scenes across domains, while the latter injects inductive bias into frozen VFMs to improve geometric feature extraction. Importantly, to mitigate computational overhead, we further propose a lightweight decision network that selectively activates MoE modules based on input complexity, balancing efficiency with accuracy. Extensive experiments demonstrate that our method exhibits state-of-the-art cross-domain and joint generalization across multiple benchmarks without dataset-specific adaptation. The code is available at \textcolor{red}{https://github.com/cocowy1/SMoE-Stereo}.
Yun Wang 0053, Longguang Wang, Chenghao Zhang 0003, Zhanjie Zhang, Ao Ma 0005, Chenyou Fan, Tin Lun Lam, Junjie Hu 0003
ICCV2
2025 Unsupervised Degradation Representation Learning for Unpaired Restoration of Images and Point Clouds
abstract
Restoration tasks in low-level vision aim to restore high-quality (HQ) data from their low-quality (LQ) observations. To circumvents the difficulty of acquiring paired data in real scenarios, unpaired approaches that aim to restore HQ data solely on unpaired data are drawing increasing interest. Since restoration tasks are tightly coupled with the degradation model, unknown and highly diverse degradations in real scenarios make learning from unpaired data quite challenging. In this paper, we propose a degradation representation learning scheme to address this challenge. By learning to distinguish various degradations in the representation space, our degradation representations can extract implicit degradation information in an unsupervised manner. Moreover, to handle diverse degradations, we develop degradation-aware (DA) convolutions with flexible adaption to various degradations to fully exploit the degrdation information in the learned representations. Based on our degradation representations and DA convolutions, we introduce a generic framework for unpaired restoration tasks. Based on our framework, we propose UnIRnet and UnPRnet for unpaired image and point cloud restoration tasks, respectively. It is demonstrated that our degradation representation learning scheme can extract discriminative representations to obtain accurate degradation information. Experiments on unpaired image and point cloud restoration tasks show that our UnIRnet and UnPRnet achieve state-of-the-art performance.
Longguang Wang, Yulan Guo, Yingqian Wang 0002, Jun-Gang Yang, Wei An 0003
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Event-Based Motion Deblurring With Blur-Aware Reconstruction Filter
abstract
Event-based motion deblurring aims at reconstructing a sharp image from a single blurry image and its corresponding events triggered during the exposure time. Existing methods learn the spatial distribution of blur from blurred images, then treat events as temporal residuals and learn blurred temporal features from them, and finally restore clear images through spatio-temporal interaction of the two features. However, due to the high coupling of detailed features such as the texture and contour of the scene with blur features, it is difficult to directly learn effective blur spatial distribution from the original blurred image. In this paper, we provide a novel perspective, i.e., employing the blur indication provided by events, to instruct the network in spatially differentiated image reconstruction. Due to the consistency between event spatial distribution and image blur, event spatial indication can learn blur spatial features more simply and directly, and serve as a complement to temporal residual guidance to improve deblurring performance. Based on the above insight, we propose an event-based motion deblurring network consisting of a Multi-Scale Event-based Double Integral (MS-EDI) module designed from temporal residual guidance, and a Blur-Aware Filter Prediction (BAFP) module to conduct filter processing directed by spatial blur indication. The network, after incorporating spatial residual guidance, has significantly enhanced its generalization ability, surpassing the best-performing image-based and event-based methods on both synthetic, semi-synthetic, and real-world datasets. In addition, our method can be extended to blurry image super-resolution and achieves impressive performance. Our code is available at:https://github.com/ChenYichen9527/MBNetnow.
Chushu Zhang, Wei An 0003, Longguang Wang, Qiang Ling 0002
IEEE Trans. Circuits Syst. Video Technol.4
2025 ARBiBench: Benchmarking and Analyzing Adversarial Robustness of Binarized Convolutional Neural Networks
abstract
Binarized convolutional neural networks (BCNNs), which restrict the weights and activations of the model to +1 or −1, provide notable reductions in memory requirements and enhanced model inference speed during deployment. Current research on BCNNs primarily revolves around addressing the performance degradation resulting from binarization. However, the investigation of the effects of extreme discretization on the robustness of BCNNs has been largely overlooked, despite its critical relevance to real-world applications. To this end, we propose ARBiBench, a comprehensive benchmark for evaluating the adversarial robustness of BCNNs in the image classification task. The key contributions of ARBiBench include: 1) systematically evaluating the robustness of seven influential BCNN methods across various architectures; 2) rigorous validation of diverse adversarial attack methods; and 3) novel empirical findings showing that BCNNs exhibit weaker robustness than full-precision networks on small datasets but surprisingly stronger robustness on large-scale datasets. Leveraging Information Bottleneck theory, we further demonstrate how data scale and model capacity collectively determine BCNNs’ adversarial robustness. These findings not only challenge conventional assumptions about BCNN security, but also provide new insights for developing robust yet efficient neural network architectures.
Li Liu 0002, Bowen Peng, Zhen Liu 0004, Longguang Wang, Yingmei Wei
IEEE Trans. Inf. Forensics Secur.6
2025 Motion and Appearance Decoupling Representation for Event Cameras
abstract
Event cameras, with high temporal resolution and high dynamic range, have shown great potential under extreme scenarios such as high-speed movement and low illumination. However, previous event representation methods typically aggregate event data into a single dense tensor, often overlooking the dynamic changes of events within a given time unit. This limitation can introduce historical artifacts and semantic inconsistencies, ultimately degrading model performance. Inspired by human visual prior, we propose a motion and appearance decoupling (MAD) event representation to disentangle the mixed spatial-temporal event tensor into two independent branches. This bio-inspired design helps the network extract discriminative temporal (i.e., motion) and spatial (i.e., appearance) information, thus reducing the network's learning burden toward complex high-level interpretation tasks. In our method, the event motion guided attention module (EMGA) is designed to achieve temporal and spatial feature interaction and fusion sequentially. Based on EMGA, three specially designed decoder heads are proposed for several representative event-based tasks (i.e., object detection, semantic segmentation, and human pose estimation). Experimental results demonstrate that our method achieves state-of-the-art performance on the above three tasks, which reveals that our method is an easy-to-implement replacement for currently event-based methods. Our code is available at: https://github.com/ChenYichen9527/MAD-representation.
Boyang Li 0007, Yingqian Wang 0002, Xinyi Ying, Longguang Wang, Chushu Zhang, Yulan Guo, Wei An 0003
IEEE Trans. Image Process.5
2025 CompletionMamba: Taming State Space Model for Point Cloud Completion
abstract
Point cloud completion aims to reconstruct complete 3D shapes from partial scans. The long-range dependencies between points and shape perception are crucial for this task. While Transformers are effective due to their global processing ability, the quadratic complexity of their attention mechanism makes them unsuitable for long sequences when computational resources are constrained. As an alternative, State Space Models (SSMs) provide a memory-efficient solution for handling long-range dependencies, yet applying them directly to unordered point clouds presents challenges because of their intrinsic causality requirements. Existing methods attempt to address this by sorting points along a single axis. This, however, often overlooks complex causal relationships in 3D space since adjacency relationships based on Euclidean distance between points in the 3D space may not be preserved by this linear arrangement. To overcome this issue, we introduce CompletionMamba, a novel SSM-based network designed to harness SSMs for capturing both global and local dependencies within a point cloud. Initially, the input point cloud is causally structured by rearranging its coordinates. Then, a local SSM framework is proposed that defines neighborhood spaces around each point based on Euclidean distance, enhancing the causal structure. Although local SSM enhances relationships in short and long distance sequences, it still lacks full shape modeling of point cloud. To address this, we propose a novel shape-aware Mamba by integrating the shape code of each 3D shape into the model, enabling shape information propagation to all points. Our experiments show that CompletionMamba achieves state-of-the-art performance on both the MVP and PCN datasets.
Zhiheng Fu, Longguang Wang, Lian Xu, Hamid Laga, Yulan Guo, Farid Boussaïd, Mohammed Bennamoun
IEEE Trans. Image Process.3
2025 ADStereo: Efficient Stereo Matching With Adaptive Downsampling and Disparity Alignment
abstract
The balance between accuracy and computational efficiency is crucial for the applications of deep learning-based stereo matching algorithms in real-world scenarios. Since matching cost aggregation is usually the most computationally expensive component, a common practice is to construct cost volumes at a low resolution for aggregation and then directly regress a high-resolution disparity map. However, current solutions often suffer from limitations such as the loss of discriminative features caused by downsampling operations that treat all pixels equally, and spatial misalignment resulting from repeated downsampling and upsampling. To overcome these challenges, this paper presents two sampling strategies: the Adaptive Downsampling Module (ADM) and the Disparity Alignment Module (DAM), to prioritize real-time inference while ensuring accuracy. The ADM leverages local features to learn adaptive weights, enabling more effective downsampling while preserving crucial structure information. On the other hand, the DAM employs a learnable interpolation strategy to predict transformation offsets of pixels, thereby mitigating the spatial misalignment issue. Building upon these modules, we introduce ADStereo, a real-time yet accurate network that achieves highly competitive performance on multiple public benchmarks. Specifically, our ADStereo runs over faster than the current state-of-the-art CREStereo (0.054s vs. ) under the same hardware while achieving comparable accuracy (1.82% vs. 1.69%) on the KITTI stereo 2015 benchmark. The codes are available at: https://github.com/cocowy1/ADStereo.
Yun Wang 0053, Kunhong Li 0001, Longguang Wang, Junjie Hu 0003, Dapeng Oliver Wu, Yulan Guo
IEEE Trans. Image Process.3
2025 Enhancing Event-Based Video Reconstruction With Bidirectional Temporal Information
abstract
Event-based video reconstruction has emerged as an appealing research direction to break through the limitations of traditional cameras to better record dynamic scenes. Most existing methods reconstruct each frame from its corresponding event subset in chronological order. Since the temporal information contained in the whole event sequence is not fully exploited, these methods suffer inferior reconstruction quality. In this paper, we propose to enhance event-based video reconstruction by leveraging the bidirectional temporal information in event sequences. The proposed model processes event sequences in a bidirectional fashion, allowing for exploiting bidirectional information in the whole sequence. Furthermore, a transformer-based temporal information fusion module is introduced to aggregate long-range information in both temporal and spatial dimensions. Additionally, we propose a new dataset for the event-based video reconstruction task which contains a variety of objects and movement patterns. Extensive experiments demonstrate that the proposed model outperforms existing state-of-the-art event-based video reconstruction methods both quantitatively and qualitatively.
Pinghai Gao, Longguang Wang, Sheng Ao, Ye Zhang 0037, Yulan Guo
IEEE Trans. Multim.2
2025 Adaptive Sparse Memory Networks for Efficient and Robust Video Object Segmentation
abstract
Recently, memory-based networks have achieved promising performance for video object segmentation (VOS). However, existing methods still suffer from unsatisfactory segmentation accuracy and inferior efficiency. The reasons are mainly twofold: 1) during memory construction, the inflexible memory storage mechanism results in a weak discriminative ability for similar appearances in complex scenarios, leading to video-level temporal redundancy, and 2) during memory reading, matching robustness and memory retrieval accuracy decrease as the number of video frames increases. To address these challenges, we propose an adaptive sparse memory network (ASM) that efficiently and effectively performs VOS by sparsely leveraging previous guidance while attending to key information. Specifically, we design an adaptive sparse memory constructor (ASMC) to adaptively memorize informative past frames according to dynamic temporal changes in video frames. Furthermore, we introduce an attentive local memory reader (ALMR) to quickly retrieve relevant information using a subset of memory, thereby reducing frame-level redundant computation and noise in a simpler and more convenient manner. To prevent key features from being discarded by the subset of memory, we further propose a novel attentive local feature aggregation (ALFA) module, which preserves useful cues by selectively aggregating discriminative spatial dependence from adjacent frames, thereby effectively increasing the receptive field of each memory frame. Extensive experiments demonstrate that our model achieves state-of-the-art performance with real-time speed on six popular VOS benchmarks. Furthermore, our ASM can be applied to existing memory-based methods as generic plugins to achieve significant performance improvements. More importantly, our method exhibits robustness in handling sparse videos with low frame rates.
Jisheng Dang, Huicheng Zheng, Xiaohao Xu, Longguang Wang, Qingyong Hu, Yulan Guo
IEEE Trans. Neural Networks Learn. Syst.4
2025 Real-World Light Field Image Super-Resolution Via Degradation Modulation
abstract
Recent years have witnessed the great advances of deep neural networks (DNNs) in light field (LF) image super-resolution (SR). However, existing DNN-based LF image SR methods are developed on a single fixed degradation (e.g., bicubic downsampling), and thus cannot be applied to super-resolve real LF images with diverse degradation. In this article, we propose a simple yet effective method for real-world LF image SR. In our method, a practical LF degradation model is developed to formulate the degradation process of real LF images. Then, a convolutional neural network is designed to incorporate the degradation prior into the SR process. By training on LF images using our formulated degradation, our network can learn to modulate different degradation while incorporating both spatial and angular information in LF images. Extensive experiments on both synthetically degraded and real-world LF images demonstrate the effectiveness of our method. Compared with existing state-of-the-art single and LF image SR methods, our method achieves superior SR performance under a wide range of degradation, and generalizes better to real LF images. Codes and models are available at https://yingqianwang.github.io/LF-DMnet/.
Yingqian Wang 0002, Zhengyu Liang, Longguang Wang, Jun-Gang Yang, Wei An 0003, Yulan Guo
IEEE Trans. Neural Networks Learn. Syst.3
2024 Pluggable Style Representation Learning for Multi-style Transfer
Hongda Liu 0001, Longguang Wang, Weijun Guan, Ye Zhang 0037, Yulan Guo
ACCV (6)2
2024 LoS: Local Structure-Guided Stereo Matching
abstract
Estimating disparities in challenging areas is difficult and limits the performance of stereo matching models. In this paper, we exploit local structure information (LSI) to better handle these areas. Specifically, our LSI comprises a series of key elements, including the slant plane (parameterised by disparity gradients), disparity offset details and neighbouring relations. This LSI empowers our method to effectively handle intricate structures, including object boundaries and curved surfaces. We bootstrap the LSI from monocular depth and subsequently refine it to bet-ter capture the underlying scene geometry constraints in an iterative manner. Building upon the LSI, we introduce the Local Structure-Guided Propagation (LSGP), which enhances the disparity initialization, optimization, and refinement processes. By combining LSGP with a Gated Re-current Unit (GRU), we present our novel stereo matching method, referred to as Local Structure-guided stereo matching (LoS). Remarkably, LoS achieves top-ranking results on four widely recognized public benchmark datasets (ETH3D, Middlebury, KITTI 15 & 12) and robust vision challenge, demonstrating the superior capabilities of our model.
Kunhong Li 0001, Longguang Wang, Ye Zhang 0037, Shunbo Zhou, Yulan Guo
CVPR2
2024 Learning Coupled Dictionaries from Unpaired Data for Image Super-Resolution
abstract
The difficulty of acquiring high-resolution (HR) and low-resolution (LR) image pairs in real scenarios limits the performance of existing learning-based image super-resolution (SR) methods in the real world. To conduct training on real-world unpaired data, current methods focus on synthesizing pseudo LR images to associate unpaired images. However, the realness and diversity of pseudo LR images are vulnerable due to the large image space. In this paper, we cir-cumvent the difficulty of image generation and propose an alternative to build the connection between unpaired images in a compact proxy space. Specifically, we first construct coupled HR and LR dictionaries, and then encode HR and LR images into a common latent code space using these dictionaries. In addition, we develop an autoencoder-based framework to couple these dictionaries during optimization by reconstructing input HR and LR images. The coupled dictionaries enable our method to employ a shal-low network architecture with only 18 layers to achieve efficient image SR. Extensive experiments show that our method (DictSR) can effectively model the LR-to-HR mapping in coupled dictionaries and produces state-of-the-art performance on benchmark datasets.
Longguang Wang, Juncheng Li 0003, Yingqian Wang 0002, Qingyong Hu, Yulan Guo
CVPR1
2024 AEDNet: Adaptive Embedding and Multiview-Aware Disentanglement for Point Cloud Completion
Zhiheng Fu, Longguang Wang, Lian Xu, Zhiyong Wang 0001, Hamid Laga, Yulan Guo, Farid Boussaïd, Mohammed Bennamoun
ECCV (11)2
2024 Distractor-Free Novel View Synthesis via Exploiting Memorization Effect in Optimization
Kunhong Li 0001, Minglin Chen, Longguang Wang, Shunbo Zhou, Yulan Guo
ECCV (54)4
2024 Learning Representations from Foundation Models for Domain Generalized Stereo Matching
Longguang Wang, Kunhong Li 0001, Yun Wang 0053, Yulan Guo
ECCV (42)2
2024 ACRF: Compressing Explicit Neural Radiance Fields via Attribute Compression
abstract
In this work, we study the problem of explicit NeRF compression. Through analyzing recent explicit NeRF models, we reformulate the task of explicit NeRF compression as 3D data compression. We further introduce our NeRF compression framework, Attributed Compression of Radiance Field (ACRF), which focuses on the compression of the explicit neural 3D representation. The neural 3D structure is pruned and converted to points with features, which are further encoded using importance-guided feature encoding. Furthermore, we employ an importance-prioritized entropy model to estimate the probability distribution of transform coefficients, which are then entropy coded with an arithmetic coder using the predicted distribution. Within this framework, we present two models, ACRF and ACRF-F, to strike a balance between compression performance and encoding time budget. Our experiments, which include both synthetic and real-world datasets such as Synthetic-NeRF and Tanks&Temples, demonstrate the superior performance of our proposed algorithm.
Guangchi Fang, Qingyong Hu, Longguang Wang, Yulan Guo
ICLR3
2024 Mask-ControlNet: Higher-Quality Image Generation with an Additional Mask Prompt
Zhiqi Huang 0005, Hui Xiong 0001, Haoyu Wang 0003, Longguang Wang
ICPR (6)4
2024 Tangram-Splatting: Optimizing 3D Gaussian Splatting Through Tangram-inspired Shape Priors
abstract
With the growth of VR and AR industry, 3D reconstruction has become a more and more important topic in multimedia. Although 3D Gaussian Splatting achieves state-of-the-art in 3D Reconstruction, a large number of Gaussians are needed to fit a 3D scene due to the Gibbs Phenomenon. The pursuit of compressing 3D Gaussian Splatting and reducing memory overhead has long been a focal point. Embarking on this trajectory, our study delves into this domain, aiming to mitigate these challenges. Inspired by the tangram, a Chinese ancient puzzle, we introduce a novel methodology (Tangram-Splatting) that leverages shape priors to optimize 3D scene fitting. Central to our approach is a pioneering technique that diversifies Gaussian function types while preserving algorithmic efficiency. Through exhaustive experimentation, we demonstrate that our method achieves a remarkable average reduction of 62.4% in memory consumption used to store optimized parameters and decreases the training time by at least 10 minutes, with only marginal sacrifices in PSNR performance, typically under 0.3 dB, and our algorithm is even better on some datasets. This reduction in memory burden is of paramount significance for real-world applications, mitigating the substantial memory footprint and transmission burden traditionally associated with such algorithms. Our algorithm underscores the profound potential of Tangram-Splatting in advancing multimedia applications.
Yi Wang 0095, Ningze Zhong, Minglin Chen, Longguang Wang, Yulan Guo
ACM Multimedia4
2024 Guest Editorial: Advanced image restoration and enhancement in the wild
abstract
Image restoration and enhancement has always been a fundamental task in computer vision and is widely used in numerous applications, such as surveillance imaging, remote sensing, and medical imaging. In recent years, remarkable progress has been witnessed with deep learning techniques. Despite the promising performance achieved on synthetic data, compelling research challenges remain to be addressed in the wild. These include: (i) degradation models for low-quality images in the real world are complicated and unknown, (ii) paired low-quality and high-quality data are difficult to acquire in the real world, and a large quantity of real data are provided in an unpaired form, (iii) it is challenging to incorporate cross-modal information provided by advanced imaging techniques (e.g. RGB-D camera) for image restoration, (iv) real-time inference on edge devices is important for image restoration and enhancement methods, and (v) it is difficult to provide the confidence or performance bounds of a learning-based method on different images/regions. This special issue invites original contributions in datasets, innovative architectures, and training methods for image restoration and enhancement to address these and other challenges. In this Special Issue, we have received 17 papers, of which 8 papers underwent the peer review process, while the rest were desk-rejected. Among these reviewed papers, 5 papers have been accepted and 3 papers have been rejected as they did not meet the criteria of IET Computer Vision. Thus, the overall submissions were of high quality, which marks the success of this Special Issue. The five eventually accepted papers can be clustered into two categories, namely video reconstruction and image super-resolution. The first category of papers aims at reconstructing high-quality videos. The papers in this category are of Zhang et al., Gu et al., and Xu et al. The second category of papers studies the task of image super-resolution. The papers in this category are of Dou et al. and Yang et al. A brief presentation of each of the paper in this special issue is as follows. Zhang et al. propose a point-image fusion network for event-based frame interpolation. Temporal information in event streams plays a critical role in this task as it provides temporal context cues complementary to images. Previous approaches commonly transform the unstructured event data to structured data formats through voxelisation and then employ advanced CNNs to extract temporal information. However, the voxelisation operation inevitably leads to information loss and introduces redundant computation. To address these limitations, the proposed method directly extracts temporal information from the events at the point level without relying on any voxelisation operation. Afterwards, a fusion module is adopted to aggregate complementary cues from both points and images for frame interpolation. Experiments on both synthetic and real-world datasets show that their method produces state-of-the-art accuracy with high efficiency. Gu et al. develop a temporal shift reconstruction network for compressive video sensing. To exploit the temporal cues between adjacent frames during the reconstruction of videos, most previous approaches commonly preform alignment between initial reconstructions. However, the estimated motions are usually too coarse to provide accurate temporal information. To remedy this, the proposed network employs stacked temporal shift reconstruction blocks to enhance the initial reconstruction progressively. Within each block, an efficient temporal shift operation is used to capture temporal structures in addition to computational overheads. Then, a bidirectional alignment module is adopted to capture the temporal dependencies in a video sequence. Different from previous methods that only extract supplementary information from the key frames, the proposed alignment module can receive temporal information from the whole video sequence via bidirectional propagations. Experiments demonstrate the superior performance of the proposed method. Qu et al. propose a lightweight video frame interpolation network with a three-scale encoding-decoding structure. Specifically, multi-scale motion information is first extracted from the input video. Then, recurrent convolutional layers are adopted to refine the resultant features. Afterwards, the resultant features are aggregated to generate high-quality interpolated frames. Experimental results on the CelebA and Helen datasets show that the proposed method outperforms state-of-the-art methods while using fewer parameters. Dou et al. introduce a decoder structure-guided CNN-Transformer network for face super-resolution. Most previous approaches follow a multi-task learning paradigm to perform landmark detection while super-resolving the low-resolution images. However, these methods require additional annotation cost, and the extracted facial prior structures are usually of low quality. To address these issues, the proposed network employs a global-local feature extraction unit to extract the global structure while capturing local texture details. In addition, a multi-state fusion module is incorporated to aggregate embeddings from different stages. Experiments show that the proposed method surpasses previous approaches by notable margins. Yang et al. study the problem of blind super-resolution and propose a method to exploit degradation information through degradation representation learning. Specifically, a generative adversarial network is employed to model the degradation process from HR images to LR images and constrain the data distribution of the synthetic LR images. Then, the learnt representation is adopted to super-resolve the input low-resolution images using a transformer-based SR network. Experiments on both synthetic and real-world datasets demonstrate the effectiveness and superiority of the proposed method. Longguang Wang received his BE and PhD degrees from Shandong University and National University of Defence Technology (NUDT) in 2015 and 2022, respectively. He is currently an assistant professor with Aviation University of Air Force. He authored more than 40 peer-reviewed journals and conference publications (including TPAMI, TIP, CVPR, ICCV, and ECCV). He has organised three workshops at CVPR 2022 and 2023. His research interests include low-level vision and 3D vision, particularly on image restoration, image enhancement, image generation, depth estimation, point cloud understanding, and network acceleration. He received the CSIG Excellent Doctoral Dissertation Nomination Award in 2022 (17 nationwide). Juncheng Li received the Ph.D. degree from the School of Computer Science and Technology, East China Normal University, in 2021. He also worked as a Postdoctoral Fellow at the Center for Mathematical Artificial Intelligence, The Chinese University of Hong Kong. He is currently an assistant professor with Shanghai University. His research interests include artificial intelligence and its applications to computer vision (e.g. image segmentation) and image processing (e.g. image super-resolution, image denoising, and image dehazing). He has published more than 25 papers in top journals and conferences, including TIP, TNNLS, TMM, ECCV, ICCV, AAAI, ACMMM, and IJCAI. He also received several premium awards, including the Shanghai Outstanding Ph.D. Graduates, CUHK Research Fellowship Scheme, and the winner of 2019 ICCV-AIM. Naoto Yokoya received the M.Eng. and Ph.D. degrees from the Department of Aeronautics and Astronautics, The University of Tokyo, Tokyo, Japan, in 2010 and 2013, respectively. From 2013 to 2017, he was an assistant professor with The University of Tokyo. From 2015 to 2017, he was an Alexander von Humboldt Fellow, working at the German Aerospace Center, Oberpfaffenhofen, Germany and at the Technical University of Munich, Munich, Germany. He is currently a lecturer with The University of Tokyo and a unit leader with the RIKEN Center for Advanced Intelligence Project, Tokyo, where he leads the Geoinformatics Unit. His research interests include image processing, data fusion, and machine learning for understanding remote sensing images with applications to disaster management. Dr. Yokoya received the First Place in the 2017 IEEE Geoscience and Remote Sensing Society (GRSS) Data Fusion Contest organised by the IEEE Image Analysis and Data Fusion Technical Committee (IADF TC). From 2019 to 2021, he was the Chair and the Co-Chair (2017–2019) of the IEEE GRSS IADF TC. Since 2018, he has been an associate editor of IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing (JSTARS). Radu Timofte received his Ph.D. degree in Electrical Engineering from the KU Leuven, Belgium, in 2013. Currently, he is a professor and holds the Chair for Computer Science IV (Computer Vision) at the University of Wurzburg, Germany. Also, he is a lecturer and a group leader at ETH Zurich, Switzerland. He is a member of the editorial board of top journals such as IEEE TPAMI, Elsevier's CVIU and NEUCOM, and SIAM's SIIMS. He regularly serves as an area chair and as a reviewer for top conferences such as CVPR, ICCV, IJCAI, and ECCV. His work received several awards. Radu Timofte is the 2022 awardee of the Alexander von Humboldt Professorship for Artificial Intelligence. He is a co-founder of Merantix and a co-organiser of NTIRE, CLIC, AIM, Mobile AI, and PIRM workshops and challenges. His current research interests include deep learning, mobile AI, visual tracking, computational photography, and image/video compression, restoration, enhancement, and manipulation. Yulan Guo received the B.E. and Ph.D. degrees from NUDT in 2008 and 2015, respectively. He has authored over 100 articles at highly referred journals and conferences. His current research interests focus on 3D vision, particularly on 3D feature learning, 3D modelling, 3D object recognition, and scene understanding. He served as an associate editor for IEEE Transactions on Image Processing, IET Computer Vision, IET Image Processing, and Computers & Graphics. He also served as an area chair for CVPR 2023/2021, ICCV 2021, and ACM Multimedia 2021. He organised several tutorials, workshops, and challenges in prestigious conferences, such as CVPR 2016, CVPR 2019, ICCV 2021, 3DV 2021, CVPR 2022, ICPR 2022, and ECCV 2022. He is a senior member of IEEE and ACM. Data sharing is not applicable to this article as no new data were created or analysed in this study. Longguang Wang received his B.E. and Ph.D. degrees from Shandong University and National University of Defense Technology in 2015 and 2022, respectively. He is currently an assistant professor with Aviation University of Air Force. He authored more than 60 peer reviewed journal and conference publications (including TPAMI, TIP, CVPR, ICCV and ECCV). He served as a reviewer for more than 10 international journals (including TPAMI and TIP) and conferences (including CVPR, ICCV and ECCV). He has organized workshops at CVPR 2022/2023/2024. His research interests include low-level vision and 3D vision, particularly on image restoration, image generation, point cloud understanding, and network acceleration. His received the CSIG Excellent Doctoral Dissertation Nomination Award in 2022 (17 nationalwide). Juncheng Li received the Ph.D. degree from the School of Computer Science and Technology, East China Normal University, in 2021. He also worked as a Postdoctoral Fellow at the Center for Mathematical Artificial Intelligence, The Chinese University of Hong Kong. He is currently an assistant professor with Shanghai University. His research interests include artificial intelligence and its applications to computer vision (e.g. image segmentation) and image processing (e.g. image super-resolution, image denoising, and image dehazing). He has published more than 25 papers in top journals and conferences, including TIP, TNNLS, TMM, ECCV, ICCV, AAAI, ACMMM and IJCAI. He also received several premium awards, including the Shanghai Outstanding Ph.D. Graduates, CUHK Research Fellowship Scheme, the winner of 2019 ICCV-AIM, etc. Meanwhile, he served as a reviewer for more than 20 international journals and conferences. Naoto Yokoya received the M.Eng. and Ph.D. degrees from the Department of Aeronautics and Astronautics, The University of Tokyo, Tokyo, Japan, in 2010 and 2013, respectively. From 2013 to 2017, he was an Assistant Professor with The University of Tokyo. From 2015 to 2017, he was an Alexander von Humboldt Fellow, working at the German Aerospace Center, Oberpfaffenhofen, Germany, and at the Technical University of Munich, Munich, Germany. He is currently a Lecturer with The University of Tokyo, and a Unit Leader with the RIKEN Center for Advanced Intelligence Project, Tokyo, where he leads the Geoinformatics Unit. His research interests include image processing, data fusion, and machine learning for understanding remote sensing images, with applications to disaster management. Dr. Yokoya received the First Place in the 2017 IEEE Geoscience and Remote Sensing Society (GRSS) Data Fusion Contest organized by the IEEE Image Analysis and Data Fusion Technical Committee (IADF TC). From 2019 to 2021, he was the Chair and the Co-Chair (2017–2019) of the IEEE GRSS IADF TC. Since 2018, he has been an Associate Editor of IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing (JSTARS). Radu Timofte received his Ph.D. degree in Electrical Engineering from the KU Leuven, Belgium, in 2013. Currently, he is a professor and holds the Chair for Computer Science IV (Computer Vision) at theUniversity of Wurzburg, Germany. He is a member of the editorial board of top journals such as IEEE TPAMI, Elsevier's CVIU and NEUCOM, and SIAM's SIIMS. He regularly serves as an area chair and as a reviewer for top conferences such as CVPR, ICCV, IJCAI and ECCV. His work received several awards. Radu Timofte is the 2022 awardee of an Alexandervon Humboldt Professorship for Artificial Intelligence. He is a co-founder of Merantix and a co-organizer of NTIRE, CLIC, AIM, Mobile AI and PIRM workshops and challenges. His current research interests include deep learning, mobile AI, visual tracking, computational photography, image/video compression, restoration, enhancement and manipulation. Yulan Guo received the B.E. and Ph.D. degrees from National University of Defense Technology (NUDT) in 2008 and 2015, respectively. He has authored over 100 articles at highly referred journals and conferences. His current research interests focus on 3D vision, particularly on 3D feature learning, 3D modeling, 3D object recognition, and scene understanding. He served as an associate editor for IEEE Transactions on Image Processing, IET Computer Vision, IET Image Processing, and Computers & Graphics. He also served as an area chair for CVPR 2023/2021, ICCV 2021, and ACM Multimedia 2021. He organized several tutorials, workshops, and challenges in prestigious conferences, such as CVPR 2016, CVPR 2019, ICCV 2021, 3DV 2021, CVPR 2022, ICPR 2022 and ECCV 2022. He is a Senior Member of IEEE and ACM.
Longguang Wang, Juncheng Li 0003, Naoto Yokoya, Radu Timofte, Yulan Guo
IET Comput. Vis.1
2024 Point Spatio-Temporal Pyramid Network for Point Cloud Video Understanding
abstract
The robustness to spatio-temporal sampling is significant for point cloud video understanding. Previous works overlook this issue and usually suffer notable performance drops when point densities and frame rates are changed. To remedy this, we propose a point spatio-temporal pyramid (PoST-Py) to improve the sampling robustness of point cloud video modeling. Specifically, we propose a pluggable PoST-Py to collect multi-scale feature maps from different layers of the backbone. Then, these features are integrated into a unified representation. This allows the model to capture multi-scale spatio-temporal information simultaneously. In addition, we employ the temporal cardinality difference to enhance the features to capture motion information. Extensive experiments show that PoST-Py achieves state-of-the-art performance, particularly with a notable improvement of over 2% under varying point sampling. This demonstrates the improved robustness of our method. The code is available athttps://github.com/JohnsonSign/PoST-Py.
Longguang Wang, Yulan Guo, Xi Zhou 0001
IEEE Signal Process. Lett.2
2024 Self-Supervised Multi-Frame Monocular Depth Estimation for Dynamic Scenes
abstract
Self-supervised multi-frame depth estimation outperforms single-frame approaches by utilizing not only appearance information, but also geometric information. A common practice for multi-frame methods is to employ feature-metric bundle adjustment (FBA) to refine depth map initialized from the single-frame prior. However, FBA cannot always provide effective residual updates due to unreliable matching costs, which are corrupted by thin texture, occlusion, and especially object motion. To tackle this problem, we propose a context-aware transformer (CAT) to refine the corrupted matching costs by leveraging the spatial context information. Specifically, the CAT adaptively aggregates matching costs according to the spatial affinity inferred from local appearance context, and produces reliable contextual costs for FBA. Moreover, we design a motion-aware regularization loss to provide supervision for regions with moving objects, making CAT competent for dynamic scenes. Extensive experiments and analyses on the KITTI and Cityscapes datasets demonstrate the effectiveness and superior generalization capability of our approach.
Guanghui Wu, Hao Liu 0061, Longguang Wang, Kunhong Li 0001, Yulan Guo, Zengping Chen
IEEE Trans. Circuits Syst. Video Technol.3
2024 Mixed-Precision Network Quantization for Infrared Small Target Segmentation
abstract
Network quantization is leveraged to reduce the model size, memory footprint, and computational cost of deep neural networks. It is achieved by representing float weights and activations with lower bit counterparts, which is essential for model deployment on resource-limited devices. However, due to the extremely small size of infrared small targets in the feature map, low-bit quantization could lead to huge information loss of small targets and thus causes severe segmentation performance degradation. To achieve low-bit quantization while maintaining the segmentation performance, we first study the quantization sensitivity of small target segmentation network and observe the sensitivity heterogeneity of different layers in the network. Specifically, feature maps in shallow layers and encoder subnetwork are more vulnerable to information loss caused by quantization as compared to deep layers and decoder subnetwork. Based on these observations, we are motivated to assign a different bitwidth for each block according to their quantization sensitivity. A simple yet effective symmetrically progressive decreasing mixed-precision quantization (SPMix-Q) method is proposed to achieve high-performance segmentation under low-bit quantization (i.e., 2.42 bits for weights and 3.82 bits for activations). The experimental results show that our SPMix-Q achieves comparable accuracy with only 1/13 model size, 1/4.6 memory footprint, and 1/29 computational cost to the full-precision counterparts. Compared with the homogeneous low-bit quantization methods, our method achieves much better performance in terms of intersection of union (IoU) on the benchmark datasets. Our mobile-system-on-a-chip (SOC) (e.g., Kyrin 980, Snapdragon 660, and Dimensity 800U) deployable android application package (APK) is available at:https://github.com/YeRen123455/SIRST-Quantization-Deployment.
Boyang Li 0007, Longguang Wang, Yingqian Wang 0002, Tianhao Wu 0014, Zaiping Lin, Wei An 0003, Yulan Guo
IEEE Trans. Geosci. Remote. Sens.2
2024 Deep Semantic Graph Matching for Large-Scale Outdoor Point Cloud Registration
abstract
Current point cloud registration methods are mainly based on local geometric information and usually ignore the semantic information contained in the scenes. In this paper, we treat the point cloud registration problem as a semantic instance matching and registration task, and propose a deep semantic graph matching method (DeepSGM) for large-scale outdoor point cloud registration. Firstly, the semantic categorical labels of 3D points are obtained using a semantic segmentation network. The adjacent points with the same category labels are then clustered together using the Euclidean clustering algorithm to obtain the semantic instances, which are represented by three kinds of attributes including spatial location information, semantic categorical information, and global geometric shape information. Secondly, the semantic adjacency graph is constructed based on the spatial adjacency relations of semantic instances. To fully explore the topological structures between semantic instances in the same scene and across different scenes, the spatial distribution features and the semantic categorical features are learned with graph convolutional networks, and the global geometric shape features are learned with a PointNet-like network. These three kinds of features are further enhanced with the self-attention and cross-attention mechanisms. Thirdly, the semantic instance matching is formulated as an optimal transport problem, and solved through an optimal matching layer. Finally, the geometric transformation matrix between two point clouds is first estimated by the SVD algorithm and then refined by the ICP algorithm. Experimental results conducted on the KITTI Odometry dataset demonstrate that the proposed method improves the registration performance and outperforms various state-of-the-art methods.
Shaocong Liu, Tao Wang 0075, Yan Zhang 0159, Ruqin Zhou, Li Li 0100, Chenguang Dai, Longguang Wang, Hanyun Wang
IEEE Trans. Geosci. Remote. Sens.8
2024 Learning Spherical Radiance Field for Efficient 360° Unbounded Novel View Synthesis
abstract
Novel view synthesis aims at rendering any posed images from sparse observations of the scene. Recently, neural radiance fields (NeRF) have demonstrated their effectiveness in synthesizing novel views of a bounded scene. However, most existing methods cannot be directly extended to 360° unbounded scenes where the camera orientations and scene depths are unconstrained with large variations. In this paper, we present a spherical radiance field (SRF) for efficient novel view synthesis in 360° unbounded scenes. Specifically, we represent a 3D scene as multiple concentric spheres with different radii. In particular, each sphere encodes its corresponding layered scene into implicit representations and is parameterized with an equirectangular projection image. A shallow multi-layer perceptron (MLP) is then used to infer the density and color from these sphere representations for volume rendering. Moreover, an occupancy grid is introduced to cache the density field and guide the ray sampling, which accelerates the training and rendering procedures by reducing the number of samples along the ray. Experiments show that our method can well fit 360° unbounded scenes and produces state-of-the-art results on three benchmark datasets with less than 30 minutes of training time on a 3090 GPU, surpassing Mip-NeRF 360 with a 400× speedup. In addition, our method achieves competitive performance in terms of both accuracy and efficiency on a bounded dataset. Project page: https://minglin-chen.github.io/SphericalRF.
Minglin Chen, Longguang Wang, Yinjie Lei, Zilong Dong, Yulan Guo
IEEE Trans. Image Process.2
2024 Beyond Appearance: Multi-Frame Spatio-Temporal Context Memory Networks for Efficient and Robust Video Object Segmentation
abstract
Current video object segmentation approaches primarily rely on frame-wise appearance information to perform matching. Despite significant progress, reliable matching becomes challenging due to rapid changes of the object's appearance over time. Moreover, previous matching mechanisms suffer from redundant computation and noise interference as the number of accumulated frames increases. In this paper, we introduce a multi-frame spatio-temporal context memory (STCM) network to exploit discriminative spatio-temporal cues in multiple adjacent frames by utilizing a multi-frame context interaction module (MCI) for memory construction. Based on the proposed MCI module, a sparse group memory reader is developed to enable efficient sparse matching during memory reading. Our proposed method is generic and achieves state-of-the-art performance with real-time speed on benchmark datasets such as DAVIS and YouTube-VOS. In addition, our model exhibits robustness to sparse videos with low frame rates.
Jisheng Dang, Huicheng Zheng, Xiaohao Xu, Longguang Wang, Yulan Guo
IEEE Trans. Image Process.4
2024 Cost Volume Aggregation in Stereo Matching Revisited: A Disparity Classification Perspective
abstract
Cost aggregation plays a critical role in existing stereo matching methods. In this paper, we revisit cost aggregation in stereo matching from disparity classification and propose a generic yet efficient Disparity Context Aggregation (DCA) module to improve the performance of CNN-based methods. Our approach is based on an insight that a coarse disparity class prior is beneficial to disparity regression. To obtain such a prior, we first classify pixels in an image into several disparity classes and treat pixels within the same class as homogeneous regions. We then generate homogeneous region representations and incorporate these representations into the cost volume to suppress irrelevant information while enhancing the matching ability for cost aggregation. With the help of homogeneous region representations, efficient and informative cost aggregation can be achieved with only a shallow 3D CNN. Our DCA module is fully-differentiable and well-compatible with different network architectures, which can be seamlessly plugged into existing networks to improve performance with small additional overheads. It is demonstrated that our DCA module can effectively exploit disparity class priors to improve the performance of cost aggregation. Based on our DCA, we design a highly accurate network named DCANet, which achieves state-of-the-art performance on several benchmarks.
Yun Wang 0053, Longguang Wang, Kunhong Li 0001, Dapeng Oliver Wu, Yulan Guo
IEEE Trans. Image Process.2
2024 Temporo-Spatial Parallel Sparse Memory Networks for Efficient Video Object Segmentation
abstract
Memory-based networks have achieved tremendous success in video object segmentation. However, these methods still suffer from unfaithful segmentation and inferior efficiency under complicated video scenarios. The reasons are mainly threefold: 1) Weak perception of fast-moving targets due to individual frame memory patterns without capturing inter-frame motion; 2) Lack of discrimination to visually similar appearances due to the limited receptive field; 3) Redundant computation caused by matching with all memorized frames. To address these issues, we propose a Temporo-Spatial Parallel Sparse Memory network (TSPSM) for efficient video object segmentation. Our TSPSM constructs a temporal memory bank and a spatial memory bank in parallel to memorize complementary discriminative object cues. The temporal bank exploits discriminative temporal motion cues, while the spatial bank mines spatial context cues between adjacent frames with large receptive fields, thereby alleviating the ambiguity caused by similar instances and fast movements. To reduce redundant computation without sacrificing performance during the matching step, we further design a parallel sparse memory reader based on the constructed informative memory banks, which efficiently retrieves relevant temporal and spatial information in a parallel way. Experiments demonstrate that our TSPSM achieves state-of-the-art performance with real-time speed on DAVIS, and YouTube-VOS benchmarks. Furthermore, extensive experiments show that the proposed TSPMC module can be applied to existing methods as a generic plugin to significantly improve performance.
Jisheng Dang, Huicheng Zheng, Bimei Wang, Longguang Wang, Yulan Guo
IEEE Trans. Intell. Transp. Syst.4
2024 DDAug: Differentiable Data Augmentation for Weakly Supervised Semantic Segmentation
abstract
Weakly supervised semantic segmentation(WSSS) with image-level labels has witnessed promising advances with the help ofclass activation maps(CAM). However, CAM is always confined to small discriminative seed regions due to its simple classification loss guided training manner. To handle this problem, recent works introduced specifically designed regularizations and modules to expand the CAM seed regions, serving as the final segmentation masks. In this paper, we surprisingly find that the classification loss could suppress the gains from these regularization and modules in the late training phase, thereby limiting the further growth of CAM, which we call as theexplicit supervision disturb(ESD) issue. Interestingly, we find that specificdata augmentation(DA) operations (e.g., CutMix) can relieve such ESD issue, and the benefits introduced by different DA operations vary a lot. To maximize the benefits, we proposedifferentiable data augmentation(DDAug) to automatically search for the proper DA policy. Specifically, we design amulti-level search spaceto sequentially sample DA operations with different properties. Extensive experiments demonstrate that the proposed DDAug can alleviate the ESD issue and introduce consistent improvements to various popular WSSS methods, achieving the state-of-the-art performance on the MS COCO 2014 and PASCAL VOC 2012 datasets.
Boyang Li 0007, Fei Zhang 0016, Longguang Wang, Yingqian Wang 0002, Ting Liu 0017, Zaiping Lin, Wei An 0003, Yulan Guo
IEEE Trans. Multim.3
2024 Heterogeneous Graph Transformer for Multiple Tiny Object Tracking in RGB-T Videos
abstract
Tracking multiple tiny objects is highly challenging due to their weak appearance and limited features. Existing multi-object tracking algorithms generally focus on singlemodality scenes, and overlook the complementary characteristics of tiny objects captured by multiple remote sensors. To enhance tracking performance by integrating complementary information from multiple sources, we propose a novel framework called HGT-Track (Heterogeneous Graph Transformer based Multi-Tiny-Object Tracking). Specifically, we first employ a Transformer-based encoder to embed images from different modalities. Subsequently, we utilize Heterogeneous Graph Transformer to aggregate spatial and temporal information from multiple modalities to generate detection and tracking features. Additionally, we introduce a target re-detection module (ReDet) to ensure tracklet continuity by maintaining consistency across different modalities. Furthermore, this paper introduces the first benchmark VT-Tiny-MOT (Visible-Thermal Tiny MultiObject Tracking) for RGB-T fused multiple tiny object tracking. Extensive experiments are conducted on VT-Tiny-MOT, and the results have demonstrated the effectiveness of our method. Compared to other state-of-the-art methods, our method achieves better performance in terms of MOTA (Multiple-Object Tracking Accuracy) and ID-F1 score. The code and dataset will be made available at https://github.com/xuqingyu26/HGTMT
Longguang Wang, Weidong Sheng, Yingqian Wang 0002, Chao Ma 0014, Wei An 0003
IEEE Trans. Multim.2
2023 PointCMP: Contrastive Mask Prediction for Self-supervised Learning on Point Cloud Videos
abstract
Self-supervised learning can extract representations of good quality from solely unlabeled data, which is ap-pealing for point cloud videos due to their high labelling cost. In this paper, we propose a contrastive mask prediction (PointCMP) framework for self-supervised learning on point cloud videos. Specifically, our PointCMP employs a two-branch structure to achieve simultaneous learning of both local and global spatiotemporal information. On top of this two-branch structure, a mutual similarity based augmentation module is developed to synthesize hard samples at the feature level. By masking dominant tokens and erasing principal channels, we generate hard samples to facilitate learning representations with better discrimi-nation and generalization performance. Extensive experiments show that our PointCMP achieves the state-of-the-art performance on benchmark datasets and outperforms existing full-supervised counterparts. Transfer learning results demonstrate the superiority of the learned representations across different datasets and tasks.
Xiaoxiao Sheng, Longguang Wang, Yulan Guo, Xi Zhou 0001
CVPR3
2023 VAPCNet: Viewpoint-Aware 3D Point Cloud Completion
abstract
Most existing learning-based 3D point cloud completion methods ignore the fact that the completion process is highly coupled with the viewpoint of a partial scan. However, the various viewpoints of incompletely scanned objects in real-world applications are normally unknown and directly estimating the viewpoint of each incomplete object is usually time-consuming and leads to huge annotation cost. In this paper, we thus propose an unsupervised viewpoint representation learning scheme for 3D point cloud completion without explicit viewpoint estimation. To be specific, we learn abstract representations of partial scans to distinguish various viewpoints in the representation space rather than the explicit estimation in the 3D space. We also introduce a Viewpoint-Aware Point cloud Completion Network (VAPCNet) with flexible adaption to various viewpoints based on the learned representations. The proposed viewpoint representation learning scheme can extract discriminative representations to obtain accurate viewpoint information. Reported experiments on two popular public datasets show that our VAPCNet achieves state-of-the-art performance for the point cloud completion task. Source code is available at https://github.com/FZH92128/VAPCNet.
Zhiheng Fu, Longguang Wang, Lian Xu, Zhiyong Wang 0001, Hamid Laga, Yulan Guo, Farid Boussaïd, Mohammed Bennamoun
ICCV2
2023 Monte Carlo Linear Clustering with Single-Point Supervision is Enough for Infrared Small Target Detection
abstract
Single-frame infrared small target (SIRST) detection aims at separating small targets from clutter backgrounds on infrared images. Recently, deep learning based methods have achieved promising performance on SIRST detection, but at the cost of a large amount of training data with expensive pixel-level annotations. To reduce the annotation burden, we propose the first method to achieve SIRST detection with single-point supervision. The core idea of this work is to recover the per-pixel mask of each target from the given single point label by using clustering approaches, which looks simple but is indeed challenging since targets are always insalient and accompanied with background clutters. To handle this issue, we introduce randomness to the clustering process by adding noise to the input images, and then obtain much more reliable pseudo masks by averaging the clustered results. Thanks to this "Monte Carlo" clustering approach, our method can accurately recover pseudo masks and thus turn arbitrary fully supervised SIRST detection networks into weakly supervised ones with only single point annotation. Experiments on four datasets demonstrate that our method can be applied to existing SIRST detection networks to achieve comparable performance with their fully-supervised counterparts, which reveals that single-point supervision is strong enough for SIRST detection. Our code will be available at: https://github.com/YeRen123455/SIRST-Single-Point-Supervision.
Boyang Li 0007, Yingqian Wang 0002, Longguang Wang, Fei Zhang 0016, Ting Liu 0017, Zaiping Lin, Wei An 0003, Yulan Guo
ICCV3
2023 Learning Non-Local Spatial-Angular Correlation for Light Field Image Super-Resolution
abstract
Exploiting spatial-angular correlation is crucial to light field (LF) image super-resolution (SR), but is highly challenging due to its non-local property caused by the disparities among LF images. Although many deep neural networks (DNNs) have been developed for LF image SR and achieved continuously improved performance, existing methods cannot well leverage the long-range spatial-angular correlation and thus suffer a significant performance drop when handling scenes with large disparity variations. In this paper, we propose a simple yet effective method to learn the non-local spatial-angular correlation for LF image SR. In our method, we adopt the epipolar plane image (EPI) representation to project the 4D spatial-angular correlation onto multiple 2D EPI planes, and then develop a Transformer network with repetitive self-attention operations to learn the spatial-angular correlation by modeling the dependencies between each pair of EPI pixels. Our method can fully incorporate the information from all angular views while achieving a global receptive field along the epipolar line. We conduct extensive experiments with insightful visualizations to validate the effectiveness of our method. Comparative results on five public datasets show that our method not only achieves state-of-the-art SR performance but also performs robust to disparity variations. Code is publicly available at https://github.com/ZhengyuLiang24/EPIT.
Zhengyu Liang, Yingqian Wang 0002, Longguang Wang, Jun-Gang Yang, Shilin Zhou 0001, Yulan Guo
ICCV3
2023 Masked Spatio-Temporal Structure Prediction for Self-supervised Learning on Point Cloud Videos
abstract
Recently, the community has made tremendous progress in developing effective methods for point cloud video understanding that learn from massive amounts of labeled data. However, annotating point cloud videos is usually notoriously expensive. Moreover, training via one or only a few traditional tasks (e.g., classification) may be insufficient to learn subtle details of the spatio-temporal structure existing in point cloud videos. In this paper, we propose a Masked Spatio-Temporal Structure Prediction (MaST-Pre) method to capture the structure of point cloud videos without human annotations. MaST-Pre is based on spatio-temporal point-tube masking and consists of two self-supervised learning tasks. First, by reconstructing masked point tubes, our method is able to capture the appearance information of point cloud videos. Second, to learn motion, we propose a temporal cardinality difference prediction task that estimates the change in the number of points within a point tube. In this way, MaST-Pre is forced to model the spatial and temporal structure in point cloud videos. Extensive experiments on MSRAction-3D, NTU-RGBD, NvGesture, and SHREC’17 demonstrate the effectiveness of the proposed method. The code is available at https://github.com/JohnsonSign/MaST-Pre.
Xiaoxiao Sheng, Hehe Fan, Longguang Wang, Yulan Guo, Xi Zhou 0001
ICCV4
2023 Point Contrastive Prediction with Semantic Clustering for Self-Supervised Learning on Point Cloud Videos
abstract
We propose a unified point cloud video self-supervised learning framework for object-centric and scene-centric data. Previous methods commonly conduct representation learning at the clip or frame level and cannot well capture fine-grained semantics. Instead of contrasting the representations of clips or frames, in this paper, we propose a unified self-supervised framework by conducting contrastive learning at the point level. Moreover, we introduce a new pretext task by achieving semantic alignment of superpoints, which further facilitates the representations to capture semantic cues at multiple scales. In addition, due to the high redundancy in the temporal dimension of dynamic point clouds, directly conducting contrastive learning at the point level usually leads to massive undesired negatives and insufficient modeling of positive representations. To remedy this, we propose a selection strategy to retain proper negatives and make use of high-similarity samples from other instances as positive supplements. Extensive experiments show that our method outperforms supervised counterparts on a wide range of downstream tasks and demonstrates the superior transferability of the learned representations.
Xiaoxiao Sheng, Gang Xiao 0002, Longguang Wang, Yulan Guo, Hehe Fan
ICCV4
2023 Disentangling Light Fields for Super-Resolution and Disparity Estimation
abstract
Light field (LF) cameras record both intensity and directions of light rays, and encode 3D scenes into 4D LF images. Recently, many convolutional neural networks (CNNs) have been proposed for various LF image processing tasks. However, it is challenging for CNNs to effectively process LF images since the spatial and angular information are highly inter-twined with varying disparities. In this paper, we propose a generic mechanism to disentangle these coupled information for LF image processing. Specifically, we first design a class of domain-specific convolutions to disentangle LFs from different dimensions, and then leverage these disentangled features by designing task-specific modules. Our disentangling mechanism can well incorporate the LF structure prior and effectively handle 4D LF data. Based on the proposed mechanism, we develop three networks (i.e., DistgSSR, DistgASR and DistgDisp) for spatial super-resolution, angular super-resolution and disparity estimation. Experimental results show that our networks achieve state-of-the-art performance on all these three tasks, which demonstrates the effectiveness, efficiency, and generality of our disentangling mechanism. Project page: https://yingqianwang.github.io/DistgLF/.
Yingqian Wang 0002, Longguang Wang, Gaochang Wu, Jun-Gang Yang, Wei An 0003, Jingyi Yu 0001, Yulan Guo
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Lightweight Pixel Difference Networks for Efficient Visual Representation Learning
abstract
Recently, there have been tremendous efforts in developing lightweight Deep Neural Networks (DNNs) with satisfactory accuracy, which can enable the ubiquitous deployment of DNNs in edge devices. The core challenge of developing compact and efficient DNNs lies in how to balance the competing goals of achieving high accuracy and high efficiency. In this paper we propose two novel types of convolutions, dubbed Pixel Difference Convolution (PDC) and Binary PDC (Bi-PDC) which enjoy the following benefits: capturing higher-order local differential information, computationally efficient, and able to be integrated with existing DNNs. With PDC and Bi-PDC, we further present two lightweight deep networks named Pixel Difference Networks (PiDiNet) and Binary PiDiNet (Bi-PiDiNet) respectively to learn highly efficient yet more accurate representations for visual tasks including edge detection and object recognition. Extensive experiments on popular datasets (BSDS500, ImageNet, LFW, YTF, etc.) show that PiDiNet and Bi-PiDiNet achieve the best accuracy-efficiency trade-off. For edge detection, PiDiNet is the first network that can be trained without ImageNet, and can achieve the human-level performance on BSDS500 at 100 FPS and with 1 M parameters. For object recognition, among existing Binary DNNs, Bi-PiDiNet achieves the best accuracy and a nearly 2× reduction of computational cost on ResNet18.
Zhuo Su 0002, Longguang Wang, Hua Zhang 0008, Zhen Liu 0004, Matti Pietikäinen, Li Liu 0002
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Exploring Fine-Grained Sparsity in Convolutional Neural Networks for Efficient Inference
abstract
Neural networks contain considerable redundant computation, which drags down the inference efficiency and hinders the deployment on resource-limited devices. In this paper, we study the sparsity in convolutional neural networks and propose a generic sparse mask mechanism to improve the inference efficiency of networks. Specifically, sparse masks are learned in both data and channel dimensions to dynamically localize and skip redundant computation at a fine-grained level. Based on our sparse mask mechanism, we develop SMPointSeg, SMSR, and SMStereo for point cloud semantic segmentation, single image super-resolution, and stereo matching tasks, respectively. It is demonstrated that our sparse masks are well compatible to different model components and network architectures to accurately localize redundant computation, with computational cost being significantly reduced for practical speedup. Extensive experiments show that our SMPointSeg, SMSR, and SMStereo achieve state-of-the-art performance on benchmark datasets in terms of both accuracy and efficiency.
Longguang Wang, Yulan Guo, Yingqian Wang 0002, Xinyi Ying, Zaiping Lin, Wei An 0003
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Learning scalable dynamic filter in convolutional networks
Shuanglin Wu, Xinyi Ying, Longguang Wang, Jun-Gang Yang, Wei An 0003
Pattern Recognit. Lett.4
2023 Not All Patches Are Equal: Hierarchical Dataset Condensation for Single Image Super-Resolution
abstract
Although the performance of single image super-resolution (SR) has been significantly improved with deep neural networks, existing methods commonly require millions of iterations for training, which not only limits their training efficiency, but also causes considerable energy consumption. In this paper, we comprehensively study the redundancy of existing training datasets and reveal that not all patches are equal for SR network training. We observe that a large percentage of patches with low textures or similar textures lead to high computation costs but make low contributions to SR performance. Then, we propose a dataset condensation method to remove these redundant patches hierarchically. Extensive experiments demonstrate that our dataset condensation method can effectively reduce the redundancy of SR datasets with a 90% condensation rate on DIV2K. With our condensed dataset, baseline networks can achieve significant improvement in terms of training efficiency while maintaining competitive accuracy. Codes are available athttps://github.com/QingtangDing/DCSR.
Qingtang Ding, Zhengyu Liang, Longguang Wang, Yingqian Wang 0002, Jun-Gang Yang
IEEE Signal Process. Lett.3
2023 RepISD-Net: Learning Efficient Infrared Small-Target Detection Network via Structural Re-Parameterization
abstract
Infrared small target detection is a challenging task for deep learning-based methods because targets tend to disappear in the deep layers. To handle this problem, existing deep neural networks usually apply various dense and skip connections for feature maintenance. Although these well-designed networks have achieved good detection performance, the complex network structures reduce their efficiency. In this paper, we propose a simple yet efficient network (RepISD-Net) for infrared small target detection. The core of our RepISD-Net is to use different network architectures but equivalent model parameters for training and inference, respectively. Specifically, in the training phase, we design a parallel multi-branch edge compensation block (ECB) to enhance the local salient features and capture finer contour characteristic of infrared small targets. In the inference phase, the multi-branch topology structures are merged into a single branch with only cascaded 3×3 convolutions for fast inference. We conduct extensive experiments on several public datasets to validate the effectiveness of our method. Experimental results demonstrate that our RepISD-Net can achieve comparable or even better detection performance with significant acceleration in inference speed as compared to state-of-the-art infrared small target detection methods. Code is submitted for review and will be released upon acceptance.
Shuanglin Wu, Longguang Wang, Yingqian Wang 0002, Jun-Gang Yang, Wei An 0003
IEEE Trans. Geosci. Remote. Sens.3
2023 Dense Nested Attention Network for Infrared Small Target Detection
abstract
Single-frame infrared small target (SIRST) detection aims at separating small targets from clutter backgrounds. With the advances of deep learning, CNN-based methods have yielded promising results in generic object detection due to their powerful modeling capability. However, existing CNN-based methods cannot be directly applied to infrared small targets since pooling layers in their networks could lead to the loss of targets in deep layers. To handle this problem, we propose a dense nested attention network (DNA-Net) in this paper. Specifically, we design a dense nested interactive module (DNIM) to achieve progressive interaction among high-level and low-level features. With the repetitive interaction in DNIM, the information of infrared small targets in deep layers can be maintained. Based on DNIM, we further propose a cascaded channel and spatial attention module (CSAM) to adaptively enhance multi-level features. With our DNA-Net, contextual information of small targets can be well incorporated and fully exploited by repetitive fusion and enhancement. Moreover, we develop an infrared small target dataset (namely, NUDT-SIRST) and propose a set of evaluation metrics to conduct comprehensive performance evaluation. Experiments on both public and our self-developed datasets demonstrate the effectiveness of our method. Compared to other state-of-the-art methods, our method achieves better performance in terms of probability of detection (${P}_{d}$), false-alarm rate (${F}_{a}$), and intersection of union ($IoU$).
Boyang Li 0007, Longguang Wang, Yingqian Wang 0002, Zaiping Lin, Wei An 0003, Yulan Guo
IEEE Trans. Image Process.3
2023 CVCNet: Learning Cost Volume Compression for Efficient Stereo Matching
abstract
State-of-the-art deep learning based stereo matching algorithms usually rely on full-size cost volumes for highly accurate disparity estimation. The full-size cost volume processes all possible disparity candidates equally without considering their different matching uncertainties. Consequently, considerable redundant computation is involved on those candidates with very low matching uncertainties, making these methods difficult to be deployed in real-time applications. To tackle this problem, we propose CVCNet featuring an adaptive disparity range prediction module (ADR) and a disparity refinement module (DRM). The ADR adaptively predicts pixel-wise disparity range to discard the “unimportant” disparity candidates. It enables our network to obtain a compressed cost volume. Besides, the DRM improves disparity range prediction and refines the predicted disparity map. With the proposed modules, our CVCNet learns to build a compressed cost volume to achieve efficient disparity estimation. Experimental results on the KITTI and SceneFlow datasets show that our method achieves state-of-the-art performance, and runs at a significant order of magnitude faster speed than existing 3D CNN based methods. Particularly, our method ranks$\mathbf {1}\mathrm{st}$on the KITTI 2012 and KITTI 2015 benchmarks among all published methods with running time shorter than 100 ms.
Yulan Guo, Yun Wang 0053, Longguang Wang, Zi Wang 0008
IEEE Trans. Multim.3
2022 Occlusion-Aware Cost Constructor for Light Field Depth Estimation
abstract
Matching cost construction is a key step in light field (LF) depth estimation, but was rarely studied in the deep learning era. Recent deep learning-based LF depth estimation methods construct matching cost by sequentially shifting each sub-aperture image (SAI) with a series of pre-defined offsets, which is complex and time-consuming. In this paper, we propose a simple and fast cost constructor to construct matching cost for LF depth estimation. Our cost constructor is composed by a series of convolutions with specifically designed dilation rates. By applying our cost constructor to SAI arrays, pixels under predefined disparities can be integrated and matching cost can be constructed without using any shifting operation. More importantly, the proposed cost constructor is occlusion-aware and can handle occlusions by dynamically modulating pixels from different views. Based on the proposed cost constructor, we develop a deep network for LF depth estimation. Our network ranks first on the commonly used 4D LF benchmark in terms of the mean square error (MSE), and achieves a faster running time than other state-of-the-art methods.
Yingqian Wang 0002, Longguang Wang, Zhengyu Liang, Jun-Gang Yang, Wei An 0003, Yulan Guo
CVPR2
2022 Decoupling Makes Weakly Supervised Local Feature Better
abstract
Weakly supervised learning can help local feature methods to overcome the obstacle of acquiring a large-scale dataset with densely labeled correspondences. However, since weak supervision cannot distinguish the losses caused by the detection and description steps, directly conducting weakly supervised learning within a joint training describe-then-detect pipeline suffers limited performance. In this paper, we propose a decoupled training describe-then-detect pipeline tailored for weakly supervised local feature learning. Within our pipeline, the detection step is decoupled from the description step and postponed until discriminative and robust descriptors are learned. In addition, we introduce a line-to-window search strategy to explicitly use the camera pose information for better descriptor learning. Extensive experiments show that our method, namely PoSFeat (Camera Pose Supervised Feature), outperforms previous fully and weakly supervised methods and achieves state-of-the-art performance on a wide range of downstream task.
Kunhong Li 0001, Longguang Wang, Li Liu 0002, Qing Ran, Kai Xu 0004, Yulan Guo
CVPR2
2022 Learnable Lookup Table for Neural Network Quantization
abstract
Neural network quantization aims at reducing bit-widths of weights and activations for memory and computational efficiency. Since a linear quantizer (i.e., round(·) function) cannot well fit the bell-shaped distributions of weights and activations, many existing methods use predefined functions (e.g., exponential function) with learnable parameters to build the quantizer for joint optimization. However, these complicated quantizers introduce considerable computational overhead during inference since activation quantization should be conducted online. In this paper, we formulate the quantization process as a simple lookup operation and propose to learn lookup tables as quantizers. Specifically, we develop differentiable lookup tables and introduce several training strategies for optimization. Our lookup tables can be trained with the network in an end-to-end manner to fit the distributions in different layers and have very small additional computational cost. Comparison with previous methods show that quantized networks using our lookup tables achieve state-of-the-art performance on image classification, image super-resolution, and point cloud classification tasks.
Longguang Wang, Yingqian Wang 0002, Li Liu 0002, Wei An 0003, Yulan Guo
CVPR1
2022 Learning Mutual Modulation for Self-supervised Cross-Modal Super-Resolution
Naoto Yokoya, Longguang Wang, Tatsumi Uezato
ECCV (19)3
2022 Parallax Attention for Unsupervised Stereo Correspondence Learning
abstract
Stereo image pairs encode 3D scene cues into stereo correspondences between the left and right images. To exploit 3D cues within stereo images, recent CNN based methods commonly use cost volume techniques to capture stereo correspondence over large disparities. However, since disparities can vary significantly for stereo cameras with different baselines, focal lengths and resolutions, the fixed maximum disparity used in cost volume techniques hinders them to handle different stereo image pairs with large disparity variations. In this paper, we propose a generic parallax-attention mechanism (PAM) to capture stereo correspondence regardless of disparity variations. Our PAM integrates epipolar constraints with attention mechanism to calculate feature similarities along the epipolar line to capture stereo correspondence. Based on our PAM, we propose a parallax-attention stereo matching network (PASMnet) and a parallax-attention stereo image super-resolution network (PASSRnet) for stereo matching and stereo image super-resolution tasks. Moreover, we introduce a new and large-scale dataset named Flickr1024 for stereo image super-resolution. Experimental results show that our PAM is generic and can effectively learn stereo correspondence under large disparity variations in an unsupervised manner. Comparative results show that our PASMnet and PASSRnet achieve the state-of-the-art performance.
Longguang Wang, Yulan Guo, Yingqian Wang 0002, Zhengfa Liang, Zaiping Lin, Jun-Gang Yang, Wei An 0003
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 Light Field Image Super-Resolution With Transformers
abstract
Light field (LF) image super-resolution (SR) aims at reconstructing high-resolution LF images from their low-resolution counterparts. Although CNN-based methods have achieved remarkable performance in LF image SR, these methods cannot fully model the non-local properties of the 4D LF data. In this paper, we propose a simple but effective Transformer-based method for LF image SR. In our method, an angular Transformer is designed to incorporate complementary information among different views, and a spatial Transformer is developed to capture both local and long-range dependencies within each sub-aperture image. With the proposed angular and spatial Transformers, the beneficial information in an LF can be fully exploited and the SR performance is boosted. We validate the effectiveness of our angular and spatial Transformers through extensive ablation studies, and compare our method to recent state-of-the-art methods on five public LF datasets. Our method achieves superior SR performance with a small model size and low computational cost. Code is available at.1
Zhengyu Liang, Yingqian Wang 0002, Longguang Wang, Jun-Gang Yang, Shilin Zhou 0001
IEEE Signal Process. Lett.3
2022 Gated Recurrent Multiattention Network for VHR Remote Sensing Image Classification
abstract
With the advances of deep learning, many recent CNN-based methods have yielded promising results for image classification. In very high-resolution (VHR) remote sensing images, the contributions of different regions to image classification can vary significantly, because informative areas are generally limited and scattered throughout the whole image. Therefore, how to pay more attention to these informative areas and better incorporate them over long distances are two main challenges to be addressed. In this article, we propose a gated recurrent multiattention neural network (GRMA-Net) to address these problems. Because informative features generally occur at multiple stages in a network (i.e., local texture features at shallow layers and global profile features at deep layers), we use multilevel attention modules to focus on informative regions to extract more discriminative features. Then, these features are arranged as spatial sequences and fed into a deep-gated recurrent unit (GRU) to capture long-range dependency and contextual relationship. We evaluate our method on the UC Merced (UCM), Aerial Image dataset (AID), NWPU-RESISC (NWPU), and Optimal-31 (Optimal) datasets. Experimental results have demonstrated the superior performance of our method as compared to other state-of-the-art methods.
Boyang Li 0007, Yulan Guo, Jun-Gang Yang, Longguang Wang, Yingqian Wang 0002, Wei An 0003
IEEE Trans. Geosci. Remote. Sens.4
2021 Exploring Sparsity in Image Super-Resolution for Efficient Inference
abstract
Current CNN-based super-resolution (SR) methods process all locations equally with computational resources being uniformly assigned in space. However, since missing details in low-resolution (LR) images mainly exist in regions of edges and textures, less computational resources are required for those flat regions. Therefore, existing CNN-based methods involve redundant computation in flat regions, which increases their computational cost and limits their applications on mobile devices. In this paper, we explore the sparsity in image SR to improve inference efficiency of SR networks. Specifically, we develop a Sparse Mask SR (SMSR) network to learn sparse masks to prune redundant computation. Within our SMSR, spatial masks learn to identify "important" regions while channel masks learn to mark redundant channels in those "unimportant" regions. Consequently, redundant computation can be accurately localized and skipped while maintaining comparable performance. It is demonstrated that our SMSR achieves state-of-the-art performance with 41%/33%/27% FLOPs being reduced for ×2/3/4 SR. Code is available at: https://github.com/LongguangWang/SMSR.
Longguang Wang, Yingqian Wang 0002, Xinyi Ying, Zaiping Lin, Wei An 0003, Yulan Guo
CVPR1
2021 Unsupervised Degradation Representation Learning for Blind Super-Resolution
abstract
Most existing CNN-based super-resolution (SR) methods are developed based on an assumption that the degradation is fixed and known (e.g., bicubic downsampling). However, these methods suffer a severe performance drop when the real degradation is different from their assumption. To handle various unknown degradations in real-world applications, previous methods rely on degradation estimation to reconstruct the SR image. Nevertheless, degradation estimation methods are usually time-consuming and may lead to SR failure due to large estimation errors. In this paper, we propose an unsupervised degradation representation learning scheme for blind SR without explicit degradation estimation. Specifically, we learn abstract representations to distinguish various degradations in the representation space rather than explicit estimation in the pixel space. Moreover, we introduce a Degradation-Aware SR (DASR) network with flexible adaption to various degradations based on the learned representations. It is demonstrated that our degradation representation learning scheme can extract discriminative representations to obtain accurate degradation information. Experiments on both synthetic and real images show that our network achieves state-of-the-art performance for the blind SR task. Code is available at: https://github.com/LongguangWang/DASR.
Longguang Wang, Yingqian Wang 0002, Jun-Gang Yang, Wei An 0003, Yulan Guo
CVPR1
2021 Learning A Single Network for Scale-Arbitrary Super-Resolution
abstract
Recently, the performance of single image super-resolution (SR) has been significantly improved with powerful networks. However, these networks are developed for image SR with specific integer scale factors (e.g., ×2/3/4), and cannot handle non-integer and asymmetric SR. In this paper, we propose to learn a scale-arbitrary image SR network from scale-specific networks. Specifically, we develop a plug-in module for existing SR networks to perform scale-arbitrary SR, which consists of multiple scale-aware feature adaption blocks and a scale-aware upsampling layer. Moreover, conditional convolution is used in our plug-in module to generate dynamic scale-aware filters, which enables our network to adapt to arbitrary scale factors. Our plug-in module can be easily adapted to existing networks to realize scale-arbitrary SR with a single model. These networks plugged with our module can produce promising results for non-integer and asymmetric SR while maintaining state-of-the-art performance for SR with integer scale factors. Besides, the additional computational and memory cost of our module is very small.
Longguang Wang, Yingqian Wang 0002, Zaiping Lin, Jun-Gang Yang, Wei An 0003, Yulan Guo
ICCV1
2021 Deep Bilateral Learning for Stereo Image Super-Resolution
abstract
Bilateral filter has demonstrated its effectiveness in many traditional methods for image restoration tasks. In this letter, we incorporate the idea of bilateral grid processing in a CNN framework and propose a bilateral stereo super-resolution network (BSSRnet). Specifically, we use a parallax-attention module to incorporate information from left and right views to learn content-aware bilateral filters. Then, these bilateral filters are used to recover missing details at different spatial locations while preserving stereo consistency. Our network is fully differentiable and is robust to both content and disparity variations. Comparative results show that our BSSRnet achieves state-of-the-art performance on the Flickr1024, Middlebury, KITTI 2012 and KITTI 2015 datasets. Source code is available at.
Longguang Wang, Yingqian Wang 0002, Weidong Sheng, Xinpu Deng
IEEE Signal Process. Lett.2
2021 Remote Sensing Image Super-Resolution Using Second-Order Multi-Scale Networks
abstract
Remotely sensed images, especially in urban areas, have highly complex spatial distribution, since the ground objects have diverse ranges of sizes and shapes. This largely increases the difficulty of super-resolution (SR) tasks. Current deep convolutional neural network (CNN)-based SR methods often show limited performance when coping with complicated images. This article develops a second-order multi-scale super-resolution network (SMSR) to explore reconstruction tasks for difficult cases. Specifically, we propose a single-path feature reuse which cleverly captures multi-scale feature information through aggregating the features learned at different depths of a single path. Further, we present a second-order learning mechanism, which double reuses small-difference and large-difference features at local and global levels, makes use of the learned multi-scale information at maximum. The proposed methods achieve multi-scale learning using small-size convolution only, resulting in a lightweight and high-performance SR network. Experimental results show the superiority of our SMSR over state-of-the-art methods in super-resolving complicated image patterns. The effectiveness of SMSR is also demonstrated through its support to object recognition task.
Longguang Wang, Xu Sun 0005, Xiuping Jia, Lianru Gao, Bing Zhang 0001
IEEE Trans. Geosci. Remote. Sens.2
2021 Light Field Image Super-Resolution Using Deformable Convolution
abstract
Light field (LF) cameras can record scenes from multiple perspectives, and thus introduce beneficial angular information for image super-resolution (SR). However, it is challenging to incorporate angular information due to disparities among LF images. In this paper, we propose a deformable convolution network (i.e., LF-DFnet) to handle the disparity problem for LF image SR. Specifically, we design an angular deformable alignment module (ADAM) for feature-level alignment. Based on ADAM, we further propose a collect-and-distribute approach to perform bidirectional alignment between the center-view feature and each side-view feature. Using our approach, angular information can be well incorporated and encoded into features of each view, which benefits the SR reconstruction of all LF images. Moreover, we develop a baseline-adjustable LF dataset to evaluate SR performance under different disparity variations. Experiments on both public and our self-developed datasets have demonstrated the superiority of our method. Our LF-DFnet can generate high-resolution images with more faithful details and achieve state-of-the-art reconstruction accuracy. Besides, our LF-DFnet is more robust to disparity variations, which has not been well addressed in literature.
Yingqian Wang 0002, Jun-Gang Yang, Longguang Wang, Xinyi Ying, Tianhao Wu 0014, Wei An 0003, Yulan Guo
IEEE Trans. Image Process.3
2020 Spatial-Angular Interaction for Light Field Image Super-Resolution
Yingqian Wang 0002, Longguang Wang, Jun-Gang Yang, Wei An 0003, Jingyi Yu 0001, Yulan Guo
ECCV (23)2
2020 DeOccNet: Learning to See Through Foreground Occlusions in Light Fields
abstract
Background objects occluded in some views of a light field (LF) camera can be seen by other views. Consequently, occluded surfaces are possible to be reconstructed from LF images. In this paper, we handle the LF de-occlusion (LF-DeOcc) problem using a deep encoder-decoder network (namely, DeOccNet). In our method, sub-aperture images (SAIs) are first given to the encoder to incorporate both spatial and angular information. The encoded representations are then used by the decoder to render an occlusion-free center-view SAI. To the best of our knowledge, DeOccNet is the first deep learning-based LF-DeOcc method. To handle the insufficiency oftraining data, we propose an LF synthesis approach to embed selected occlusion masks into existing LF images. Besides, several synthetic and real-world LFs are developed for performance evaluation. Experimental results show that, after training on the generated data, our DeOccNet can effectively remove foreground occlusions and achieves superior performance as compared to other state-of-the-art methods. Source codes are available at: https://github.com/YingqianWang/DeOccNet.
Yingqian Wang 0002, Tianhao Wu 0014, Jun-Gang Yang, Longguang Wang, Wei An 0003, Yulan Guo
WACV4
2020 A Stereo Attention Module for Stereo Image Super-Resolution
abstract
In stereo image super-resolution (SR), exploiting both intra-view and cross-view information is significant but challenging. As existing single image SR (SISR) methods are powerful in intra-view information exploitation, in this letter, we propose a generic stereo attention module (SAM) to extend arbitrary SISR networks for stereo image SR. Specifically, we apply two identical pretrained SISR networks to stereo images. The extracted stereo features at different stages are fed to SAMs to interact cross-view information. Finally, the intra-view and cross-view information is incorporated by SISR networks for stereo image SR. Experiments on the KITTI2012, KITTI2015 and Middlebury datasets have demonstrated the effectiveness of our scheme. Using SAM, we can exploit cross-view information while maintaining the superiority of intra-view information exploitation, resulting in notable performance gain to SISR networks. Moreover, SRResNet equipped with our SAM outperforms the state-of-the-art stereo SR methods. Source code is available at https://github.com/XinyiYing/SAM.
Xinyi Ying, Yingqian Wang 0002, Longguang Wang, Weidong Sheng, Wei An 0003, Yulan Guo
IEEE Signal Process. Lett.3
2020 Deformable 3D Convolution for Video Super-Resolution
abstract
The spatio-temporal information among video sequences is significant for video super-resolution (SR). However, the spatio-temporal information cannot be fully used by existing video SR methods since spatial feature extraction and temporal motion compensation are usually performed sequentially. In this paper, we propose a deformable 3D convolution network (D3Dnet) to incorporate spatio-temporal information from both spatial and temporal dimensions for video SR. Specifically, we introduce deformable 3D convolution (D3D) to integrate deformable convolution with 3D convolution, obtaining both superior spatio-temporal modeling capability and motion-aware modeling flexibility. Extensive experiments have demonstrated the effectiveness of D3D in exploiting spatio-temporal information. Comparative results show that our network achieves state-of-the-art SR performance. Code is available at: https://github.com/XinyiYing/D3Dnet.
Xinyi Ying, Longguang Wang, Yingqian Wang 0002, Weidong Sheng, Wei An 0003, Yulan Guo
IEEE Signal Process. Lett.2
2020 Deep Video Super-Resolution Using HR Optical Flow Estimation
abstract
Video super-resolution (SR) aims at generating a sequence of high-resolution (HR) frames with plausible and temporally consistent details from their low-resolution (LR) counterparts. The key challenge for video SR lies in the effective exploitation of temporal dependency between consecutive frames. Existing deep learning based methods commonly estimate optical flows between LR frames to provide temporal dependency. However, the resolution conflict between LR optical flows and HR outputs hinders the recovery of fine details. In this paper, we propose an end-to-end video SR network to super-resolve both optical flows and images. Optical flow SR from LR frames provides accurate temporal dependency and ultimately improves video SR performance. Specifically, we first propose an optical flow reconstruction network (OFRnet) to infer HR optical flows in a coarse-to-fine manner. Then, motion compensation is performed using HR optical flows to encode temporal dependency. Finally, compensated LR inputs are fed to a super-resolution network (SRnet) to generate SR results. Extensive experiments have been conducted to demonstrate the effectiveness of HR optical flows for SR performance improvement. Comparative results on the Vid4 and DAVIS-10 datasets show that our network achieves the state-of-the-art performance.
Longguang Wang, Yulan Guo, Li Liu 0002, Zaiping Lin, Xinpu Deng, Wei An 0003
IEEE Trans. Image Process.1
2019 Learning Parallax Attention for Stereo Image Super-Resolution
abstract
Stereo image pairs can be used to improve the performance of super-resolution (SR) since additional information is provided from a second viewpoint. However, it is challenging to incorporate this information for SR since disparities between stereo images vary significantly. In this paper, we propose a parallax-attention stereo superresolution network (PASSRnet) to integrate the information from a stereo image pair for SR. Specifically, we introduce a parallax-attention mechanism with a global receptive field along the epipolar line to handle different stereo images with large disparity variations. We also propose a new and the largest dataset for stereo image SR (namely, Flickr1024). Extensive experiments demonstrate that the parallax-attention mechanism can capture correspondence between stereo images to improve SR performance with a small computational and memory cost. Comparative results show that our PASSRnet achieves the state-of-the-art performance on the Middlebury, KITTI 2012 and KITTI 2015 datasets.
Longguang Wang, Yingqian Wang 0002, Zhengfa Liang, Zaiping Lin, Jun-Gang Yang, Wei An 0003, Yulan Guo
CVPR1
2018 Learning for Video Super-Resolution Through HR Optical Flow Estimation
Longguang Wang, Yulan Guo, Zaiping Lin, Xinpu Deng, Wei An 0003
ACCV (1)1