Shuai Liu 0009

dblp:76/5789-0009 · DBLP profile ↗
← Back
28ranked-venue papers
11as first author
27since 2021 · last 2026
0000-0001-5002-4754ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 6 first-author · 17 since 2021Artificial intelligence and machine learning · 15 · 7 first-author · 15 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Unsupervised Large-Scale Point Cloud Registration via Spherical Projection Consistency
abstract
In real-world applications, large-scale outdoor Li- DAR point cloud registration faces significant challenges, including data sparsity, occlusions, and reliance on expensive ground-truth pose labels. To address these issues, this paper proposes an unsupervised registration framework based on spherical projection consistency. Specifically, both the input point cloud and its spatially transformed counterpart are projected into range images, and their spatial consistency is exploited as a supervision signal. Furthermore, a multi-scale patch-topatch feature fusion module is introduced to effectively integrate features from point clouds and range images, thereby enhancing feature discriminability. In addition, a dynamic masking strategy is applied in the range image domain to mitigate the impact of sparsity variations and further strengthen the spatial consistency of the supervision signal. Extensive experiments on the KITTI and NuScenes datasets demonstrate that the proposed method outperforms state-of-the-art unsupervised approaches, validating its effectiveness and robustness in large-scale outdoor scenarios.
Yejun Shou, Shuai Liu 0009, Zhijie Xu, Yanlong Cao
IEEE Signal Process. Lett.2
2025 FSHNet: Fully Sparse Hybrid Network for 3D Object Detection
abstract
Fully sparse 3D detectors have recently gained significant attention due to their efficiency in long-range detection. However, sparse 3D detectors extract features only from non-empty voxels, which impairs long-range interactions and causes the center feature missing. The former weakens the feature extraction capability, while the latter hinders network optimization. To address these challenges, we introduce the Fully Sparse Hybrid Network (FSHNet). FSHNet incorporates a proposed SlotFormer block to enhance the long-range feature extraction capability of existing sparse encoders. The SlotFormer divides sparse voxels using a slot partition approach, which, compared to traditional window partition, provides a larger receptive field. Additionally, we propose a dynamic sparse label assignment strategy to deeply optimize the network by providing more high-quality positive samples. To further enhance performance, we introduce a sparse upsampling module to refine downsampled voxels, preserving fine-grained details crucial for detecting small objects. Extensive experiments on the Waymo, nuScenes, and Argoverse2 benchmarks demonstrate the effectiveness of FSHNet. The code is available at https://github.com/Say2L/FSHNet.
Shuai Liu 0009, Mingyue Cui, Boyang Li 0009, Quanmin Liang, Tinghe Hong, Yunxiao Shan, Kai Huang 0001
CVPR1
2025 TrackFusion: Enhancing Multi-Object Tracking With Temporal Trajectory Modeling and Frame-Integrated Detection
abstract
Although MOTIP is the SOTA multi-object tracking method, there are still some issues that limit its performance. First, MOTIP still has defects in temporal information modeling, which leads to the failure to fully utilize the historical information of the tracked target and affects the correlation performance of the model. Second, in MOT, objects in consecutive video frames usually have temporal continuity and spatial consistency. Therefore, the object information of the previous frame can effectively assist the detection of the current frame. However, MOTIP performs independent detection between each frame, which does not fully utilize the correlation information between frames, resulting in suboptimal model performance. To address the above problems, we propose TrackFusion, which optimizes model performance from the perspective of trajectory modeling and inter-frame joint detection. First, we extract embeddings in video sequences through a Transformer-based detector, then combine the embeddings of the same object in different frames into sequences and input them into the trajectory modeling module for sequence association. This strategy effectively enhances the association ability. Thanks to these improvements, TrackFusion’s HOTA on the DanceTrack test set reached 68.6%, an increase of 1.1% compared to MOTIP’s 67.5%.
Shuai Liu 0009, Bingyang Wang, Jiaojiao Dai, Jinqing Qi, Huchuan Lu, You He 0002
ICASSP2
2025 Efficient Event Camera Data Pretraining with Adaptive Prompt Fusion
Quanmin Liang, Shuai Liu 0009, Xinzi Cao, Jinyi Lu, Feidiao Yang, Wei Zhang 0161, Kai Huang 0001, Yonghong Tian 0001
ICCV3
2025 Fine-Structure Preserved Real-World Image Super-Resolution Via Transfer Vae Training
Qiaosi Yi, Shuai Liu 0009, Rongyuan Wu, Lingchen Sun, Yuhui Wu 0001, Lei Zhang 0006
ICCV2
2025 ROAD-6: A Diverse Dataset for Unexpected Hazard Recognition in Autonomous Vehicles
Shehzad Ali, Md Tanvir Islam, Minh-Son Dao, Ikhyun Lee, Shuai Liu 0009, Khan Muhammad 0001
ICMR5
2025 ESOD: Event-Based Small Object Detection
abstract
Event-based object detection plays a crucial role in scenarios involving high-speed motion, extreme lighting conditions, and high-frequency detection. However, existing methods fail to address the challenges posed by small objects, including discriminative feature deficiency, the loss of critical information, and the inherent sparsity of event data. Moreover, the lack of benchmark datasets has significantly hindered progress in this field. To tackle these issues, we propose the Fully Deformable Detection Network (FDDNet), a lightweight framework that dynamically adapts to extract key features. First, we introduce a Long-Term Deformable Temporal Receptive Module (LDTR), which aligns critical features across consecutive event streams and leverages a State Space Model for long-range temporal modeling, enhancing the detection of high-speed small objects. Second, to address the sparsity of event data and the concentration of key features along object edges, we design a Sparse Feature Aggregation Block (SFAB) within the backbone and a coarse-to-fine deformable detection head, enabling hierarchical feature refinement from local to global, and improving the detection quality of sparse targets. Finally, to mitigate the lack of event-based small object datasets, we develop a high-quality, annotation-free data acquisition method and collect a real-world benchmark dataset for validation. Extensive experiments demonstrate that our approach achieves state-of-the-art (SOTA) performance on event-based small object detection tasks, with a mAP of 37.4% (+2.4%) on our benchmark and runs at 88 FPS, showcasing both accuracy and real-time capability. Our code and Supplement are available at https://github.com/Lqm26/ESOD.
Quanmin Liang, Jinyi Lu, Shuai Liu 0009, Yinzheng Zhao, Wei Zhang 0161, Kai Huang 0001, Yonghong Tian 0001
ACM Multimedia4
2025 SURDS: Benchmarking Spatial Understanding and Reasoning in Driving Scenarios with Vision Language Models
abstract
Accurate spatial reasoning in outdoor environments—covering geometry, object pose, and inter-object relationships—is fundamental to downstream tasks such as mapping, motion forecasting, and high-level planning in autonomous driving. We introduce SURDS, a large-scale benchmark designed to systematically evaluate the spatial reasoning capabilities of vision language models (VLMs). Built on the nuScenes dataset, SURDS comprises 41,080 vision–question–answer training instances and 9,250 evaluation samples, spanning six spatial categories: orientation, depth estimation, pixel-level localization, pairwise distance, lateral ordering, and front–behind relations. We benchmark leading general-purpose VLMs, including GPT, Gemini, and Qwen, revealing persistent limitations in fine-grained spatial understanding. To address these deficiencies, we go beyond static evaluation and explore whether alignment techniques can improve spatial reasoning performance. Specifically, we propose a reinforcement learning–based alignment scheme leveraging spatially grounded reward signals—capturing both perception-level accuracy (location) and reasoning consistency (logic). We further incorporate final-answer correctness and output-format rewards to guide fine-grained policy adaptation. Our GRPO-aligned variant achieves overall score of 40.80 in SURDS benchmark. Notably, it outperforms proprietary systems such as GPT-4o (13.30) and Gemini-2.0-flash (35.71). To our best knowledge, this is the first study to demonstrate that reinforcement learning–based alignment can significantly and consistently enhance the spatial reasoning capabilities of VLMs in real-world driving contexts. We release the SURDS benchmark, evaluation toolkit, and GRPO alignment code through: https://github.com/XiandaGuo/Drive-MLLM.
Xianda Guo, Ruijun Zhang, Yiqun Duan, Dujun Nie, Wenke Huang 0003, Chenming Zhang, Shuai Liu 0009, Hao Zhao 0002, Long Chen 0005
NeurIPS8
2025 GaussianFusion: Gaussian-Based Multi-Sensor Fusion for End-to-End Autonomous Driving
abstract
Multi-sensor fusion is crucial for improving the performance and robustness of end-to-end autonomous driving systems. Existing methods predominantly adopt either attention-based flatten fusion or bird’s eye view fusion through geometric transformations. However, these approaches often suffer from limited interpretability or dense computational overhead. In this paper, we introduce GaussianFusion, a Gaussian-based multi-sensor fusion framework for end-to-end autonomous driving. Our method employs intuitive and compact Gaussian representations as intermediate carriers to aggregate information from diverse sensors. Specifically, we initialize a set of 2D Gaussians uniformly across the driving scene, where each Gaussian is parameterized by physical attributes and equipped with explicit and implicit features. These Gaussians are progressively refined by integrating multi-modal features. The explicit features capture rich semantic and spatial information about the traffic scene, while the implicit features provide complementary cues beneficial for trajectory planning. To fully exploit rich spatial and semantic information in Gaussians, we design a cascade planning head that iteratively refines trajectory predictions through interactions with Gaussians. Extensive experiments on the NAVSIM and Bench2Drive benchmarks demonstrate the effectiveness and robustness of the proposed GaussianFusion framework. The source code is included in the supplementary material and will be released publicly.
Shuai Liu 0009, Quanmin Liang, Zefeng Li, Boyang Li 0009, Kai Huang 0001
NeurIPS1
2025 TeamFed: Teamwork Principles-Inspired Federated Learning for 3D Object Detection
Siheng Ren, Boyang Li 0009, Shuai Liu 0009, Jiahui Liao, Mingyue Cui, Kai Huang 0001
PRCV (11)3
2025 No-Reference Image Quality Assessment: Past, Present, and Future
abstract
ABSTRACT No‐reference image quality assessment (NR‐IQA) has garnered significant attention due to its critical role in various image processing applications. This survey provides a comprehensive and systematic review of NR‐IQA methods, datasets, and challenges, offering new perspectives and insights for the field. Specifically, we propose a novel taxonomy for NR‐IQA methods based on distortion scenarios and design principles, which distinguishes this work from previous surveys. Representative methods within each category are thoroughly examined, with a focus on their strengths, limitations, and performance characteristics. Additionally, we review 20 widely used NR‐IQA datasets that serve as benchmarks for evaluating these methods, providing detailed information on the number of images, distortion types, and distortion levels for each dataset. Furthermore, we identify and discuss key challenges currently faced by NR‐IQA methods, such as handling diverse and complex distortions, ensuring generalisation across datasets and devices, and achieving real‐time performance. We also suggest potential future research directions to address these issues. In summary, this survey offers a comprehensive and systematic examination of NR‐IQA methods, datasets, and challenges, offering valuable insights and guidance for researchers and practitioners working in the NR‐IQA domain.
Qingyu Mao, Shuai Liu 0009, Qilei Li, Gwanggil Jeon, Hyunbum Kim, David Camacho
Expert Syst. J. Knowl. Eng.2
2025 Robust manipulated media localization and detection based on high frequency and texture features
abstract
Advances in facial manipulation techniques have resulted in the increasing trend of realistic and indistinguishable identity swap media, which mislead the viewers and accompanied by severe security concerns. While current deepfake detectors demonstrate strong performance under high-quality conditions, they still face notable limitations. This article proposes a novel framework mining high frequency and degraded texture features for locating manipulated traces and improving the generalization ability. To improve the universality of the proposed detector, we design the Multi-feature Mining Stream for capturing the global and subtle discrepancies of undegraded images. Moreover, the Encoder-Decoder Structure is introduced for gaining high localization accuracy and full resolution manipulated regions. This work attempts to solve the tampered region localization issue and achieve face forgery image detection at the meantime. This contributes to help the model perform a more effective differentiation between real and fake content when confronted with high- or low-quality compressed images. Comprehensive experiments on the popular FaceForensics++, Celeb-DF, and DFDC datasets demonstrate the superior performance and robustness of our proposed framework, in particular, achieving performance improvements ranging from 1% to 10% in comparison with the most recent related work.
Shuai Liu 0009, Shengfa Miao, Huasong Yi, Xin Jin 0005, Yuru Kou, Hanxian Duan
Discov. Comput.2
2025 Turbid Underwater Image Enhancement With Illumination-Constrained and Structure-Preserved Retinex Model
abstract
Turbid underwater images often suffer from color distortion, contrast degradation, and detail loss. To improve the visual quality of these images, this paper proposes an illumination-constrained, structure-preserved retinex variational model. The proposed approach consists of three main components: a nonlinear model based on the classical retinex theory to represent the multiple adverse deformations of turbid underwater images; an adaptive channel compensation method to correct the color cast; and an illumination-constrained structure-preserved variational retinex model that simultaneously estimates a smooth illumination component and a detail display reflection component and uniformly predicts the noise pattern of preprocessed underwater images. Specifically, an adaptive weight matrix is proposed to reveal the structural details in reflectance. The overall smoothness of illumination is constrain by exponential guided filtering and l1/2 norm. The total intensity of the noise pattern is constrained by l2 norm. To solve the resulting optimization problem, we employ alternating direction minimization of logless transformations of Lagrange multipliers. Extensive experiments demonstrate the effectiveness of the proposed method in improving the quality of turbid underwater images. Beyond subjective visual observations, the method also exhibits competitive performance in objective image quality evaluations.
Shuai Liu 0009, Yuchao Zheng 0001, Jianru Li, Huimin Lu 0001, Zhengxiang Shen, Zhanshan Wang 0002
IEEE Trans. Circuits Syst. Video Technol.1
2025 Real-Time Semantic Segmentation via a Densely Aggregated Bilateral Network
abstract
With the growing demands of applications on online devices, the speed-accuracy trade-off is critical in the semantic segmentation system. Recently, the bilateral segmentation network has shown promising capacity to achieve the balance between favorable accuracy and fast speed, and has become the mainstream backbone in real-time semantic segmentation. Segmentation of target objects relies on high-level semantics, whereas it requires detailed low-level features to model specific local patterns for accurate location. However, the lightweight backbone of bilateral architecture limits the extraction of semantic context and spatial details. And the late fusion of the bilateral streams incurs the insufficient aggregation of semantic context and spatial details. In this article, we propose a densely aggregated bilateral network (DAB-Net) for real-time semantic segmentation. In the context path, a patchwise context enhancement (PCE) module is proposed to efficiently capture the local semantic contextual information from spatialwise and channelwise, respectively. Meanwhile, a context-guided spatial path (CGSP) is designed to exploit more spatial information by encoding finer details from the raw image and the transition from the context path. Finally, with multiple interactions between bilateral branches, the intertwined outputs from bilateral streams are combined in a unified decoder for a final interaction to further enhance the feature representation, which generates the final segmentation prediction. Experimental results on three public benchmarks demonstrate that our proposed method achieves higher accuracy with a limited decay in speed, which performs favorably against state-of-the-art real-time approaches and runs at 31.1 frames/s (FPS) on the high resolution of . The source code is released at https://github.com/isyangshu/DABNet.
Shu Yang 0004, Lu Zhang 0053, Shuai Liu 0009, Huchuan Lu, Hao Chen 0011
IEEE Trans. Neural Networks Learn. Syst.3
2024 Self-Supervised High Dynamic Range Imaging with Multi-Exposure Images in Dynamic Scenes
abstract
Merging multi-exposure images is a common approach for obtaining high dynamic range (HDR) images, with the primary challenge being the avoidance of ghosting artifacts in dynamic scenes. Recent methods have proposed using deep neural networks for deghosting. However, the methods typically rely on sufficient data with HDR ground-truths, which are difficult and costly to collect. In this work, to eliminate the need for labeled data, we propose SelfHDR, a self-supervised HDR reconstruction method that only requires dynamic multi-exposure images during training. Specifically, SelfHDR learns a reconstruction network under the supervision of two complementary components, which can be constructed from multi-exposure images and focus on HDR color as well as structure, respectively. The color component is estimated from aligned multi-exposure images, while the structure one is generated through a structure-focused network that is supervised by the color component and an input reference (\eg, medium-exposure) image. During testing, the learned reconstruction network is directly deployed to predict an HDR image. Experiments on real-world images demonstrate our SelfHDR achieves superior results against the state-of-the-art self-supervised methods, and comparable performance to supervised ones. Codes are available at https://github.com/cszhilu1998/SelfHDR
Zhilu Zhang 0001, Shuai Liu 0009, Xiaotao Wang, Wangmeng Zuo
ICLR3
2024 DCDet: Dynamic Cross-based 3D Object Detector
Shuai Liu 0009, Boyang Li 0009, Zhiyu Fang, Kai Huang 0001
IJCAI1
2024 FFAM: Feature Factorization Activation Map for Explanation of 3D Detectors
abstract
LiDAR-based 3D object detection has made impressive progress recently, yet most existing models are black-box, lacking interpretability. Previous explanation approaches primarily focus on analyzing image-based models and are not readily applicable to LiDAR-based 3D detectors. In this paper, we propose a feature factorization activation map (FFAM) to generate high-quality visual explanations for 3D detectors. FFAM employs non-negative matrix factorization to generate concept activation maps and subsequently aggregates these maps to obtain a global visual explanation. To achieve object-specific visual explanations, we refine the global visual explanation using the feature gradient of a target object. Additionally, we introduce a voxel upsampling strategy to align the scale between the activation map and input point cloud. We qualitatively and quantitatively analyze FFAM with multiple detectors on several datasets. Experimental results validate the high-quality visual explanations produced by FFAM. The code is available at \url{https://anonymous.4open.science/r/FFAM-B9AF}.
Shuai Liu 0009, Boyang Li 0009, Zhiyu Fang, Mingyue Cui, Kai Huang 0001
NeurIPS1
2024 MCDC-Net: Multi-scale forgery image detection network based on central difference convolution
abstract
Abstract Generative Adversarial Networks (GANs) emerged thanks to the development of deep neural networks. Forgery images generated by various variants of GANs are widely spread on the Internet, which may be damage personal credibility and cause huge property losses. Thus, numerous methods are proposed to detect forgery images, but most of them are designed to detect forgery faces. Therefore, a method to detect forgery images of various scenes is proposed. In this work, central difference convolution and vanilla convolution (CDC‐Mix) are mixed after considering the depth and width features of neural networks and analyzing the influence of attention on network performance. Based on CDC‐Mix, a separable convolution (SeparableCDC‐Mix) is proposed. The proposed method consists of three parts: (1) CDC‐Mix and SeparableCDC‐Mix are used to extract the gradient information and texture features; (2) CDCM is used to extract the multi‐scale information of the image; (3) multi‐scale fusion module (MS‐Fusion) is used to fuse the multi‐scale information from different locations of the network. A large number of experiments have been carried out on several datasets generated by GAN, and the experimental results show that the proposed method has a great improvement compared with the existing advanced methods.
Defen He, Xin Jin 0005, Zien Cheng, Shuai Liu 0009, Shaowen Yao 0001, Wei Zhou 0011
IET Image Process.5
2024 NIV-SSD: Neighbor IoU-voting single-stage object detector from point cloud
Shuai Liu 0009, Di Wang 0011, Quan Wang 0006, Kai Huang 0001
Neurocomputing1
2024 Efficient Adaptive Feature Fusion Network for Remote-Sensing Image Super-Resolution
abstract
Image super-resolution is a fundamental low-level vision task aimed at recovering high-resolution images with fine details. Deep learning has significantly enhanced the performance of super-resolution techniques for remote sensing imagery. However, increasing the depth of networks and the size of their parameters has resulted in substantial computational and storage burdens. To address this challenge, we propose an adaptive approach that learns both local and global information for each region. We introduce a lightweight hybrid model named the Efficient Adaptive Feature Fusion Network, which combines CNNs and Transformers to fully exploit the texture information in remote sensing images. This model leverages local details and long-range dependencies within images in an adaptive manner to achieve superior super-resolution. Specifically, a set of Transformers is employed to model the self-similarity between pixels and perform dense texture pattern predictions at each pixel, while a set of CNNs captures local details within the images. The computed global and local features serve as inputs to the proposed Adaptive Contextual Fusion Block, which learns to fuse local and global information across different regions to generate robust image super-resolution features. We conduct extensive experimental evaluations of the proposed method on the UCMerced and AID datasets, demonstrating its outstanding performance in terms of PSNR and SSIM metrics. Comprehensive experiments validate the effectiveness of our approach, showing that the proposed method achieves an excellent balance between performance and complexity.
Shuai Hao 0007, Shuai Liu 0009, Xu Jia 0012, Huchuan Lu, You He 0002
IEEE Signal Process. Lett.2
2023 Physics-Guided ISO-Dependent Sensor Noise Modeling for Extreme Low-Light Photography
abstract
Although deep neural networks have achieved astonishing performance in many vision tasks, existing learningbased methods are far inferior to the physical model-based solutions in extreme low-light sensor noise modeling. To tap the potential of learning-based sensor noise modeling, we investigate the noise formation in a typical imaging process and propose a novel physics-guided ISO-dependent sensor noise modeling approach. Specifically, we build a normalizing flow-based framework to represent the complex noise characteristics of CMOS camera sensors. Each component of the noise model is dedicated to a particular kind of noise under the guidance of physical models. Moreover, we take into consideration of the ISO dependence in the noise model, which is not completely considered by the existing learning-based methods. For training the proposed noise model, a new dataset is further collected with paired noisy-clean images, as well as flat-field and bias frames covering a wide range of ISO settings. Compared to existing methods, the proposed noise model is equipped with a flexible structure and accurate modeling capabilities, which is beneficial for better denoising performance in extreme low-light scenes. The dataset and code are available at https://github.com/happycaoyue/LLD.
Yue Cao 0009, Ming Liu 0018, Shuai Liu 0009, Xiaotao Wang, Wangmeng Zuo
CVPR3
2023 Beyond Image Borders: Learning Feature Extrapolation for Unbounded Image Composition
abstract
For improving image composition and aesthetic quality, most existing methods modulate the captured images by striking out redundant content near the image borders. However, such image cropping methods are limited in the range of image views. Some methods have been suggested to extrapolate the images and predict cropping boxes from the extrapolated image. Nonetheless, the synthesized extrapolated regions may be included in the cropped image, making the image composition result not real and potentially with degraded image quality. In this paper, we circumvent this issue by presenting a joint framework for both unbounded recommendation of camera view and image composition (i.e., UNIC). In this way, the cropped image is a sub-image of the image acquired by the predicted camera view, and thus can be guaranteed to be real and consistent in image quality. Specifically, our framework takes the current camera preview frame as input and provides a recommendation for view adjustment, which contains operations unlimited by the image borders, such as zooming in or out and camera movement. To improve the prediction accuracy of view adjustment prediction, we further extend the field of view by feature extrapolation. After one or several times of view adjustments, our method converges and results in both a camera view and a bounding box showing the image composition recommendation. Extensive experiments are conducted on the datasets constructed upon existing image cropping datasets, showing the effectiveness of our UNIC in unbounded recommendation of camera view and image composition. The source code, dataset, and pre-trained models is available at https://github.com/liuxiaoyu1104/UNIC.
Xiaoyu Liu 0006, Ming Liu 0018, Junyi Li 0005, Shuai Liu 0009, Xiaotao Wang, Wangmeng Zuo
ICCV4
2023 Cross-Modal Enhancement Network for Multimodal Sentiment Analysis
abstract
Multimodal sentiment analysis (MSA) plays an important role in many applications, such as intelligent question-answering, computer-assisted psychotherapy and video understanding, and has attracted considerable attention in recent years. It leverages multimodal signals including verbal language, facial gestures, and acoustic behaviors to identify sentiments in videos. Language modality typically outperforms nonverbal modalities in MSA. Therefore, strengthening the significance of language in MSA will be a vital way to promote recognition accuracy. Considering that the meaning of a sentence often varies in different nonverbal contexts, combining nonverbal information with text representations is conducive to understanding the exact emotion conveyed by an utterance. In this paper, we propose a Cross-modal Enhancement Network (CENet) model to enhance text representations by integrating visual and acoustic information into a language model. Specifically, it embeds a Cross-modal Enhancement (CE) module, which enhances each word representation according to long-range emotional cues implied in unaligned nonverbal data, into a transformer-based pre-trained language model. Moreover, a feature transformation strategy is introduced for acoustic and visual modalities to reduce the distribution differences between the initial representations of verbal and nonverbal modalities, thereby facilitating the fusion of distinct modalities. Extensive experiments on benchmark datasets demonstrate the significant gains of CENet over state-of-the-art methods.
Di Wang 0011, Shuai Liu 0009, Quan Wang 0006, Yumin Tian, Lihuo He, Xinbo Gao 0001
IEEE Trans. Multim.2
2022 Multi-Object Tracking Meets Moving UAV
abstract
Multi-object tracking in unmanned aerial vehicle (UAV) videos is an important vision task and can be applied in a wide range of applications. However, conventional multi-object trackers do not work well on UAV videos due to the challenging factors of irregular motion caused by moving camera and view change in 3D directions. In this paper, we propose a UAVMOT network specially for multi-object tracking in UAV views. The UAVMOT introduces an ID feature update module to enhance the object's feature association. To better handle the complex motions under UAV views, we develop an adaptive motion filter module. In addition, a gradient balanced focal loss is used to tackle the imbalance categories and small objects detection problem. Experimental results on the VisDrone2019 and UAVDT datasets demonstrate that the proposed UAVMOT achieves considerable improvement against the state-of-the-art tracking methods on UAV videos.
Shuai Liu 0009, Xin Li 0034, Huchuan Lu, You He 0002
CVPR1
2022 Multiple Feature Mining Based on Local Correlation and Frequency Information for Face Forgery Detection
abstract
As facial image manipulation techniques developed, deep fake detection attracted extensive attentions. Although researchers have made remarkable progresses in deepfake detection recently, which is still suffering from two limitations: a) current detectors achieve high accuracy in the high-quality videos and images, but it is hard to capture local and subtle artifacts in the low-quality and high-compression media; b) few of deep fake detection methods gain satisfying performance under cross-database scenario, because detector overfit to specific color textures producing by same manipulation algorithm. Inspired the above issues, this paper proposes a novel framework fusing local related features and frequency information to mine the forgery patterns. Firstly, we design multi-feature enhancement module, which amplifies implicit local disc repancies and capture spatial correlation from three shallow feature layers and high-level semantic layer guided by attention maps. Secondly, dual frequency decomposition module is proposed for disassembling high-frequency and low-frequency features, the forgery artifacts are exposed after dual cross attention block processing in the frequency spectrum. Features from the two streams are fused to the classification for the final result. Comprehensive experiments demonstrate the superior performance of our proposed approach in the low-quality benchmark database and cross-dataset sce-nario.
Shuai Liu 0009, Xin Jin 0005, Zhenli He, Wei Zhou 0011, Shaowen Yao 0001, Qiannian Wang
ICTAI1
2022 Center-Boundary Dual Attention for Oriented Object Detection in Remote Sensing Images
abstract
Recently, anchor-free object detectors have shown promising performance in oriented object detection on remote sensing images. However, the objects in remote sensing images always have large variations in arbitrary orientations, sizes, and aspect ratios, which makes the existing anchor-free methods hard to obtain satisfactory results. In this article, we propose a novel anchor-free detector, center-boundary dual attention (CBDA) network (CBDA-Net), for fast and accurate oriented object detection on remote sensing images. In CBDA-Net, we construct a CBDA module, which utilizes a dual attention mechanism to extract attention features on the center and boundary regions of objects. The CBDA module can learn more essential features for rotating objects and reduce the interference from complex background. Besides, to resolve the influence of object aspect ratio on angle errors, we propose an aspect ratio weighted angle loss (arwLoss), where diffident penalties are assigned on the angle loss based on the aspect ratios of objects. This loss construction is effective in improving the detection accuracy of oriented objects, especially for slender objects. We conduct extensive experiments on two publish benchmarks, i.e., DOTA and HRSC2016. The experimental results demonstrate that our CBDA-Net achieves favorable performance against other anchor-free state of the arts with a real-time speed of 50 FPS.
Shuai Liu 0009, Lu Zhang 0053, Huchuan Lu, You He 0002
IEEE Trans. Geosci. Remote. Sens.1
2021 Polar Ray: A Single-stage Angle-free Detector for Oriented Object Detection in Aerial Images
abstract
Oriented bounding boxes are widely used for object detection in aerial images. Existing oriented object detection methods typically follow the general object detection paradigm by adding an extra rotation angle on the horizontal bounding boxes. However, the angular periodicity incurs the difficulty in angle regression and rotation sensitivity on bounding boxes. In this paper, we propose a new anchor-free oriented object detector, Polar Ray Network (PRNet), where object keypoints are represented by polar coordinates without angle regression. Our PRNet learns a set of polar rays from the object center to boundary with predefined equal-distributed angles. We introduce a dynamic PointConv module to optimize the regression of polar ray by incorporating object corner features. Furthermore, a classification feature guidance module is presented to improve the classification accuracy by incorporating more spatial contents from polar rays. Experimental results on two public datasets, i.e., DOTA and HRSC2016, demonstrate that the proposed PRNet significantly outperforms existing anchor-free detectors, and shows highly competitiveness with the state-of-the-art two-stage anchor-based methods.
Shuai Liu 0009, Lu Zhang 0053, Shuai Hao 0007, Huchuan Lu, You He 0002
ACM Multimedia1
2019 Adaptive deep residual network for single image super-resolution
abstract
In recent years, deep learning has achieved great success in the field of image processing. In the single image super-resolution (SISR) task, the convolutional neural network (CNN) extracts the features of the image through deeper layers, and has achieved impressive results. In this paper, we propose a single image super-resolution model based on Adaptive Deep Residual named as ADR-SR, which uses the Input Output Same Size (IOSS) structure, and releases the dependence of upsampling layers compared with the existing SR methods. Specifically, the key element of our model is the Adaptive Residual Block (ARB), which replaces the commonly used constant factor with an adaptive residual factor. The experiments prove the effectiveness of our ADR-SR model, which can not only reconstruct images with better visual effects, but also get better objective performances.
Shuai Liu 0009, Ruipeng Gang, Chenghua Li, Ruixia Song
Comput. Vis. Media1