Xiang Li 0041

dblp:40/1491-41 · DBLP profile ↗
← Back
91ranked-venue papers
11as first author
71since 2021 · last 2026
0000-0002-4996-7365ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 76 · 9 first-author · 59 since 2021Graphics, computer vision, multimedia, augmented reality and games · 48 · 5 first-author · 35 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SM3Det: A Unified Model for Multi-Modal Remote Sensing Object Detection
abstract
With the rapid advancement of remote sensing technology, high-resolution multi-modal imagery is now more widely accessible. Conventional object detection models are trained on a single dataset, often restricted to a specific imaging modality and annotation format. However, such an approach overlooks the valuable shared knowledge across multi-modalities and limits the model’s applicability in more versatile scenarios. This paper introduces a new task called Multi-Modal Datasets and Multi-Task Object Detection (M2Det) for remote sensing, designed to accurately detect horizontal or oriented objects from any sensor modality. This task poses challenges due to 1) the trade-offs involved in managing multi-modal modelling and 2) the complexities of multi-task optimization. To address these, we establish a benchmark dataset and propose a unified model, SM3Det (Single Model for Multi-Modal datasets and Multi-Task object Detection). SM3Det leverages a grid-level sparse MoE backbone to enable joint knowledge learning while preserving distinct feature representations for different modalities. Furthermore, we propose a novel consistency and synchronization optimization mechanism, allowing it to effectively handle varying levels of learning difficulty across modalities and tasks. Extensive experiments demonstrate SM3Det's effectiveness and generalizability, consistently outperforming the combination of specialized models on individual datasets.
Yuxuan Li 0004, Xiang Li 0041, Yimian Dai, Qibin Hou, Ming-Ming Cheng, Jian Yang 0003
AAAI2
2026 DenoDet V2: Phase-Amplitude Cross Denoising for SAR Object Detection
abstract
One of the primary challenges in Synthetic Aperture Radar (SAR) object detection lies in the pervasive influence of coherent noise. As a common practice, most existing methods, whether handcrafted approaches or deep learning-based methods, employ the analysis or enhancement of object spatial-domain characteristics to achieve implicit denoising. In this paper, we propose DenoDet V2, which explores a completely novel and different perspective to deconstruct and modulate the features in the transform domain via a carefully designed attention architecture. Compared to DenoDet V1, DenoDet V2 is a major advancement that exploits the complementary nature of amplitude and phase information through a band-wise mutual modulation mechanism, which enables a reciprocal enhancement between phase and amplitude spectra. Extensive experiments on various SAR datasets demonstrate the state-of-the-art performance of DenoDet V2. Notably, DenoDet V2 achieves a significant 0.8% improvement on SARDet-100K dataset compared to DenoDet V1, while reducing the model complexity by half.
Kang Ni, Minrui Zou, Yuxuan Li 0004, Xiang Li 0041, Kehua Guo, Ming-Ming Cheng, Yimian Dai
AAAI4
2026 SpatioTemporal Difference Network for Video Depth Super-Resolution
abstract
Depth super-resolution has achieved impressive performance, and the incorporation of multi-frame information further enhances reconstruction quality. Nevertheless, statistical analyses reveal that video depth super-resolution remains affected by pronounced long-tailed distributions, with the long-tailed effects primarily manifesting in spatial non-smooth regions and temporal variation zones. To address these challenges, we propose a novel SpatioTemporal Difference Network (STDNet) comprising two core branches: a spatial difference branch and a temporal difference branch. In the spatial difference branch, we introduce a spatial difference mechanism to mitigate the long-tailed issues in spatial non-smooth regions. This mechanism dynamically aligns RGB features with learned spatial difference representations, enabling intra-frame RGB-D aggregation for depth calibration. In the temporal difference branch, we further design a temporal difference strategy that preferentially propagates temporal variation information from adjacent RGB and depth frames to the current depth frame, leveraging temporal difference representations to achieve precise motion compensation in temporal long-tailed areas. Extensive experimental results across multiple datasets demonstrate the effectiveness of our STDNet, outperforming existing approaches.
Zhengxue Wang, Xiang Li 0041, Zhiqiang Yan 0001, Jian Yang 0003
AAAI3
2026 Strip R-CNN: Large Strip Convolution for Remote Sensing Object Detection
abstract
In this paper, we show that current approaches using large square kernels or transformer-based global modeling aggregate contextual information uniformly across spatial dimensions, leading to feature dilution and localization errors for elongated targets. To mitigate this issue, we propose Strip R-CNN, the first work to systematically explore large strip convolutions for remote sensing object detection. Our key insight is that strip convolutions enable directional feature aggregation along the dominant spatial dimension of slender objects, reducing background interference while preserving essential geometric information. We design two core components: (i) StripNet, a backbone network employing sequential orthogonal large strip convolutions to capture anisotropic spatial patterns, and (ii) Strip Head, which enhances localization precision by incorporating strip convolutions into the detection head. Unlike previous large-kernel approaches that suffer from computational redundancy and isotropic limitations, our method achieves superior performance with remarkable efficiency. Extensive experiments on multiple benchmarks (DOTA, FAIR1M, HRSC2016, and DIOR) demonstrate significant improvements, with our 30M parameter model achieving 82.75% mAP on DOTA-v1.0, establishing a new state-of-the-art record while providing new insights into anisotropic feature learning for remote sensing applications.
Xinbin Yuan, Zhaohui Zheng 0003, Yuxuan Li 0004, Xialei Liu, Li Liu 0004, Xiang Li 0041, Qibin Hou, Ming-Ming Cheng
AAAI6
2026 S2I-DiT: Unlocking the semantic-to-image transferability by fine-tuning large diffusion transformer models
Enze Xie, Chongjian Ge, Xiang Li 0041, Lingyu Si, Changwen Zheng, Zhenguo Li
Pattern Recognit.4
2026 DTSI: Towards faster convergence of query-based detectors for rotated dense aerial images
Xiang Li 0041, Yuncong Yao, Wankou Yang
Pattern Recognit.3
2026 PanoKernel: Large Distortion-Aware Kernel for Panoramic Depth Perception
Zhiqiang Yan 0001, Zhijie Shen, Jiayi Yuan 0003, Xiang Li 0041, Jun Li 0027, Jian Yang 0003
IEEE Trans. Intell. Transp. Syst.4
2025 Multi-clue Consistency Learning to Bridge Gaps Between General and Oriented Object in Semi-supervised Detection
abstract
While existing semi-supervised object detection (SSOD) methods perform well in general scenes, they encounter challenges in handling oriented objects in aerial images. We experimentally find three gaps between general and oriented object detection in semi-supervised learning: 1) Sampling inconsistency: the common center sampling is not suitable for oriented objects with larger aspect ratios when selecting positive labels from labeled data. 2) Assignment inconsistency: balancing the precision and localization quality of oriented pseudo-boxes poses greater challenges which introduces more noise when selecting positive labels from unlabeled data. 3) Confidence inconsistency: there exists more mismatch between the predicted classification and localization qualities when considering oriented objects, affecting the selection of pseudo-labels. Therefore, we propose a Multi-clue Consistency Learning (MCL) framework to bridge gaps between general and oriented objects in semi-supervised detection. Specifically, considering various shapes of rotated objects, the Gaussian Center Assignment is specially designed to select the pixel-level positive labels from labeled data. We then introduce the Scale-aware Label Assignment to select pixel-level pseudo-labels instead of unreliable pseudo-boxes, which is a divide-and-rule strategy suited for objects with various scales. The Consistent Confidence Soft Label is adopted to further boost the detector by maintaining the alignment of the predicted results. Comprehensive experiments on DOTA-v1.5 and DOTA-v1.0 benchmarks demonstrate that our proposed MCL can achieve state-of-the-art performance in the semi-supervised oriented object detection task.
Chunyan Xu, Xiang Li 0041, YuXuan Li, Ziqi Gu, Zhen Cui 0001
AAAI3
2025 From Words to Worth: Newborn Article Impact Prediction with LLM
abstract
Predicting the future impact of newly published articles is pivotal for advancing scientific discovery in an era of unprecedented scholarly expansion. This paper introduces a promising approach, leveraging the capabilities of LLMs to predict the future impact of newborn articles solely based on titles and abstracts. Breaking away from traditional methods heavily reliant on external data, we propose fine-tuning the LLM to uncover the intrinsic semantic patterns shared by highly impactful articles from a vast collection of text-score pairs. These semantic features are further utilized to predict the proposed indicator, TNCSIsp, which incorporates favorable normalization properties across value, field, and time. To facilitate parameter-efficient fine-tuning of the LLM, we have also meticulously curated a dataset containing over 12,000 entries, each annotated with titles, abstracts, and their corresponding TNCSIsp values. Experimental results reveal an MAE of 0.216 and an NDCG@20 of 0.901, setting new benchmarks in predicting the impact of newborn articles. Finally, we present a real-world application example for predicting the impact of newborn journal articles to demonstrate its noteworthy practical value. Overall, our findings challenge existing paradigms and propose a shift towards a more content-focused prediction of academic impact, offering new insights for article impact prediction.
Penghai Zhao, Kairan Dou, Jinyu Tian 0006, Ying Tai, Jian Yang 0003, Ming-Ming Cheng, Xiang Li 0041
AAAI8
2025 Wolf in Sheep's Clothing: Understanding and Detecting Mobile Cloaking in Blackhat SEO
Yan Li 0195, Zhenrui Zhang, Zhiyu Lang, Xiang Li 0041, Donghong Sun
ACNS (2)4
2025 InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption
abstract
Text-to-video generation has evolved rapidly in recent years, delivering remarkable results. Training typically relies on video-caption paired data, which plays a crucial role in enhancing generation performance. However, current video captions often suffer from insufficient details, hallucinations and imprecise motion depiction, affecting the fidelity and consistency of generated videos. In this work, we propose a novel instance-aware structured caption framework, termed InstanceCap, to achieve instance-level and fine-grained video caption for the first time. Based on this scheme, we design an auxiliary models cluster to convert original video into instances to enhance instance fidelity. Video instances are further used to refine dense prompts into structured phrases, achieving concise yet precise descriptions. Furthermore, a 22K InstanceVid dataset is curated for training, and an enhancement pipeline that tailored to InstanceCap structure is proposed for inference. Experimental results demonstrate that our proposed InstanceCap significantly outperform previous models, ensuring high fidelity between captions and videos while reducing hallucinations.
Tiehan Fan, Kepan Nan, Rui Xie 0005, Penghao Zhou, Zhenheng Yang, Chaoyou Fu, Xiang Li 0041, Jian Yang 0003, Ying Tai
CVPR7
2025 RSAR: Restricted State Angle Resolver and Rotated SAR Benchmark
abstract
Rotated object detection has made significant progress in the optical remote sensing. However, advancements in the Synthetic Aperture Radar (SAR) field are laggard behind, primarily due to the absence of a large-scale dataset. Annotating such a dataset is inefficient and costly. A promising solution is to employ a weakly supervised model (e.g., trained with available horizontal boxes only) to generate pseudo-rotated boxes for reference before manual calibration. Unfortunately, the existing weakly supervised models exhibit limited accuracy in predicting the object’s angle. Previous works attempt to enhance angle prediction by using angle resolvers that decouple angles into cosine and sine encodings. In this work, we first reevaluate these resolvers from a unified perspective of dimension mapping and expose that they share the same shortcomings: these methods overlook the unit cycle constraint inherent in these encodings, easily leading to prediction biases. To address this issue, we propose the Unit Cycle Resolver (UCR), which incorporates a unit circle constraint loss to improve angle prediction accuracy. Our approach can effectively improve the performance of existing state-of-the-art weakly supervised methods and even surpasses fully supervised models on existing optical benchmarks (i.e., DOTA-v1.0). With the aid of UCR, we further annotate and introduce RSAR, the largest multi-class rotated SAR object detection dataset to date. Extensive experiments on both RSAR and optical datasets demonstrate that our UCR enhances angle prediction accuracy. Our dataset and code can be found at: https://github.com/zhasion/RSAR.
Xin Zhang 0170, Xue Yang 0005, Yuxuan Li 0004, Jian Yang 0003, Ming-Ming Cheng, Xiang Li 0041
CVPR6
2025 DISTA-Net: Dynamic Closely-Spaced Infrared Small Target Unmixing
abstract
Resolving closely-spaced small targets in dense clusters presents a significant challenge in infrared imaging, as the overlapping signals hinder precise determination of their quantity, sub-pixel positions, and radiation intensities. While deep learning has advanced the field of infrared small target detection, its application to closely-spaced infrared small targets has not yet been explored. This gap exists primarily due to the complexity of separating superimposed characteristics and the lack of an open-source infrastructure. In this work, we propose the Dynamic Iterative Shrinkage Thresholding Network (DISTA-Net), which reconceptualizes traditional sparse reconstruction within a dynamic framework. DISTA-Net adaptively generates convolution weights and thresholding parameters to tailor the reconstruction process in real time. To the best of our knowledge, DISTA-Net is the first deep learning model designed specifically for the unmixing of closely-spaced infrared small targets, achieving superior sub-pixel detection accuracy. Moreover, we have established the first open-source ecosystem to foster further research in this field. This ecosystem comprises three key components: (1) CSIST-100K, a publicly available benchmark dataset; (2) CSO-mAP, a custom evaluation metric for sub-pixel detection; and (3) GrokCSO, an open-source toolkit featuring DISTA-Net and other models. Our code and dataset are available at https://github.com/GrokCV/GrokCSO.
Shengdong Han, Shangdong Yang, Yuxuan Li 0004, Xin Zhang 0170, Xiang Li 0041, Jian Yang 0003, Ming-Ming Cheng, Yimian Dai
ICCV5
2025 Advancing Textual Prompt Learning with Anchored Attributes
Zheng Li 0028, Yibing Song, Ming-Ming Cheng, Xiang Li 0041, Jian Yang 0003
ICCV4
2025 OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation
abstract
Text-to-video (T2V) generation has recently garnered significant attention thanks to the large multi-modality model Sora. However, T2V generation still faces two important challenges: 1) Lacking a precise open sourced high-quality dataset. The previously popular video datasets, e.g.WebVid-10M and Panda-70M, overly emphasized large scale, resulting in the inclusion of many low-quality videos and short, imprecise captions. Therefore, it is challenging but crucial to collect a precise high-quality dataset while maintaining a scale of millions for T2V generation. 2) Ignoring to fully utilize textual information. Recent T2V methods have focused on vision transformers, using a simple cross attention module for video generation, which falls short of making full use of semantic information from text tokens. To address these issues, we introduce OpenVid-1M, a precise high-quality dataset with expressive captions. This open-scenario dataset contains over 1 million text-video pairs, facilitating research on T2V generation. Furthermore, we curate 433K 1080p videos from OpenVid-1M to create OpenVidHD-0.4M, advancing high-definition video generation. Additionally, we propose a novel Multi-modal Video Diffusion Transformer (MVDiT) capable of mining both structure information from visual tokens and semantic information from text tokens. Extensive experiments and ablation studies verify the superiority of OpenVid-1M over previous datasets and the effectiveness of our MVDiT.
Kepan Nan, Rui Xie 0005, Penghao Zhou, Tiehan Fan, Zhenheng Yang, Xiang Li 0041, Jian Yang 0003, Ying Tai
ICLR7
2025 Rethinking Point Cloud Data Augmentation: Topologically Consistent Deformation
abstract
Data augmentation has been widely used in machine learning. Its main goal is to transform and expand the original data using various techniques, creating a more diverse and enriched training dataset. However, due to the disorder and irregularity of point clouds, existing methods struggle to enrich geometric diversity and maintain topological consistency, leading to imprecise point cloud understanding. In this paper, we propose SinPoint, a novel method designed to preserve the topological structure of the original point cloud through a homeomorphism. It utilizes the Sine function to generate smooth displacements. This simulates object deformations, thereby producing a rich diversity of samples. In addition, we propose a Markov chain Augmentation Process to further expand the data distribution by combining different basic transformations through a random process. Our extensive experiments demonstrate that our method consistently outperforms existing Mixup and Deformation methods on various benchmark point cloud datasets, improving performance for shape classification and part segmentation tasks. Specifically, when used with PointNet++ and DGCNN, our method achieves a state-of-the-art accuracy of 90.2 in shape classification with the real-world ScanObjectNN dataset. We release the code at https://github.com/CSBJian/SinPoint.
Jian Bi, Qianliang Wu, Xiang Li 0041, Shuo Chen 0003, Jianjun Qian, Lei Luo 0001, Jian Yang 0003
ICML3
2025 Deep Height Decoupling for Precise Vision-Based 3D Occupancy Prediction
abstract
The task of vision-based 3D occupancy prediction aims to reconstruct 3D geometry and estimate its semantic classes from 2D color images, where the 2D-to-3D view transformation is an indispensable step. Most previous methods conduct forward projection, such as BEVPooling and VoxelPooling, both of which map the 2D image features into 3D grids. However, the current grid representing features within a certain height range usually introduces many confusing features that belong to other height ranges. To address this challenge, we present Deep Height Decoupling (DHD), a novel framework that incorporates explicit height prior to filter out the confusing features. Specifically, DHD first predicts height maps via explicit supervision. Based on the height distribution statistics, DHD designs Mask Guided Height Sampling (MGHS) to adaptively decouple the height map into multiple binary masks. MGHS projects the 2D image features into multiple subspaces, where each grid contains features within reasonable height ranges. Finally, a Synergistic Feature Aggregation (SFA) module is deployed to enhance the feature representation through channel and spatial affinities, enabling further occupancy refinement. On the popular Occ3D-nuScenes benchmark, our method achieves state-of-the-art performance even with minimal input frames. Source code is released at https://github.com/yanzq95/DHD.
Zhiqiang Yan 0001, Zhengxue Wang, Xiang Li 0041, Le Hui, Jian Yang 0003
ICRA4
2025 Category-Aware 3D Object Composition with Disentangled Texture and Shape Multi-view Diffusion
Zeren Xiong, Zedong Zhang, Xiang Li 0041, Ying Tai, Jian Yang 0003, Jun Li 0027
ACM Multimedia4
2025 Revisiting End-to-End Learning with Slide-level Supervision in Computational Pathology
abstract
Pre-trained encoders for offline feature extraction followed by multiple instance learning (MIL) aggregators have become the dominant paradigm in computational pathology (CPath), benefiting cancer diagnosis and prognosis. However, performance limitations arise from the absence of encoder fine-tuning for downstream tasks and disjoint optimization with MIL. While slide-level supervised end-to-end (E2E) learning is an intuitive solution to this issue, it faces challenges such as high computational demands and suboptimal results. These limitations motivate us to revisit E2E learning. We argue that prior work neglects inherent E2E optimization challenges, leading to performance disparities compared to traditional two-stage methods. In this paper, we pioneer the elucidation of optimization challenge caused by sparse-attention MIL and propose a novel MIL called ABMILX. ABMILX mitigates this problem through global correlation-based attention refinement and multi-head mechanisms. With the efficient multi-scale random patch sampling strategy, an E2E trained ResNet with ABMILX surpasses SOTA foundation models under the two-stage paradigm across multiple challenging benchmarks, while remaining computationally efficient ($<$ 10 RTX3090 GPU hours). We demonstrate the potential of E2E learning in CPath and calls for greater research focus in this area. The code is https://github.com/DearCaat/E2E-WSI-ABMILX.
Fengtao Zhou, Xiang Li 0041, Ming-Ming Cheng
NeurIPS6
2025 See through the Dark: Learning Illumination-affined Representations for Nighttime Occupancy Prediction
abstract
Occupancy prediction aims to estimate the 3D spatial distribution of occupied regions along with their corresponding semantic labels. Existing vision-based methods perform well on daytime benchmarks but struggle in nighttime scenarios due to limited visibility and challenging lighting conditions. To address these challenges, we propose LIAR, a novel framework that learns illumination-affined representations. LIAR first introduces Selective Low-light Image Enhancement (SLLIE), which leverages the illumination priors from daytime scenes to adaptively determine whether a nighttime image is genuinely dark or sufficiently well-lit, enabling more targeted global enhancement. Building on the illumination maps generated by SLLIE, LIAR further incorporates two illumination-aware components: 2D Illumination-guided Sampling (2D-IGS) and 3D Illumination-driven Projection (3D-IDP), to respectively tackle local underexposure and overexposure. Specifically, 2D-IGS modulates feature sampling positions according to illumination maps, assigning larger offsets to darker regions and smaller ones to brighter regions, thereby alleviating feature degradation in underexposed areas. Subsequently, 3D-IDP enhances semantic understanding in overexposed regions by constructing illumination intensity fields and supplying refined residual queries to the BEV context refinement process. Extensive experiments on both real and synthetic datasets demonstrate the superior performance of LIAR under challenging nighttime scenarios. The source code and pretrained models are available [here](https://github.com/yanzq95/LIAR).
Zhiqiang Yan 0001, Yigong Zhang, Xiang Li 0041, Jian Yang 0003
NeurIPS4
2025 Representation Entanglement for Generation: Training Diffusion Transformers Is Much Easier Than You Think
abstract
REPA and its variants effectively mitigate training challenges in diffusion models by incorporating external visual representations from pretrained models, through alignment between the noisy hidden projections of denoising networks and foundational clean image representations. We argue that the external alignment, which is absent during the entire denoising inference process, falls short of fully harnessing the potential of discriminative representations. In this work, we propose a straightforward method called $\textit{$\textbf{R}$epresentation $\textbf{E}$ntanglement for $\textbf{G}$eneration}$ ($\textbf{REG}$), which entangles low-level image latents with a single high-level class token from pretrained foundation models for denoising. REG acquires the capability to produce coherent image-class pairs directly from pure noise, substantially improving both generation quality and training efficiency. This is accomplished with negligible additional inference overhead, requiring only one single additional token for denoising (<0.5\% increase in FLOPs and latency). The inference process concurrently reconstructs both image latents and their corresponding global semantics, where the acquired semantic knowledge actively guides and enhances the image generation process. On ImageNet 256$\times$256, SiT-XL/2 + REG demonstrates remarkable convergence acceleration, achieving $\textbf{63}\times$ and $\textbf{23}\times$ faster training than SiT-XL/2 and SiT-XL/2 + REPA, respectively. More impressively, SiT-L/2 + REG trained for merely 400K iterations outperforms SiT-XL/2 + REPA trained for 4M iterations ($\textbf{10}\times$ longer). Code is available at: https://github.com/Martinser/REG.
Ruijing Shi, Shanghua Gao, Zhenyuan Chen, Lei Wang 0118, Zhaowei Chen, Hongcheng Gao, Jian Yang 0003, Ming-Ming Cheng, Xiang Li 0041
NeurIPS12
2025 SlimHead: Rethinking the Efficiency Bottleneck in Dense Object Detection
Zhaohui Zheng 0003, Ping Wang 0072, Le Zhang 0001, Xiang Li 0041, Qibin Hou, Ming-Ming Cheng
PRCV (16)5
2025 LSKNet: A Foundation Lightweight Backbone for Remote Sensing
Yuxuan Li 0004, Xiang Li 0041, Yimian Dai, Qibin Hou, Li Liu 0004, Yongxiang Liu, Ming-Ming Cheng, Jian Yang 0003
Int. J. Comput. Vis.2
2025 RigNet++: Semantic Assisted Repetitive Image Guided Network for Depth Completion
Zhiqiang Yan 0001, Xiang Li 0041, Le Hui, Zhenyu Zhang 0005, Jun Li 0027, Jian Yang 0003
Int. J. Comput. Vis.2
2025 PointOBB-v3: Expanding Performance Boundaries of Single Point-Supervised Oriented Object Detection
Peiyuan Zhang, Xue Yang 0005, Yi Yu 0010, Qingyun Li, Yue Zhou 0005, Xiaosong Jia, Jingdong Chen, Xiang Li 0041, Junchi Yan, Yansheng Li 0001
Int. J. Comput. Vis.10
2025 YOLO-MS: Rethinking Multi-Scale Representation Learning for Real-Time Object Detection
abstract
We aim at providing the object detection community with an efficient and performant object detector, termed YOLO-MS. The core design is based on a series of investigations on how multi-branch features of the basic block and convolutions with different kernel sizes affect the detection performance of objects at different scales. The outcome is a new strategy that can significantly enhance multi-scale feature representations of real-time object detectors. To verify the effectiveness of our work, we train our YOLO-MS on the MS COCO dataset from scratch without relying on any other large-scale datasets, like ImageNet or pre-trained weights. Without bells and whistles, our YOLO-MS outperforms the recent state-of-the-art real-time object detectors, including YOLO-v7, RTMDet, and YOLO-v8. Taking the XS version of YOLO-MS as an example, it can achieve an AP score of 42+% on MS COCO, which is about 2% higher than RTMDet with the same model size. Furthermore, our work can also serve as a plug-and-play module for other YOLO models. Typically, our method significantly advances the APs, APl, and AP of YOLOv8-N from 18%+, 52%+, and 37%+ to 20%+, 55%+, and 40%+, respectively, with even fewer parameters and MACs.
Xinbin Yuan, Jiabao Wang 0005, Xiang Li 0041, Qibin Hou, Ming-Ming Cheng
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Tri-Perspective View Decomposition for Geometry Aware Depth Completion and Super-Resolution
abstract
Depth completion and super-resolution are crucial tasks for comprehensive RGB-D scene understanding, as they involve reconstructing the precise 3D geometry of a scene from sparse or low-resolution depth measurements. However, most existing methods either rely solely on 2D depth representations or directly incorporate raw 3D point clouds for compensation, which are still insufficient to capture the fine-grained 3D geometry of the scene. In this paper, we introduce Tri-Perspective View Decomposition (TPVD) frameworks that can explicitly model 3D geometry. To this end, (1) TPVD ingeniously decomposes the original 3D point cloud into three 2D views, one of which corresponds to the sparse or low-resolution depth input. (2) For sufficient geometric interaction, TPV Fusion is designed to update the 2D TPV features through recurrent 2D-3D-2D aggregation. (3) By adaptively searching for TPV affinitive neighbors, two additional refinement heads are developed for these two tasks to further improve the geometric consistency. Meanwhile, we build novel datasets named TOFDC for depth completion and TOFDSR for depth super-resolution. Both datasets are acquired using time-of-flight (TOF) sensors and color cameras on smartphones. Extensive experiments on TOFDC, KITTI, NYUv2, SUN RGBD, VKITTI, TOFDSR, RGB-D-D, Lu, and Middlebury datasets indicate that our TPVD outperforms previous depth completion and super-resolution methods, reaching the state of the art.
Zhiqiang Yan 0001, Kun Wang 0042, Xiang Li 0041, Guangwei Gao, Jun Li 0027, Jian Yang 0003
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Fine-Grained Visual Text Prompting
abstract
Vision-Language Models (VLMs), such as CLIP, excel in zero-shot image-level visual understanding but struggle with object-based tasks requiring precise localization and recognition. Visual prompts, like colorful boxes or circles, are suggested to enhance local perception. However, these methods often include irrelevant and noisy pixels, leading to suboptimal performance. The design of better visual prompts and their collaboration with text prompting remains underexplored. This paper introduces Fine-Grained Visual Text Prompting (FGVTP), a new zero-shot framework for object-based tasks using precise semantic masks and reinforced image-text alignment. FGVTP comprises Fine-Grained Visual Prompting (FGVP) and Consistency-Enhanced Text Prompting (CETP). Specifically, we carefully study visual prompting designs by exploring more visual markings that vary in shape and form. FGVP uses semantic masks from a segmenter like the Segment Anything Model (SAM) and employs background blurring (Blur Reverse Mask) to highlight targets while maintaining spatial coherence. Further, CETP enhances image-text alignment by prompting captions based on FGVP-processed images. As a result, FGVTP achieves superior zero-shot referring expression comprehension on RefCOCO/+/g benchmarks, outperforming previous SOTA methods by 5.8% on average. Part detection experiments conducted on the PACO dataset further validate the preponderance of FGVTP over existing works. Code is available at https://github.com/ylingfeng/FGVP.
Lingfeng Yang, Xiang Li 0041, Yueze Wang, Jian Yang 0003
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Non-Aligned Supervision for Real Image Dehazing
abstract
Removing haze from real-world images is challenging due to unpredictable weather conditions, resulting in the misalignment of hazy and clear image pairs. In this paper, we propose an innovative dehazing framework that operates under non-aligned supervision. This framework is grounded in the atmospheric scattering model, and consists of three interconnected networks: dehazing, airlight, and transmission networks. In particular, we explore a non-alignment scenario that a clear reference image, unaligned with the input hazy image, is utilized to supervise the dehazing network. To implement this, we present a multi-scale reference loss that compares the feature representations between the referred image and the dehazed output. Our scenario makes it easier to collect hazy/clear image pairs in real-world environments, even under conditions of misalignment and shift views. To showcase the effectiveness of our scenario, we have collected a new hazy dataset including 415 image pairs captured by mobile Phone in both rural and urban areas, called "Phone-Hazy". Furthermore, we introduce a self-attention network based on mean and variance for modeling real infinite airlight, using the dark channel prior as positional guidance. Experimental results demonstrate the superior performance of our framework over existing state-of-the-art techniques in the real-world image dehazing task. Phone-Hazy and code will be available at https://fanjunkai1.github.io/projectpage/NSDNet/index.html.
Junkai Fan, Xiang Li 0041, Jianjun Qian, Jun Li 0027, Jian Yang 0003
IEEE Trans. Circuits Syst. Video Technol.2
2025 MoCoLSK: Modality-Conditioned High-Resolution Downscaling for Land Surface Temperature
abstract
Land surface temperature (LST) is a critical parameter for environmental studies, but directly obtaining high spatial resolution LST data remains challenging due to the spatiotemporal tradeoff in satellite remote sensing. Guided LST downscaling has emerged as an alternative solution to overcome these limitations, but current methods often neglect spatial nonstationarity, and there is a lack of an open-source ecosystem for deep learning methods. In this article, we propose the modality-conditioned large selective kernel (MoCoLSK) network, a novel architecture that dynamically fuses multimodal data through modality-conditioned projections. MoCoLSK achieves a confluence of dynamic receptive field adjustment and multimodal feature fusion, leading to enhanced LST prediction accuracy. Furthermore, we establish the GrokLST Project, a comprehensive open-source ecosystem featuring the GrokLST dataset, a high-resolution (HR) benchmark, and the GrokLST toolkit, an open-source PyTorch-based toolkit encapsulating MoCoLSK alongside 40+ state-of-the-art approaches. Extensive experimental results validate MoCoLSK’s effectiveness in capturing complex dependencies and subtle variations within multispectral data, outperforming existing methods in LST downscaling. Our code, dataset, and toolkit are available athttps://github.com/GrokCV/GrokLST.
Qun Dai, Chunyang Yuan, Yimian Dai, Yuxuan Li 0004, Xiang Li 0041, Kang Ni, Jianhui Xu, Xiangbo Shu, Jian Yang 0003
IEEE Trans. Geosci. Remote. Sens.5
2025 MSOD: A Large-Scale Multiscene Dataset and a Novel Diagonal-Geometry Loss for SAR Object Detection
abstract
Synthetic Aperture Radar (SAR) has attracted significant attention due to its excellent all-weather imaging capabilities. However, SAR image object detection methods face two major challenges: 1) Most existing datasets are small in volume and single in category and scene. 2) Existing IoU-based loss functions cannot fully capture the relationship between prediction and target bounding boxes. To further advance the development of the SAR object detection method, we construct a large-scale multi-scene SAR object detection dataset called MSOD. It comprises three distinct scenarios, containing 40K images and about 1M instances of interest classified into six categories. In addition, we propose a novel diagonal-based similarity loss, Diagonal-Geometry IoU (DGIoU), to optimize the performance of SAR object detection by measuring the similarity between the diagonal of the prediction and target boxes. Specifically, we equivalently represent a rectangular box as a diagonal, and then define DGIoU based on the similarity of a set of sampling points between the diagonals of the predicted box and the target box. DGIoU effectively characterizes the difference between the predicted box and the target box, particularly in box inclusion and separation cases, resulting in improved localization accuracy. Numerous experimental results demonstrate that MSOD is closer to practical application and more challenging than existing SAR image datasets, and serves as a strong benchmark for evaluating the effectiveness of various IoU loss functions. The dataset and code are available at: https://github.com/wchao0601/MSOD-DGIoU.
Wenxuan Fang 0001, Xiang Li 0041, Jian Yang 0003, Lei Luo 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 AltNeRF: Learning Robust Neural Radiance Field via Alternating Depth-Pose Optimization
abstract
Neural Radiance Fields (NeRF) have shown promise in generating realistic novel views from sparse scene images. However, existing NeRF approaches often encounter challenges due to the lack of explicit 3D supervision and imprecise camera poses, resulting in suboptimal outcomes. To tackle these issues, we propose AltNeRF---a novel framework designed to create resilient NeRF representations using self-supervised monocular depth estimation (SMDE) from monocular videos, without relying on known camera poses. SMDE in AltNeRF masterfully learns depth and pose priors to regulate NeRF training. The depth prior enriches NeRF's capacity for precise scene geometry depiction, while the pose prior provides a robust starting point for subsequent pose refinement. Moreover, we introduce an alternating algorithm that harmoniously melds NeRF outputs into SMDE through a consistence-driven mechanism, thus enhancing the integrity of depth priors. This alternation empowers AltNeRF to progressively refine NeRF representations, yielding the synthesis of realistic novel views. Extensive experiments showcase the compelling capabilities of AltNeRF in generating high-fidelity and robust novel views that closely resemble reality.
Kun Wang 0042, Zhiqiang Yan 0001, Huang Tian, Zhenyu Zhang 0005, Xiang Li 0041, Jun Li 0027, Jian Yang 0003
AAAI5
2024 PromptKD: Unsupervised Prompt Distillation for Vision-Language Models
abstract
Prompt learning has emerged as a valuable technique in enhancing vision-language models (VLMs) such as CLIP for downstream tasks in specific domains. Existing work mainly focuses on designing various learning forms of prompts, neglecting the potential of prompts as effective distillers for learning from larger teacher models. In this paper, we introduce an unsupervised domain prompt distillation framework, which aims to transfer the knowledge of a larger teacher model to a lightweight target model through prompt-driven imitation using unlabeled domain images. Specifically, our framework consists of two distinct stages. In the initial stage, we pre-train a large CLIP teacher model using domain (few-shot) labels. After pretraining, we leverage the unique decoupled-modality characteristics of CLIP by pre-computing and storing the text features as class vectors only once through the teacher text encoder. In the subsequent stage, the stored class vectors are shared across teacher and student image encoders for calculating the predicted logits. Further, we align the logits of both the teacher and student models via KL divergence, encouraging the student image encoder to generate similar probability distributions to the teacher through the learnable prompts. The proposed prompt distillation process eliminates the reliance on labeled data, enabling the algorithm to leverage a vast amount of unlabeled images within the domain. Finally, the well-trained student image encoders and pre-stored text features (class vectors) are utilized for inference. To our best knowledge, we are the first to (1) perform unsupervised domain-specific prompt-driven knowledge distillation for CLIP, and (2) establish a practical pre-storing mechanism of text features as shared class vectors between teacher and student. Extensive experiments on 11 datasets demonstrate the effectiveness of our method. Code is publicly available at https://github.com/zhengli97/PromptKD.
Zheng Li 0028, Xiang Li 0041, Xin Zhang 0170, Weiqiang Wang 0002, Shuo Chen 0003, Jian Yang 0003
CVPR2
2024 CrossKD: Cross-Head Knowledge Distillation for Object Detection
abstract
Knowledge Distillation (KD) has been validated as an effective model compression technique for learning compact object detectors. Existing state-of-the-art KD methods for object detection are mostly based on feature imitation. In this paper, we present a general and effective prediction mimicking distillation scheme, called CrossKD, which delivers the intermediate features of the student's detection head to the teacher's detection head. The resulting crosshead predictions are then forced to mimic the teacher's predictions. This manner relieves the student's head from receiving contradictory supervision signals from the annotations and the teacher's predictions, greatly improving the student's detection performance. Moreover, as mimicking the teacher's predictions is the target of KD, CrossKD offers more task-oriented information in contrast with feature imitation. On MS COCO, with only prediction mimicking losses applied, our CrossKD boosts the average precision of GFL ResNet-50 with 1 × training schedule from 40.2 to 43.7, outperforming all existing KD methods. In addition, our method also works well when distilling detectors with heterogeneous backbones.
Jiabao Wang 0005, Zhaohui Zheng 0003, Xiang Li 0041, Ming-Ming Cheng, Qibin Hou
CVPR4
2024 Distilling Knowledge from Large-Scale Image Models for Object Detection
Wenhai Wang, Xiang Li 0041, Jian Yang 0003, Jifeng Dai, Yu Qiao 0001, Shanshan Zhang 0001
ECCV (84)3
2024 Cascade Prompt Learning for Vision-Language Model Adaptation
Xin Zhang 0170, Zheng Li 0028, Zhaowei Chen, Jiajun Liang, Jian Yang 0003, Xiang Li 0041
ECCV (50)7
2024 SARDet-100K: Towards Open-Source Benchmark and ToolKit for Large-Scale SAR Object Detection
abstract
Synthetic Aperture Radar (SAR) object detection has gained significant attention recently due to its irreplaceable all-weather imaging capabilities. However, this research field suffers from both limited public datasets (mostly comprising <2K images with only mono-category objects) and inaccessible source code. To tackle these challenges, we establish a new benchmark dataset and an open-source method for large-scale SAR object detection. Our dataset, SARDet-100K, is a result of intense surveying, collecting, and standardizing 10 existing SAR detection datasets, providing a large-scale and diverse dataset for research purposes. To the best of our knowledge, SARDet-100K is the first COCO-level large-scale multi-class SAR object detection dataset ever created. With this high-quality dataset, we conducted comprehensive experiments and uncovered a crucial challenge in SAR object detection: the substantial disparities between the pretraining on RGB datasets and finetuning on SAR datasets in terms of both data domain and model structure. To bridge these gaps, we propose a novel Multi-Stage with Filter Augmentation (MSFA) pretraining framework that tackles the problems from the perspective of data input, domain transition, and model migration. The proposed MSFA method significantly enhances the performance of SAR object detection models while demonstrating exceptional generalizability and flexibility across diverse models. This work aims to pave the way for further advancements in SAR object detection. The dataset and code is available at \url{https://github.com/zcablii/SARDet_100K}.
Yuxuan Li 0004, Xiang Li 0041, Qibin Hou, Li Liu 0004, Ming-Ming Cheng, Jian Yang 0003
NeurIPS2
2024 DCDepth: Progressive Monocular Depth Estimation in Discrete Cosine Domain
abstract
In this paper, we introduce DCDepth, a novel framework for the long-standing monocular depth estimation task. Moving beyond conventional pixel-wise depth estimation in the spatial domain, our approach estimates the frequency coefficients of depth patches after transforming them into the discrete cosine domain. This unique formulation allows for the modeling of local depth correlations within each patch. Crucially, the frequency transformation segregates the depth information into various frequency components, with low-frequency components encapsulating the core scene structure and high-frequency components detailing the finer aspects. This decomposition forms the basis of our progressive strategy, which begins with the prediction of low-frequency components to establish a global scene context, followed by successive refinement of local details through the prediction of higher-frequency components. We conduct comprehensive experiments on NYU-Depth-V2, TOFDC, and KITTI datasets, and demonstrate the state-of-the-art performance of DCDepth. Code is available at https://github.com/w2kun/DCDepth.
Kun Wang 0042, Zhiqiang Yan 0001, Junkai Fan, Wanlu Zhu, Xiang Li 0041, Jun Li 0027, Jian Yang 0003
NeurIPS5
2024 Novel Object Synthesis via Adaptive Text-Image Harmony
abstract
In this paper, we study an object synthesis task that combines an object text with an object image to create a new object image. However, most diffusion models struggle with this task, \textit{i.e.}, often generating an object that predominantly reflects either the text or the image due to an imbalance between their inputs. To address this issue, we propose a simple yet effective method called Adaptive Text-Image Harmony (ATIH) to generate novel and surprising objects. First, we introduce a scale factor and an injection step to balance text and image features in cross-attention and to preserve image information in self-attention during the text-image inversion diffusion process, respectively. Second, to better integrate object text and image, we design a balanced loss function with a noise parameter, ensuring both optimal editability and fidelity of the object image. Third, to adaptively adjust these parameters, we present a novel similarity score function that not only maximizes the similarities between the generated object image and the input text/image but also balances these similarities to harmonize text and image integration. Extensive experiments demonstrate the effectiveness of our approach, showcasing remarkable object creations such as colobus-glass jar. https://xzr52.github.io/ATIH/
Zeren Xiong, Zedong Zhang, Shuo Chen 0003, Xiang Li 0041, Gan Sun, Jian Yang 0003, Jun Li 0027
NeurIPS5
2024 Zone Evaluation: Revealing Spatial Bias in Object Detection
abstract
A fundamental limitation of object detectors is that they suffer from "spatial bias", and in particular perform less satisfactorily when detecting objects near image borders. For a long time, there has been a lack of effective ways to measure and identify spatial bias, and little is known about where it comes from and what degree it is. To this end, we present a new zone evaluation protocol, extending from the traditional evaluation to a more generalized one, which measures the detection performance over zones, yielding a series of Zone Precisions (ZPs). For the first time, we provide numerical results, showing that the object detectors perform quite unevenly across the zones. Surprisingly, the detector's performance in the 96% border zone of the image does not reach the AP value (Average Precision, commonly regarded as the average detection performance in the entire image zone). To better understand spatial bias, a series of heuristic experiments are conducted. Our investigation excludes two intuitive conjectures about spatial bias that the object scale and the absolute positions of objects barely influence the spatial bias. We find that the key lies in the human-imperceptible divergence in data patterns between objects in different zones, thus eventually forming a visible performance gap between the zones. With these findings, we finally discuss a future direction for object detection, namely, spatial disequilibrium problem, aiming at pursuing a balanced detection ability over the entire image zone. By broadly evaluating 10 popular object detectors and 5 detection datasets, we shed light on the spatial bias of object detectors. We hope this work could raise a focus on detection robustness.
Zhaohui Zheng 0003, Qibin Hou, Xiang Li 0041, Ping Wang 0072, Ming-Ming Cheng
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Dual teachers for self-knowledge distillation
Zheng Li 0028, Xiang Li 0041, Lingfeng Yang, Renjie Song, Jian Yang 0003
Pattern Recognit.2
2024 Pick of the Bunch: Detecting Infrared Small Targets Beyond Hit-Miss Trade-Offs via Selective Rank-Aware Attention
abstract
Infrared small target detection faces the inherent challenge of precisely localizing dim targets amidst complex background clutter. Traditional approaches struggle to balance detection precision and false alarm rates. To break this dilemma, we propose SeRankDet, a deep network that achieves high accuracy beyond the conventional hit-miss trade-off, by following the “Pick of the Bunch” principle. At its core lies our selective rank-aware attention (SeRank) module, employing a nonlinear Top-K selection process that preserves the most salient responses, preventing target signal dilution while maintaining constant complexity. Furthermore, we replace the static concatenation typical in U-Net structures with our large selective feature fusion (LSFF) module, a dynamic fusion strategy that empowers SeRankDet with adaptive feature integration, enhancing its ability to discriminate true targets from false alarms. The network’s discernment is further refined by our dilated difference convolution (DDC) module, which merges differential convolution aimed at amplifying subtle target characteristics with dilated convolution to expand the receptive field, thereby substantially improving target-background separation. Despite its lightweight architecture, the proposed SeRankDet sets new benchmarks in state-of-the-art performance across multiple public datasets. The code is available athttps://github.com/GrokCV/SeRankDet.
Yimian Dai, Peiwen Pan, Yulei Qian, Yuxuan Li 0004, Xiang Li 0041, Jian Yang 0003, Huan Wang 0013
IEEE Trans. Geosci. Remote. Sens.5
2024 Learning Complementary Correlations for Depth Super-Resolution With Incomplete Data in Real World
abstract
Depth information is a significant ingredient to visually perceive the physical world. However, mainstream depth sensors, e.g., time-of-flight (ToF) cameras, often measure incomplete and low-resolution depth data, resulting in low-quality visual perception. In this article, we try to address a potentially valuable task, i.e., depth super-resolution (DSR) with incomplete data, which recovers dense and high-resolution depth map from incomplete and low-resolution one. To tackle this task, we introduce a novel incomplete DSR (IDSR) framework, including a primary branch for DSR to recover high-frequency details, and an auxiliary branch for depth completion (DC) to fill missing pixels. More importantly, we propose two modules, joint correlation learning (JCL) and iterative-cross (IC), to enhance the learning of complementary information flows between the two branches. The former module aims to learn the correlative relationships of the two branches, whilst the latter module adequately fuses higher level representations for more precise predictions. Extensive experiments show that our framework is effective and achieves the state-of-the-art performance on the real-world RGB-D-D and the synthetic NYUv2 datasets.
Zhiqiang Yan 0001, Kun Wang 0042, Xiang Li 0041, Zhenyu Zhang 0005, Jun Li 0027, Jian Yang 0003
IEEE Trans. Neural Networks Learn. Syst.3
2023 Curriculum Temperature for Knowledge Distillation
abstract
Most existing distillation methods ignore the flexible role of the temperature in the loss function and fix it as a hyper-parameter that can be decided by an inefficient grid search. In general, the temperature controls the discrepancy between two distributions and can faithfully determine the difficulty level of the distillation task. Keeping a constant temperature, i.e., a fixed level of task difficulty, is usually sub-optimal for a growing student during its progressive learning stages. In this paper, we propose a simple curriculum-based technique, termed Curriculum Temperature for Knowledge Distillation (CTKD), which controls the task difficulty level during the student's learning career through a dynamic and learnable temperature. Specifically, following an easy-to-hard curriculum, we gradually increase the distillation loss w.r.t. the temperature, leading to increased distillation difficulty in an adversarial manner. As an easy-to-use plug-in technique, CTKD can be seamlessly integrated into existing knowledge distillation frameworks and brings general improvements at a negligible additional computation cost. Extensive experiments on CIFAR-100, ImageNet-2012, and MS-COCO demonstrate the effectiveness of our method.
Zheng Li 0028, Xiang Li 0041, Lingfeng Yang, Borui Zhao, Renjie Song, Lei Luo 0001, Jun Li 0027, Jian Yang 0003
AAAI2
2023 DesNet: Decomposed Scale-Consistent Network for Unsupervised Depth Completion
abstract
Unsupervised depth completion aims to recover dense depth from the sparse one without using the ground-truth annotation. Although depth measurement obtained from LiDAR is usually sparse, it contains valid and real distance information, i.e., scale-consistent absolute depth values. Meanwhile, scale-agnostic counterparts seek to estimate relative depth and have achieved impressive performance. To leverage both the inherent characteristics, we thus suggest to model scale-consistent depth upon unsupervised scale-agnostic frameworks. Specifically, we propose the decomposed scale-consistent learning (DSCL) strategy, which disintegrates the absolute depth into relative depth prediction and global scale estimation, contributing to individual learning benefits. But unfortunately, most existing unsupervised scale-agnostic frameworks heavily suffer from depth holes due to the extremely sparse depth input and weak supervisory signal. To tackle this issue, we introduce the global depth guidance (GDG) module, which attentively propagates dense depth reference into the sparse target via novel dense-to-sparse attention. Extensive experiments show the superiority of our method on outdoor KITTI, ranking 1st and outperforming the best KBNet more than 12% in RMSE. Additionally, our approach achieves state-of-the-art performance on indoor NYUv2 benchmark as well.
Zhiqiang Yan 0001, Kun Wang 0042, Xiang Li 0041, Zhenyu Zhang 0005, Jun Li 0027, Jian Yang 0003
AAAI3
2023 Recurrent Structure Attention Guidance for Depth Super-resolution
abstract
Image guidance is an effective strategy for depth super-resolution. Generally, most existing methods employ hand-crafted operators to decompose the high-frequency (HF) and low-frequency (LF) ingredients from low-resolution depth maps and guide the HF ingredients by directly concatenating them with image features. However, the hand-designed operators usually cause inferior HF maps (e.g., distorted or structurally missing) due to the diverse appearance of complex depth maps. Moreover, the direct concatenation often results in weak guidance because not all image features have a positive effect on the HF maps. In this paper, we develop a recurrent structure attention guided (RSAG) framework, consisting of two important parts. First, we introduce a deep contrastive network with multi-scale filters for adaptive frequency-domain separation, which adopts contrastive networks from large filters to small ones to calculate the pixel contrasts for adaptive high-quality HF predictions. Second, instead of the coarse concatenation guidance, we propose a recurrent structure attention block, which iteratively utilizes the latest depth estimation and the image features to jointly select clear patterns and boundaries, aiming at providing refined guidance for accurate depth recovery. In addition, we fuse the features of HF maps to enhance the edge structures in the decomposed LF maps. Extensive experiments show that our approach obtains superior performance compared with state-of-the-art depth super-resolution methods. Our code is available at: https://github.com/Yuanjiayii/DSR-RSAG.
Jiayi Yuan 0003, Haobo Jiang, Xiang Li 0041, Jianjun Qian, Jun Li 0027, Jian Yang 0003
AAAI3
2023 Structure Flow-Guided Network for Real Depth Super-resolution
abstract
Real depth super-resolution (DSR), unlike synthetic settings, is a challenging task due to the structural distortion and the edge noise caused by the natural degradation in real-world low-resolution (LR) depth maps. These defeats result in significant structure inconsistency between the depth map and the RGB guidance, which potentially confuses the RGB-structure guidance and thereby degrades the DSR quality. In this paper, we propose a novel structure flow-guided DSR framework, where a cross-modality flow map is learned to guide the RGB-structure information transferring for precise depth upsampling. Specifically, our framework consists of a cross-modality flow-guided upsampling network (CFUNet) and a flow-enhanced pyramid edge attention network (PEANet). CFUNet contains a trilateral self-attention module combining both the geometric and semantic correlations for reliable cross-modality flow learning. Then, the learned flow maps are combined with the grid-sampling mechanism for coarse high-resolution (HR) depth prediction. PEANet targets at integrating the learned flow map as the edge attention into a pyramid network to hierarchically learn the edge-focused guidance feature for depth edge refinement. Extensive experiments on real and synthetic DSR datasets verify that our approach achieves excellent performance compared to state-of-the-art methods. Our code is available at: https://github.com/Yuanjiayii/DSR-SFG.
Jiayi Yuan 0003, Haobo Jiang, Xiang Li 0041, Jianjun Qian, Jun Li 0027, Jian Yang 0003
AAAI3
2023 Large Selective Kernel Network for Remote Sensing Object Detection
abstract
Recent research on remote sensing object detection has largely focused on improving the representation of oriented bounding boxes but has overlooked the unique prior knowledge presented in remote sensing scenarios. Such prior knowledge can be useful because tiny remote sensing objects may be mistakenly detected without referencing a sufficiently long-range context, which can vary for different objects. This paper considers these priors and proposes the lightweight Large Selective Kernel Network (LSKNet). LSKNet can dynamically adjust its large spatial receptive field to better model the ranging context of various objects in remote sensing scenarios. To our knowledge, large and selective kernel mechanisms have not been previously explored in remote sensing object detection. Without bells and whistles, our lightweight LSKNet sets new state-of-the-art scores on standard benchmarks, i.e., HRSC2016 (98.46% mAP), DOTA-v1.0 (81.85% mAP), and FAIR1M-v1.0 (47.87% mAP).
Yuxuan Li 0004, Qibin Hou, Zhaohui Zheng 0003, Ming-Ming Cheng, Jian Yang 0003, Xiang Li 0041
ICCV6
2023 Creative Birds: Self-Supervised Single-View 3D Style Transfer
abstract
In this paper, we propose a novel method for single-view 3D style transfer that generates a unique 3D object with both shape and texture transfer. Our focus lies primarily on birds, a popular subject in 3D reconstruction, for which no existing single-view 3D transfer methods have been developed. The method we propose seeks to generate a 3D mesh shape and texture of a bird from two single-view images. To achieve this, we introduce a novel shape transfer generator that comprises a dual residual gated network (DRGNet), and a multi-layer perceptron (MLP). DRGNet extracts the features of source and target images using a shared coordinate gate unit, while the MLP generates spatial coordinates for building a 3D mesh. We also introduce a semantic UV texture transfer module that implements textural style transfer using semantic UV segmentation, which ensures consistency in the semantic meaning of the transferred regions. This module can be widely adapted to many existing approaches. Finally, our method constructs a novel 3D bird using a differentiable renderer. Experimental results on the CUB dataset verify that our method achieves state-of-the-art performance on the single-view 3D style transfer task. Code is available at https://github.com/wrk226/creative_birds.
Renke Wang, Guimin Que, Shuo Chen 0003, Xiang Li 0041, Jun Li 0027, Jian Yang 0003
ICCV4
2023 ADNet: Lane Shape Prediction via Anchor Decomposition
abstract
In this paper, we revisit the limitations of anchor-based lane detection methods, which have predominantly focused on fixed anchors that stem from the edges of the image, disregarding their versatility and quality. To overcome the inflexibility of anchors, we decompose them into learning the heat map of starting points and their associated directions. This decomposition removes the limitations on the starting point of anchors, making our algorithm adaptable to different lane types in various datasets. To enhance the quality of anchors, we introduce the Large Kernel Attention (LKA) for Feature Pyramid Network (FPN). This significantly increases the receptive field, which is crucial in capturing the sufficient context as lane lines typically run throughout the entire image. We have named our proposed system the Anchor Decomposition Network (ADNet). Additionally, we propose the General Lane IoU (GLIoU) loss, which significantly improves the performance of ADNet in complex scenarios. Experimental results on three widely used lane detection benchmarks, VIL-100, CU-Lane, and TuSimple, demonstrate that our approach outperforms the state-of-the-art methods on VIL-100 and exhibits competitive accuracy on CULane and TuSimple. Code and models will be released on https://github.com/Sephirex-X/ADNet.
Lingyu Xiao, Xiang Li 0041, Wankou Yang
ICCV2
2023 Distortion and Uncertainty Aware Loss for Panoramic Depth Completion
abstract
Standard MSE or MAE loss function is commonly used in limited field-of-vision depth completion, treating each pixel equally under a basic assumption that all pixels have same contribution during optimization. Recently, with the rapid rise of panoramic photography, panoramic depth completion (PDC) has raised increasing attention in 3D computer vision. However, the assumption is inapplicable to panoramic data due to its latitude-wise distortion and high uncertainty nearby textures and edges. To handle these challenges, we propose distortion and uncertainty aware loss (DUL) that consists of a distortion-aware loss and an uncertainty-aware loss. The distortion-aware loss is designed to tackle the panoramic distortion caused by equirectangular projection, whose coordinate transformation relation is used to adaptively calculate the weight of the latitude-wise distortion, distributing uneven importance instead of the equal treatment for each pixel. The uncertainty-aware loss is presented to handle the inaccuracy in non-smooth regions. Specifically, we characterize uncertainty into PDC solutions under Bayesian deep learning framework, where a novel consistent uncertainty estimation constraint is designed to learn the consistency between multiple uncertainty maps of a single panorama. This consistency constraint allows model to produce more precise uncertainty estimation that is robust to feature deformation. Extensive experiments show the superiority of our method over standard loss functions, reaching the state of the art.
Zhiqiang Yan 0001, Xiang Li 0041, Kun Wang 0042, Shuo Chen 0003, Jun Li 0027, Jian Yang 0003
ICML2
2023 Fine-Grained Visual Prompting
abstract
Vision-Language Models (VLMs), such as CLIP, have demonstrated impressive zero-shot transfer capabilities in image-level visual perception. However, these models have shown limited performance in instance-level tasks that demand precise localization and recognition. Previous works have suggested that incorporating visual prompts, such as colorful boxes or circles, can improve the ability of models to recognize objects of interest. Nonetheless, compared to language prompting, visual prompting designs are rarely explored. Existing approaches, which employ coarse visual cues such as colorful boxes or circles, often result in sub-optimal performance due to the inclusion of irrelevant and noisy pixels. In this paper, we carefully study the visual prompting designs by exploring more fine-grained markings, such as segmentation masks and their variations. In addition, we introduce a new zero-shot framework that leverages pixel-level annotations acquired from a generalist segmentation model for fine-grained visual prompting. Consequently, our investigation reveals that a straightforward application of blur outside the target mask, referred to as the Blur Reverse Mask, exhibits exceptional effectiveness. This proposed prompting strategy leverages the precise mask annotations to reduce focus on weakly related regions while retaining spatial coherence between the target and the surrounding background. Our **F**ine-**G**rained **V**isual **P**rompting (**FGVP**) demonstrates superior performance in zero-shot comprehension of referring expressions on the RefCOCO, RefCOCO+, and RefCOCOg benchmarks. It outperforms prior methods by an average margin of 3.0\% to 4.6\%, with a maximum improvement of 12.5\% on the RefCOCO+ testA subset. The part detection experiments conducted on the PACO dataset further validate the preponderance of FGVP over existing visual prompting techniques. Code is available at https://github.com/ylingfeng/FGVP.
Lingfeng Yang, Yueze Wang, Xiang Li 0041, Jian Yang 0013
NeurIPS3
2023 APF-GAN: Exploring asymmetric pre-training and fine-tuning strategy for conditional generative adversarial network
abstract
The use of generative adversarial network (GAN)based models for the conditional generation of image semantic segmentation has shown promising results in recent years.However, there are still some limitations, including limited diversity of image style, distortion of detailed texture, unbalanced color tone, and lengthy training time.To address these issues, we propose an asymmetric pre-training and fine-tuning (APF)-GAN model.In the pretraining phase, we introduce a progressive growing mechanism for pix2pix conditional GAN frameworks to efficiently generate high-quality images with details.Subsequently, in the fine-tuning phase, we introduce novel semantic spatially-guided noise to improve the robustness of the model and increase style diversity.The proposed algorithm outperformed the high-performance GauGAN model and won the championship of the Second Jittor Artificial Intelligence Challenge.Our model was implemented in the Jittor framework and is available at https:// github.com/zcablii/jittor-Torile-PG_SPADE. APF-GAN Asymmetric pre-training and fine-tuning strategyIn this study, we propose an asymmetric pretraining and fine-tuning strategy for the conditional generative adversarial network (APF-GAN) model
Yuxuan Li 0004, Lingfeng Yang, Xiang Li 0041
Comput. Vis. Media3
2023 Denseformer: A dense transformer framework for person re-identification
abstract
Abstract Transformer has shown its effectiveness and advantage in many computer vision tasks, for example, image classification and object re‐identification (ReID). However, existing vision transformers are stacked layer by layer, lacking direct information exchange among every layer. Inspired by DenseNet, we propose a dense transformer framework (termed Denseformer) that connects each layer to every other layer through class tokens. We demonstrate that Denseformer can consistently achieve better performance on person ReID tasks across datasets (Market‐1501, DukeMTMC, MSMT17, and Occluded‐Duke), only at a negligible increase of computation. We show that Denseformer has several compelling advantages: it pays more attention to the main parts of human bodies and obtains discriminative global features.
Haoyan Ma, Xiang Li 0041, Xia Yuan, Chunxia Zhao
IET Comput. Vis.2
2023 Two-phase self-supervised pretraining for object re-identification
Haoyan Ma, Xiang Li 0041, Xia Yuan, Chunxia Zhao
Knowl. Based Syst.2
2023 Boundary-restricted metric learning
Shuo Chen 0003, Chen Gong 0002, Xiang Li 0041, Jian Yang 0003, Gang Niu 0001, Masashi Sugiyama
Mach. Learn.3
2023 Generalized Focal Loss: Towards Efficient Representation Learning for Dense Object Detection
abstract
Object detection is a fundamental computer vision task that simultaneously predicts the category and localization of the targets of interest. Recently one-stage (also termed "dense") detectors have gained much attention over two-stage ones due to their simple pipeline and friendly application to end devices. Dense object detectors basically formulate object detection as dense classification and localization (i.e., bounding box regression). The classification is usually optimized by Focal Loss and the box location is commonly learned under Dirac delta distribution. A recent trend for dense detectors is to introduce an individual prediction branch to estimate the quality of localization, which facilitates the classification to improve detection performance. This paper delves into the representations of the above three fundamental elements: quality estimation, classification and localization. Three problems are discovered in existing practices, including (1) the inconsistent usage of the quality estimation and classification between training and inference, (2) the inflexible Dirac delta distribution for localization, and (3) the deficient and implicit guidance for accurate quality estimation. To address these problems, we design new representations for these elements. Specifically, we merge the quality estimation into the class prediction vector to form a joint representation, use a vector to represent arbitrary distribution of box locations, and extract discriminant feature descriptors from the distribution vector for more reliable quality estimation. The improved representations eliminate the inconsistency risk and accurately depict the flexible distribution in real data, but contain continuous labels, which is beyond the scope of Focal Loss. We then propose Generalized Focal Loss (GFocal) that generalizes Focal Loss from its discrete form to the continuous version for successful optimization. Extensive experiments demonstrate the effectiveness of our method, without sacrificing the efficiency both in training and inference. Based on GFocal, we construct a considerably fast and lightweight detector termed NanoDet under mobile settings, which is 1.8 AP higher, 2x faster and 6x smaller than scaled YoloV4-Tiny.
Xiang Li 0041, Chengqi Lv, Wenhai Wang, Lingfeng Yang, Jian Yang 0003
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 One-Stage Cascade Refinement Networks for Infrared Small Target Detection
abstract
Single-frame infrared small target (SIRST) detection has been a challenging task due to a lack of inherent characteristics, imprecise bounding box regression, a scarcity of real-world datasets, and sensitive localization evaluation. In this article, we propose a comprehensive solution to these challenges. First, we find that the existing anchor-free label assignment method is prone to mislabeling small targets as background, leading to their omission by detectors. To overcome this issue, we propose an all-scale pseudobox-based label assignment scheme that relaxes the constraints on the scale and decouples the spatial assignment from the size of the ground-truth target. Second, motivated by the structured prior of feature pyramids, we introduce the one-stage cascade refinement network (OSCAR), which uses the high-level head as soft proposal for the low-level refinement head. This allows OSCAR to process the same target in a cascade coarse-to-fine manner. Finally, we present a new research benchmark for infrared small target detection, consisting of the SIRST-V2 dataset of real-world, high-resolution single-frame targets, the normalized contrast evaluation metric, and the DeepInfrared toolkit for detection. We conduct extensive ablation studies to evaluate the components of OSCAR and compare its performance to state-of-the-art model- and data-driven methods on the SIRST-V2 benchmark. Our results demonstrate that a top-down cascade refinement framework can improve the accuracy of infrared small target detection without sacrificing efficiency. The DeepInfrared toolkit, dataset, and trained models are available athttps://github.com/YimianDai/open-deepinfrared.
Yimian Dai, Xiang Li 0041, Fei Zhou 0006, Yulei Qian, Yaohong Chen, Jian Yang 0003
IEEE Trans. Geosci. Remote. Sens.2
2022 Knowledge Distillation for Object Detection via Rank Mimicking and Prediction-Guided Feature Imitation
abstract
Knowledge Distillation (KD) is a widely-used technology to inherit information from cumbersome teacher models to compact student models, consequently realizing model compression and acceleration. Compared with image classification, object detection is a more complex task, and designing specific KD methods for object detection is non-trivial. In this work, we elaborately study the behaviour difference between the teacher and student detection models, and obtain two intriguing observations: First, the teacher and student rank their detected candidate boxes quite differently, which results in their precision discrepancy. Second, there is a considerable gap between the feature response differences and prediction differences between teacher and student, indicating that equally imitating all the feature maps of the teacher is the sub-optimal choice for improving the student's accuracy. Based on the two observations, we propose Rank Mimicking (RM) and Prediction-guided Feature Imitation (PFI) for distilling one-stage detectors, respectively. RM takes the rank of candidate boxes from teachers as a new form of knowledge to distill, which consistently outperforms the traditional soft label distillation. PFI attempts to correlate feature differences with prediction differences, making feature imitation directly help to improve the student's accuracy. On MS COCO and PASCAL VOC benchmarks, extensive experiments are conducted on various detectors with different backbones to validate the effectiveness of our method. Specifically, RetinaNet with ResNet50 achieves 40.4% mAP on MS COCO, which is 3.5% higher than its baseline, and also outperforms previous KD methods.
Xiang Li 0041, Shanshan Zhang 0001, Yichao Wu, Ding Liang
AAAI2
2022 Spatial Group-Wise Enhance: Enhancing Semantic Feature Learning in CNN
Yuxuan Li 0004, Xiang Li 0041, Jian Yang 0003
ACCV (5)2
2022 Dynamic MLP for Fine-Grained Image Classification by Leveraging Geographical and Temporal Information
abstract
Fine-grained image classification is a challenging computer vision task where various species share similar visual appearances, resulting in misclassification if merely based on visual clues. Therefore, it is helpful to leverage additional information, e.g., the locations and dates for data shooting, which can be easily accessible but rarely exploited. In this paper, we first demonstrate that existing multimodal methods fuse multiple features only on a single dimension, which essentially has insufficient help in feature discrimination. To fully explore the potential of multimodal information, we propose a dynamic MLP on top of the image representation, which interacts with multimodal features at a higher and broader dimension. The dynamic MLP is an efficient structure parameterized by the learned embeddings of variable locations and dates. It can be regarded as an adaptive nonlinear projection for generating more discriminative image representations in visual tasks. To our best knowledge, it is the first attempt to explore the idea of dynamic networks to exploit multimodal information in fine-grained image classification tasks. Extensive experiments demonstrate the effectiveness of our method. The t-SNE algorithm visually indicates that our technique improves the recognizability of image representations that are visually similar but with different categories. Furthermore, among published works across multiple fine-grained datasets, dynamic MLP consistently achieves SOTA results11https://paperswithcode.com/dataset/inaturalist and takes third place in the iNaturalist challenge at FGVC822https://www.kaggle.com/c/inaturalist-2021/leaderboard. Code is available at httpsr//glthub.com/megvii-research/DynamicMLPForFinegrained.
Lingfeng Yang, Xiang Li 0041, Renjie Song, Borui Zhao, Juntian Tao, Jiajun Liang, Jian Yang 0003
CVPR2
2022 PseCo: Pseudo Labeling and Consistency Training for Semi-Supervised Object Detection
Xiang Li 0041, Yichao Wu, Ding Liang, Shanshan Zhang 0001
ECCV (9)2
2022 Multi-modal Masked Pre-training for Monocular Panoramic Depth Completion
Zhiqiang Yan 0001, Xiang Li 0041, Kun Wang 0042, Zhenyu Zhang 0005, Jun Li 0027, Jian Yang 0003
ECCV (1)2
2022 RigNet: Repetitive Image Guided Network for Depth Completion
Zhiqiang Yan 0001, Kun Wang 0042, Xiang Li 0041, Zhenyu Zhang 0005, Jun Li 0027, Jian Yang 0003
ECCV (27)3
2022 PPT: Anomaly Detection Dataset of Printed Products with Templates
abstract
Visual anomaly detection has been an active topic in industrial applications. In particular, it aims to classify anomalies and precisely locate defective areas in the printed products. To the best of our knowledge, there is no anomaly detection dataset for industrial printings. In this paper, we are the first to introduce a Printed Products with Templates (PPT) dataset, which contains large templates and sliced images collected from industry scene images. PPT is a challenging dataset with more variable surface defects and more disturbing background than existing related benchmarks. Furthermore, we propose a template matching method for anomaly detection of printed products, which consists of a fast template matching block with a convolutional operation using the test sliced image as its kernel, and a prediction network for generating an anomaly map of the test sliced image. Experimental results show that our method achieves state-of-the-art performance compared to the related anomaly detection approaches.
Huang Tian, Xiang Li 0041, Lingfeng Yang, Jun Li 0027, Jian Yang 0003, Weidong Du
ICIP2
2022 DTG-SSOD: Dense Teacher Guidance for Semi-Supervised Object Detection
abstract
The Mean-Teacher (MT) scheme is widely adopted in semi-supervised object detection (SSOD). In MT, sparse pseudo labels, offered by the final predictions of the teacher (e.g., after Non Maximum Suppression (NMS) post-processing), are adopted for the dense supervision for the student via hand-crafted label assignment. However, the "sparse-to-dense'' paradigm complicates the pipeline of SSOD, and simultaneously neglects the powerful direct, dense teacher supervision. In this paper, we attempt to directly leverage the dense guidance of teacher to supervise student training, i.e., the "dense-to-dense'' paradigm. Specifically, we propose the Inverse NMS Clustering (INC) and Rank Matching (RM) to instantiate the dense supervision, without the widely used, conventional sparse pseudo labels. INC leads the student to group candidate boxes into clusters in NMS as the teacher does, which is implemented by learning grouping information revealed in NMS procedure of the teacher. After obtaining the same grouping scheme as the teacher via INC, the student further imitates the rank distribution of the teacher over clustered candidates through Rank Matching. With the proposed INC and RM, we integrate Dense Teacher Guidance into Semi-Supervised Object Detection (termed "DTG-SSOD''), successfully abandoning sparse pseudo labels and enabling more informative learning on unlabeled data. On COCO benchmark, our DTG-SSOD achieves state-of-the-art performance under various labelling ratios. For example, under 10% labelling ratio, DTG-SSOD improves the supervised baseline from 26.9 to 35.9 mAP, outperforming the previous best method Soft Teacher by 1.9 points.
Xiang Li 0041, Yichao Wu, Ding Liang, Shanshan Zhang 0001
NeurIPS2
2022 RecursiveMix: Mixed Learning with History
abstract
Mix-based augmentation has been proven fundamental to the generalization of deep vision models. However, current augmentations only mix samples from the current data batch during training, which ignores the possible knowledge accumulated in the learning history. In this paper, we propose a recursive mixed-sample learning paradigm, termed ``RecursiveMix'' (RM), by exploring a novel training strategy that leverages the historical input-prediction-label triplets. More specifically, we iteratively resize the input image batch from the previous iteration and paste it into the current batch while their labels are fused proportionally to the area of the operated patches. Furthermore, a consistency loss is introduced to align the identical image semantics across the iterations, which helps the learning of scale-invariant feature representations. Based on ResNet-50, RM largely improves classification accuracy by $\sim$3.2% on CIFAR-100 and $\sim$2.8% on ImageNet with negligible extra computation/storage costs. In the downstream object detection task, the RM-pretrained model outperforms the baseline by 2.1 AP points and surpasses CutMix by 1.4 AP points under the ATSS detector on COCO. In semantic segmentation, RM also surpasses the baseline and CutMix by 1.9 and 1.1 mIoU points under UperNet on ADE20K, respectively. Codes and pretrained models are available at https://github.com/implus/RecursiveMix.
Lingfeng Yang, Xiang Li 0041, Borui Zhao, Renjie Song, Jian Yang 0003
NeurIPS2
2022 PAN++: Towards Efficient and Accurate End-to-End Spotting of Arbitrarily-Shaped Text
abstract
Scene text detection and recognition have been well explored in the past few years. Despite the progress, efficient and accurate end-to-end spotting of arbitrarily-shaped text remains challenging. In this work, we propose an end-to-end text spotting framework, termed PAN++, which can efficiently detect and recognize text of arbitrary shapes in natural scenes. PAN++ is based on the kernel representation that reformulates a text line as a text kernel (central region) surrounded by peripheral pixels. By systematically comparing with existing scene text representations, we show that our kernel representation can not only describe arbitrarily-shaped text but also well distinguish adjacent text. Moreover, as a pixel-based representation, the kernel representation can be predicted by a single fully convolutional network, which is very friendly to real-time applications. Taking the advantages of the kernel representation, we design a series of components as follows: 1) a computationally efficient feature enhancement network composed of stacked Feature Pyramid Enhancement Modules (FPEMs); 2) a lightweight detection head cooperating with Pixel Aggregation (PA); and 3) an efficient attention-based recognition head with Masked RoI. Benefiting from the kernel representation and the tailored components, our method achieves high inference speed while maintaining competitive accuracy. Extensive experiments show the superiority of our method. For example, the proposed PAN++ achieves an end-to-end text spotting F-measure of 64.9 at 29.2 FPS on the Total-Text dataset, which significantly outperforms the previous best method. Code will be available at: git.io/PAN.
Wenhai Wang, Enze Xie, Xiang Li 0041, Xuebo Liu 0001, Ding Liang, Zhibo Yang 0003, Tong Lu 0002, Chunhua Shen
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 CBi-GNN: Cross-Scale Bilateral Graph Neural Network for 3D Object Detection
abstract
3D object detection from LiDAR point clouds is a challenging task, since the point clouds are irregular and sparse. Existing one-stage methods mainly predict the 3D bounding box of 3D objects by extracting deep down-scaled features of point clouds from low-level (high-resolution, HR) feature maps to high-level (low-resolution, LR). Nonetheless, most of these methods ignore geometric context information of the down-scaled feature maps across scales, especially only using the LR feature will result in incomplete structure and less location accuracy of 3D objects. In this paper, we propose a novel cross-scale graph network-based one-stage 3D object detector to fully exploit the geometric contexts of the voxels between the down-scaled feature maps. Specifically, we first employ a 3D sparse convolution neural network to form different resolutions of feature maps of voxels. We then dynamically construct a cross-scale bilateral graph to search the neighbor non-empty voxels in the HR feature map with a fixed radius for each non-empty voxel in the LR feature map. In the constructed graph, we present a bilateral attention mechanism (i.e., self-attention and spatial attention) in the HR feature map and encode each non-empty voxel in the LR feature map by aggregating the HR features to obtain the attention features. In addition, we design a non-local part pooling operation to improve the score of the detected bounding box of 3D objects. Finally, we formulate a multi-task loss to train our network for regression of the 3D bounding box of the 3D objects. Experiments on the challenging KITTI’s 3D/BEV benchmark show that our proposed detector outperforms all one-stage 3D object detectors and is comparable to two-stage 3D object detectors. Our code is available athttps://github.com/csjxchen/CBi-GNN.
Jiaxin Chen 0001, Xiang Li 0041, Jin Xie 0001, Jun Li 0027, Jianjun Qian, Jian Yang 0003
IEEE Trans. Intell. Transp. Syst.2
2021 Generalized Focal Loss V2: Learning Reliable Localization Quality Estimation for Dense Object Detection
abstract
Localization Quality Estimation (LQE) is crucial and popular in the recent advancement of dense object detectors since it can provide accurate ranking scores that benefit the Non-Maximum Suppression processing and improve detection performance. As a common practice, most existing methods predict LQE scores through vanilla convolutional features shared with object classification or bounding box regression. In this paper, we explore a completely novel and different perspective to perform LQE – based on the learned distributions of the four parameters of the bounding box. The bounding box distributions are inspired and introduced as "General Distribution" in GFLV1, which describes the uncertainty of the predicted bounding boxes well. Such a property makes the distribution statistics of a bounding box highly correlated to its real localization quality. Specifically, a bounding box distribution with a sharp peak usually corresponds to high localization quality, and vice versa. By leveraging the close correlation between distribution statistics and the real localization quality, we develop a considerably lightweight Distribution-Guided Quality Predictor (DGQP) for reliable LQE based on GFLV1, thus producing GFLV2. To our best knowledge, it is the first attempt in object detection to use a highly relevant, statistical representation to facilitate LQE. Extensive experiments demonstrate the effectiveness of our method. Notably, GFLV2 (ResNet101) achieves 46.2 AP at 14.6 FPS, surpassing the previous state-of-the-art ATSS baseline (43.6 AP at 14.6 FPS) by absolute 2.6 AP on COCO test-dev, without sacrificing the efficiency both in training and inference.
Xiang Li 0041, Wenhai Wang, Xiaolin Hu 0001, Jun Li 0027, Jinhui Tang 0001, Jian Yang 0003
CVPR1
2021 Regularizing Nighttime Weirdness: Efficient Self-supervised Monocular Depth Estimation in the Dark
abstract
Monocular depth estimation aims at predicting depth from a single image or video. Recently, self-supervised methods draw much attention since they are free of depth annotations and achieve impressive performance on several daytime benchmarks. However, they produce weird outputs in more challenging nighttime scenarios because of low visibility and varying illuminations, which bring weak textures and break brightness-consistency assumption, respectively. To address these problems, in this paper we propose a novel framework with several improvements: (1) we introduce Priors-Based Regularization to learn distribution knowledge from unpaired depth maps and prevent model from being incorrectly trained; (2) we leverage Mapping-Consistent Image Enhancement module to enhance image visibility and contrast while maintaining brightness consistency; and (3) we present Statistics-Based Mask strategy to tune the number of removed pixels within textureless regions, using dynamic statistics. Experimental results demonstrate the effectiveness of each component. Mean-while, our framework achieves remarkable improvements and state-of-the-art results on two nighttime datasets. Code is available at https://github.com/w2kun/RNW.
Kun Wang 0042, Zhenyu Zhang 0005, Zhiqiang Yan 0001, Xiang Li 0041, Baobei Xu, Jun Li 0027, Jian Yang 0003
ICCV4
2020 Understanding the Disharmony between Weight Normalization Family and Weight Decay
abstract
The merits of fast convergence and potentially better performance of the weight normalization family have drawn increasing attention in recent years. These methods use standardization or normalization that changes the weight W to W′, which makes W′ independent to the magnitude of W. Surprisingly, W must be decayed during gradient descent, otherwise we will observe a severe under-fitting problem, which is very counter-intuitive since weight decay is widely known to prevent deep networks from over-fitting. Moreover, if we substitute (e.g., weight normalization) W′ = W∥W∥ in the original loss function ∑i L(ƒ(xi; W′),yi) + ½λ∥W′∥2, it is observed that the regularization term ½λ∥W′∥2 will be canceled as a constant ½ λ in the optimization objective. Therefore, to decay W, we need to explicitly append: ½λ∥W∥2. In this paper, we theoretically prove that ½λ∥W∥2 improves optimization only by modulating the effective learning rate and fairly has no influence on generalization when the weight normalization family is compositely employed. Furthermore, we also expose several serious problems when introducing weight decay term to weight normalization family, including the missing of global minimum, training instability and sensitivity of initialization. To address these problems, we propose an Adaptive Weight Shrink (AWS) scheme, which gradually shrinks the weights during optimization by a dynamic coefficient proportional to the magnitude of the parameter. This simple yet effective method appropriately controls the effective learning rate, which significantly improves the training stability and makes optimization more robust to initialization.
Xiang Li 0041, Shuo Chen 0003, Jian Yang 0003
AAAI1
2020 Generalized Focal Loss: Learning Qualified and Distributed Bounding Boxes for Dense Object Detection
abstract
One-stage detector basically formulates object detection as dense classification and localization (i.e., bounding box regression). The classification is usually optimized by Focal Loss and the box location is commonly learned under Dirac delta distribution. A recent trend for one-stage detectors is to introduce an \emph{individual} prediction branch to estimate the quality of localization, where the predicted quality facilitates the classification to improve detection performance. This paper delves into the \emph{representations} of the above three fundamental elements: quality estimation, classification and localization. Two problems are discovered in existing practices, including (1) the inconsistent usage of the quality estimation and classification between training and inference, and (2) the inflexible Dirac delta distribution for localization. To address the problems, we design new representations for these elements. Specifically, we merge the quality estimation into the class prediction vector to form a joint representation, and use a vector to represent arbitrary distribution of box locations. The improved representations eliminate the inconsistency risk and accurately depict the flexible distribution in real data, but contain \emph{continuous} labels, which is beyond the scope of Focal Loss. We then propose Generalized Focal Loss (GFL) that generalizes Focal Loss from its discrete form to the \emph{continuous} version for successful optimization. On COCO {\tt test-dev}, GFL achieves 45.0\% AP using ResNet-101 backbone, surpassing state-of-the-art SAPD (43.5\%) and ATSS (43.6\%) with higher or comparable inference speed.
Xiang Li 0041, Wenhai Wang, Shuo Chen 0003, Xiaolin Hu 0001, Jun Li 0027, Jinhui Tang 0001, Jian Yang 0003
NeurIPS1
2020 Joint Task-Recursive Learning for RGB-D Scene Understanding
abstract
RGB-D scene understanding under monocular camera is an emerging and challenging topic with many potential applications. In this paper, we propose a novel Task-Recursive Learning (TRL) framework to jointly and recurrently conduct three representative tasks therein containing depth estimation, surface normal prediction and semantic segmentation. TRL recursively refines the prediction results through a series of task-level interactions, where one-time cross-task interaction is abstracted as one network block of one time stage. In each stage, we serialize multiple tasks into a sequence and then recursively perform their interactions. To adaptively enhance counterpart patterns, we encapsulate interactions into a specific Task-Attentional Module (TAM) to mutually-boost the tasks from each other. Across stages, the historical experiences of previous states of tasks are selectively propagated into the next stages by using Feature-Selection unit (FS-Unit), which takes advantage of complementary information across tasks. The sequence of task-level interactions is also evolved along a coarse-to-fine scale space such that the required details may be refined progressively. Finally the task-abstracted sequence problem of multi-task prediction is framed into a recursive network. Extensive experiments on NYU-Depth v2 and SUN RGB-D datasets demonstrate that our method can recursively refines the results of the triple tasks and achieves state-of-the-art performance.
Zhenyu Zhang 0005, Zhen Cui 0001, Chunyan Xu, Zequn Jie, Xiang Li 0041, Jian Yang 0003
IEEE Trans. Pattern Anal. Mach. Intell.5
2020 Line-CNN: End-to-End Traffic Line Detection With Line Proposal Unit
abstract
The task of traffic line detection is a fundamental yet challenging problem. Previous approaches usually conduct traffic line detection via a two-stage way, namely the line segment detection followed by a segment clustering, which is very likely to ignore the global semantic information of an entire line. To address the problem, we propose an end-to-end system called Line-CNN (L-CNN), in which the key component is a novel line proposal unit (LPU). The LPU utilizes line proposals as references to locate accurate traffic curves, which forces the system to learn the global feature representation of the entire traffic lines. We benchmark the proposed L-CNN on two public datasets including MIKKI and TuSimple, and the results suggest that L-CNN outperforms the state-of-the-art methods. In addition, L-CNN can run at approximately 30 f/s on a Titan X GPU, which indicates the practicability and effectiveness of L-CNN for real-time intelligent self-driving systems.
Xiang Li 0041, Jun Li 0027, Xiaolin Hu 0001, Jian Yang 0003
IEEE Trans. Intell. Transp. Syst.1
2020 Toward Making Unsupervised Graph Hashing Discriminative
abstract
Recently, hashing has attracted much attention in visual information retrieval due to its low storage cost and fast query speed. The goal of hashing is to map original high-dimensional data into a low-dimensional binary-code space where the similar data points are assigned similar hash codes and dissimilar points are far away from each other. Existing unsupervised hashing methods mainly focus on recovering the pairwise similarity of the original data in hash space, but do not take specific measures to make the generated binary codes to be discriminative. To address this problem, this paper proposes a novel unsupervised hashing method, named “Discriminative Unsupervised Graph Hashing” (DUGH), which takes both similarity and dissimilarity of original data into consideration to learn discriminative binary codes. In particular, a probabilistic model is utilized to learn the encoding of original data in low-dimensional space, which models the original neighbor structure through both positive and negative edges in the KNN graph and then maximizes the likelihood of observing these edges. To efficiently and accurately measure the neighbor structure for largescale datasets, we propose an effective KNN graph construction algorithm based on the random projection tree and neighbor exploring techniques. The experimental results on one synthetic dataset and four typical real-world image datasets demonstrate that the proposed method significantly outperforms the state-of-the-art unsupervised hashing methods.
Chao Ma 0005, Chen Gong 0002, Xiang Li 0041, Xiaolin Huang, Wei Liu 0005, Jie Yang 0002
IEEE Trans. Multim.3
2019 Inter-Class Angular Loss for Convolutional Neural Networks
abstract
Convolutional Neural Networks (CNNs) have shown great power in various classification tasks and have achieved remarkable results in practical applications. However, the distinct learning difficulties in discriminating different pairs of classes are largely ignored by the existing networks. For instance, in CIFAR-10 dataset, distinguishing cats from dogs is usually harder than distinguishing horses from ships. By carefully studying the behavior of CNN models in the training process, we observe that the confusion level of two classes is strongly correlated with their angular separability in the feature space. That is, the larger the inter-class angle is, the lower the confusion will be. Based on this observation, we propose a novel loss function dubbed “Inter-Class Angular Loss” (ICAL), which explicitly models the class correlation and can be directly applied to many existing deep networks. By minimizing the proposed ICAL, the networks can effectively discriminate the examples in similar classes by enlarging the angle between their corresponding class vectors. Thorough experimental results on a series of vision and nonvision datasets confirm that ICAL critically improves the discriminative ability of various representative deep neural networks and generates superior performance to the original networks with conventional softmax loss.
Le Hui, Xiang Li 0041, Chen Gong 0002, Joey Tianyi Zhou, Jian Yang 0003
AAAI2
2019 Understanding the Disharmony Between Dropout and Batch Normalization by Variance Shift
abstract
This paper first answers the question ``why do the two most powerful techniques Dropout and Batch Normalization (BN) often lead to a worse performance when they are combined together in many modern neural networks, but cooperate well sometimes as in Wide ResNet (WRN)?'' in both theoretical and empirical aspects. Theoretically, we find that Dropout shifts the variance of a specific neural unit when we transfer the state of that network from training to test. However, BN maintains its statistical variance, which is accumulated from the entire learning procedure, in the test phase. The inconsistency of variances in Dropout and BN (we name this scheme ``variance shift'') causes the unstable numerical behavior in inference that leads to erroneous predictions finally. Meanwhile, the large feature dimension in WRN further reduces the ``variance shift'' to bring benefits to the overall performance. Thorough experiments on representative modern convolutional networks like DenseNet, ResNet, ResNeXt and Wide ResNet confirm our findings. According to the uncovered mechanism, we get better understandings in the combination of these two techniques and summarize guidelines for better practices.
Xiang Li 0041, Shuo Chen 0003, Xiaolin Hu 0001, Jian Yang 0003
CVPR1
2019 Selective Kernel Networks
abstract
In standard Convolutional Neural Networks (CNNs), the receptive fields of artificial neurons in each layer are designed to share the same size. It is well-known in the neuroscience community that the receptive field size of visual cortical neurons are modulated by the stimulus, which has been rarely considered in constructing CNNs. We propose a dynamic selection mechanism in CNNs that allows each neuron to adaptively adjust its receptive field size based on multiple scales of input information. A building block called Selective Kernel (SK) unit is designed, in which multiple branches with different kernel sizes are fused using softmax attention that is guided by the information in these branches. Different attentions on these branches yield different sizes of the effective receptive fields of neurons in the fusion layer. Multiple SK units are stacked to a deep network termed Selective Kernel Networks (SKNets). On the ImageNet and CIFAR benchmarks, we empirically show that SKNet outperforms the existing state-of-the-art architectures with lower model complexity. Detailed analyses show that the neurons in SKNet can capture target objects with different scales, which verifies the capability of neurons for adaptively adjusting their receptive field sizes according to the input. The code and models are available at https://github.com/implus/SKNet.
Xiang Li 0041, Wenhai Wang, Xiaolin Hu 0001, Jian Yang 0003
CVPR1
2019 Shape Robust Text Detection With Progressive Scale Expansion Network
abstract
Scene text detection has witnessed rapid progress especially with the recent development of convolutional neural networks. However, there still exists two challenges which prevent the algorithm into industry applications. On the one hand, most of the state-of-art algorithms require quadrangle bounding box which is in-accurate to locate the texts with arbitrary shape. On the other hand, two text instances which are close to each other may lead to a false detection which covers both instances. Traditionally, the segmentation-based approach can relieve the first problem but usually fail to solve the second challenge. To address these two challenges, in this paper, we propose a novel Progressive Scale Expansion Network (PSENet), which can precisely detect text instances with arbitrary shapes. More specifically, PSENet generates the different scale of kernels for each text instance, and gradually expands the minimal scale kernel to the text instance with the complete shape. Due to the fact that there are large geometrical margins among the minimal scale kernels, our method is effective to split the close text instances, making it easier to use segmentation-based methods to detect arbitrary-shaped text instances. Extensive experiments on CTW1500, Total-Text, ICDAR 2015 and ICDAR 2017 MLT validate the effectiveness of PSENet. Notably, on CTW1500, a dataset full of long curve texts, PSENet achieves a F-measure of 74.3% at 27 FPS, and our best F-measure (82.2%) outperforms state-of-art algorithms by 6.6%. The code will be released in the future.
Wenhai Wang, Enze Xie, Xiang Li 0041, Wenbo Hou, Tong Lu 0002, Gang Yu 0002, Shuai Shao 0005
CVPR3
2018 Stacked Conditional Generative Adversarial Networks for Jointly Learning Shadow Detection and Shadow Removal
abstract
Understanding shadows from a single image consists of two types of task in previous studies, containing shadow detection and shadow removal. In this paper, we present a multi-task perspective, which is not embraced by any existing work, to jointly learn both detection and removal in an end-to-end fashion that aims at enjoying the mutually improved benefits from each other. Our framework is based on a novel STacked Conditional Generative Adversarial Network (ST-CGAN), which is composed of two stacked CGANs, each with a generator and a discriminator. Specifically, a shadow image is fed into the first generator which produces a shadow detection mask. That shadow image, concatenated with its predicted mask, goes through the second generator in order to recover its shadow-free image consequently. In addition, the two corresponding discriminators are very likely to model higher level relationships and global scene characteristics for the detected shadow region and reconstruction via removing shadows, respectively. More importantly, for multi-task learning, our design of stacked paradigm provides a novel view which is notably different from the commonly used one as the multi-branch version. To fully evaluate the performance of our proposed framework, we construct the first large-scale benchmark with 1870 image triplets (shadow image, shadow mask image, and shadow-free image) under 135 scenes. Extensive experimental results consistently show the advantages of STC-GAN over several representative state-of-the-art methods on two large-scale publicly available datasets and our newly released one.
Xiang Li 0041, Jian Yang 0003
CVPR2
2018 Joint Task-Recursive Learning for Semantic Segmentation and Depth Estimation
Zhenyu Zhang 0005, Zhen Cui 0001, Chunyan Xu, Zequn Jie, Xiang Li 0041, Jian Yang 0003
ECCV (10)5
2018 Discrete Locally-Linear Preserving Hashing
abstract
Recently, hashing has attracted considerable attention for nearest neighbor search due to its fast query speed and low storage cost. However, existing unsupervised hashing algorithms have two problems in common. Firstly, the widely utilized anchor graph construction algorithm has inherent limitations in local weight estimation. Secondly, the locally linear structure in the original feature space is seldom taken into account for binary encoding. Therefore, in this paper, we propose a novel unsupervised hashing method, dubbed “dis-crete locally-linear preserving hashing”, which effectively calculates the adjacent matrix while preserving the locally linear structure in the obtained hash space. Specifically, a novel local anchor embedding algorithm is adopted to construct the approximate adjacent matrix. After that, we directly minimize the reconstruction error with the discrete constrain to learn the binary codes. Experimental results on two typical image datasets indicate that the proposed method significantly outperforms the state-of-the-art unsupervised methods.
Xiang Li 0041, Chao Ma 0005, Jie Yang 0002, Xiaolin Huang
ICIP1
2018 Teach to Hash: A Deep Supervised Hashing Framework with Data Selection
Xiang Li 0041, Chao Ma 0005, Jie Yang 0002, Yu Qiao 0003
ICONIP (1)1
2018 SESR: Single Image Super Resolution with Recursive Squeeze and Excitation Networks
abstract
Single image super resolution is a very important computer vision task, with a wide range of applications. In recent years, the depth of the super-resolution model has been constantly increasing, but with a small increase in performance, it has brought a huge amount of computation and memory consumption. In this work, in order to make the super resolution models more effective, we proposed a novel single image super resolution method via recursive squeeze and excitation networks (SESR). By introducing the squeeze and excitation module, our SESR can model the interdependencies and relationships between channels and that makes our model more efficiency. In addition, the recursive structure and progressive reconstruction method in our model minimized the layers and parameters and enabled SESR to simultaneously train multi-scale super resolution in a single model. After evaluating on four benchmark test sets, our model is proved to be above the state-of-the-art methods in terms of speed and accuracy.
Xiang Li 0041, Jian Yang 0003, Ying Tai
ICPR2
2018 Unsupervised Multi-Domain Image Translation with Domain-Specific Encoders/Decoders
abstract
Unsupervised Image-to-Image Translation achieves spectacularly advanced developments nowadays. However, recent approaches mainly focus on one model with two domains, which may face heavy burdens with the large cost of training time and the huge model parameters, under such a requirement that domains are freely transferred to each other in a general setting. To address this problem, we propose a novel and unified framework named Domain-Bank, which consists of a globally shared auto-encoder and n domain-specific encoders/decoders, assuming that there is a universal shared-latent space can be projected. Thus, we not only reduce the parameters of the model but also have a huge reduction of the time budgets. Besides the high efficiency, we show the comparable (or even better) image translation results over state-of-the-arts on various challenging unsupervised image translation tasks, including face image translation and painting style translation. We also apply the proposed framework to the domain adaptation task and achieve state-of-the-art performance on digit benchmark datasets.
Le Hui, Xiang Li 0041, Jiaxin Chen 0001, Jian Yang 0003
ICPR2
2018 Adversarial Metric Learning
abstract
In the past decades, intensive efforts have been put to design various loss functions and metric forms for metric learning problem. These improvements have shown promising results when the test data is similar to the training data. However, the trained models often fail to produce reliable distances on the ambiguous test pairs due to the different samplings between training set and test set. To address this problem, the Adversarial Metric Learning (AML) is proposed in this paper, which automatically generates adversarial pairs to remedy the sampling bias and facilitate robust metric learning. Specifically, AML consists of two adversarial stages, i.e. confusion and distinguishment. In confusion stage, the ambiguous but critical adversarial data pairs are adaptively generated to mislead the learned metric. In distinguishment stage, a metric is exhaustively learned to try its best to distinguish both adversarial pairs and original training pairs. Thanks to the challenges posed by the confusion stage in such competing process, the AML model is able to grasp plentiful difficult knowledge that has not been contained by the original training pairs, so the discriminability of AML can be significantly improved. The entire model is formulated into optimization framework, of which the global convergence is theoretically proved. The experimental results on toy data and practical datasets clearly demonstrate the superiority of AML to representative state-of-the-art metric learning models.
Shuo Chen 0003, Chen Gong 0002, Jian Yang 0003, Xiang Li 0041, Yang Wei 0003, Jun Li 0027
IJCAI4
2018 Mixed Link Networks
abstract
On the basis of the analysis by revealing the equivalence of modern networks, we find that both ResNet and DenseNet are essentially derived from the same "dense topology", yet they only differ in the form of connection: addition (dubbed "inner link") vs. concatenation (dubbed "outer link"). However, both forms of connections have the superiority and insufficiency. To combine their advantages and avoid certain limitations on representation learning, we present a highly efficient and modularized Mixed Link Network (MixNet) which is equipped with flexible inner link and outer link modules. Consequently, ResNet, DenseNet and Dual Path Network (DPN) can be regarded as a special case of MixNet, respectively. Furthermore, we demonstrate that MixNets can achieve superior efficiency in parameter over the state-of-the-art architectures on many competitive datasets like CIFAR-10/100, SVHN and ImageNet.
Wenhai Wang, Xiang Li 0041, Tong Lu 0002, Jian Yang 0003
IJCAI2
2018 Densely Connected Bidirectional LSTM with Applications to Sentence Classification
Zixiang Ding, Jianfei Yu, Xiang Li 0041, Jian Yang 0003
NLPCC (2)4
2017 A Point and Line Features Based Method for Disturbed Surface Motion Estimation
Xiang Li 0041, Yue Zhou 0005
ICONIP (3)1
2016 LightRNN: Memory and Computation-Efficient Recurrent Neural Networks
abstract
Recurrent neural networks (RNNs) have achieved state-of-the-art performances in many natural language processing tasks, such as language modeling and machine translation. However, when the vocabulary is large, the RNN model will become very big (e.g., possibly beyond the memory capacity of a GPU device) and its training will become very inefficient. In this work, we propose a novel technique to tackle this challenge. The key idea is to use 2-Component (2C) shared embedding for word representations. We allocate every word in the vocabulary into a table, each row of which is associated with a vector, and each column associated with another vector. Depending on its position in the table, a word is jointly represented by two components: a row vector and a column vector. Since the words in the same row share the row vector and the words in the same column share the column vector, we only need $2 \sqrt{|V|}$ vectors to represent a vocabulary of $|V|$ unique words, which are far less than the $|V|$ vectors required by existing approaches. Based on the 2-Component shared embedding, we design a new RNN algorithm and evaluate it using the language modeling task on several benchmark datasets. The results show that our algorithm significantly reduces the model size and speeds up the training process, without sacrifice of accuracy (it achieves similar, if not better, perplexity as compared to state-of-the-art language models). Remarkably, on the One-Billion-Word benchmark Dataset, our algorithm achieves comparable perplexity to previous language models, whilst reducing the model size by a factor of 40-100, and speeding up the training process by a factor of 2. We name our proposed algorithm \emph{LightRNN} to reflect its very small model size and very high training speed.
Xiang Li 0041, Tao Qin 0001, Jian Yang 0003, Tie-Yan Liu
NIPS1