EDBT 2026 Demo / reviewers in the wild / expert
Yuxuan Li 0004
dblp:17/7624-4
· DBLP profile ↗
17ranked-venue papers
7as first author
16since 2021 · last 2026
0000-0002-0613-3969ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 5 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SM3Det: A Unified Model for Multi-Modal Remote Sensing Object DetectionabstractWith the rapid advancement of remote sensing technology, high-resolution multi-modal imagery is now more widely accessible. Conventional object detection models are trained on a single dataset, often restricted to a specific imaging modality and annotation format. However, such an approach overlooks the valuable shared knowledge across multi-modalities and limits the model’s applicability in more versatile scenarios. This paper introduces a new task called Multi-Modal Datasets and Multi-Task Object Detection (M2Det) for remote sensing, designed to accurately detect horizontal or oriented objects from any sensor modality. This task poses challenges due to 1) the trade-offs involved in managing multi-modal modelling and 2) the complexities of multi-task optimization. To address these, we establish a benchmark dataset and propose a unified model, SM3Det (Single Model for Multi-Modal datasets and Multi-Task object Detection). SM3Det leverages a grid-level sparse MoE backbone to enable joint knowledge learning while preserving distinct feature representations for different modalities. Furthermore, we propose a novel consistency and synchronization optimization mechanism, allowing it to effectively handle varying levels of learning difficulty across modalities and tasks. Extensive experiments demonstrate SM3Det's effectiveness and generalizability, consistently outperforming the combination of specialized models on individual datasets. Yuxuan Li 0004, Xiang Li 0041, Yimian Dai, Qibin Hou, Ming-Ming Cheng, Jian Yang 0003 |
AAAI | 1 |
| 2026 | DenoDet V2: Phase-Amplitude Cross Denoising for SAR Object DetectionabstractOne of the primary challenges in Synthetic Aperture Radar (SAR) object detection lies in the pervasive influence of coherent noise. As a common practice, most existing methods, whether handcrafted approaches or deep learning-based methods, employ the analysis or enhancement of object spatial-domain characteristics to achieve implicit denoising. In this paper, we propose DenoDet V2, which explores a completely novel and different perspective to deconstruct and modulate the features in the transform domain via a carefully designed attention architecture. Compared to DenoDet V1, DenoDet V2 is a major advancement that exploits the complementary nature of amplitude and phase information through a band-wise mutual modulation mechanism, which enables a reciprocal enhancement between phase and amplitude spectra. Extensive experiments on various SAR datasets demonstrate the state-of-the-art performance of DenoDet V2. Notably, DenoDet V2 achieves a significant 0.8% improvement on SARDet-100K dataset compared to DenoDet V1, while reducing the model complexity by half. Kang Ni, Minrui Zou, Yuxuan Li 0004, Xiang Li 0041, Kehua Guo, Ming-Ming Cheng, Yimian Dai |
AAAI | 3 |
| 2026 | Strip R-CNN: Large Strip Convolution for Remote Sensing Object DetectionabstractIn this paper, we show that current approaches using large square kernels or transformer-based global modeling aggregate contextual information uniformly across spatial dimensions, leading to feature dilution and localization errors for elongated targets. To mitigate this issue, we propose Strip R-CNN, the first work to systematically explore large strip convolutions for remote sensing object detection. Our key insight is that strip convolutions enable directional feature aggregation along the dominant spatial dimension of slender objects, reducing background interference while preserving essential geometric information. We design two core components: (i) StripNet, a backbone network employing sequential orthogonal large strip convolutions to capture anisotropic spatial patterns, and (ii) Strip Head, which enhances localization precision by incorporating strip convolutions into the detection head. Unlike previous large-kernel approaches that suffer from computational redundancy and isotropic limitations, our method achieves superior performance with remarkable efficiency. Extensive experiments on multiple benchmarks (DOTA, FAIR1M, HRSC2016, and DIOR) demonstrate significant improvements, with our 30M parameter model achieving 82.75% mAP on DOTA-v1.0, establishing a new state-of-the-art record while providing new insights into anisotropic feature learning for remote sensing applications. Xinbin Yuan, Zhaohui Zheng 0003, Yuxuan Li 0004, Xialei Liu, Li Liu 0004, Xiang Li 0041, Qibin Hou, Ming-Ming Cheng |
AAAI | 3 |
| 2025 | RSAR: Restricted State Angle Resolver and Rotated SAR BenchmarkabstractRotated object detection has made significant progress in the optical remote sensing. However, advancements in the Synthetic Aperture Radar (SAR) field are laggard behind, primarily due to the absence of a large-scale dataset. Annotating such a dataset is inefficient and costly. A promising solution is to employ a weakly supervised model (e.g., trained with available horizontal boxes only) to generate pseudo-rotated boxes for reference before manual calibration. Unfortunately, the existing weakly supervised models exhibit limited accuracy in predicting the object’s angle. Previous works attempt to enhance angle prediction by using angle resolvers that decouple angles into cosine and sine encodings. In this work, we first reevaluate these resolvers from a unified perspective of dimension mapping and expose that they share the same shortcomings: these methods overlook the unit cycle constraint inherent in these encodings, easily leading to prediction biases. To address this issue, we propose the Unit Cycle Resolver (UCR), which incorporates a unit circle constraint loss to improve angle prediction accuracy. Our approach can effectively improve the performance of existing state-of-the-art weakly supervised methods and even surpasses fully supervised models on existing optical benchmarks (i.e., DOTA-v1.0). With the aid of UCR, we further annotate and introduce RSAR, the largest multi-class rotated SAR object detection dataset to date. Extensive experiments on both RSAR and optical datasets demonstrate that our UCR enhances angle prediction accuracy. Our dataset and code can be found at: https://github.com/zhasion/RSAR. Xin Zhang 0170, Xue Yang 0005, Yuxuan Li 0004, Jian Yang 0003, Ming-Ming Cheng, Xiang Li 0041 |
CVPR | 3 |
| 2025 | DISTA-Net: Dynamic Closely-Spaced Infrared Small Target UnmixingabstractResolving closely-spaced small targets in dense clusters presents a significant challenge in infrared imaging, as the overlapping signals hinder precise determination of their quantity, sub-pixel positions, and radiation intensities. While deep learning has advanced the field of infrared small target detection, its application to closely-spaced infrared small targets has not yet been explored. This gap exists primarily due to the complexity of separating superimposed characteristics and the lack of an open-source infrastructure. In this work, we propose the Dynamic Iterative Shrinkage Thresholding Network (DISTA-Net), which reconceptualizes traditional sparse reconstruction within a dynamic framework. DISTA-Net adaptively generates convolution weights and thresholding parameters to tailor the reconstruction process in real time. To the best of our knowledge, DISTA-Net is the first deep learning model designed specifically for the unmixing of closely-spaced infrared small targets, achieving superior sub-pixel detection accuracy. Moreover, we have established the first open-source ecosystem to foster further research in this field. This ecosystem comprises three key components: (1) CSIST-100K, a publicly available benchmark dataset; (2) CSO-mAP, a custom evaluation metric for sub-pixel detection; and (3) GrokCSO, an open-source toolkit featuring DISTA-Net and other models. Our code and dataset are available at https://github.com/GrokCV/GrokCSO. Shengdong Han, Shangdong Yang, Yuxuan Li 0004, Xin Zhang 0170, Xiang Li 0041, Jian Yang 0003, Ming-Ming Cheng, Yimian Dai |
ICCV | 3 |
| 2025 | Unbiased Region-Language Alignment for Open-Vocabulary Dense Prediction
Yuxuan Li 0004, Quansheng Zeng 0001, Wenhai Wang, Qibin Hou, Ming-Ming Cheng |
ICCV | 2 |
| 2025 | LSKNet: A Foundation Lightweight Backbone for Remote Sensing
Yuxuan Li 0004, Xiang Li 0041, Yimian Dai, Qibin Hou, Li Liu 0004, Yongxiang Liu, Ming-Ming Cheng, Jian Yang 0003 |
Int. J. Comput. Vis. | 1 |
| 2025 | MoCoLSK: Modality-Conditioned High-Resolution Downscaling for Land Surface TemperatureabstractLand surface temperature (LST) is a critical parameter for environmental studies, but directly obtaining high spatial resolution LST data remains challenging due to the spatiotemporal tradeoff in satellite remote sensing. Guided LST downscaling has emerged as an alternative solution to overcome these limitations, but current methods often neglect spatial nonstationarity, and there is a lack of an open-source ecosystem for deep learning methods. In this article, we propose the modality-conditioned large selective kernel (MoCoLSK) network, a novel architecture that dynamically fuses multimodal data through modality-conditioned projections. MoCoLSK achieves a confluence of dynamic receptive field adjustment and multimodal feature fusion, leading to enhanced LST prediction accuracy. Furthermore, we establish the GrokLST Project, a comprehensive open-source ecosystem featuring the GrokLST dataset, a high-resolution (HR) benchmark, and the GrokLST toolkit, an open-source PyTorch-based toolkit encapsulating MoCoLSK alongside 40+ state-of-the-art approaches. Extensive experimental results validate MoCoLSK’s effectiveness in capturing complex dependencies and subtle variations within multispectral data, outperforming existing methods in LST downscaling. Our code, dataset, and toolkit are available athttps://github.com/GrokCV/GrokLST. Qun Dai, Chunyang Yuan, Yimian Dai, Yuxuan Li 0004, Xiang Li 0041, Kang Ni, Jianhui Xu, Xiangbo Shu, Jian Yang 0003 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | SARDet-100K: Towards Open-Source Benchmark and ToolKit for Large-Scale SAR Object DetectionabstractSynthetic Aperture Radar (SAR) object detection has gained significant attention recently due to its irreplaceable all-weather imaging capabilities. However, this research field suffers from both limited public datasets (mostly comprising <2K images with only mono-category objects) and inaccessible source code. To tackle these challenges, we establish a new benchmark dataset and an open-source method for large-scale SAR object detection. Our dataset, SARDet-100K, is a result of intense surveying, collecting, and standardizing 10 existing SAR detection datasets, providing a large-scale and diverse dataset for research purposes. To the best of our knowledge, SARDet-100K is the first COCO-level large-scale multi-class SAR object detection dataset ever created. With this high-quality dataset, we conducted comprehensive experiments and uncovered a crucial challenge in SAR object detection: the substantial disparities between the pretraining on RGB datasets and finetuning on SAR datasets in terms of both data domain and model structure. To bridge these gaps, we propose a novel Multi-Stage with Filter Augmentation (MSFA) pretraining framework that tackles the problems from the perspective of data input, domain transition, and model migration. The proposed MSFA method significantly enhances the performance of SAR object detection models while demonstrating exceptional generalizability and flexibility across diverse models. This work aims to pave the way for further advancements in SAR object detection. The dataset and code is available at \url{https://github.com/zcablii/SARDet_100K}. Yuxuan Li 0004, Xiang Li 0041, Qibin Hou, Li Liu 0004, Ming-Ming Cheng, Jian Yang 0003 |
NeurIPS | 1 |
| 2024 | Pick of the Bunch: Detecting Infrared Small Targets Beyond Hit-Miss Trade-Offs via Selective Rank-Aware AttentionabstractInfrared small target detection faces the inherent challenge of precisely localizing dim targets amidst complex background clutter. Traditional approaches struggle to balance detection precision and false alarm rates. To break this dilemma, we propose SeRankDet, a deep network that achieves high accuracy beyond the conventional hit-miss trade-off, by following the “Pick of the Bunch” principle. At its core lies our selective rank-aware attention (SeRank) module, employing a nonlinear Top-K selection process that preserves the most salient responses, preventing target signal dilution while maintaining constant complexity. Furthermore, we replace the static concatenation typical in U-Net structures with our large selective feature fusion (LSFF) module, a dynamic fusion strategy that empowers SeRankDet with adaptive feature integration, enhancing its ability to discriminate true targets from false alarms. The network’s discernment is further refined by our dilated difference convolution (DDC) module, which merges differential convolution aimed at amplifying subtle target characteristics with dilated convolution to expand the receptive field, thereby substantially improving target-background separation. Despite its lightweight architecture, the proposed SeRankDet sets new benchmarks in state-of-the-art performance across multiple public datasets. The code is available athttps://github.com/GrokCV/SeRankDet. Yimian Dai, Peiwen Pan, Yulei Qian, Yuxuan Li 0004, Xiang Li 0041, Jian Yang 0003, Huan Wang 0013 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Large Selective Kernel Network for Remote Sensing Object DetectionabstractRecent research on remote sensing object detection has largely focused on improving the representation of oriented bounding boxes but has overlooked the unique prior knowledge presented in remote sensing scenarios. Such prior knowledge can be useful because tiny remote sensing objects may be mistakenly detected without referencing a sufficiently long-range context, which can vary for different objects. This paper considers these priors and proposes the lightweight Large Selective Kernel Network (LSKNet). LSKNet can dynamically adjust its large spatial receptive field to better model the ranging context of various objects in remote sensing scenarios. To our knowledge, large and selective kernel mechanisms have not been previously explored in remote sensing object detection. Without bells and whistles, our lightweight LSKNet sets new state-of-the-art scores on standard benchmarks, i.e., HRSC2016 (98.46% mAP), DOTA-v1.0 (81.85% mAP), and FAIR1M-v1.0 (47.87% mAP). Yuxuan Li 0004, Qibin Hou, Zhaohui Zheng 0003, Ming-Ming Cheng, Jian Yang 0003, Xiang Li 0041 |
ICCV | 1 |
| 2023 | APF-GAN: Exploring asymmetric pre-training and fine-tuning strategy for conditional generative adversarial networkabstractThe use of generative adversarial network (GAN)based models for the conditional generation of image semantic segmentation has shown promising results in recent years.However, there are still some limitations, including limited diversity of image style, distortion of detailed texture, unbalanced color tone, and lengthy training time.To address these issues, we propose an asymmetric pre-training and fine-tuning (APF)-GAN model.In the pretraining phase, we introduce a progressive growing mechanism for pix2pix conditional GAN frameworks to efficiently generate high-quality images with details.Subsequently, in the fine-tuning phase, we introduce novel semantic spatially-guided noise to improve the robustness of the model and increase style diversity.The proposed algorithm outperformed the high-performance GauGAN model and won the championship of the Second Jittor Artificial Intelligence Challenge.Our model was implemented in the Jittor framework and is available at https:// github.com/zcablii/jittor-Torile-PG_SPADE. APF-GAN Asymmetric pre-training and fine-tuning strategyIn this study, we propose an asymmetric pretraining and fine-tuning strategy for the conditional generative adversarial network (APF-GAN) model Yuxuan Li 0004, Lingfeng Yang, Xiang Li 0041 |
Comput. Vis. Media | 1 |
| 2023 | A Spatial Hierarchical Reasoning Network for Remote Sensing Visual Question AnsweringabstractFor visual question answering on remote sensing (RSVQA), current methods scarcely consider geospatial objects typically with large-scale differences and positional sensitive properties. Besides, modeling and reasoning the relationships between entities have rarely been explored, which leads to one-sided and inaccurate answer predictions. In this article, a novel method called spatial hierarchical reasoning network (SHRNet) is proposed, which endows a remote sensing (RS) visual question answering (VQA) system with enhanced visual–spatial reasoning capability. Specifically, a hash-based spatial multiscale visual representation module is first designed to encode multiscale visual features embedded with spatial positional information. Then, spatial hierarchical reasoning is conducted to learn the high-order inner group object relations across multiple scales under the guidance of linguistic cues. Finally, a visual-question (VQ) interaction module is employed to learn an effective image–text joint embedding for the final answer predicting. Experimental results on three public RS VQA datasets confirm the effectiveness and superiority of our model SHRNet. Zixiao Zhang, Licheng Jiao, Lingling Li 0002, Xu Liu 0006, Puhua Chen, Fang Liu 0001, Yuxuan Li 0004, Zhicheng Guo |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2022 | Spatial Group-Wise Enhance: Enhancing Semantic Feature Learning in CNN
Yuxuan Li 0004, Xiang Li 0041, Jian Yang 0003 |
ACCV (5) | 1 |
| 2022 | Laplacian Feature Pyramid Network for Object Detection in VHR Optical Remote Sensing ImagesabstractExcept for multiscale features, high-frequency features are also crucial for the identification of many objects in object detection for very high resolution optical remote sensing (VHR-ORS) images but have not been considered yet. Due to the fact that the Laplacian pyramid consists of high-frequency information at each level, we propose a Laplacian feature pyramid (FP) network (LFPN) considering both low-frequency features and high-frequency features based on FP structure to improve the object detection performance of VHR-ORS images. FP-based structures are efficient to represent multiscale features. But, in general, FP-based structures, high-frequency features are not specially considered. Such high-frequency features are important to distinguish many ground objects with sufficient details. For example, texture features are critical to distinguish basketball_court and tennis_court. The construction of LFPN consists of a bottom-up pathway, Laplacian pathway, and a fusion pathway, which generate low-frequency pyramid, high-frequency pyramid, and compound pyramid, respectively. The bottom-up pathway follows the computation flow of the backbone convolutional neural networks (CNNs) which is similar to general FP-based structures. The Laplacian pathway extracts the high-frequency features of objects through a trainable Laplacian operator. Finally, the low-frequency and high-frequency FPs are fused to generate the compound pyramid in efficient ways. To evaluate the performance of LFPN, we embed LFPN into both two-stage object detection (T-LFPN) systems and single-stage object detection (S-LFPN) systems to conduct experiments. Experiments on a public challenging ten-class data set NWPU VHR-10 demonstrate the superior performance of LFPN in both T-LFPN and S-LFPN systems and state-of-the-art performance of LFPN-based detectors. Licheng Jiao, Yuxuan Li 0004, Zhongjian Huang, Haoran Wang 0008 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Sparse Learning-Based Correlation Filter for Robust Trackingabstract-norm based sparse response regularization term to restrain unexpected crests in response for CF framework. CF trackers learn online to regress the region of interest into a Gaussian response. However, due to the uncertain transformations of tracked object, there are many unexpected crests in the response map. When the response of tracked object is corrupted by other crests, the tracker will lost the object. Therefore, the sparse response is used to increase the robustness to transformations of tracked object. Since the novel term is directly incorporated into the objective function of the CF framework, it can be used to improve the performance of many methods which are based on this framework. Moreover, from the solutions we derive, the new method will not increase the computational complexity. Through the experiments on benchmarks of OTB-100, TempleColor, VOT2016 and VOT2017, the proposed regularization term can improve the tracking performance of various CF trackers, including those based on standard discriminative CF framework and those based on context-aware CF framework. We also embed the sparse response regularization term in the state-of-the-art integrated tracker MCCT to test its generalization performance. Although MCCT is an expert integrated tracker and owns an exquisite algorithm for selecting experts, the experimental results show that our method can still improve its long-term tracking performance without increasing computational complexity. Licheng Jiao, Yuxuan Li 0004, Jia Liu 0020 |
IEEE Trans. Image Process. | 3 |
| 2019 | Weak Moving Object Detection In Optical Remote Sensing Video With Motion-Drive Fusion NetworkabstractObject detection in optical remote sensing video (ORSV) is a new trend which makes it possible for obtaining richer information in more complicated and diverse situations. However, the small objects are blurred in the videos captured from optical sensor assembled in satellite, limited by the devices and natural weather. The concept of weak object is defined in this situation that the objects are extremely small and hardly detected with only one static image. Therefore, we propose a simple but efficient method for weak moving object detection in ORSV by combining the temporal information from neighbor frames and spatial features from image pixels. First, we compute the difference map between two adjacent frames, and stack it with original RGB channel so that a (1+3)-channel input data is made. Then, a motion-drive D-RGB (difference map with RGB image) fusion network is developed to obtain the feature map of this (1+3)-channel data. To adapt unusual scale in ORSV images, based on statistical prior objects size, we change the size of anchor box in original Faster R-CNN. The proposed method is demonstrated to improve the mean average precision on detecting weak moving remote sensing objects. Yuxuan Li 0004, Licheng Jiao, Xu Tang 0004, Xiangrong Zhang |
IGARSS | 1 |