EDBT 2026 Demo / reviewers in the wild / expert
Chang Xu 0027
dblp:97/2966-27
· DBLP profile ↗
15ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0002-3078-0496ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 8 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Oriented Tiny Object Detection: A Dataset, Benchmark, and Dynamic Unbiased LearningabstractDetecting oriented tiny objects, which are limited in appearance information yet prevalent in real-world applications, remains an intricate and under-explored problem. To address this, we systematically introduce a new dataset, a benchmark, and a dynamic coarse-to-fine learning scheme in this study. Our proposed dataset, AI-TOD-R, features the smallest object sizes among all oriented object detection datasets. Based on AI-TOD-R, we present a benchmark spanning a broad range of detection paradigms, including both fully-supervised and label-efficient approaches. Through investigation, we identify a learning bias presents across various learning pipelines: confident objects become increasingly confident, while vulnerable oriented tiny objects are further marginalized, hindering their detection performance. To mitigate this issue, we propose a Dynamic Coarse-to-Fine Learning (DCFL) scheme towards unbiased learning. DCFL dynamically updates prior positions to better align with the limited areas of oriented tiny objects, and it assigns samples in a way that balances both quantity and quality across different object shapes, thus mitigating biases in prior settings and sample selection. Extensive experiments across 10 challenging object detection datasets demonstrate that DCFL achieves state-of-the-art accuracy, high efficiency, and remarkable versatility. Chang Xu 0027, Ruixiang Zhang, Wen Yang 0001, Jian Ding 0001, Gui-Song Xia |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | DN-TOD: Robust tiny object detection amidst label noise
Chang Xu 0027, Wen Yang 0001, Ruixiang Zhang, Yan Zhang 0115, Gui-Song Xia |
Pattern Recognit. | 2 |
| 2025 | SMARTIES: Spectrum-Aware Multi-Sensor Auto-Encoder for Remote Sensing Images
Gencer Sumbul, Chang Xu 0027, Emanuele Dalsasso, Devis Tuia |
ICCV | 2 |
| 2025 | ReFocal: Addressing Learning Imbalances for Accurate Tiny Object Detection in Aerial ImageryabstractTiny objects in aerial imagery usually exhibit an extremely limited number of pixels, significantly affecting the object detection model’s learning process. While existing research has attempted to improve tiny objects’ positive sample quantity for scale-balanced learning, the primary focus lies on the object level. We argue that mitigating learning imbalance requires a comprehensive consideration encompassing object-level, sample-level, and feature-level improvements. To this end, we propose ReFocal, a learning strategy comprised of ReFocal Loss and ReFocal feature pyramid network (FPN), to mitigate imbalances across these three levels. ReFocal Loss utilizes a magnitude factor to regulate the learning magnitude of objects with varying sample counts and a novel focal rate adjuster to differentiate sample quality at the sample level, enabling the detector to prioritize high-quality samples within each object. ReFocal FPN employs a refocusing mechanism to dynamically enhance detailed information in high-level feature maps without introducing additional computational cost, thus addressing the feature-level imbalance. Extensive experiments on AI-TOD-v2 and TinyPerson datasets demonstrate the superiority of our proposed method over previous single-stage methods, particularly for very tiny objects. Zijuan Chen, Chang Xu 0027, Wen Yang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2025 | Detecting Every Object From EventsabstractObject detection is critical in autonomous driving, and it is more practical yet challenging to localize objects of unknown categories: an endeavour known as Class-Agnostic Object Detection (CAOD). Existing studies on CAOD predominantly rely on RGB cameras, but these frame-based sensors usually have high latency and limited dynamic range, leading to safety risks under extreme conditions like fast-moving objects, overexposure, and darkness. In this study, we turn to the event-based vision, featured by its sub-millisecond latency and high dynamic range, for robust CAOD. We propose Detecting Every Object in Events (DEOE), an approach aimed at achieving high-speed, class-agnostic object detection in event-based vision. Built upon the fast event-based backbone: recurrent vision transformer, we jointly consider the spatial and temporal consistencies to identify potential objects. The discovered potential objects are assimilated as soft positive samples to avoid being suppressed as backgrounds. Moreover, we introduce a disentangled objectness head to separate the foreground-background classification and novel object discovery tasks, enhancing the model's generalization in localizing novel objects while maintaining a strong ability to filter out the background. Extensive experiments confirm the superiority of our proposed DEOE in both open-set and closed-set settings, outperforming strong baseline methods. Haitian Zhang, Chang Xu 0027, Xinya Wang, Bingde Liu, Guang Hua 0001, Lei Yu 0006, Wen Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Unsupervised Multiview UAV Image Geolocalization via Iterative RenderingabstractUnmanned Aerial Vehicle (UAV) Cross-View Geo-Localization (CVGL) poses significant challenges due to the substantial view discrepancies between oblique UAV images and overhead satellite images. Existing methods heavily rely on supervised learning with labeled datasets to extract viewpoint-invariant features for cross-view retrieval. However, these approaches are computationally expensive, prone to overfitting region-specific cues, and exhibit limited generalizability to new regions. To overcome this issue, we propose an unsupervised solution that lifts the scene representation to 3D space from UAV observations for satellite image generation, providing a robust representation against view distortion. By generating orthogonal images that closely resemble satellite views, our method reduces view discrepancies in feature representation and mitigates shortcuts in region-specific image pairing. To further align the perspective of the rendered image with the real one, we design an iterative camera pose updating mechanism that progressively modulates the rendered query image with potential satellite targets, eliminating spatial offsets relative to the reference images. Additionally, this iterative refinement strategy enhances cross-view feature invariance through view-consistent fusion across iterations. As such, our unsupervised paradigm naturally avoids the problem of region-specific overfitting, enabling generic CVGL for UAV images without feature fine-tuning or data-driven training. Experiments on the University-1652 and SUES-200 datasets demonstrate that our approach significantly improves geo-localization accuracy while maintaining robustness across diverse regions. Notably, without model fine-tuning or paired training, our method achieves competitive performance with recent supervised methods. Haoyuan Li 0005, Chang Xu 0027, Wen Yang 0001, Li Mi, Huai Yu, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | Minimizing Sample Redundancy for Label-Efficient Object Detection in Aerial ImagesabstractObjects in aerial images tend to be densely scattered and appear in arbitrary orientations, making the annotation process quite costly. To reduce the annotation cost, existing methods propose randomly annotating a proportion of images or objects for aerial object detection with fewer label usage. These approaches, however, can lead to redundancy in labels and inherit the biases associated with the imbalance in datasets. To minimize sample redundancy and alleviate data imbalance, we propose a novel labeling pattern that acquires heterogeneous object labels in a class-orthogonal manner, preserving a broader diversity of samples for each category with less annotation effort. To improve data utility, we design a Dynamic Multi-View Learning (DML) strategy to overcome the sample quantity-quality dilemma in current pseudo-labeling methods—a high pseudo-label threshold reduces sample quantity, while low thresholds compromise sample quality. First, DML separates model predictions into multiple hierarchies for finer screening, mitigating the suppression of unlabeled objects in binary pseudo-label strategies. With this separation, DML learns to construct a new view by injecting high-quality samples and masking low-quality regions in this view, simultaneously expanding sample quantity while ensuring sample quality. Unlike previous methods that mine pseudo labels solely from unlabelled regions, DML releases this constraint by learning to expand high-quality samples with a dynamic view. Extensive experiments on five benchmark datasets validate our method’s state-of-the-art accuracy and label efficiency. Notably, with approximately 5% DOTA-v2.0 annotations, DML achieves nearly 90% of the fully supervised performance. The codes will be available at https://github.com/ZhangRuixiang-WHU/ALOD_DML/. Ruixiang Zhang, Chang Xu 0027, Wen Yang 0001, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | ConGeo: Robust Cross-View Geo-Localization Across Ground View Variations
Li Mi, Chang Xu 0027, Javiera Castillo-Navarro, Syrielle Montariol, Wen Yang 0001, Antoine Bosselut, Devis Tuia |
ECCV (14) | 2 |
| 2024 | Decoupling Representation for Nighttime Aerial TrackingabstractNighttime aerial tracking is an indispensable step towards around-the-clock real-world applications. However, RGBbased tracking algorithms face significant challenges at night due to their vulnerability to illumination. Observing that different feature channels have varying sensitivity to illumination, we propose to decouple the representation for illuminationsensitive and illumination-insensitive embeddings. We devise a Nighttime aerial tracking scheme via Decouple Representations, termed NiDR, where the Illumination-Invariant Embedding (IIE) module and the Illumination-Sensitive Embedding (ISE) module are designed to decouple representations. We achieve this semantic decoupling by utilizing a pair of normlight and low-light images and regulating the reconstruction and consistency relations between features. Experiments on UAVDark135 exhibit the remarkable performance of NiDR under challenging nighttime scenarios, surpassing the secondbest competitor by a large margin of 3.1% on precision. Xu Lei 0002, Yan Zhang 0115, Chang Xu 0027, Wen Yang 0001, Wensheng Cheng |
IGARSS | 3 |
| 2024 | NiDR: Nighttime Aerial Tracking via Decoupled RepresentationsabstractVanilla aerial trackers exhibit sensitivity to low-light conditions (e.g., nighttime aerial tracking scenario). To mitigate this, existing methods incorporate the light enhancement method as a preprocessing for aerial tracking. Despite the advancements, these approaches are restricted to the disparity in task objectives between the enhancer and tracker. Motivated by the observation that feature channels exhibit varying sensitivity to illumination, we propose to decouple the feature representation into two distinct parts: 1) illumination-invariant feature embedding and 2) illumination-sensitive feature embedding. The former, realized by the illumination invariant embedding (IIE) module, enhances features that remain invariant to illumination changes. Meanwhile, the latter, facilitated by the illumination sensitive embedding (ISE) module, aims to mitigate the negative impact of illumination-sensitive features on tracking performance. Building upon this decoupling strategy, we introduce NiDR, a simple yet effective nighttime aerial tracker. The proposed NiDR exhibits strong performance on three nighttime aerial tracking benchmarks (i.e., UAVDark135, NAT2021, and DarkTrack2021). Notably, it outperforms previous competitors by large margins, e.g., 3.1 points on the UAVDark135 and 2.0 points on the Darktrack2021 in terms of precision for nighttime scenarios. Xu Lei 0002, Yan Zhang 0115, Chang Xu 0027, Wensheng Cheng, Wen Yang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Learning Cross-View Visual Geo-Localization Without Ground TruthabstractCross-view geo-localization (CVGL) involves determining the geographical location of a query image by matching it with a corresponding GPS-tagged reference image. Current state-of-the-art methods predominantly rely on training models with labeled paired images, incurring substantial annotation costs and training burdens. In this study, we investigate the adaptation of frozen models for CVGL without requiring ground-truth pair labels. We observe that training on unlabeled cross-view images presents significant challenges, including establishing relationships within unlabeled data and reconciling view discrepancies between uncertain queries and references. To address these challenges, we propose a self-supervised learning framework to train a learnable adapter for a frozen foundation model (FM). This adapter is designed to map feature distributions from diverse views into a uniform space using unlabeled data exclusively. To establish relationships within unlabeled data, we introduce an expectation-maximization (EM)-based pseudolabeling module, which iteratively estimates matching between cross-view features and optimizes the adapter. To maintain the robustness of the FM’s representation, we incorporate an information consistency module with a reconstruction loss, ensuring that adapted features retain strong discriminative ability across views. Experimental results demonstrate that our proposed method achieves significant improvements over vanilla FMs and competitive accuracy compared to supervised methods while necessitating fewer training parameters and relying solely on unlabeled data. Evaluation of our adaptation for task-specific models further highlights its broad applicability. Particularly, on the University-1652 dataset, our method outperforms the FM baseline by a substantial margin, achieving about 39 points improvement in Recall@1 and more than 34 points increase in average precision (AP). The project is available athttps://collebt.github.io/EM-CVGL. Haoyuan Li 0005, Chang Xu 0027, Wen Yang 0001, Huai Yu, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Learning Cross-Modality High-Resolution Representation for Thermal Small-Object DetectionabstractThermal infrared (TIR) object detection plays a crucial role in diverse around-the-clock applications, such as search and rescue operations and wildlife protection. Achieving rapid and robust detection of small objects from an aerial perspective is particularly significant in these scenarios. However, the task is compounded by two interrelated challenges, rendering it even more tricky. For one, small objects only occupy a few pixels and contain limited information. For another, TIR sensors are typically low-resolution (LR) due to inherent challenges associated with the imaging mechanism of the TIR spectrum. In contrast, high-resolution (HR) RGB sensors are readily available due to their cost-effectiveness and widespread application. Recognizing the importance of HR information, especially in the context of small object detection, we propose a cross-modality high-resolution knowledge distillation framework (CMHRD), which leverages knowledge from the HR-RGB modality and provides a novel strategy for TIR small object detection. The proposed framework introduces three key components: a super-resolution generative distillation loss for cross-modal high-resolution representation learning, a cross-modality affinity distillation loss to extract scene-level cross-modality information, and a response distillation loss aimed at mimicking the HR prediction. To facilitate research on small object detection with HR-RGB and LR-TIR data, we have curated and annotated two datasets, namely NOAA-Seal and VTUAV-det-small. Experimental results on the NOAA-Seal demonstrate that CMHRD yields significant improvements, achieving a remarkable 6.39 mAP50 increase over a strong baseline without introducing additional computational cost during inference. Experiments on single-category dataset VTUAV-det-small and multi-category dataset RTDOD also show consistent improvements brought by CMHRD. The project is available at https://github.com/NNNNerd/CMHRD. Yan Zhang 0115, Xu Lei 0002, Chang Xu 0027, Wen Yang 0001, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Dynamic Coarse-to-Fine Learning for Oriented Tiny Object DetectionabstractDetecting arbitrarily oriented tiny objects poses intense challenges to existing detectors, especially for label assignment. Despite the exploration of adaptive label assignment in recent oriented object detectors, the extreme geometry shape and limited feature of oriented tiny objects still induce severe mismatch and imbalance issues. Specifically, the position prior, positive sample feature, and instance are mismatched, and the learning of extreme-shaped objects is biased and unbalanced due to little proper feature supervision. To tackle these issues, we propose a dynamic prior along with the coarse-to-fine assigner, dubbed DCFL. For one thing, we model the prior, label assignment, and object representation all in a dynamic manner to alleviate the mismatch issue. For another, we leverage the coarse prior matching and finer posterior constraint to dynamically assign labels, providing appropriate and relatively balanced supervision for diverse instances. Extensive experiments on six datasets show substantial improvements to the baseline. Notably, we obtain the state-of-the-art performance for one-stage detectors on the DOTA-v1.5, DOTA-v2.0, and DIOR-R datasets under single-scale training and testing. Codes are available at https://github.com/Chasel-Tsui/mmrotate-dcfl. Chang Xu 0027, Jian Ding 0001, Jinwang Wang, Wen Yang 0001, Huai Yu, Lei Yu 0006, Gui-Song Xia |
CVPR | 1 |
| 2023 | A3Track: Achieving Precise Target Tracking in Aerial Images With Receptive Field AlignmentabstractTracking arbitrary objects in aerial images presents formidable challenges to existing trackers. Among these challenges, the large scale variation and arbitrary geometry shape of visual targets are pronounced, resulting in two-fold mismatch issues between the feature receptive field and the tracking target. For one, there is a mismatch between the prior receptive field center and arbitrary-shaped targets. For another, the single receptive field mismatches the significantly scale-varied targets in the aerial imagery. To handle these challenges, we propose to Achieve precise Aerial tracking with receptive field Alignment, dubbed A3Track. The proposed A3Track is comprised of two modules: a Receptive Field Alignment (RFA) module and a Pyramid Receptive Field (PRF) module. First of all, we transform and update the receptive field center progressively, which drives the feature sampling location onto the targets’ main body, thus gradually yielding precise feature representation for arbitrary-shaped targets. We term this progressively updating process as the Receptive Field Alignment. Moreover, the PRF module constructs a set of pyramid features for the target, providing a multi-scale receptive field to handle the large scale variation of tracking objects. On four benchmarks, the new tracker A3Track achieves leading performance compared with existing methods and shows consistent improvements over baselines. The project is available at: https://chnleixu.github.io/A3Track-web/. Xu Lei 0002, Chang Xu 0027, Wensheng Cheng, Wen Yang 0001, Gui-Song Xia |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | RFLA: Gaussian Receptive Field Based Label Assignment for Tiny Object Detection
Chang Xu 0027, Jinwang Wang, Wen Yang 0001, Huai Yu, Lei Yu 0006, Gui-Song Xia |
ECCV (9) | 1 |