EDBT 2026 Demo / reviewers in the wild / expert
Yidan Zhang 0002
dblp:11/8540-2
· DBLP profile ↗
16ranked-venue papers
2as first author
16since 2021 · last 2026
0000-0002-7466-0234ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 13 · 2 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CFVBench: A Comprehensive Video Benchmark for Fine-grained Multimodal Retrieval-Augmented GenerationabstractMultimodal Retrieval-Augmented Generation (MRAG) enables Multimodal Large Language Models (MLLMs) to generate responses with external multimodal evidence. Recently, numerous video-based MRAG benchmarks have been proposed to evaluate model capabilities across retrieval and generation stages in MRAG. However, existing benchmarks remain limited in modality coverage and format diversity, often focusing on single- or limited-modality tasks, or coarse-grained scene understanding. To address these gaps, we introduce CFVBench, a large-scale, manually verified benchmark constructed from 599 publicly available videos, yielding 5,360 open-ended QA pairs. CFVBench, spans high-density formats and domains such as chart-heavy reports, news broadcasts, and software tutorials, requiring models to retrieve and reason over long temporal video spans while maintaining fine-grained multimodal information. Using CFVBench, we systematically evaluate 7 retrieval methods and 14 widely-used MLLMs, revealing a critical bottleneck: current models (even GPT5 or Gemini) struggle to capture transient yet essential fine-grained multimodal details. To mitigate this, we propose Adaptive Visual Refinement (AVR), a plug-and-paly framework that adaptively increases frame sampling density and selectively invokes external tools when necessary. Experiments show that AVR consistently enhances fine-grained multimodal comprehension and improves performance across all evaluated MLLMs. Kaiwen Wei, Ruida Liu, Changzai Pan, Yidan Zhang 0002, Peijin Wang, Yingchao Feng |
WWW | 11 |
| 2026 | Lightweight stereo image super-resolution via adaptive pruning and bridge distillation
Zhe Zhang 0041, Bingzheng Liu, Lei Chen 0091, Pengzhi Li, Yidan Zhang 0002, Jianjun Lei 0001 |
Knowl. Based Syst. | 5 |
| 2025 | MARA: A Multimodal Adaptive Retrieval-Augmented Framework for Document Question AnsweringabstractRetrieval-based multimodal document QA aims to identify and integrate relevant information from visually rich documents with complex multimodal structures. While retrieval-augmented generation (RAG) has shown strong performance in text-based QA, its extensions to multimodal documents remain underexplored and face significant limitations. Specifically, current approaches rely on query-agnostic document representations that overlook salient content and use static top-k evidence selection, which fails to adapt to the uncertain distribution of relevant information. To address these limitations, we propose the Multimodal Adaptive Retrieval-Augmented (MARA) framework, which introduces query-adaptive mechanisms to both retrieval and generation. MARA consists of two components: a Query-Aligned Region Encoder that builds multi-level document representations and reweights them based on query relevance to improve retrieval precision; and a Self-Reflective Evidence Controller that monitors evidence sufficiency during generation and adaptively incorporates content from lower-ranked sources using a sliding-window strategy. Experiments on six multimodal QA benchmarks demonstrate that MARA consistently improves retrieval relevance and answer quality over existing SOTA method. Haoquan Zhai, Yuchen Li 0006, Hengyi Cai, Peirong Zhang 0002, Yidan Zhang 0002, Lei Wang 0265, Chunle Wang, Yingyan Hou, Shuaiqiang Wang, Dawei Yin 0001 |
ACM Multimedia | 6 |
| 2025 | Referring Multi-Object Tracking in Satellite Videos: A New Benchmark and BaselineabstractReferring multi-object tracking (RMOT), which aims to track one or more objects in a video based on a natural language query, is increasingly crucial for a wide range of real-world applications. However, the study of RMOT in satellite video (RMOT-SV) scenarios remains limited, largely due to the high cost of data acquisition and the difficulty of annotation. To address this gap, we introduce RefSat, the first dataset for benchmarking RMOT-SV. RefSat comprises 212 top-down viewpoint video clips, totaling 31,129 frames, collected from a variety of publicly available satellite video datasets. By combining manual annotations of object appearance and position with automatic motion estimation, we build a semi-automatic pipeline that generates high-quality natural language descriptions covering object attributes and motion trajectories, resulting in over 4,000 objects paired with carefully designed textual queries. RefSat features satellite-specific challenges such as small object sizes, cloud occlusions, and motion-referenced semantics. To address these, we introduce RSRefTrack, a tailored baseline designed for small object perception and motion-aware grounding, which outperforms existing state-of-the-art RMOT methods on the RefSat benchmark. Project page: https://github.com/Zhang-Peirong/RefSat Peirong Zhang 0002, Yidan Zhang 0002, Hanru Shi, Dianyu Wang, Lei Wang 0265 |
ACM Multimedia | 2 |
| 2025 | Language-Guided Object Localization via Refined Spotting Enhancement in Remote Sensing ImageryabstractLanguage-guided remote sensing image object localization uses intuitive natural language interactions to locate objects of interest within satellite or drone imagery, and has a wide range of practical applications. Early research on this task typically used discriminative models that relied on predefined task heads, limited to locating a single object and lacking flexibility. Recent years, multi-modal large language models based generative models leverage their language understanding ability to comprehend more complex language references and provide flexible outputs. However, these models tend to be too large and have limited accuracy in spotting dense, small objects in remote sensing scenarios. The reasons are: 1) Generative models treating continuous coordinate prediction as a token classification problem which fails to reflect the actual localization gap. 2) Current methods typically use CLIP pre-trained encoders that align text-image pairs at a global level, may not capture the fine-grained semantic object information necessary for accurate localization. To address these shortcomings, we introduce a lightweight generative model named LM-RSE (Localization Model with Refined Spotting Enhancement). Refined Spotting Enhancement includes the refinement of bounding box outputs and the integration of fine-grained semantic features. Specifically, we design a Bounding Box Refinement (BBR) approach that includes a special token, a Bounding Box decoder, and a custom regression loss function to refine the spotting precision and optimize the training process. Additionally, we propose a Fine-Grained Semantic Integration (FGSI) strategy, integrates a fine-grained image encoder, a Vision-Language Semantic Processing (VLSP) layer and a two-stage, full-parameter training strategy, all working together to effectively enhance the granularity of spotting. Building on this, we further refine the language-guided object localization task into two types: expression-guided and class-guided. For the former, we utilize the RSVGD dataset for evaluation and achieved state-of-the-art performance; for the latter, our evaluation results surpassed those of GeoChat. Our code and checkpoints will be released at https://github.com/Zhang-Peirong/LM-RSE. Peirong Zhang 0002, Yidan Zhang 0002, Yingyan Hou, Lei Wang 0265 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Photogrammetry for Unconstrained Optical Satellite Imagery With Combined Neural Radiance FieldsabstractWe introduce a novel method tailored for unconstrained multiview optical satellite photogrammetry in time-varying illumination and reflection conditions. Our approach uses continuous radiance fields to represent surface radiance and albedo based on radiometry principles, integrating both static and transient components for satellite photogrammetry. In addition, an innovative self-supervised mechanism is introduced to optimize the learning process which leverages dark regions’ accentuation, transient and static composition, and shadow regularization. Evaluations on multidate WorldView-3 images affirm that our model consistently surpasses the state-of-the-art techniques. Xiaohe Li, Zide Fan, Yidan Zhang 0002, Yunping Ge, Lijie Wen 0001 |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | Few-Shot Incremental Object Detection in Aerial Imagery via Dual-Frequency PromptabstractRecently, there has been a growing interest in few-shot incremental object detection (FSIOD). It learns new tasks with limited data while mitigating catastrophic forgetting on previous tasks. However, existing FSIOD methods experience parameter changes after training on new tasks, causing a parameter competition issue among tasks. Additionally, the background information differs among various tasks, and using a common background weight for all tasks results in the background shift. Constrained by these two issues, existing methods only alleviate catastrophic forgetting and cannot wholly prevent the performance decline on previous tasks. Especially for complex remote sensing images with messy background, the models trained on new tasks exhibit noticeable performance drops on previous tasks. In this paper, we propose a novel FSIOD method via dual-frequency prompt to address these challenges, named FSIOD-DFP. It can completely eliminate catastrophic forgetting while mitigating over-fitting. Specifically, a dual-frequency prompt generator is designed to tackle the parameter competition issue. It decouples the frequency components of images to produce prompts that modify the images to adapt to the base model trained on previous tasks. Compared to traditional prompts, our generator introduces fewer parameters to address over-fitting for limited data and allows freezing the base model to maintain the performance of previous data. Besides, a self-regularization loss is introduced to guide the prompt-modified images to leverage the knowledge of the base model effectively. Furthermore, we propose a task-decoupled detection head to address the background shift problem. It separates the detection heads for new and previous tasks to resolve the conflict in the background between different tasks. In FSIOD-DFP, only a prompt generator and a novel detection head are added and fine-tuned when learning a new task. Experiments on three remote sensing object detection datasets demonstrate that our method achieves state-of-the-art performance on both new and previous tasks in all few-shot incremental settings. Wenhui Diao, Junxi Li, Yidan Zhang 0002, Peijin Wang, Xian Sun 0001, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Confidence-Weighted Dual-Teacher Networks With Biased Contrastive Learning for Semi-Supervised Semantic Segmentation in Remote Sensing ImagesabstractSemantic segmentation of remote sensing images is vital in remote sensing technology. High-quality models for this task require a vast amount of images, manual annotation is a process that is time-consuming and labor-intensive. Consequently, this has catalyzed the emergence of semi-supervised semantic segmentation methods. However, the complexity of foreground categories in remote sensing images poses a challenge to maintaining prediction consistency. Moreover, inherent characteristics such as intra-class variations and inter-class similarities result in a certain degree of confusion among features of different classes in the feature space. This impacts the final classification results. In order to improve the model’s consistency and optimize the classification of categories based on features, this paper proposes a new semi-supervised semantic segmentation framework that combines consistency regularization and contrastive learning. In terms of consistency regularization, the proposed method incorporates dual teacher networks, introduces ClassMix for image augmentation, and utilizes confidence levels to integrate the predictions from these networks. By introducing perturbations at both the network and image levels, while simultaneously maintaining consistency, the predictive prowess and generalization ability of the model are enhanced. For contrastive learning, Postive-Unlabeled Learning (PU-Learning) is employed to improve the problem of mis-sampling when selecting features. At the same time, higher biased weights are allocated to more challenging negative samples, thereby elevating the complexity of feature learning and enhancing the discriminative capability of the final feature representation space. Our extensive experiments on the ISPRS Vaihingen dataset and the challenging iSAID dataset have served to underscore the superior performance of our approach. Zide Fan, Xiyu Qi, Yidan Zhang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Negative-Core Sample Knowledge Distillation for Oriented Object Detection in Remote Sensing ImageabstractKnowledge distillation (KD) has been one of the most effective methods for enhancing the performance of lightweight detectors, crucial for remote sensing edge intelligence models. However, many mainstream distillation methods that are centered around the paradigm of distilling positive samples show weak exploitation of the student’s potential. This arises due to these methods overlooking the core teacher-student difference in remote sensing scenarios with vast and object-similar backgrounds. In this article, from the point of distillation sample and knowledge hierarchy, we design a negative-core sample knowledge distillation (NSD) method for improving the performance of the lightweight object detection model. Specifically, a negative-core sample (NCS) is innovatively employed to transfer effective background discrimination knowledge for bridging the core difference. KD for NCS across four levels—pixel, logit, box, and angle—are customized to fully leverage the teacher’s insights. Category direction estimation (CE) is incorporated into the angle KD to convey NCS-oriented knowledge more effectively. Extensive experiments conducted on multiple remote sensing datasets achieve state-of-the-art (SOTA) performance, demonstrating the effectiveness of the proposed NSD. Codes are available athttps://github.com/Changan00/NSD. Yidan Zhang 0002, Feilong Huang, Xiyu Qi |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | ST-Net: Scattering Topology Network for Aircraft Classification in High-Resolution SAR ImagesabstractAircraft classification in synthetic aperture radar (SAR) images plays a considerable role in global region management and surveillance. Recently, deep learning has been applied to solve the classification problem and made significant progress. Due to the imaging variability at different angles and component scattering discreteness in SAR images, previous works have had difficulty in achieving desirable classification results. To address these issues, we study the positional and semantic relationship between the scattering points and propose an innovative scattering topology network (ST-Net) in this article. First, considering the diversity of imaging results caused by different target attitude angles, we extract and transform the scattering cluster centers to update the information of various categories. It can guide the model to strengthen the discriminative features and mitigate the impact of imaging variability on classification performance. Second, a novel scattering topology module (STM) is introduced to model the spatial relationships and semantic information interaction of discrete scattering points. In this process, the topology relations and scattering characteristics are enhanced for further accurate classification. Third, context attention excitation (CAE) is designed to capture significant global and semantic information, which is conducive to suppressing background interference and reducing category confusion. In conclusion, the ST-Net is presented with the SAR imaging mechanism and the topology geometric representation of aircraft. We construct the SAR aircraft category dataset (SAR-ACD) and conduct extensive experiments on it to show the effectiveness of ST-Net, which illustrates that our method achieves superior classification performance. Yuzhuo Kang, Zhirui Wang 0003, Haoyu Zuo, Yidan Zhang 0002, Zhujun Yang, Xian Sun 0001, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | CFTracker: Multi-Object Tracking With Cross-Frame Connections in Satellite VideosabstractMulti-object tracking (MOT) in satellite videos is an essential topic with many applications, such as traffic monitoring and disaster response. However, many multi-object trackers that perform well in natural scenes show weak generalization in satellite videos due to low object discrimination caused by low spatial resolution and the widespread indistinguishable background, such as clouds and reflections. In this paper, we design a novel multi-object tracking framework called CFTracker (Cross-Frame Tracker) for satellite videos from the point of both network structure and training method. On the one hand, in network structure design, a cross-frame feature update module (CFU) is proposed to enhance object recognition and reduce the response to background noises by using rich temporal semantic information. On the other hand, we reveal that the picture-pair training approach used by the mainstream MOT network is not entirely conducive to the network learning temporal semantic information. To better grasp the cross-frame feature connections and output time-consistent motion predictions, we train CFTracker by a novel cross-frame training flow (CT). Experiments demonstrate the effectiveness of our CFTracker and obtain state-of-the-art tracking accuracy and precision of 72.9% score on the AIR-MOT dataset and 57.1% score on the VISO dataset. The code will be available online. Lingyu Kong, Yidan Zhang 0002, Wenhui Diao, Zining Zhu 0003, Lei Wang 0077 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Few-Shot Object Detection in Aerial Imagery Guided by Text-Modal KnowledgeabstractFew-shot object detection (FSOD) has received numerous attention due to the difficulty and time-consuming of labeling objects. Recent researches achieve excellent performance in a natural scene by only using a few instances of novel classes to fine-tune the last prediction layer of the model well-trained on plentiful base data. However, compared with natural scene objects with a single direction and small size variety, the direction and size of the objects in remote sensing images (RSIs) vary greatly. The methods proposed for the natural scene cannot be directly applied to RSIs. In this article, we first propose a strong baseline for RSIs. It fine-tunes all detector components acting on high-level features and effectively improves the performance of novel classes. Further analyzing the results of the baseline, we find that the error for novel classes is mainly concentrated in classification. It misclassifies novel classes as confusable base classes or backgrounds due to the difficulty in extracting generalized information from limited instances. As is well-known, text-modal knowledge can highly summarize the generalized and unique characteristics of categories. Thus, we introduce text-modal descriptions for each category and propose an FSOD method guided by TExt-MOdal knowledge, called TEMO. Specifically, a text-modal knowledge extractor and a cross-modal assembly module are proposed to extract text features and fuse the text-modal features into visual-modal features. The fused features greatly reduce the classification confusion of novel classes. Furthermore, we introduce a mask strategy and a separation loss to avoid over-fitting and ambiguity of text-modal features. Experimental results on detection in optical remote sensing images (DIOR), Northwestern Polytechnical University (NWPU), and fine-grained object recognition in high-resolution remote sensing imagery (FAIR1M) illustrate that our TEMO achieves state-of-the-art performance in all settings. Xian Sun 0001, Wenhui Diao, Yongqiang Mao, Junxi Li, Yidan Zhang 0002, Peijin Wang, Kun Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | PICS: Paradigms Integration and Contrastive Selection for Semisupervised Remote Sensing Images Semantic SegmentationabstractRemote sensing images semantic segmentation is a fundamental yet challenging task, which has long relied heavily on sufficient pixelwise annotations. Semisupervised learning is proposed to address the problem of high dependence on labeled data by exploiting more learnable samples generated from the large amounts of accessible unlabeled data. However, affected by the complexity and diversity of remote sensing images, various misclassifications often occur and lead to errors accumulation during model training. Errors accumulation will destroy the consistency of model training and lead to degradation of final segmentation performance. In this article, in order to further alleviate the damage caused by the errors to the consistency of model training and improve final segmentation accuracy, we propose a novel semisupervised segmentation framework, paradigms integration and contrastive selection (PICS). First, multiple proven semisupervised paradigms are integrated to generate pseudolabeled samples with less noise. Second, a loss-based contrastive selection method is explored to distinguish generated samples that contain different degrees of inevitable misclassification, thereby further maintaining the approximation of the generated samples and the ground truth in the sample space. By generating and selecting high-quality pseudolabeled samples for selective self-training, we can better guarantee consistency during model training and obtain better segmentation results. Extensive experiments over the ISPRS Vaihingen, Potsdam, and the challenging iSAID benchmarks demonstrate that our method yields significant accuracy boosting on the segmentation results and achieves on-par performance with the state of the arts. Xiyu Qi, Yongqiang Mao, Yidan Zhang 0002, Lei Wang 0077 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Bridging the Gap Between Cumbersome and Light Detectors via Layer-Calibration and Task-Disentangle Distillation in Remote Sensing ImageryabstractWith urgent application requirements, such as satellite in-orbit processing and unmanned aerial vehicle tracking, knowledge distillation (KD) following the teacher–student teaching mechanism has shown great potential to obtain lightweight detectors. However, compact students have limited accuracy due to the interference of large-scale variations and blurred boundaries in remote sensing objects. Specifically, previous methods mostly force teacher–student responses from the layer of the same depth and scale to align. Stereotyped manual interlayer associations may cause discriminative features of multiscale objects to be incorrectly bundled. Furthermore, the regression branch follows the identical distillation paradigm as the classification branch, resulting in ambiguous object bounding box deviations. To solve the above two issues, we propose an effective KD framework called layer-calibration and task-disentangle distillation (LTD). First, the cross-layer calibration distillation (CCD) structure is innovatively proposed. It adaptively binds a student layer with several related target layers, rather than a fixed layer in the teacher model. Appropriate and clear knowledge of large and small objects is transmitted. Since the CCD structure requires explicit global inner product computation between multiple layers, the local implicit calibration (LIC) module is further proposed to reduce distilled convergence difficulty. Second, the task-aware spatial disentangle distillation (TASD) structure is devised to transfer task-decoupled semantics and localization knowledge in a divide-and-conquer manner, alleviating objects’ localization imprecision. Experiments demonstrate that our LTD achieves state-of-the-art performance on several datasets and is a plug-and-play approach to most detectors. The code will be available soon. Yidan Zhang 0002, Xian Sun 0001, Junxi Li, Yongqiang Mao, Lei Wang 0077 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | SIL-LAND: Segmentation Incremental Learning in Aerial Imagery via LAbel Number Distribution ConsistencyabstractSegmentation incremental learning has received a lot of attention in recent years due to the ability to overcome the problem of catastrophic forgetting. Our study found that differences in label number distribution affect the performance of segmentation incremental learning. Because the labels for pixels of the old category are marked as background when the model is trained on the new tasks, the label number distribution is inconsistent with static learning that is considered to be the upper bound on incremental learning, which hinders the mitigation of the catastrophic forgetting problem. In response to the above problems, we propose an incremental learning method named SIL-LAND, which improves the accuracy by making the label number distribution of our method close to that of static learning. From the perspective of high-level semantic labels, we propose the prototype update mechanism for the problem that non-adaptive representative prototypes ignore the sample diversity of semantic categories in remote sensing images. By compensating for the difference in label number distribution at the feature level, the distance between the prototype and the actual class center is reduced; Aiming at the lack of semantic consistency between feature vectors and prototypes, we propose a similarity measure module to increase the intra-class similarity between the prototype and corresponding feature vectors. From the perspective of one-hot labels, we propose label reconstruction, including foreground screening and background padding to make the number distribution of one-hot labels as close as possible to that of static learning. A series of experimental results demonstrate the effectiveness of our method. Junxi Li, Wenhui Diao, Peijin Wang, Yidan Zhang 0002, Zhujun Yang, Guangluan Xu, Xian Sun 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Learning Efficient and Accurate Detectors With Dynamic Knowledge Distillation in Remote Sensing ImageryabstractDeep convolutional neural networks (CNNs) have brought a tremendous increase in detection accuracy, but too cumbersome model makes them hard to deploy on low computation edge devices, such as satellites and unmanned aerial vehicles. A promising method to tackle this problem is knowledge distillation (KD), which makes models lightweight with satisfactory accuracy. For remote sensing images, the objects are usually environment-related and located in a cluttered scene. The features that objects’ semantic information relies on are tangled. However, existing distillation methods only imitate feature distribution derived from regions, including objects resulting in poor performance. Furthermore, masses of instances generated by teachers are blindly inherited, even if some of them are outliers. In this article, we propose a general and effective KD framework called dynamic knowledge distillation (DKD). First, our framework leverages the dynamic global distillation (GD) module to discover valuable regions from the foreground and background for multiscale features imitation, avoiding ignoring the potential geographical spatial relationship. Second, we propose a dynamic instance selection distillation (ISD) module to give students the ability of self-judgment through the magnitude of detection loss. Third, toward more accurate handling of hard samples in regression, a training-status-aware loss is tailored to guide students mine knowledge about objects with large aspect ratio or small size. Extensive experiments are conducted to show the effectiveness of DKD framework. The detection results on DOTA and NWPU VHR-10 dataset illustrate that our method is suitable for single-stage, two-stage and even anchor-free detectors. It shows the state-of-the-art performance. The code will be publicly available. Yidan Zhang 0002, Xian Sun 0001, Wenhui Diao, Kun Fu 0001, Lei Wang 0077 |
IEEE Trans. Geosci. Remote. Sens. | 1 |