EDBT 2026 Demo / reviewers in the wild / expert
Peirong Zhang 0002
dblp:306/3180-2
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2025
0009-0003-4171-3234ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Question answering and dialogue systems · 30% Video understanding and tracking · 30% Segmentation and scene understanding · 30% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 4 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Question answering and dialogue systems › machine reading comprehension
document question answering |
0.9 | 1 | 2025 | MARA: A Multimodal Adaptive Retrieval-Augmented Framework for Document Question Answering · ACM Multimedia 2025 |
Computer vision › Video understanding and tracking › multi-object tracking
referring multi-object tracking |
0.9 | 1 | 2025 | Referring Multi-Object Tracking in Satellite Videos: A New Benchmark and Baseline · ACM Multimedia 2025 |
Information retrieval
retrieval-augmented generation |
0.9 | 1 | 2025 | MARA: A Multimodal Adaptive Retrieval-Augmented Framework for Document Question Answering · ACM Multimedia 2025 |
Computer vision › Vision and language
visual grounding |
0.3 | 1 | 2025 | Referring Multi-Object Tracking in Satellite Videos: A New Benchmark and Baseline · ACM Multimedia 2025 |
Methods — techniques the papers use, named apart from their topics
self-reflective evidence control · 1.7query-adaptive retrieval · 1.7semi-automatic annotation · 0.9motion estimation · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MARA: A Multimodal Adaptive Retrieval-Augmented Framework for Document Question AnsweringabstractRetrieval-based multimodal document QA aims to identify and integrate relevant information from visually rich documents with complex multimodal structures. While retrieval-augmented generation (RAG) has shown strong performance in text-based QA, its extensions to multimodal documents remain underexplored and face significant limitations. Specifically, current approaches rely on query-agnostic document representations that overlook salient content and use static top-k evidence selection, which fails to adapt to the uncertain distribution of relevant information. To address these limitations, we propose the Multimodal Adaptive Retrieval-Augmented (MARA) framework, which introduces query-adaptive mechanisms to both retrieval and generation. MARA consists of two components: a Query-Aligned Region Encoder that builds multi-level document representations and reweights them based on query relevance to improve retrieval precision; and a Self-Reflective Evidence Controller that monitors evidence sufficiency during generation and adaptively incorporates content from lower-ranked sources using a sliding-window strategy. Experiments on six multimodal QA benchmarks demonstrate that MARA consistently improves retrieval relevance and answer quality over existing SOTA method. Haoquan Zhai, Yuchen Li 0006, Hengyi Cai, Peirong Zhang 0002, Yidan Zhang 0002, Lei Wang 0265, Chunle Wang, Yingyan Hou, Shuaiqiang Wang, Dawei Yin 0001 |
ACM Multimedia | 5 |
| 2025 | Referring Multi-Object Tracking in Satellite Videos: A New Benchmark and BaselineabstractReferring multi-object tracking (RMOT), which aims to track one or more objects in a video based on a natural language query, is increasingly crucial for a wide range of real-world applications. However, the study of RMOT in satellite video (RMOT-SV) scenarios remains limited, largely due to the high cost of data acquisition and the difficulty of annotation. To address this gap, we introduce RefSat, the first dataset for benchmarking RMOT-SV. RefSat comprises 212 top-down viewpoint video clips, totaling 31,129 frames, collected from a variety of publicly available satellite video datasets. By combining manual annotations of object appearance and position with automatic motion estimation, we build a semi-automatic pipeline that generates high-quality natural language descriptions covering object attributes and motion trajectories, resulting in over 4,000 objects paired with carefully designed textual queries. RefSat features satellite-specific challenges such as small object sizes, cloud occlusions, and motion-referenced semantics. To address these, we introduce RSRefTrack, a tailored baseline designed for small object perception and motion-aware grounding, which outperforms existing state-of-the-art RMOT methods on the RefSat benchmark. Project page: https://github.com/Zhang-Peirong/RefSat Peirong Zhang 0002, Yidan Zhang 0002, Hanru Shi, Dianyu Wang, Lei Wang 0265 |
ACM Multimedia | 1 |
| 2025 | Language-Guided Object Localization via Refined Spotting Enhancement in Remote Sensing ImageryabstractLanguage-guided remote sensing image object localization uses intuitive natural language interactions to locate objects of interest within satellite or drone imagery, and has a wide range of practical applications. Early research on this task typically used discriminative models that relied on predefined task heads, limited to locating a single object and lacking flexibility. Recent years, multi-modal large language models based generative models leverage their language understanding ability to comprehend more complex language references and provide flexible outputs. However, these models tend to be too large and have limited accuracy in spotting dense, small objects in remote sensing scenarios. The reasons are: 1) Generative models treating continuous coordinate prediction as a token classification problem which fails to reflect the actual localization gap. 2) Current methods typically use CLIP pre-trained encoders that align text-image pairs at a global level, may not capture the fine-grained semantic object information necessary for accurate localization. To address these shortcomings, we introduce a lightweight generative model named LM-RSE (Localization Model with Refined Spotting Enhancement). Refined Spotting Enhancement includes the refinement of bounding box outputs and the integration of fine-grained semantic features. Specifically, we design a Bounding Box Refinement (BBR) approach that includes a special token, a Bounding Box decoder, and a custom regression loss function to refine the spotting precision and optimize the training process. Additionally, we propose a Fine-Grained Semantic Integration (FGSI) strategy, integrates a fine-grained image encoder, a Vision-Language Semantic Processing (VLSP) layer and a two-stage, full-parameter training strategy, all working together to effectively enhance the granularity of spotting. Building on this, we further refine the language-guided object localization task into two types: expression-guided and class-guided. For the former, we utilize the RSVGD dataset for evaluation and achieved state-of-the-art performance; for the latter, our evaluation results surpassed those of GeoChat. Our code and checkpoints will be released at https://github.com/Zhang-Peirong/LM-RSE. Peirong Zhang 0002, Yidan Zhang 0002, Yingyan Hou, Lei Wang 0265 |
IEEE Trans. Geosci. Remote. Sens. | 1 |