Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Peirong Zhang 0002

dblp:306/3180-2 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2025
0009-0003-4171-3234ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Question answering and dialogue systems · 30% Video understanding and tracking · 30% Segmentation and scene understanding · 30%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems › machine reading comprehension
document question answering
0.912025
MARA: A Multimodal Adaptive Retrieval-Augmented Framework for Document Question Answering · ACM Multimedia 2025
Computer vision › Video understanding and tracking › multi-object tracking
referring multi-object tracking
0.912025
Referring Multi-Object Tracking in Satellite Videos: A New Benchmark and Baseline · ACM Multimedia 2025
Information retrieval
retrieval-augmented generation
0.912025
MARA: A Multimodal Adaptive Retrieval-Augmented Framework for Document Question Answering · ACM Multimedia 2025
Computer vision › Vision and language
visual grounding
0.312025
Referring Multi-Object Tracking in Satellite Videos: A New Benchmark and Baseline · ACM Multimedia 2025

Methods — techniques the papers use, named apart from their topics

self-reflective evidence control · 1.7query-adaptive retrieval · 1.7semi-automatic annotation · 0.9motion estimation · 0.9
YearPublicationVenuePosition
2025 MARA: A Multimodal Adaptive Retrieval-Augmented Framework for Document Question Answering
abstract
Retrieval-based multimodal document QA aims to identify and integrate relevant information from visually rich documents with complex multimodal structures. While retrieval-augmented generation (RAG) has shown strong performance in text-based QA, its extensions to multimodal documents remain underexplored and face significant limitations. Specifically, current approaches rely on query-agnostic document representations that overlook salient content and use static top-k evidence selection, which fails to adapt to the uncertain distribution of relevant information. To address these limitations, we propose the Multimodal Adaptive Retrieval-Augmented (MARA) framework, which introduces query-adaptive mechanisms to both retrieval and generation. MARA consists of two components: a Query-Aligned Region Encoder that builds multi-level document representations and reweights them based on query relevance to improve retrieval precision; and a Self-Reflective Evidence Controller that monitors evidence sufficiency during generation and adaptively incorporates content from lower-ranked sources using a sliding-window strategy. Experiments on six multimodal QA benchmarks demonstrate that MARA consistently improves retrieval relevance and answer quality over existing SOTA method.
Haoquan Zhai, Yuchen Li 0006, Hengyi Cai, Peirong Zhang 0002, Yidan Zhang 0002, Lei Wang 0265, Chunle Wang, Yingyan Hou, Shuaiqiang Wang, Dawei Yin 0001
ACM Multimedia5
2025 Referring Multi-Object Tracking in Satellite Videos: A New Benchmark and Baseline
abstract
Referring multi-object tracking (RMOT), which aims to track one or more objects in a video based on a natural language query, is increasingly crucial for a wide range of real-world applications. However, the study of RMOT in satellite video (RMOT-SV) scenarios remains limited, largely due to the high cost of data acquisition and the difficulty of annotation. To address this gap, we introduce RefSat, the first dataset for benchmarking RMOT-SV. RefSat comprises 212 top-down viewpoint video clips, totaling 31,129 frames, collected from a variety of publicly available satellite video datasets. By combining manual annotations of object appearance and position with automatic motion estimation, we build a semi-automatic pipeline that generates high-quality natural language descriptions covering object attributes and motion trajectories, resulting in over 4,000 objects paired with carefully designed textual queries. RefSat features satellite-specific challenges such as small object sizes, cloud occlusions, and motion-referenced semantics. To address these, we introduce RSRefTrack, a tailored baseline designed for small object perception and motion-aware grounding, which outperforms existing state-of-the-art RMOT methods on the RefSat benchmark. Project page: https://github.com/Zhang-Peirong/RefSat
Peirong Zhang 0002, Yidan Zhang 0002, Hanru Shi, Dianyu Wang, Lei Wang 0265
ACM Multimedia1
2025 Language-Guided Object Localization via Refined Spotting Enhancement in Remote Sensing Imagery
abstract
Language-guided remote sensing image object localization uses intuitive natural language interactions to locate objects of interest within satellite or drone imagery, and has a wide range of practical applications. Early research on this task typically used discriminative models that relied on predefined task heads, limited to locating a single object and lacking flexibility. Recent years, multi-modal large language models based generative models leverage their language understanding ability to comprehend more complex language references and provide flexible outputs. However, these models tend to be too large and have limited accuracy in spotting dense, small objects in remote sensing scenarios. The reasons are: 1) Generative models treating continuous coordinate prediction as a token classification problem which fails to reflect the actual localization gap. 2) Current methods typically use CLIP pre-trained encoders that align text-image pairs at a global level, may not capture the fine-grained semantic object information necessary for accurate localization. To address these shortcomings, we introduce a lightweight generative model named LM-RSE (Localization Model with Refined Spotting Enhancement). Refined Spotting Enhancement includes the refinement of bounding box outputs and the integration of fine-grained semantic features. Specifically, we design a Bounding Box Refinement (BBR) approach that includes a special token, a Bounding Box decoder, and a custom regression loss function to refine the spotting precision and optimize the training process. Additionally, we propose a Fine-Grained Semantic Integration (FGSI) strategy, integrates a fine-grained image encoder, a Vision-Language Semantic Processing (VLSP) layer and a two-stage, full-parameter training strategy, all working together to effectively enhance the granularity of spotting. Building on this, we further refine the language-guided object localization task into two types: expression-guided and class-guided. For the former, we utilize the RSVGD dataset for evaluation and achieved state-of-the-art performance; for the latter, our evaluation results surpassed those of GeoChat. Our code and checkpoints will be released at https://github.com/Zhang-Peirong/LM-RSE.
Peirong Zhang 0002, Yidan Zhang 0002, Yingyan Hou, Lei Wang 0265
IEEE Trans. Geosci. Remote. Sens.1