VLDB 2026 Research / reviewers in the wild / expert
Enzhuo Zhang
dblp:401/6384
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Vision and language · 87% Efficient and distributed learning · 13% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
vision-language model |
0.9 | 1 | 2025 | Quality-Driven Curation of Remote Sensing Vision-Language Data via Learned Scoring Models · NeurIPS 2025 |
Computer vision › Vision and language › vision-language model › vision-language model adaptation
vision-language model fine-tuning |
0.9 | 1 | 2025 | Quality-Driven Curation of Remote Sensing Vision-Language Data via Learned Scoring Models · NeurIPS 2025 |
Machine learning › Efficient and distributed learning
data selection |
0.3 | 1 | 2025 | Quality-Driven Curation of Remote Sensing Vision-Language Data via Learned Scoring Models · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 0.9best-of-n test-time scaling · 0.9CLIP scoring · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Quality-Driven Curation of Remote Sensing Vision-Language Data via Learned Scoring ModelsabstractVision-Language Models (VLMs) have demonstrated great potential in interpreting remote sensing (RS) images through language-guided semantic. However, the effectiveness of these VLMs critically depends on high-quality image-text training data that captures rich semantic relationships between visual content and language descriptions. Unlike natural images, RS lacks large-scale interleaved image-text pairs from web data, making data collection challenging. While current approaches rely primarily on rule-based methods or flagship VLMs for data synthesis, a systematic framework for automated quality assessment of such synthetically generated RS vision-language data is notably absent. To fill this gap, we propose a novel score model trained on large-scale RS vision-language preference data for automated quality assessment. Our empirical results demonstrate that fine-tuning CLIP or advanced VLMs (e.g., Qwen2-VL) with the top 30% of data ranked by our score model achieves superior accuracy compared to both full-data fine-tuning and CLIP-score-based ranking approaches. Furthermore, we demonstrate applications of our scoring model for reinforcement learning (RL) training and best-of-N (BoN) test-time scaling, enabling significant improvements in VLM performance for RS tasks. Our code, model, and dataset are publicly available. Dilxat Muhtar, Enzhuo Zhang, Zhenshi Li, Yanglangxing He, Pengfeng Xiao, Xueliang Zhang 0002 |
NeurIPS | 2 |