EDBT 2026 Demo / reviewers in the wild / expert
Michael Ogezi
dblp:349/0307
· DBLP profile ↗
1ranked-venue papers
1as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Knowledge representation and reasoning · 44% Vision and language · 44% 3D vision · 13% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Knowledge representation and reasoning
spatial reasoning |
0.9 | 1 | 2025 | SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data · ACL (1) 2025 |
Computer vision › Vision and language
visual question answering |
0.9 | 1 | 2025 | SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data · ACL (1) 2025 |
Computer vision › 3D vision › 3d scene understanding
spatial relation understanding |
0.3 | 1 | 2025 | SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data · ACL (1) 2025 |
Methods — techniques the papers use, named apart from their topics
vision-language model fine-tuning · 0.9synthetic data generation · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic DataabstractVision-language models (VLMs) work well in tasks ranging from image captioning to visual question answering (VQA), yet they struggle with spatial reasoning, a key skill for understanding our physical world that humans excel at.We find that spatial relations are generally rare in widely used VL datasets, with only a few being well represented while most form a long tail of underrepresented relations.This gap leaves VLMs ill-equipped to handle diverse spatial relationships.To bridge it, we construct a synthetic VQA dataset focused on spatial reasoning generated from hyperdetailed image descriptions in Localized Narratives, DOCCI, and PixMo-Cap.Our dataset consists of 455k samples containing 3.4 million QA pairs.Trained on this dataset, our Spatial-Reasoning Enhanced (SpaRE) VLMs show strong improvements on spatial reasoning benchmarks, achieving up to a 49% performance gain on the What's Up benchmark, while maintaining strong results on general tasks.Our work narrows the gap between human and VLM spatial reasoning and makes VLMs more capable in real-world tasks such as robotics and navigation.We plan to share our code and dataset in due course. Michael Ogezi, Freda Shi |
ACL (1) | 1 |