VLDB 2026 Research / reviewers in the wild / expert
Freya Tan
dblp:417/6828
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Segmentation and scene understanding · 44% Vision and language · 44% Image recognition and object detection · 13% |
Topics — the 3 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Segmentation and scene understanding
scene understanding |
1.0 | 1 | 2026 | MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes · AAAI 2026 |
Computer vision › Vision and language › multimodal reasoning
vision-language model reasoning |
1.0 | 1 | 2026 | MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes · AAAI 2026 |
Computer vision › Image recognition and object detection › object detection › category-specific object detection
person detection |
0.3 | 1 | 2026 | MINGLE: VLMs for Semantically Complex Region Detection in Urban Scenes · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
vision-language model · 1.0spatial aggregation · 1.0depth estimation · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MINGLE: VLMs for Semantically Complex Region Detection in Urban ScenesabstractUnderstanding group-level social interactions in public spaces is crucial for urban planning, informing the design of socially vibrant and inclusive environments. Detecting such interactions from images involves interpreting subtle visual cues such as relations, proximity and co-movement – semantically complex signals that go beyond traditional object detection. To address this challenge, we introduce a social group region detection task, which requires inferring and spatially grounding visual regions defined by abstract interpersonal relations. We propose MINGLE (Modeling INterpersonal Group-Level Engagement), a modular three-stage pipeline that integrates: (1) off-the-shelf human detection and depth estimation, (2) VLM-based reasoning to classify pairwise social affiliation, and (3) a lightweight spatial aggregation algorithm to localize socially connected groups. To support this task and encourage future research, we present a new dataset of 100K urban street-view images annotated with bounding boxes and labels for both individuals and socially interacting groups. The annotations combine human-created labels and outputs from the MINGLE pipeline, ensuring semantic richness and broad coverage of real world scenarios. Liu Liu 0018, Alexandra Schild, Marco Cipriano, Fatimeh Al Ghannam, Freya Tan, Gerard de Melo, Andres Sevtsuk |
AAAI | 5 |