EDBT 2026 Demo / reviewers in the wild / expert
Ahmed Bourouis
dblp:365/5952
· DBLP profile ↗
1ranked-venue papers
1as first author
1since 2021 · last 2024
0009-0007-8827-387XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Segmentation and scene understanding · 67% Vision and language · 33% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.8 | 1 | 2024 | Open Vocabulary Semantic Scene Sketch Understanding · CVPR 2024 |
Computer vision › Segmentation and scene understanding › image segmentation
sketch segmentation |
0.8 | 1 | 2024 | Open Vocabulary Semantic Scene Sketch Understanding · CVPR 2024 |
Computer vision › Vision and language
vision-language model |
0.8 | 1 | 2024 | Open Vocabulary Semantic Scene Sketch Understanding · CVPR 2024 |
Methods — techniques the papers use, named apart from their topics
visual prompt tuning · 0.8vision transformer · 0.8cross-attention · 0.8CLIP · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Open Vocabulary Semantic Scene Sketch UnderstandingabstractWe study the underexplored but fundamental problem of machine understanding of abstract freehand scene sketches. We introduce a sketch encoder that ensures a semantically-aware feature space, which we evaluate by testing its performance on a semantic sketch segmentation task. To train our model, we rely only on bitmap sketches accompanied by brief captions, avoiding the need for pixel-level annotations. To generalize to a large set of sketches and categories, we build upon a vision transformer encoder pre-trained with the CLIP model. We freeze the text encoder and perform visual-prompt tuning of the visual encoder branch while introducing a set of critical modifications. First, we augment the classical key-query (k-q) self-attention blocks with value-value (v-v) self-attention blocks. Central to our model is a two-level hierarchical training that enables efficient semantic disentanglement: The first level ensures holistic scene sketch encoding, and the second level focuses on individual categories. In the second level of the hierarchy, we introduce cross-attention between the text and vision branches. Our method outperforms zero-shot CLIP segmentation results by 37 points, reaching a pixel accuracy of 85.5% on the FS-COCO sketch dataset. Finally, we conduct a user study that allows us to identify further improvements needed over our method to reconcile machine and human understanding of freehand scene sketches. Ahmed Bourouis, Judith Ellen Fan, Yulia Gryaditskaya |
CVPR | 1 |