VLDB 2026 Research / reviewers in the wild / expert
Gal Oren 0002
dblp:183/0735-2
· DBLP profile ↗
1ranked-venue papers
0as first author
0since 2021 · last 2019
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 1Graphics, computer vision, multimedia, augmented reality and games · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Vision and language · 100% |
Topics — the 2 heaviest of 2, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › image captioning
discriminative caption generation |
0.4 | 1 | 2019 | Joint Optimization for Cooperative Image Captioning · ICCV 2019 |
Computer vision › Vision and language
image captioning |
0.4 | 1 | 2019 | Joint Optimization for Cooperative Image Captioning · ICCV 2019 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 0.4partial-sampling straight-through · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Joint Optimization for Cooperative Image CaptioningabstractWhen describing images with natural language, descriptions can be made more informative if tuned for downstream tasks. This can be achieved by training two networks: a "speaker" that generates sentences given an image and a "listener" that uses them to perform a task. Unfortunately, training multiple networks jointly to communicate, faces two major challenges. First, the descriptions generated by a speaker network are discrete and stochastic, making optimization very hard and inefficient. Second, joint training usually causes the vocabulary used during communication to drift and diverge from natural language. To address these challenges, we present an effective optimization technique based on partial-sampling from a multinomial distribution combined with straight-through gradient updates, which we name PSST for Partial-Sampling Straight-Through. We then show that the generated descriptions can be kept close to natural by constraining them to be similar to human descriptions. Together, this approach creates descriptions that are both more discriminative and more natural than previous approaches. Evaluations on the COCO benchmark show that PSST improve the recall@10 from 60% to 86% maintaining comparable language naturalness. Human evaluations show that it also increases naturalness while keeping the discriminative power of generated captions. Gilad Vered, Gal Oren 0002, Yuval Atzmon, Gal Chechik |
ICCV | 2 |