EDBT 2026 Demo / reviewers in the wild / expert
Ge Ya Luo
dblp:380/4100
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Generative modeling · 70% Reinforcement learning · 23% Trustworthy machine learning · 7% | |
| Computer graphics and multimedia
1 paper |
Multimedia systems and quality of experience · 100% |
Topics — the 5 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling › diffusion model
image editing |
0.9 | 1 | 2025 | The Promise of RL for Autoregressive Image Editing · NeurIPS 2025 |
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
0.9 | 1 | 2025 | The Promise of RL for Autoregressive Image Editing · NeurIPS 2025 |
Machine learning › Generative modeling
video generation |
0.9 | 1 | 2025 | Beyond FVD: An Enhanced Evaluation Metrics for Video Generation Distribution Quality · ICLR 2025 |
Machine learning › Generative modeling › generative model evaluation
video generation evaluation |
0.9 | 1 | 2025 | Beyond FVD: An Enhanced Evaluation Metrics for Video Generation Distribution Quality · ICLR 2025 |
Multimedia systems and quality of experience
video quality assessment |
0.9 | 1 | 2025 | Beyond FVD: An Enhanced Evaluation Metrics for Video Generation Distribution Quality · ICLR 2025 |
Methods — techniques the papers use, named apart from their topics
maximum mean discrepancy · 1.7joint embedding predictive architecture · 1.7fréchet video distance · 1.7supervised fine-tuning · 0.9multimodal LLM verifier · 0.9chain-of-thought · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Beyond FVD: An Enhanced Evaluation Metrics for Video Generation Distribution QualityabstractThe Fréchet Video Distance (FVD) is a widely adopted metric for evaluating video generation distribution quality. However, its effectiveness relies on critical assumptions. Our analysis reveals three significant limitations: (1) the non-Gaussianity of the Inflated 3D Convnet (I3D) feature space; (2) the insensitivity of I3D features to temporal distortions; (3) the impractical sample sizes required for reliable estimation. These findings undermine FVD's reliability and show that FVD falls short as a standalone metric for video generation evaluation. After extensive analysis of a wide range of metrics and backbone architectures, we propose JEDi, the JEPA Embedding Distance,
based on features derived from a Joint Embedding Predictive Architecture, measured using Maximum Mean Discrepancy with polynomial kernel. Our experiments on multiple open-source datasets show clear evidence that it is a superior alternative to the widely used FVD metric, requiring only 16% of the samples to reach its steady value, while increasing alignment with human evaluation by 34%, on average.
Project page: https://oooolga.github.io/JEDi.github.io/. Ge Ya Luo, Gian Mario Favero, Zhi Hao Luo, Alexia Jolicoeur-Martineau, Christopher Joseph Pal |
ICLR | 1 |
| 2025 | The Promise of RL for Autoregressive Image EditingabstractWhile image generation techniques are now capable of producing high-quality images that respect prompts which span multiple sentences, the task of text-guided image editing remains a challenge. Even edit requests that consist of only a few words often fail to be executed correctly. We explore three strategies to enhance performance on a wide range of image editing tasks: supervised fine-tuning (SFT), reinforcement learning (RL), and Chain-of-Thought (CoT) reasoning. In order to study all these components in one consistent framework, we adopt an autoregressive multimodal model that processes textual and visual tokens in a unified manner.
We find RL combined with a large multi-modal LLM verifier to be the most effective of these strategies.
As a result, we release **EARL**: **E**diting with **A**utoregression and **RL**, a strong RL-based image editing model that performs competitively on a diverse range of edits compared to strong baselines, despite using much less training data. Thus, EARL pushes the frontier of autoregressive multimodal models on image editing. We release our code, training data, and trained models at [https://github.com/mair-lab/EARL](https://github.com/mair-lab/EARL). Saba Ahmadi, Rabiul Awal, Ankur Sikarwar, Amirhossein Kazemnejad, Ge Ya Luo, Juan A. Rodríguez, Sai Rajeswar, Siva Reddy, Christopher Joseph Pal, Benno Krojer, Aishwarya Agrawal |
NeurIPS | 5 |