VLDB 2026 Research / reviewers in the wild / expert
Akash Gokul
dblp:267/8682
· DBLP profile ↗
8ranked-venue papers
0as first author
8since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Generative modeling · 74% Language models and text generation · 12% Learning paradigms · 6% | |
| Computer graphics and multimedia
2 papers |
Visual content generation and editing · 84% Audio and music processing · 16% | |
| Software engineering, system software, and programming languages
1 paper |
Program synthesis and code generation · 100% |
Topics — the 23 heaviest of 23, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
3.8 | 5 | 2025 | LaViDa: A Large Diffusion Language Model for Multimodal Understanding · NeurIPS 2025 OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows · CVPR 2025 Aligning Diffusion Models by Optimizing Human Utility · NeurIPS 2024 |
Machine learning › Generative modeling › multimodal generation
any-to-any generation |
0.9 | 1 | 2025 | OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows · CVPR 2025 |
Machine learning › Generative modeling › diffusion model
diffusion transformer |
0.9 | 1 | 2025 | Reflect-DiT: Inference-Time Scaling for Text-to-Image Diffusion Transformers via In-Context Reflection · ICCV 2025 |
Machine learning › Generative modeling › diffusion model › discrete diffusion model
discrete diffusion language model |
0.9 | 1 | 2025 | LaViDa: A Large Diffusion Language Model for Multimodal Understanding · NeurIPS 2025 |
Machine learning › Generative modeling › multimodal generation
multimodal generative model |
0.9 | 1 | 2025 | OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows · CVPR 2025 |
Computer vision › Vision and language
multimodal understanding |
0.9 | 1 | 2025 | LaViDa: A Large Diffusion Language Model for Multimodal Understanding · NeurIPS 2025 |
Machine learning › Generative modeling › diffusion model
rectified flow |
0.9 | 1 | 2025 | OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows · CVPR 2025 |
Natural language and speech › Language models and text generation
test-time scaling |
0.9 | 1 | 2025 | Reflect-DiT: Inference-Time Scaling for Text-to-Image Diffusion Transformers via In-Context Reflection · ICCV 2025 |
Machine learning › Generative modeling › diffusion model › text-to-image generation
text-to-image diffusion model |
0.9 | 1 | 2025 | Reflect-DiT: Inference-Time Scaling for Text-to-Image Diffusion Transformers via In-Context Reflection · ICCV 2025 |
Natural language and speech › Language models and text generation › alignment
preference alignment |
0.8 | 1 | 2024 | Aligning Diffusion Models by Optimizing Human Utility · NeurIPS 2024 |
Machine learning › Generative modeling › diffusion model › text-to-image generation
text-to-image diffusion model alignment |
0.8 | 1 | 2024 | Aligning Diffusion Models by Optimizing Human Utility · NeurIPS 2024 |
Program synthesis and code generation › generative programming
modular code generation |
0.8 | 1 | 2024 | CodeChain: Towards Modular Code Generation Through Chain of Self-revisions with Representative Sub-modules · ICLR 2024 |
Machine learning › Generative modeling › diffusion model › guided diffusion
classifier guidance |
0.7 | 1 | 2023 | End-to-End Diffusion Latent Optimization Improves Classifier Guidance · ICCV 2023 |
Machine learning › Generative modeling › diffusion model
diffusion inversion |
0.7 | 1 | 2023 | EDICT: Exact Diffusion Inversion via Coupled Transformations · CVPR 2023 |
Machine learning › Generative modeling
latent space optimization |
0.7 | 1 | 2023 | End-to-End Diffusion Latent Optimization Improves Classifier Guidance · ICCV 2023 |
Visual content generation and editing
image editing |
0.7 | 1 | 2023 | EDICT: Exact Diffusion Inversion via Coupled Transformations · CVPR 2023 |
Visual content generation and editing › image editing
real image editing |
0.7 | 1 | 2023 | EDICT: Exact Diffusion Inversion via Coupled Transformations · CVPR 2023 |
Machine learning › Learning paradigms › continual learning
catastrophic forgetting |
0.5 | 1 | 2021 | Remembering for the Right Reasons: Explanations Reduce Catastrophic Forgetting · ICLR 2021 |
Machine learning › Learning paradigms
continual learning |
0.5 | 1 | 2021 | Remembering for the Right Reasons: Explanations Reduce Catastrophic Forgetting · ICLR 2021 |
Machine learning › Trustworthy machine learning
interpretability |
0.5 | 1 | 2021 | Remembering for the Right Reasons: Explanations Reduce Catastrophic Forgetting · ICLR 2021 |
Natural language and speech › Language models and text generation
text generation |
0.3 | 1 | 2025 | OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows · CVPR 2025 |
Audio and music processing
sound synthesis |
0.3 | 1 | 2025 | OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows · CVPR 2025 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.2 | 1 | 2023 | End-to-End Diffusion Latent Optimization Improves Classifier Guidance · ICCV 2023 |
Methods — techniques the papers use, named apart from their topics
multimodal transformer · 1.7classifier-free guidance · 1.7affine coupling layer · 1.3timestep shifting · 0.9prefix KV cache · 0.9inference-time scaling · 0.9in-context reflection · 0.9complementary masking · 0.9human utility maximization · 0.8clustering of sub-modules · 0.8chain-of-thought prompting · 0.8binary feedback · 0.8denoising diffusion implicit model · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | OmniFlow: Any-to-Any Generation with Multi-Modal Rectified FlowsabstractWe introduce OmniFlow, a novel generative model designed for any-to-any generation tasks such as text-to-image, text-to-audio, and audio-to-image synthesis. OmniFlow advances the rectified flow (RF) framework used in text-to-image models to handle the joint distribution of multiple modalities. It outperforms previous any-to-any models on a wide range of tasks, such as text-to-image and text-to-audio synthesis. Our work offers three key contributions: First, we extend RF to a multi-modal setting and introduce a novel guidance mechanism, enabling users to flexibly control the alignment between different modalities in the generated outputs. Second, we propose a novel architecture that extends the text-to-image MMDiT architecture of Stable Diffusion 3 and enables audio and text generation. The extended modules can be efficiently pretrained individually and merged with the vanilla text-to-image MMDiT for fine-tuning. Lastly, we conduct a comprehensive study of the design choices of rectified flow transformers for large-scale audio and text generation, providing valuable insights into optimizing performance across various modalities. Code is available at https://github.com/jacklishufan/OmniFlows. Konstantinos Kallidromitis, Akash Gokul, Zichun Liao, Yusuke Kato, Kazuki Kozuka, Aditya Grover |
CVPR | 3 |
| 2025 | Reflect-DiT: Inference-Time Scaling for Text-to-Image Diffusion Transformers via In-Context Reflection
Konstantinos Kallidromitis, Akash Gokul, Arsh Koneru, Yusuke Kato, Kazuki Kozuka, Aditya Grover |
ICCV | 3 |
| 2025 | LaViDa: A Large Diffusion Language Model for Multimodal UnderstandingabstractModern Vision-Language Models (VLMs) can solve a wide range of tasks requiring visual reasoning. In real-world scenarios, desirable properties for VLMs include fast inference and controllable generation (e.g., constraining outputs to adhere to a desired format). However, existing autoregressive (AR) VLMs like LLaVA struggle in these aspects. Discrete diffusion models (DMs) offer a promising alternative, enabling parallel decoding for faster inference and bidirectional context for controllable generation through text-infilling. While effective in language-only settings, DMs' potential for multimodal tasks is underexplored. We introduce LaViDa, a family of VLMs built on DMs. We build LaViDa by equipping DMs with a vision encoder and jointly fine-tune the combined parts for multimodal instruction following. To address challenges encountered, LaViDa incorporates novel techniques such as complementary masking for effective training, prefix KV cache for efficient inference, and timestep shifting for high-quality sampling. Experiments show that LaViDa achieves competitive or superior performance to AR VLMs on multi-modal benchmarks such as MMMU, while offering unique advantages of DMs, including flexible speed-quality tradeoff, controllability, and bidirectional reasoning. On COCO captioning, LaViDa surpasses Open-LLaVa-Next-8B by +4.1 CIDEr with 1.92x speedup. On bidirectional tasks, it achieves +59% improvement on Constrained Poem Completion. These results demonstrate LaViDa as a strong alternative to AR VLMs. Code and models is available at https://github.com/jacklishufan/LaViDa Konstantinos Kallidromitis, Hritik Bansal, Akash Gokul, Yusuke Kato, Kazuki Kozuka, Jason Kuen, Kai-Wei Chang 0001, Aditya Grover |
NeurIPS | 4 |
| 2024 | CodeChain: Towards Modular Code Generation Through Chain of Self-revisions with Representative Sub-modulesabstractLarge Language Models (LLMs) have already become quite proficient at solving simpler programming tasks like those in HumanEval or MBPP benchmarks. However, solving more complex and competitive programming tasks is still quite challenging for these models - possibly due to their tendency to generate solutions as monolithic code blocks instead of decomposing them into logical sub-tasks and sub-modules. On the other hand, experienced programmers instinctively write modularized code with abstraction for solving complex tasks, often reusing previously developed modules. To address this gap, we propose CodeChain, a novel framework for inference that elicits modularized code generation through a chain of self-revisions, each being guided by some representative sub-modules generated in previous iterations. Concretely, CodeChain first instructs the LLM to generate modularized codes through chain-of-thought prompting. Then it applies a chain of self-revisions by iterating the two steps: 1) extracting and clustering the generated sub-modules and selecting the cluster representatives as the more generic and re-usable implementations, and 2) augmenting the original chain-of-thought prompt with these selected module-implementations and instructing the LLM to re-generate new modularized solutions. We find that by naturally encouraging the LLM to reuse the previously developed and verified sub-modules, CodeChain can significantly boost both modularity as well as correctness of the generated solutions, achieving relative pass@1 improvements of 35\% on APPS and 76\% on CodeContests. It is shown to be effective on both OpenAI LLMs as well as open-sourced LLMs like WizardCoder. We also conduct comprehensive ablation studies with different methods of prompting, number of clusters, model sizes, program qualities, etc., to provide useful insights that underpin CodeChain's success. Hung Le 0003, Hailin Chen, Amrita Saha, Akash Gokul, Doyen Sahoo, Shafiq R. Joty |
ICLR | 4 |
| 2024 | Aligning Diffusion Models by Optimizing Human UtilityabstractWe present Diffusion-KTO, a novel approach for aligning text-to-image diffusion models by formulating the alignment objective as the maximization of expected human utility. Unlike previous methods, Diffusion-KTO does not require collecting pairwise preference data nor training a complex reward model. Instead, our objective uses per-image binary feedback signals, e.g. likes or dislikes, to align the model with human preferences. After fine-tuning using Diffusion-KTO, text-to-image diffusion models exhibit improved performance compared to existing techniques, including supervised fine-tuning and Diffusion-DPO, both in terms of human judgment and automatic evaluation metrics such as PickScore and ImageReward. Overall, Diffusion-KTO unlocks the potential of leveraging readily available per-image binary preference signals and broadens the applicability of aligning text-to-image diffusion models with human preferences. Konstantinos Kallidromitis, Akash Gokul, Yusuke Kato, Kazuki Kozuka |
NeurIPS | 3 |
| 2023 | EDICT: Exact Diffusion Inversion via Coupled TransformationsabstractFinding an initial noise vector that produces an input image when fed into the diffusion process (known as inversion) is an important problem in denoising diffusion models (DDMs), with applications for real image editing. The standard approach for real image editing with inversion uses denoising diffusion implicit models (DDIMs [29]) to deterministically noise the image to the intermediate state along the path that the denoising would follow given the original conditioning. However, DDIM inversion for real images is unstable as it relies on local linearization assumptions, which result in the propagation of errors, leading to incorrect image reconstruction and loss of content. To alleviate these problems, we propose Exact Diffusion Inversion via Coupled Transformations (EDICT), an inversion method that draws inspiration from affine coupling layers. EDICT enables mathematically exact inversion of real and model-generated images by maintaining two coupled noise vectors which are used to invert each other in an alternating fashion. Using Stable Diffusion [25], a state-of-the-art latent diffusion model, we demonstrate that EDICT successfully reconstructs real images with high fidelity. On complex image datasets like MS-COCO, EDICT reconstruction significantly outperforms DDIM, improving the mean square error of reconstruction by a factor of two. Using noise vectors inverted from real images, EDICT enables a wide range of image edits—from local and global semantic edits to image stylization—while maintaining fidelity to the original image structure. EDICT requires no model training/finetuning, prompt tuning, or extra data and can be combined with any pretrained DDM. Bram Wallace, Akash Gokul |
CVPR | 2 |
| 2023 | End-to-End Diffusion Latent Optimization Improves Classifier GuidanceabstractClassifier guidance—using the gradients of an image classifier to steer the generations of a diffusion model—has the potential to dramatically expand the creative control over image generation and editing. However, currently classifier guidance requires either training new noise-aware models to obtain accurate gradients or using a one-step denoising approximation of the final generation, which leads to misaligned gradients and sub-optimal control. We highlight this approximation’s shortcomings and propose a novel guidance method: Direct Optimization of Diffusion Latents (DOODL), which enables plug-and-play guidance by optimizing diffusion latents w.r.t. the gradients of a pre-trained classifier on the true generated pixels, using an invertible diffusion process to achieve memory-efficient backpropagation. Showcasing the potential of more precise guidance, DOODL outperforms one-step classifier guidance on computational and human evaluation metrics across different forms of guidance: using CLIP guidance to improve generations of complex prompts from DrawBench, using fine-grained visual classifiers to expand the vocabulary of Stable Diffusion, enabling image-conditioned generation with a CLIP visual encoder, and improving image aesthetics using an aesthetic scoring network. Bram Wallace, Akash Gokul, Stefano Ermon |
ICCV | 2 |
| 2021 | Remembering for the Right Reasons: Explanations Reduce Catastrophic Forgetting
Sayna Ebrahimi, Suzanne Petryk, Akash Gokul, William Gan, Joseph Gonzalez 0001, Marcus Rohrbach, Trevor Darrell |
ICLR | 3 |