VLDB 2026 Research / reviewers in the wild / expert
Yonatan Dukler
dblp:242/3844
· DBLP profile ↗
8ranked-venue papers
5as first author
6since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Deep learning architectures and training · 31% Trustworthy machine learning · 16% Vision and language · 11% | |
| Computer graphics and multimedia
1 paper |
Image and video processing · 100% |
Topics — the 22 heaviest of 25, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training › transformer
vision transformer |
0.9 | 2 | 2023 | Learning Expressive Prompting With Residuals for Vision Transformers · CVPR 2023 Your representations are in the network: composable and parallel adaptation for large scale models · NeurIPS 2023 |
Machine learning › Trustworthy machine learning › hallucination
hallucination evaluation |
0.8 | 1 | 2024 | THRONE: An Object-Based Hallucination Benchmark for the Free-Form Generations of Large Vision-Language Models · CVPR 2024 |
Machine learning › Deep learning architectures and training
memory mechanism |
0.8 | 1 | 2024 | B'MOJO: Hybrid State Space Realizations of Foundation Models with Eidetic and Fading Memory · NeurIPS 2024 |
Machine learning › Deep learning architectures and training
sequence modeling |
0.8 | 1 | 2024 | B'MOJO: Hybrid State Space Realizations of Foundation Models with Eidetic and Fading Memory · NeurIPS 2024 |
Computer vision › Vision and language › vision-language model
vision-language model evaluation |
0.8 | 1 | 2024 | THRONE: An Object-Based Hallucination Benchmark for the Free-Form Generations of Large Vision-Language Models · CVPR 2024 |
Machine learning › Trustworthy machine learning › hallucination
vision-language model hallucination |
0.8 | 1 | 2024 | THRONE: An Object-Based Hallucination Benchmark for the Free-Form Generations of Large Vision-Language Models · CVPR 2024 |
Machine learning › Kernel, tree and ensemble methods
ensemble learning |
0.7 | 1 | 2023 | SAFE: Machine Unlearning With Shard Graphs · ICCV 2023 |
Machine learning › Trustworthy machine learning
machine unlearning |
0.7 | 1 | 2023 | SAFE: Machine Unlearning With Shard Graphs · ICCV 2023 |
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning |
0.7 | 1 | 2023 | Your representations are in the network: composable and parallel adaptation for large scale models · NeurIPS 2023 |
Computer vision › Vision and language › vision-language model
prompt learning |
0.7 | 1 | 2023 | Learning Expressive Prompting With Residuals for Vision Transformers · CVPR 2023 |
Machine learning › Transfer learning and domain adaptation › parameter-efficient transfer learning
visual prompt tuning |
0.7 | 1 | 2023 | Learning Expressive Prompting With Residuals for Vision Transformers · CVPR 2023 |
Machine learning › Optimization for machine learning › convergence analysis
convergence analysis of neural network training |
0.4 | 1 | 2020 | Optimization Theory for ReLU Neural Networks Trained with Normalization Layers · ICML 2020 |
Machine learning › Deep learning architectures and training
loss landscape |
0.4 | 1 | 2020 | Optimization Theory for ReLU Neural Networks Trained with Normalization Layers · ICML 2020 |
Machine learning › Optimization for machine learning
non-convex optimization |
0.4 | 1 | 2020 | Optimization Theory for ReLU Neural Networks Trained with Normalization Layers · ICML 2020 |
Machine learning › Deep learning architectures and training › normalization
normalization layers |
0.4 | 1 | 2020 | Optimization Theory for ReLU Neural Networks Trained with Normalization Layers · ICML 2020 |
Machine learning › Deep learning architectures and training › normalization
weight normalization |
0.4 | 1 | 2020 | Optimization Theory for ReLU Neural Networks Trained with Normalization Layers · ICML 2020 |
Machine learning › Generative modeling
generative adversarial network |
0.4 | 1 | 2019 | Wasserstein of Wasserstein Loss for Learning Generative Models · ICML 2019 |
Machine learning › Deep learning architectures and training › regularization › gradient regularization
gradient penalty |
0.4 | 1 | 2019 | Wasserstein of Wasserstein Loss for Learning Generative Models · ICML 2019 |
Machine learning › Generative modeling › generative adversarial network
Wasserstein GAN |
0.4 | 1 | 2019 | Wasserstein of Wasserstein Loss for Learning Generative Models · ICML 2019 |
Natural language and speech › Language models and text generation
language modeling |
0.2 | 1 | 2024 | B'MOJO: Hybrid State Space Realizations of Foundation Models with Eidetic and Fading Memory · NeurIPS 2024 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.2 | 1 | 2024 | THRONE: An Object-Based Hallucination Benchmark for the Free-Form Generations of Large Vision-Language Models · CVPR 2024 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.2 | 1 | 2023 | Learning Expressive Prompting With Residuals for Vision Transformers · CVPR 2023 |
Methods — techniques the papers use, named apart from their topics
stochastic realization theory · 0.8state space model · 0.8pseudo-labeling · 0.8data augmentation · 0.8shard graph · 0.7residual tokens · 0.7output tokens · 0.7ensemble · 0.7cross-attention modules · 0.7adapter · 0.7wasserstein distance · 0.4riemannian gradient penalty · 0.4convolution · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | THRONE: An Object-Based Hallucination Benchmark for the Free-Form Generations of Large Vision-Language ModelsabstractMitigating hallucinations in large vision-language models (LVLMs) remains an open problem. Recent benchmarks do not address hallucinations in open-ended free-form responses, which we term “Type I hallucinations”. Instead, they focus on hallucinations responding to very specific question formats-typically a multiple-choice response regarding a particular object or attribute-which we term “Type II hallucinations”. Additionally, such benchmarks often require external API calls to models which are subject to change. In practice, we observe that a reduction in Type II hallucinations does not lead to a reduction in Type I hallucinations but rather that the two forms of halluci-nations are often anti-correlated. To address this, we propose THRONE, a novel object-based automatic framework for quantitatively evaluating Type I hallucinations in LVLM free-form outputs. We use public language models (LMs) to identify hallucinations in LVLM responses and compute informative metrics. By evaluating a large selection of recent LVLMs using public datasets, we show that an improvement in existing metrics do not lead to a reduction in Type I hallucinations, and that established benchmarks for measuring Type I hallucinations are incomplete. Finally, we provide a simple and effective data augmentation method to reduce Type I and Type II hallucinations as a strong baseline. Prannay Kaul, Zhizhong Li 0001, Hao Yang 0043, Yonatan Dukler, Ashwin Swaminathan, C. J. Taylor, Stefano Soatto |
CVPR | 4 |
| 2024 | B'MOJO: Hybrid State Space Realizations of Foundation Models with Eidetic and Fading MemoryabstractWe describe a family of architectures to support transductive inference by allowing memory to grow to a finite but a-priori unknown bound while making efficient use of finite resources for inference. Current architectures use such resources to represent data either eidetically over a finite span ('context' in Transformers), or fading over an infinite span (in State Space Models, or SSMs). Recent hybrid architectures have combined eidetic and fading memory, but with limitations that do not allow the designer or the learning process to seamlessly modulate the two, nor to extend the eidetic memory span. We leverage ideas from Stochastic Realization Theory to develop a class of models called B'MOJO to seamlessly combine eidetic and fading memory within an elementary composable module. The overall architecture can be used to implement models that can access short-term eidetic memory 'in-context,' permanent structural memory 'in-weights,' fading memory 'in-state,' and long-term eidetic memory 'in-storage' by natively incorporating retrieval from an asynchronously updated memory. We show that Transformers, existing SSMs such as Mamba, and hybrid architectures such as Jamba are special cases of B'MOJO and describe a basic implementation that can be stacked and scaled efficiently in hardware. We test B'MOJO on transductive inference tasks, such as associative recall, where it outperforms existing SSMs and Hybrid models; as a baseline, we test ordinary language modeling where B'MOJO achieves perplexity comparable to similarly-sized Transformers and SSMs up to 1.4B parameters, while being up to 10% faster to train. Finally, we test whether models trained inductively on a-priori bounded sequences (up to 8K tokens) can still perform transductive inference on sequences many-fold longer. B'MOJO's ability to modulate eidetic and fading memory results in better inference on longer sequences tested up to 32K tokens, four-fold the length of the longest sequences seen during training. Luca Zancato, Arjun Seshadri, Yonatan Dukler, Aditya Golatkar, Yantao Shen 0002, Benjamin Bowman, Matthew Trager, Alessandro Achille, Stefano Soatto |
NeurIPS | 3 |
| 2023 | Learning Expressive Prompting With Residuals for Vision TransformersabstractPrompt learning is an efficient approach to adapt transformers by inserting learnable set of parameters into the input and intermediate representations of a pre-trained model. In this work, we present Expressive Prompts with Residuals (EXPRES) which modifies the prompt learning paradigm specifically for effective adaptation of vision transformers (ViT). Our method constructs downstream representations via learnable “output” tokens (shal-low prompts), that are akin to the learned class tokens of the ViT. Further for better steering of the downstream representation processed by the frozen transformer, we introduce residual learnable tokens that are added to the output of various computations. We apply EXPRES for image classification and few-shot semantic segmentation, and show our method is capable of achieving state of the art prompt tuning on 3/3 categories of the VTAB benchmark. In addition to strong performance, we observe that our approach is an order of magnitude more prompt efficient than existing visual prompting baselines. We analytically show the computational benefits of our approach over weight space adaptation techniques like finetuning. Lastly we systematically corroborate the architectural design of our method via a series of ablation experiments. Rajshekhar Das, Yonatan Dukler, Avinash Ravichandran, Ashwin Swaminathan |
CVPR | 2 |
| 2023 | SAFE: Machine Unlearning With Shard GraphsabstractWe present Synergy Aware Forgetting Ensemble (SAFE), a method to adapt large models on a diverse collection of data while minimizing the expected cost to remove the influence of training samples from the trained model. This process, also known as selective forgetting or unlearning, is often conducted by partitioning a dataset into shards, training fully independent models on each, then ensembling the resulting models. Increasing the number of shards reduces the expected cost to forget but at the same time it increases inference cost and reduces the final accuracy of the model since synergistic information between samples is lost during the independent model training. Rather than treating each shard as independent, SAFE introduces the notion of a shard graph, which allows incorporating limited information from other shards during training, trading off a modest increase in expected forgetting cost with a significant increase in accuracy, all while still attaining complete removal of residual influence after forgetting. SAFE uses a lightweight system of adapters which can be trained while reusing most of the computations. This allows SAFE to be trained on shards an order-of-magnitude smaller than current state-of-the-art methods (thus reducing the forgetting costs) while also maintaining high accuracy, as we demonstrate empirically on fine-grained computer vision datasets. Yonatan Dukler, Benjamin Bowman, Alessandro Achille, Aditya Golatkar, Ashwin Swaminathan, Stefano Soatto |
ICCV | 1 |
| 2023 | Your representations are in the network: composable and parallel adaptation for large scale modelsabstractWe present a framework for transfer learning that efficiently adapts a large base-model by learning lightweight cross-attention modules attached to its intermediate activations.
We name our approach InCA (Introspective-Cross-Attention) and show that it can efficiently survey a network’s representations and identify strong performing adapter models for a downstream task.
During training, InCA enables training numerous adapters efficiently and in parallel, isolated from the frozen base model. On the ViT-L/16 architecture, our experiments show that a single adapter, 1.3% of the full model, is able to reach full fine-tuning accuracy on average across 11 challenging downstream classification tasks.
Compared with other forms of parameter-efficient adaptation, the isolated nature of the InCA adaptation is computationally desirable for large-scale models. For instance, we adapt ViT-G/14 (1.8B+ parameters) quickly with 20+ adapters in parallel on a single V100 GPU (76% GPU memory reduction) and exhaustively identify its most useful representations.
We further demonstrate how the adapters learned by InCA can be incrementally modified or combined for flexible learning scenarios and our approach achieves state of the art performance on the ImageNet-to-Sketch multi-task benchmark. Yonatan Dukler, Alessandro Achille, Hao Yang 0043, Varsha Vivek, Luca Zancato, Benjamin Bowman, Avinash Ravichandran, Charless C. Fowlkes, Ashwin Swaminathan, Stefano Soatto |
NeurIPS | 1 |
| 2022 | DIVA: Dataset Derivative of a Learning Task
Yonatan Dukler, Alessandro Achille, Giovanni Paolini, Avinash Ravichandran, Marzia Polito, Stefano Soatto |
ICLR | 1 |
| 2020 | Optimization Theory for ReLU Neural Networks Trained with Normalization LayersabstractThe current paradigm of deep neural networks has been successful in part due to the use of normalization layers. Normalization layers like Batch Normalization, Layer Normalization and Weight Normalization are ubiquitous in practice as they improve the generalization performance and training speed of neural networks significantly. Nonetheless, the vast majority of current deep learning theory and non-convex optimization literature focuses on the un-normalized setting. We bridge this gap by providing the first global convergence result for 2 layer non-linear neural networks with ReLU activations trained with a normalization layer, namely Weight Normalization. The analysis shows how the introduction of normalization layers changes the optimization landscape and in some settings enables faster convergence as compared with un-normalized neural networks. Yonatan Dukler, Quanquan Gu, Guido Montúfar |
ICML | 1 |
| 2019 | Wasserstein of Wasserstein Loss for Learning Generative ModelsabstractThe Wasserstein distance serves as a loss function for unsupervised learning which depends on the choice of a ground metric on sample space. We propose to use the Wasserstein distance itself as the ground metric on the sample space of images. This ground metric is known as an effective distance for image retrieval, that correlates with human perception. We derive the Wasserstein ground metric on pixel space and define a Riemannian Wasserstein gradient penalty to be used in the Wasserstein Generative Adversarial Network (WGAN) framework. The new gradient penalty is computed efficiently via convolutions on the $L^2$ gradients with negligible additional computational cost. The new formulation is more robust to the natural variability of the data and provides for a more continuous discriminator in sample space. Yonatan Dukler, Wuchen Li, Alex Tong Lin, Guido Montúfar |
ICML | 1 |