EDBT 2026 Demo / reviewers in the wild / expert
Amrutha Saseendran
dblp:289/0537
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Language models and text generation · 25% Generative modeling · 23% Vision and language · 22% |
Topics — the 23 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
autoencoder |
1.1 | 2 | 2022 | Trading off Image Quality for Robustness is not Necessary with Regularized Deterministic Autoencoders · NeurIPS 2022 Shape your Space: A Gaussian Mixture Regularization Approach to Deterministic Autoencoders · NeurIPS 2021 |
Machine learning › Deep learning architectures and training › autoencoder
deterministic autoencoder |
1.1 | 2 | 2022 | Trading off Image Quality for Robustness is not Necessary with Regularized Deterministic Autoencoders · NeurIPS 2022 Shape your Space: A Gaussian Mixture Regularization Approach to Deterministic Autoencoders · NeurIPS 2021 |
Machine learning › Transfer learning and domain adaptation
cross-modal transfer |
0.9 | 1 | 2025 | Diffusion Instruction Tuning · ICML 2025 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | Diffusion Instruction Tuning · ICML 2025 |
Natural language and speech › Language models and text generation › large language model inference
inference-time techniques |
0.9 | 1 | 2025 | Balancing Act: Diversity and Consistency in Large Language Model Ensembles · ICLR 2025 |
Natural language and speech › Language models and text generation
instruction tuning |
0.9 | 1 | 2025 | Diffusion Instruction Tuning · ICML 2025 |
Computer vision › Vision and language › visual grounding
language-guided segmentation |
0.9 | 1 | 2025 | Segment Anyword: Mask Prompt Inversion for Open-Set Grounded Segmentation · ICML 2025 |
Natural language and speech › Language models and text generation › large language model
large language model ensemble |
0.9 | 1 | 2025 | Balancing Act: Diversity and Consistency in Large Language Model Ensembles · ICLR 2025 |
Natural language and speech › Language models and text generation › LLM agents › LLM collaboration
mixture of agents |
0.9 | 1 | 2025 | Balancing Act: Diversity and Consistency in Large Language Model Ensembles · ICLR 2025 |
Computer vision › Segmentation and scene understanding › open-world segmentation
open-set segmentation |
0.9 | 1 | 2025 | Segment Anyword: Mask Prompt Inversion for Open-Set Grounded Segmentation · ICML 2025 |
Natural language and speech › Language models and text generation › decoding
self-consistency decoding |
0.9 | 1 | 2025 | Balancing Act: Diversity and Consistency in Large Language Model Ensembles · ICLR 2025 |
Computer vision › Vision and language
vision-language model |
0.9 | 1 | 2025 | Diffusion Instruction Tuning · ICML 2025 |
Computer vision › Vision and language › vision-language model
prompt learning |
0.8 | 1 | 2024 | An Image is Worth Multiple Words: Discovering Object Level Concepts using Multi-Concept Prompt Learning · ICML 2024 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.8 | 1 | 2024 | An Image is Worth Multiple Words: Discovering Object Level Concepts using Multi-Concept Prompt Learning · ICML 2024 |
Machine learning › Generative modeling
variational autoencoder |
0.7 | 2 | 2022 | Shape your Space: A Gaussian Mixture Regularization Approach to Deterministic Autoencoders · NeurIPS 2021 Trading off Image Quality for Robustness is not Necessary with Regularized Deterministic Autoencoders · NeurIPS 2022 |
Machine learning › Trustworthy machine learning › robustness
adversarial robustness |
0.6 | 1 | 2022 | Trading off Image Quality for Robustness is not Necessary with Regularized Deterministic Autoencoders · NeurIPS 2022 |
Machine learning › Trustworthy machine learning
robustness |
0.6 | 1 | 2022 | Trading off Image Quality for Robustness is not Necessary with Regularized Deterministic Autoencoders · NeurIPS 2022 |
Machine learning › Generative modeling › image generation
conditional image generation |
0.5 | 1 | 2021 | Multi-Class Multi-Instance Count Conditioned Adversarial Image Generation · ICCV 2021 |
Machine learning › Generative modeling › diffusion model
controllable generation |
0.5 | 1 | 2021 | Multi-Class Multi-Instance Count Conditioned Adversarial Image Generation · ICCV 2021 |
Machine learning › Generative modeling
generative adversarial network |
0.5 | 1 | 2021 | Multi-Class Multi-Instance Count Conditioned Adversarial Image Generation · ICCV 2021 |
Machine learning › Generative modeling
latent space regularization |
0.5 | 1 | 2021 | Shape your Space: A Gaussian Mixture Regularization Approach to Deterministic Autoencoders · NeurIPS 2021 |
Machine learning › Trustworthy machine learning › adversarial machine learning
adversarial vulnerability |
0.2 | 1 | 2022 | Trading off Image Quality for Robustness is not Necessary with Regularized Deterministic Autoencoders · NeurIPS 2022 |
Computer vision › Image recognition and object detection
object counting |
0.1 | 1 | 2021 | Multi-Class Multi-Instance Count Conditioned Adversarial Image Generation · ICCV 2021 |
Methods — techniques the papers use, named apart from their topics
adversarial training · 1.1supervised fine-tuning · 0.9stable diffusion · 0.9prompt regularization · 0.9mixture refinement · 0.9ensemble gating · 0.9diffusion model · 0.9cross-validation · 0.9cross-attention · 0.9attention alignment · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Balancing Act: Diversity and Consistency in Large Language Model EnsemblesabstractEnsembling strategies for Large Language Models (LLMs) have demonstrated significant potential in improving performance across various tasks by combining the strengths of individual models. However, identifying the most effective ensembling method remains an open challenge, as neither maximizing output consistency through self-consistency decoding nor enhancing model diversity via frameworks like "Mixture of Agents" has proven universally optimal. Motivated by this, we propose a unified framework to examine the trade-offs between task performance, model diversity, and output consistency in ensembles. More specifically, we introduce a consistency score that defines a gating mechanism for mixtures of agents and an algorithm for mixture refinement to investigate these trade-offs at the semantic and model levels, respectively. We incorporate our insights into a novel inference-time LLM ensembling strategy called the Dynamic Mixture of Agents (DMoA) and demonstrate that it achieves a new state-of-the-art result in the challenging Big Bench Hard mixed evaluations benchmark. Our analysis reveals that cross-validation bias can enhance performance, contingent on the expertise of the constituent models. We further demonstrate that distinct reasoning tasks—such as arithmetic reasoning, commonsense reasoning, and instruction following—require different model capabilities, leading to inherent task-dependent trade-offs that DMoA balances effectively. Ahmed Abdulaal, Nina Montaña Brown, Aryo Pradipta Gema, Daniel C. Castro, Daniel C. Alexander, Philip Teare, Tom Diethe, Dino Oglic, Amrutha Saseendran |
ICLR | 10 |
| 2025 | Diffusion Instruction TuningabstractWe introduce Lavender, a simple supervised fine-tuning (SFT) method that boosts the performance of advanced vision-language models (VLMs) by leveraging state-of-the-art image generation models such as Stable Diffusion. Specifically, Lavender aligns the text-vision attention in the VLM transformer with the equivalent used by Stable Diffusion during SFT, instead of adapting separate encoders. This alignment enriches the model’s visual understanding and significantly boosts performance across in- and out-of-distribution tasks. Lavender requires just 0.13 million training examples—2.5% of typical large-scale SFT datasets—and fine-tunes on standard hardware (8 GPUs) in a single day. It consistently improves state-of-the-art open-source multimodal LLMs (e.g., Llama-3.2-11B, MiniCPM-Llama3-v2.5), achieving up to 30% gains and a 68% boost on challenging out-of-distribution medical QA tasks. By efficiently transferring the visual expertise of image generators with minimal supervision, Lavender offers a scalable solution for more accurate vision-language systems. Code, training data, and models are available on the project page. Ryutaro Tanno, Amrutha Saseendran, Tom Diethe, Philip Teare |
ICML | 3 |
| 2025 | Segment Anyword: Mask Prompt Inversion for Open-Set Grounded SegmentationabstractOpen-set image segmentation poses a significant challenge because existing methods often demand extensive training or fine-tuning and generally struggle to segment unified objects consistently across diverse text reference expressions. Motivated by this, we propose Segment Anyword, a novel training-free visual concept prompt learning approach for open-set language grounded segmentation that relies on token-level cross-attention maps from a frozen diffusion model to produce segmentation surrogates or *mask prompts*, which are then refined into targeted object masks. Initial prompts typically lack coherence and consistency as the complexity of the image-text increases, resulting in suboptimal mask fragments. To tackle this issue, we further introduce a novel linguistic-guided visual prompt regularization that binds and clusters visual prompts based on sentence dependency and syntactic structural information, enabling the extraction of robust, noise-tolerant mask prompts, and significant improvements in segmentation accuracy. The proposed approach is effective, generalizes across different open-set segmentation tasks, and achieves state-of-the-art results of 52.5 (+6.8 relative) mIoU on Pascal Context 59, 67.73 (+25.73 relative) cIoU on gRefCOCO, and 67.4 (+1.1 relative to fine-tuned methods) mIoU on GranDf, which is the most complex open-set grounded segmentation task in the field. Amrutha Saseendran, Xilin He, Fariba Yousefi, Nikolay Burlutskiy, Dino Oglic, Tom Diethe, Philip Teare, Huiyu Zhou 0001 |
ICML | 2 |
| 2024 | An Image is Worth Multiple Words: Discovering Object Level Concepts using Multi-Concept Prompt LearningabstractTextural Inversion, a prompt learning method, learns a singular text embedding for a new "word" to represent image style and appearance, allowing it to be integrated into natural language sentences to generate novel synthesised images. However, identifying multiple unknown object-level concepts within one scene remains a complex challenge. While recent methods have resorted to cropping or masking individual images to learn multiple concepts, these techniques often require prior knowledge of new concepts and are labour-intensive. To address this challenge, we introduce *Multi-Concept Prompt Learning (MCPL)*, where multiple unknown "words" are simultaneously learned from a single sentence-image pair, without any imagery annotations. To enhance the accuracy of word-concept correlation and refine attention mask boundaries, we propose three regularisation techniques: *Attention Masking*, *Prompts Contrastive Loss*, and *Bind Adjective*. Extensive quantitative comparisons with both real-world categories and biomedical images demonstrate that our method can learn new semantically disentangled concepts. Our approach emphasises learning solely from textual embeddings, using less than 10% of the storage space compared to others. The project page, code, and data are available at [https://astrazeneca.github.io/mcpl.github.io](https://astrazeneca.github.io/mcpl.github.io). Ryutaro Tanno, Amrutha Saseendran, Tom Diethe, Philip Teare |
ICML | 3 |
| 2022 | Trading off Image Quality for Robustness is not Necessary with Regularized Deterministic AutoencodersabstractThe susceptibility of Variational Autoencoders (VAEs) to adversarial attacks indicates the necessity to evaluate the robustness of the learned representations along with the generation performance. The vulnerability of VAEs has been attributed to the limitations associated with their variational formulation. Deterministic autoencoders could overcome the practical limitations associated with VAEs and offer a promising alternative for image generation applications. In this work, we propose an adversarially robust deterministic autoencoder with superior performance in terms of both generation and robustness of the learned representations. We introduce a regularization scheme to incorporate adversarially perturbed data points to the training pipeline without increasing the computational complexity or compromising the generation fidelity by leveraging a loss based on the two-point Kolmogorov–Smirnov test between representations. We conduct extensive experimental studies on popular image benchmark datasets to quantify the robustness of the proposed approach based on the adversarial attacks targeted at VAEs. Our empirical findings show that the proposed method achieves significant performance in both robustness and fidelity when compared to the robust VAE models. Amrutha Saseendran, Kathrin Skubch, Stefan Falkner, Margret Keuper |
NeurIPS | 1 |
| 2021 | Multi-Class Multi-Instance Count Conditioned Adversarial Image GenerationabstractImage generation has rapidly evolved in recent years. Modern architectures for adversarial training allow to generate even high resolution images with remarkable quality. At the same time, more and more effort is dedicated towards controlling the content of generated images. In this paper, we take one further step in this direction and propose a conditional generative adversarial network (GAN) that generates images with a defined number of objects from given classes. This entails two fundamental abilities (1) being able to generate high-quality images given a complex constraint and (2) being able to count object instances per class in a given image. Our proposed model modularly extends the successful StyleGAN2 architecture with a count-based conditioning as well as with a regression subnetwork to count the number of generated objects per class during training. In experiments on three different datasets, we show that the proposed model learns to generate images according to the given multiple-class count condition even in the presence of complex backgrounds. In particular, we propose a new dataset, CityCount, which is derived from the Cityscapes street scenes dataset, to evaluate our approach in a challenging and practically relevant scenario. An implementation is available at https://github.com/boschresearch/MCCGAN. Amrutha Saseendran, Kathrin Skubch, Margret Keuper |
ICCV | 1 |
| 2021 | Shape your Space: A Gaussian Mixture Regularization Approach to Deterministic AutoencodersabstractVariational Autoencoders (VAEs) are powerful probabilistic models to learn representations of complex data distributions. One important limitation of VAEs is the strong prior assumption that latent representations learned by the model follow a simple uni-modal Gaussian distribution. Further, the variational training procedure poses considerable practical challenges. Recently proposed regularized autoencoders offer a deterministic autoencoding framework, that simplifies the original VAE objective and is significantly easier to train. Since these models only provide weak control over the learned latent distribution, they require an ex-post density estimation step to generate samples comparable to those of VAEs. In this paper, we propose a simple and end-to-end trainable deterministic autoencoding framework, that efficiently shapes the latent space of the model during training and utilizes the capacity of expressive multi-modal latent distributions. The proposed training procedure provides direct evidence if the latent distribution adequately captures complex aspects of the encoded data. We show in experiments the expressiveness and sample quality of our model in various challenging continuous and discrete domains. An implementation is available at https://github.com/boschresearch/GMM_DAE. Amrutha Saseendran, Kathrin Skubch, Stefan Falkner, Margret Keuper |
NeurIPS | 1 |