EDBT 2026 Demo / reviewers in the wild / expert
Renjith Prasad Kaippilly Mana
dblp:402/3423
· DBLP profile ↗
2ranked-venue papers
2as first author
2since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Vision and language · 33% Generative modeling · 17% Language models and text generation · 17% | |
| Computer graphics and multimedia
1 paper |
Multimedia analysis and retrieval · 100% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
diffusion model |
1.0 | 1 | 2026 | DETONATE - A Benchmark for Text-to-Image Alignment and Kernelized Direct Preference Optimization · AAAI 2026 |
Natural language and speech › Language models and text generation › preference optimization
direct preference optimization |
1.0 | 1 | 2026 | DETONATE - A Benchmark for Text-to-Image Alignment and Kernelized Direct Preference Optimization · AAAI 2026 |
Machine learning › Trustworthy machine learning
fairness |
1.0 | 1 | 2026 | DETONATE - A Benchmark for Text-to-Image Alignment and Kernelized Direct Preference Optimization · AAAI 2026 |
Computer vision › Vision and language › cross-modal alignment › image-text alignment
text-to-image alignment |
1.0 | 1 | 2026 | DETONATE - A Benchmark for Text-to-Image Alignment and Kernelized Direct Preference Optimization · AAAI 2026 |
Natural language and speech › Question answering and dialogue systems
multimodal dialogue system |
0.9 | 1 | 2025 | Pic2Prep: A Multimodal Conversational Agent for Cooking Assistance · AAAI 2025 |
Computer vision › Vision and language › vision-language generation
recipe generation from images |
0.9 | 1 | 2025 | Pic2Prep: A Multimodal Conversational Agent for Cooking Assistance · AAAI 2025 |
Multimedia analysis and retrieval
food computing |
0.9 | 1 | 2025 | Pic2Prep: A Multimodal Conversational Agent for Cooking Assistance · AAAI 2025 |
Methods — techniques the papers use, named apart from their topics
stable diffusion · 1.7mistral · 1.7cookgen · 1.7BLIP · 1.7kernel methods · 1.0heavy-tailed self-regularization · 1.0divergence measure · 1.0direct preference optimization · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DETONATE - A Benchmark for Text-to-Image Alignment and Kernelized Direct Preference OptimizationabstractAlignment is crucial for text-to-image (T2I) models to ensure that the generated images faithfully capture user intent while maintaining safety and fairness. Direct Preference Optimization (DPO) has emerged as a key alignment technique for large language models (LLMs), and its influence is now extending to T2I systems. This paper introduces DPO-Kernels for T2I models, a novel extension of DPO that enhances alignment across three key dimensions: (i) Hybrid Loss, which integrates embedding-based objectives with the traditional probability-based loss to improve optimization; (ii) Kernelized Representations, leveraging Radial Basis Function (RBF), Polynomial, and Wavelet kernels to enable richer feature transformations, ensuring better separation between safe and unsafe inputs; and (iii) Divergence Selection, expanding beyond DPO’s default Kullback–Leibler (KL) regularizer by incorporating alternative divergence measures such as Wasserstein and Rényi divergences to enhance stability and robustness in alignment training. We introduce DETONATE, the first large-scale benchmark of its kind, comprising approximately 100K curated image pairs, categorized as chosen and rejected. This benchmark encapsulates three critical axes of social bias and discrimination: Race, Gender, and Disability. The prompts are sourced from the hate speech datasets, while the images are generated using state-of-the-art T2I models, including Stable Diffusion 3.5 Large (SD-3.5), Stable Diffusion XL (SD-XL), and Midjourney. Furthermore, to evaluate alignment beyond surface metrics, we introduce the Alignment Quality Index (AQI) for T2I systems: a novel geometric measure that quantifies latent space separability of safe/unsafe image activations, revealing hidden model vulnerabilities. While alignment techniques often risk overfitting, we empirically demonstrate that DPO-Kernels preserve strong generalization bounds using the theory of Heavy-Tailed Self-Regularization (HT-SR). Renjith Prasad Kaippilly Mana, Abhilekh Borah, Hasnat Md Abdullah, Chathurangi Shyalika, Ritvik Garimella, Rajarshi Roy 0007, Harshul Raj Surana, Nasrin Imanpour, Suranjana Trivedy, Amit P. Sheth, Amitava Das 0001 |
AAAI | 1 |
| 2025 | Pic2Prep: A Multimodal Conversational Agent for Cooking AssistanceabstractAs the demand for healthier, personalized culinary experiences grows, so does the need for advanced food computation models that offer more than basic nutritional insights. However, current food computation models lack the depth to provide actionable insights like ingredient substitution or alternative cooking actions to suit users’ dietary goals. To address this, we introduce and demonstrate Pic2Prep, a multimodal conversational system that generates detailed cooking instructions, actions and ingredient lists from both images and text provided by users. The system is developed using a novel dataset generated through Stable Diffusion, where the input consists of recipe titles and ingredient lists from the Recipe1M dataset to create synthesized food images with variations. This dataset is used to fine-tune the Bootstrapping Language-Image Pre-training (BLIP) model to extract cooking instructions and ingredients from food images. Pic2Prep also employs the CookGen model, a small-scale custom generative model to derive specific cooking actions from cooking instructions. A custom mapper, trained on the Mistral model, links these actions to the corresponding ingredients, creating a comprehensive understanding of the cooking process. The system features an interactive user interface that allows users to input images and ask targeted questions, receiving real-time responses. Renjith Prasad Kaippilly Mana, Chathurangi Shyalika, Revathy Venkataramanan, Darssan Eswaramoorthi, Amit P. Sheth |
AAAI | 1 |