EDBT 2026 Demo / reviewers in the wild / expert
Jayakumar Subramanian
dblp:202/5957
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Reinforcement learning · 28% Generative modeling · 24% Language models and text generation · 18% |
Topics — the 14 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
offline reinforcement learning |
1.4 | 2 | 2025 | Offline RL by Reward-Weighted Fine-Tuning for Conversation Optimization · NeurIPS 2025 Medical Dead-ends and Learning to Identify High-Risk States and Treatments · NeurIPS 2021 |
Natural language and speech › Language models and text generation › text generation › surface realization
linearization |
1.0 | 1 | 2026 | Lizard: An Efficient Linearization Framework for Large Language Models · ACL (1) 2026 |
Natural language and speech › Language models and text generation › language modeling
long-context language modeling |
1.0 | 1 | 2026 | Lizard: An Efficient Linearization Framework for Large Language Models · ACL (1) 2026 |
Machine learning › Generative modeling › diffusion model › diffusion model adaptation
diffusion model alignment |
0.9 | 1 | 2025 | Measuring And Improving Engagement of Text-to-Image Generation Models · ICLR 2025 |
Machine learning › Reinforcement learning › reinforcement learning from human feedback
reward alignment |
0.9 | 1 | 2025 | Measuring And Improving Engagement of Text-to-Image Generation Models · ICLR 2025 |
Machine learning › Generative modeling › diffusion model › diffusion model adaptation
reward fine-tuning |
0.9 | 1 | 2025 | Offline RL by Reward-Weighted Fine-Tuning for Conversation Optimization · NeurIPS 2025 |
Machine learning › Generative modeling › diffusion model
text-to-image generation |
0.9 | 1 | 2025 | Measuring And Improving Engagement of Text-to-Image Generation Models · ICLR 2025 |
Computer vision › Vision and language
vision-language model |
0.9 | 1 | 2025 | Measuring And Improving Engagement of Text-to-Image Generation Models · ICLR 2025 |
Machine learning › Trustworthy machine learning › interpretability
explainable reinforcement learning |
0.7 | 1 | 2023 | Explaining RL Decisions with Trajectories · ICLR 2023 |
Machine learning › Trustworthy machine learning
interpretability |
0.7 | 1 | 2023 | Explaining RL Decisions with Trajectories · ICLR 2023 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning
approximate planning |
0.6 | 1 | 2022 | Approximate Information State for Approximate Planning and Reinforcement Learning in Partially Observed Systems · J. Mach. Learn. Res. 2022 |
Machine learning › Reinforcement learning › policy optimization
policy gradient |
0.6 | 1 | 2022 | Approximate Information State for Approximate Planning and Reinforcement Learning in Partially Observed Systems · J. Mach. Learn. Res. 2022 |
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › search control
dead-end detection |
0.5 | 1 | 2021 | Medical Dead-ends and Learning to Identify High-Risk States and Treatments · NeurIPS 2021 |
Machine learning › Reinforcement learning
dynamic programming |
0.2 | 1 | 2022 | Approximate Information State for Approximate Planning and Reinforcement Learning in Partially Observed Systems · J. Mach. Learn. Res. 2022 |
Methods — techniques the papers use, named apart from their topics
supervised fine-tuning · 1.7linear attention · 1.0reinforcement learning · 0.9direct preference optimization · 0.9trajectory-based explanation · 0.7policy gradient · 0.6information state · 0.6approximate dynamic programming · 0.6deep neural network · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Lizard: An Efficient Linearization Framework for Large Language ModelsabstractChien Van Nguyen, Huy Huu Nguyen, Ruiyi Zhang, Hanieh Deilamsalehy, Puneet Mathur, Viet Dac Lai, Haoliang Wang, Jayakumar Subramanian, Ryan A. Rossi, Trung Bui, Nikos Vlassis, Franck Dernoncourt, Thien Huu Nguyen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Chien Van Nguyen, Huy Huu Nguyen, Ruiyi Zhang 0002, Hanieh Deilamsalehy, Puneet Mathur, Viet Dac Lai, Jayakumar Subramanian, Ryan Rossi, Trung Bui, Nikos Vlassis, Franck Dernoncourt, Thien Huu Nguyen |
ACL (1) | 8 |
| 2025 | AsyncVoice Agent: Real-Time Explanation for LLM Planning and ReasoningabstractEffective human-AI collaboration on complex reasoning tasks requires that users understand and interact with the model’s process, not just receive an output. However, the monolithic text from methods like Chain-of-Thought (CoT) prevents this, as current interfaces lack real-time verbalization and robust user barge-in. We present AsyncVoice Agent, a system whose asynchronous architecture decouples a streaming LLM backend from a conversational voice frontend. This design allows narration and inference to run in parallel, empowering users to interrupt, query, and steer the model’s reasoning process at any time. Objective benchmarks show this approach reduces interaction latency by more than 600 $\times$ compared to monolithic baselines while ensuring high fidelity and competitive task accuracy. By enabling a two-way dialogue with a model’s thought process, AsyncVoice Agent offers a new paradigm for building more effective, steerable, and trustworthy human-AI systems for high-stakes tasks.1 Yueqian Lin, Zhengmian Hu, Jayakumar Subramanian, Qinsi Wang, Nikos Vlassis, Hai Li 0001, Yiran Chen 0001 |
ASRU | 3 |
| 2025 | Measuring And Improving Engagement of Text-to-Image Generation ModelsabstractRecent advances in text-to-image generation have achieved impressive aesthetic quality, making these models usable for both personal and commercial purposes. However, in the fields of marketing and advertising, images are often created to be more engaging, as reflected in user behaviors such as increasing clicks, likes, and purchases, in addition to being aesthetically pleasing. To this end, we introduce the challenge of optimizing the image generation process for improved viewer engagement. In order to study image engagement and utility in real-world marketing scenarios, we collect *EngagingImageNet*, the first large-scale dataset of images, along with associated user engagement metrics. Further, we find that existing image evaluation metrics like aesthetics, CLIPScore, PickScore, ImageReward, *etc.* are unable to capture viewer engagement. To address the lack of reliable metrics for assessing image utility, we use the *EngagingImageNet* dataset to train *EngageNet*, an engagement-aware Vision Language Model (VLM) that predicts viewer engagement of images by leveraging contextual information about the tweet content, enterprise details, and posting time. We then explore methods to enhance the engagement of text-to-image models, making initial strides in this direction. These include conditioning image generation on improved prompts, supervised fine-tuning of stable diffusion on high-performing images, and reinforcement learning to align stable diffusion with *EngageNet*-based reward signals, all of which lead to the generation of images with higher viewer engagement. Finally, we propose the *Engagement Arena*, to benchmark text-to-image models based on their ability to generate engaging images, using *EngageNet* as the evaluator, thereby encouraging the research community to measure further advances in the engagement of text-to-image modeling. These contributions provide a new pathway for advancing utility-driven image generation, with significant implications for the commercial application of image generation. We have released our code and dataset on [behavior-in-the-wild.github.io/image-engagement](https://behavior-in-the-wild.github.io/image-engagement). Varun Khurana, Yaman Singla, Jayakumar Subramanian, Changyou Chen, Rajiv Ratn Shah, Balaji Krishnamurthy |
ICLR | 3 |
| 2025 | Offline RL by Reward-Weighted Fine-Tuning for Conversation OptimizationabstractOffline reinforcement learning (RL) is a variant of RL where the policy is learned from a previously collected dataset of trajectories and rewards. In our work, we propose a practical approach to offline RL with large language models (LLMs). We recast the problem as reward-weighted fine-tuning, which can be solved using similar techniques to supervised fine-tuning (SFT). To showcase the value of our approach, we apply it to learning short-horizon question-answering policies of a fixed length, where the agent reasons about potential answers or asks clarifying questions. Our work stands in a stark contrast to state-of-the-art methods in this domain, based on SFT and direct preference optimization, which have additional hyper-parameters and do not directly optimize for rewards. We compare to them empirically, and report major gains in both optimized rewards and language quality. Subhojyoti Mukherjee, Viet Dac Lai, Raghavendra Addanki, Ryan Rossi, Seunghyun Yoon 0002, Trung Bui, Anup B. Rao, Jayakumar Subramanian, Branislav Kveton |
NeurIPS | 8 |
| 2023 | Explaining RL Decisions with Trajectories
Shripad V. Deshmukh, Arpan Dasgupta, Balaji Krishnamurthy, Chirag Agarwal, Georgios Theocharous, Jayakumar Subramanian |
ICLR | 7 |
| 2022 | Approximate Information State for Approximate Planning and Reinforcement Learning in Partially Observed SystemsabstractWe propose a theoretical framework for approximate planning and learning in partially observed systems. Our framework is based on the fundamental notion of information state. We provide two definitions of information state---i) a function of history which is sufficient to compute the expected reward and predict its next value; ii) a function of the history which can be recursively updated and is sufficient to compute the expected reward and predict the next observation. An information state always leads to a dynamic programming decomposition. Our key result is to show that if a function of the history (called AIS) approximately satisfies the properties of the information state, then there is a corresponding approximate dynamic program. We show that the policy computed using this is approximately optimal with bounded loss of optimality. We show that several approximations in state, observation and action spaces in literature can be viewed as instances of AIS. In some of these cases, we obtain tighter bounds. A salient feature of AIS is that it can be learnt from data. We present AIS based multi-time scale policy gradient algorithms and detailed numerical experiments with low, moderate and high dimensional environments. Jayakumar Subramanian, Amit Sinha, Raihan Seraj, Aditya Mahajan |
J. Mach. Learn. Res. | 1 |
| 2021 | Medical Dead-ends and Learning to Identify High-Risk States and TreatmentsabstractMachine learning has successfully framed many sequential decision making problems as either supervised prediction, or optimal decision-making policy identification via reinforcement learning. In data-constrained offline settings, both approaches may fail as they assume fully optimal behavior or rely on exploring alternatives that may not exist. We introduce an inherently different approach that identifies "dead-ends" of a state space. We focus on patient condition in the intensive care unit, where a "medical dead-end" indicates that a patient will expire, regardless of all potential future treatment sequences. We postulate "treatment security" as avoiding treatments with probability proportional to their chance of leading to dead-ends, present a formal proof, and frame discovery as an RL problem. We then train three independent deep neural models for automated state construction, dead-end discovery and confirmation. Our empirical results discover that dead-ends exist in real clinical data among septic patients, and further reveal gaps between secure treatments and those administered. Mehdi Fatemi, Taylor W. Killian, Jayakumar Subramanian, Marzyeh Ghassemi |
NeurIPS | 3 |