EDBT 2026 Demo / reviewers in the wild / expert
Lajanugen Logeswaran
dblp:157/3603
· DBLP profile ↗
23ranked-venue papers
9as first author
16since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 8 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Scalable Video-to-Dataset Generation for Cross-Platform Mobile AgentsabstractRecent advancements in Large Language Models (LLMs) and Vision-Language Models (VLMs) have sparked significant interest in developing GUI visual agents. We introduce MONDAY (Mobile OS Navigation Task Dataset for Agents from YouTube), a large-scale dataset of 313K annotated frames from 20K instructional videos capturing diverse real-world mobile OS navigation across multiple platforms. Models that include MONDAY in their pretraining phases demonstrate robust cross-platform generalization capabilities, consistently outperforming models trained on existing single OS datasets while achieving an average performance gain of 18.11%p on an unseen mobile OS platform. To enable continuous dataset expansion as mobile platforms evolve, we present an automated framework that leverages publicly available video content to create comprehensive task datasets without manual annotation. Our framework comprises robust OCR-based scene detection (95.04% F1-score), near-perfect UI element detection (99.87% hit ratio), and novel multi-step action identification to extract reliable action sequences across diverse interface configurations. We contribute both the MONDAY dataset and our automated collection framework to facilitate future research in mobile OS navigation, available at https://github.com/runamu/monday. Yunseok Jang 0001, Yeda Song, Sungryull Sohn, Lajanugen Logeswaran, Tiange Luo, Dong-Ki Kim, Kyunghoon Bae, Honglak Lee |
CVPR | 4 |
| 2025 | Revisiting LLM Value Probing Strategies: Are They Robust and Expressive?abstractThe value orientation of Large Language Models (LLMs) has been extensively studied, as it can shape user experiences across demographic groups.However, two key challenges remain:(1) the lack of systematic comparison across value probing strategies, despite the Multiple Choice Question (MCQ) setting being vulnerable to perturbations, and (2) the uncertainty over whether probed values capture in-context information or predict models' real-world actions.In this paper, we systematically compare three widely used value probing methods: token likelihood, sequence perplexity, and text generation.Our results show that all three methods exhibit large variances under non-semantic perturbations in prompts and option formats, with sequence perplexity being the most robust overall.We further introduce two tasks to assess expressiveness: demographic prompting, testing whether probed values adapt to cultural context; and value-action agreement, testing the alignment of probed values with value-based actions.We find that demographic context has little effect on the text generation method, and probed values only weakly correlate with action preferences across all methods.Our work highlights the instability and the limited expressive power of current value probing methods, calling for more reliable LLM value representations. Mehar Singh, Lajanugen Logeswaran, Moontae Lee, Honglak Lee, Rada Mihalcea |
EMNLP | 3 |
| 2025 | Visual Test-Time Scaling for GUI Agent Grounding
Tiange Luo, Lajanugen Logeswaran, Justin Johnson 0001, Honglak Lee |
ICCV | 2 |
| 2025 | MLRC-Bench: Can Language Agents Solve Machine Learning Research Challenges?abstractWe introduce MLRC-Bench, a benchmark designed to quantify how effectively language agents can tackle challenging Machine Learning (ML) Research Competitions, with a focus on open research problems that demand novel methodologies.Unlike prior work, e.g., AI Scientist, which evaluates the end-to-end agentic pipeline by using LLM-as-a-judge, MLRC-Bench measures the key steps of proposing and implementing novel research methods and evaluates them with rigorous protocol and objective metrics.Our curated suite of 7 competition tasks reveals significant challenges for LLM agents. Even the best-performing tested agent (gemini-exp-1206 under MLAB) closes only 9.3% of the gap between baseline and top human participant scores.Furthermore, our analysis reveals a misalignment between the LLM-judged innovation and their actual performance on cutting-edge ML research problems. MLRC-Bench is a dynamic benchmark, which is designed to continually grow with new ML competitions to encourage rigorous and objective evaluations of AI’s research capabilities. Our leaderboard and code are publicly available at https://huggingface.co/spaces/launch/MLRC_Bench. Muhammad Khalifa, Shitanshu Bhushan, Grant D. Murphy, Lajanugen Logeswaran, Jaekyeom Kim, Moontae Lee, Honglak Lee, Lu Wang 0008 |
NeurIPS | 5 |
| 2024 | Code Models are Zero-shot Precondition ReasonersabstractLajanugen Logeswaran, Sungryull Sohn, Yiwei Lyu, Anthony Liu, Dong-Ki Kim, Dongsub Shim, Moontae Lee, Honglak Lee. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Lajanugen Logeswaran, Sungryull Sohn, Yiwei Lyu 0001, Anthony Z. Liu, Dong-Ki Kim, Dongsub Shim, Moontae Lee, Honglak Lee |
NAACL-HLT | 1 |
| 2024 | Understanding the Capabilities and Limitations of Large Language Models for Cultural CommonsenseabstractSiqi Shen, Lajanugen Logeswaran, Moontae Lee, Honglak Lee, Soujanya Poria, Rada Mihalcea. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Lajanugen Logeswaran, Moontae Lee, Honglak Lee, Soujanya Poria, Rada Mihalcea |
NAACL-HLT | 2 |
| 2024 | You don't need a personality test to know these models are unreliable: Assessing the Reliability of Large Language Models on Psychometric InstrumentsabstractBangzhao Shu, Lechen Zhang, Minje Choi, Lavinia Dunagan, Lajanugen Logeswaran, Moontae Lee, Dallas Card, David Jurgens. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Bangzhao Shu, Minje Choi, Lavinia Dunagan, Lajanugen Logeswaran, Moontae Lee, Dallas Card, David Jurgens |
NAACL-HLT | 5 |
| 2024 | AutoGuide: Automated Generation and Selection of Context-Aware Guidelines for Large Language Model AgentsabstractRecent advances in large language models (LLMs) have empowered AI agents capable of performing various sequential decision-making tasks. However, effectively guiding LLMs to perform well in unfamiliar domains like web navigation, where they lack sufficient knowledge, has proven to be difficult with the demonstration-based in-context learning paradigm. In this paper, we introduce a novel framework, called AutoGuide, which addresses this limitation by automatically generating context-aware guidelines from offline experiences. Importantly, each context-aware guideline is expressed in concise natural language and follows a conditional structure, clearly describing the context where it is applicable. As a result, our guidelines facilitate the provision of relevant knowledge for the agent's current decision-making process, overcoming the limitations of the conventional demonstration-based learning paradigm. Our evaluation demonstrates that AutoGuide significantly outperforms competitive baselines in complex benchmark domains, including real-world web navigation. Dong-Ki Kim, Jaekyeom Kim, Sungryull Sohn, Lajanugen Logeswaran, Kyunghoon Bae, Honglak Lee |
NeurIPS | 5 |
| 2023 | Learning Compositional Tasks from Language InstructionsabstractThe ability to combine learned knowledge and skills to solve novel tasks is a key aspect of generalization in humans that allows us to understand and perform tasks described by novel language utterances. While progress has been made in supervised learning settings, no work has yet studied compositional generalization of a reinforcement learning agent following natural language instructions in an embodied environment. We develop a set of tasks in a photo-realistic simulated kitchen environment that allow us to study the degree to which a behavioral policy captures the systematicity in language by studying its zero-shot generalization performance on held out natural language instructions. We show that our agent which leverages a novel additive action-value decomposition in tandem with attention based subgoal prediction is able to exploit composition in text instructions to generalize to unseen tasks. Lajanugen Logeswaran, Wilka Carvalho, Honglak Lee |
AAAI | 1 |
| 2023 | Knowledge Unlearning for Mitigating Privacy Risks in Language ModelsabstractJoel Jang, Dongkeun Yoon, Sohee Yang, Sungmin Cha, Moontae Lee, Lajanugen Logeswaran, Minjoon Seo. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Joel Jang, Dongkeun Yoon, Sohee Yang, Sungmin Cha, Moontae Lee, Lajanugen Logeswaran, Minjoon Seo |
ACL (1) | 6 |
| 2023 | Few-shot Reranking for Multi-hop QA via Language Model PromptingabstractWe study few-shot reranking for multi-hop QA (MQA) with open-domain questions.To alleviate the need for a large number of labeled question-document pairs for retriever training, we propose PROMPTRANK, which relies on language model prompting for multi-hop path reranking.PROMPTRANK first constructs an instruction-based prompt that includes a candidate document path and then computes the relevance score between a given question and the path based on the conditional likelihood of the question given the path prompt according to a language model.PROMPTRANK yields strong retrieval performance on HotpotQA with only 128 training examples compared to state-of-theart methods trained on thousands of examples -73.6 recall@10 by PROMPTRANK vs. 77.8 by PathRetriever (Asai et al., 2020) and 77.5 by multi-hop dense retrieval (Xiong et al., 2021). Muhammad Khalifa, Lajanugen Logeswaran, Moontae Lee, Honglak Lee, Lu Wang 0008 |
ACL (1) | 2 |
| 2023 | A Picture is Worth a Thousand Words: Language Models Plan from PixelsabstractPlanning is an important capability of artificial agents that perform long-horizon tasks in real-world environments.In this work, we explore the use of pre-trained language models (PLMs) to reason about plan sequences from text instructions in embodied visual environments.Prior PLM based approaches for planning either assume observations are available in the form of text (e.g., provided by a captioning model), reason about plans from the instruction alone, or incorporate information about the visual environment in limited ways (such as a pre-trained affordance function).In contrast, we show that PLMs can accurately plan even when observations are directly encoded as input prompts for the PLM.We show that this simple approach outperforms prior approaches in experiments on the ALFWorld and VirtualHome benchmarks. Anthony Z. Liu, Lajanugen Logeswaran, Sungryull Sohn, Honglak Lee |
EMNLP | 2 |
| 2023 | TOD-Flow: Modeling the Structure of Task-Oriented DialoguesabstractTask-Oriented Dialogue (TOD) systems have become crucial components in interactive artificial intelligence applications.While recent advances have capitalized on pre-trained language models (PLMs), they exhibit limitations regarding transparency and controllability.To address these challenges, we propose a novel approach focusing on inferring the TOD-Flow graph from dialogue data annotated with dialog acts, uncovering the underlying task structure in the form of a graph.The inferred TOD-Flow graph can be easily integrated with any dialogue model to improve its prediction performance, transparency, and controllability.Our TOD-Flow graph learns what a model can, should, and should not predict, effectively reducing the search space and providing a rationale for the model's prediction.We show that the proposed TOD-Flow graph better resembles human-annotated graphs compared to prior approaches.Furthermore, when combined with several dialogue policies and end-to-end dialogue models, we demonstrate that our approach significantly improves dialog act classification and end-to-end response generation performance in the Multi-WOZ and SGD benchmarks.Code available Sungryull Sohn, Yiwei Lyu 0001, Anthony Z. Liu, Lajanugen Logeswaran, Dong-Ki Kim, Dongsub Shim, Honglak Lee |
EMNLP | 4 |
| 2023 | Merging Generated and Retrieved Knowledge for Open-Domain QAabstractOpen-domain question answering (QA) systems are often built with retrieval modules.However, retrieving passages from a given source is known to suffer from insufficient knowledge coverage.Alternatively, prompting large language models (LLMs) to generate contextual passages based on their parametric knowledge has been shown to improve QA performance.Yet, LLMs tend to "hallucinate" content that conflicts with the retrieved knowledge.Based on the intuition that answers supported by both sources are more likely to be correct, we propose COMBO, a Compatibility-Oriented knowledge Merging for Better Open-domain QA framework, to effectively leverage the two sources of information.Concretely, we match LLM-generated passages with retrieved counterparts into compatible pairs, based on discriminators trained with silver compatibility labels.Then a Fusionin-Decoder-based (Izacard and Grave, 2021b) reader model handles passage pairs to arrive at the final answer.Experiments show that COMBO outperforms competitive baselines on three out of four tested open-domain QA benchmarks.Further analysis reveals that our proposed framework demonstrates greater efficacy in scenarios with a higher degree of knowledge conflicts. Muhammad Khalifa, Lajanugen Logeswaran, Moontae Lee, Honglak Lee, Lu Wang 0008 |
EMNLP | 3 |
| 2023 | Exploring the Benefits of Training Expert Language Models over Instruction TuningabstractRecently, Language Models (LMs) instruction-tuned on multiple tasks, also known as multitask-prompted fine-tuning (MT), have shown capabilities to generalize to unseen tasks. Previous work has shown that scaling the number of finetuning datasets and instructions is the key component in making stronger MT LMs. In this work, we report surprising findings that show an expert LM trained on just a single task can outperform an MT LM trained with 300+ different tasks on 11 different unseen datasets and on 13 datasets of the BIG-bench benchmark by an average of 3.20% and 1.29%, respectively. This finding casts doubt on the previously held belief that simply scaling the number of tasks makes stronger MT LMs. Leveraging this finding, we further show that this distributed approach of training multiple expert LMs instead of a single MT LM for zero-shot inference possesses many benefits including (1) avoiding negative task transfer that often occurs during instruction tuning, (2) being able to continually learn new tasks without having to re-train on previous tasks to avoid catastrophic forgetting, and (3) showing compositional capabilities when merging individual experts together. Joel Jang, Seungone Kim, Seonghyeon Ye, Doyoung Kim 0001, Lajanugen Logeswaran, Moontae Lee, Kyungjae Lee 0002, Minjoon Seo |
ICML | 5 |
| 2022 | Few-shot Subgoal Planning with Language ModelsabstractLajanugen Logeswaran, Yao Fu, Moontae Lee, Honglak Lee. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Lajanugen Logeswaran, Moontae Lee, Honglak Lee |
NAACL-HLT | 1 |
| 2019 | Zero-Shot Entity Linking by Reading Entity DescriptionsabstractWe present the zero-shot entity linking task, where mentions must be linked to unseen entities without in-domain labeled data.The goal is to enable robust transfer to highly specialized domains, and so no metadata or alias tables are assumed.In this setting, entities are only identified by text descriptions, and models must rely strictly on language understanding to resolve the new entities.First, we show that strong reading comprehension models pre-trained on large unlabeled data can be used to generalize to unseen entities.Second, we propose a simple and effective adaptive pre-training strategy, which we term domainadaptive pre-training (DAP), to address the domain shift problem associated with linking unseen entities in a new domain.We present experiments on a new dataset that we construct for this task and show that DAP improves over strong pre-training baselines, including BERT. Lajanugen Logeswaran, Ming-Wei Chang, Kenton Lee, Kristina Toutanova, Jacob Devlin, Honglak Lee |
ACL (1) | 1 |
| 2018 | Sentence Ordering and Coherence Modeling using Recurrent Neural NetworksabstractModeling the structure of coherent texts is a key NLP problem. The task of coherently organizing a given set of sentences has been commonly used to build and evaluate models that understand such structure. We propose an end-to-end unsupervised deep learning approach based on the set-to-sequence framework to address this problem. Our model strongly outperforms prior methods in the order discrimination task and a novel task of ordering abstracts from scientific articles. Furthermore, our work shows that useful text representations can be obtained by learning to order sentences. Visualizing the learned sentence representations shows that the model captures high-level logical structure in paragraphs. Our representations perform comparably to state-of-the-art pre-training methods on sentence similarity and paraphrase detection tasks. Lajanugen Logeswaran, Honglak Lee, Dragomir R. Radev |
AAAI | 1 |
| 2018 | An efficient framework for learning sentence representations
Lajanugen Logeswaran, Honglak Lee |
ICLR (Poster) | 1 |
| 2018 | Content preserving text generation with attribute controlsabstractIn this work, we address the problem of modifying textual attributes of sentences. Given an input sentence and a set of attribute labels, we attempt to generate sentences that are compatible with the conditioning information. To ensure that the model generates content compatible sentences, we introduce a reconstruction loss which interpolates between auto-encoding and back-translation loss components. We propose an adversarial loss to enforce generated samples to be attribute compatible and realistic. Through quantitative, qualitative and human evaluations we demonstrate that our model is capable of generating fluent sentences that better reflect the conditioning information compared to prior methods. We further demonstrate that the model is capable of simultaneously controlling multiple attributes. Lajanugen Logeswaran, Honglak Lee, Samy Bengio |
NeurIPS | 1 |
| 2016 | Performance, Resource, and Cost Aware Resource Provisioning in the CloudabstractCurrent cloud users pay for statically configured VM sizes irrespective of usage. It is more favorable for users to consume (and be billed for) just the right amount of resources necessary to satisfy the performance requirement of their applications. We take a novel perspective to enable such resource usage, where we assume that the cloud operator exposes a small, dynamic fraction of its infrastructure, corresponding resource specifications, and constraints to each application. We then propose a dynamic and computationally efficient reconfiguration scheme which comprises an Application Performance Model, a Cost Model, and a Reconfiguration algorithm. The performance model estimates application performance given specific resources. Cost model assigns a value to resource candidates made available to the application considering the lease expense, reconfiguration penalty, and operating income. A reconfiguration algorithm, assisted by the cost model, then makes optimal reconfiguration decisions. Simulation results for RUBiS and filebench-fileserver applications and Worldcup workloads show significant cost savings while meeting performance targets compared to rule-based scaling. Lajanugen Logeswaran, H. M. N. Dilum Bandara, H. S. Bhathiya |
CLOUD | 1 |
| 2016 | Generative Adversarial Text to Image SynthesisabstractAutomatic synthesis of realistic images from text would be interesting and useful, but current AI systems are still far from this goal. However, in recent years generic and powerful recurrent neural network architectures have been developed to learn discriminative text feature representations. Meanwhile, deep convolutional generative adversarial networks (GANs) have begun to generate highly compelling images of specific categories such as faces, album covers, room interiors and flowers. In this work, we develop a novel deep architecture and GAN formulation to effectively bridge these advances in text and image modeling, translating visual concepts from characters to pixels. We demonstrate the capability of our model to generate plausible images of birds and flowers from detailed text descriptions. Scott E. Reed, Zeynep Akata, Xinchen Yan, Lajanugen Logeswaran, Bernt Schiele, Honglak Lee |
ICML | 4 |
| 2014 | Solving Jigsaw Puzzles using Paths and Cycles
Lajanugen Logeswaran |
BMVC | 1 |