VLDB 2026 Research / reviewers in the wild / expert
Megh Thakkar
dblp:92/6840
· DBLP profile ↗
13ranked-venue papers
2as first author
13since 2021 · last 2025
0009-0004-8224-622XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 2 first-author · 11 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | How to Train Your LLM Web Agent: A Statistical DiagnosisabstractLarge language model (LLM) agents for web interfaces have advanced rapidly, yet open-source systems still lag behind proprietary agents. Bridging this gap is key to enabling customizable, efficient, and privacy-preserving agents. Two challenges hinder progress: the reproducibility issues in RL and LLM agent training, where results often depend on sensitive factors like seeds and decoding parameters, and the focus of prior work on single-step tasks, overlooking the complexities of web-based, multi-step decision-making.
We address these gaps by providing a statistically driven study of training LLM agents for web tasks. Our two-stage pipeline combines imitation learning from a Llama 3.3 70B teacher with on-policy fine-tuning via Group Relative Policy Optimization (GRPO) on a Llama 3.1 8B student. Through 240 configuration sweeps and rigorous bootstrapping, we chart the first compute allocation curve for open-source LLM web agents. Our findings show that dedicating one-third of compute to teacher traces and the rest to RL improves MiniWoB++ success by 6 points and closes 60\% of the gap to GPT-4o on WorkArena, while cutting GPU costs by 45\%. We introduce a principled hyperparameter sensitivity analysis, offering actionable guidelines for robust and cost-effective agent training. Dheeraj Vattikonda, Santhoshi Ravichandran, Emiliano Penaloza, Hadi Nekoei, Thibault Le Sellier de Chezelles, Megh Thakkar, Nicolas Angelard-Gontier, Miguel Muñoz-Mármol, Sahar Omidi Shayegan, Stefania Raimondo, Steve (Xue) Liu, Alexandre Drouin, Alexandre Piché, Alexandre Lacoste, Massimo Caccia |
NeurIPS | 6 |
| 2024 | A Deep Dive into the Trade-Offs of Parameter-Efficient Preference Alignment TechniquesabstractMegh Thakkar, Quentin Fournier, Matthew Riemer, Pin-Yu Chen, Amal Zouaq, Payel Das, Sarath Chandar. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Megh Thakkar, Quentin Fournier, Matthew Riemer, Amal Zouaq, Sarath Chandar |
ACL (1) | 1 |
| 2024 | WorkArena++: Towards Compositional Planning and Reasoning-based Common Knowledge Work TasksabstractThe ability of large language models (LLMs) to mimic human-like intelligence has led to a surge in LLM-based autonomous agents. Though recent LLMs seem capable of planning and reasoning given user instructions, their effectiveness in applying these capabilities for autonomous task solving remains underexplored. This is especially true in enterprise settings, where automated agents hold the promise of a high impact. To fill this gap, we propose WorkArena++, a novel benchmark consisting of 682 tasks corresponding to realistic workflows routinely performed by knowledge workers. WorkArena++ is designed to evaluate the planning, problem-solving, logical/arithmetic reasoning, retrieval, and contextual understanding abilities of web agents. Our empirical studies across state-of-the-art LLMs and vision-language models (VLMs), as well as human workers, reveal several challenges for such models to serve as useful assistants in the workplace. In addition to the benchmark, we provide a mechanism to effortlessly generate thousands of ground-truth observation/action traces, which can be used for fine-tuning existing models. Overall, we expect this work to serve as a useful resource to help the community progress towards capable autonomous agents. The benchmark can be found at https://github.com/ServiceNow/WorkArena. Léo Boisvert, Megh Thakkar, Maxime Gasse, Massimo Caccia, Thibault Le Sellier de Chezelles, Quentin Cappart, Nicolas Chapados, Alexandre Lacoste, Alexandre Drouin |
NeurIPS | 2 |
| 2023 | Towards Robust Low-Resource Fine-Tuning with Multi-View Compressed RepresentationsabstractDue to the huge amount of parameters, finetuning of pretrained language models (PLMs) is prone to overfitting in the low resource scenarios.In this work, we present a novel method that operates on the hidden representations of a PLM to reduce overfitting.During fine-tuning, our method inserts random autoencoders between the hidden layers of a PLM, which transform activations from the previous layers into multi-view compressed representations before feeding them into the upper layers.The autoencoders are plugged out after fine-tuning, so our method does not add extra parameters or increase computation cost during inference.Our method demonstrates promising performance improvement across a wide range of sequenceand token-level low-resource NLP tasks.Our code is available at https://github.com/DAMO- NLP-SG/MVCR. Xingxuan Li, Megh Thakkar, Xin Li 0056, Shafiq R. Joty, Luo Si, Lidong Bing |
ACL (1) | 3 |
| 2023 | Randomized Smoothing with Masked Inference for Adversarially Robust Text ClassificationsabstractLarge-scale pre-trained language models have shown outstanding performance in a variety of NLP tasks.However, they are also known to be significantly brittle against specifically crafted adversarial examples, leading to increasing interest in probing the adversarial robustness of NLP systems.We introduce RSMI, a novel two-stage framework that combines randomized smoothing (RS) with masked inference (MI) to improve the adversarial robustness of NLP systems.RS transforms a classifier into a smoothed classifier to obtain robust representations, whereas MI forces a model to exploit the surrounding context of a masked token in an input sequence.RSMI improves adversarial robustness by 2 to 3 times over existing state-of-the-art methods on benchmark datasets.We also perform in-depth qualitative analysis to validate the effectiveness of the different stages of RSMI and probe the impact of its components through extensive ablations.By empirically proving the stability of RSMI, we put it forward as a practical method to robustly train large-scale NLP models.Our code and datasets are available at https://github.com/Han8931/rsmi_nlp. Han Cheol Moon, Shafiq R. Joty, Megh Thakkar |
ACL (1) | 4 |
| 2023 | Self-Influence Guided Data Reweighting for Language Model Pre-trainingabstractLanguage Models (LMs) pre-trained with selfsupervision on large text corpora have become the default starting point for developing models for various NLP tasks.Once the pre-training corpus has been assembled, all data samples in the corpus are treated with equal importance during LM pre-training.However, due to varying levels of relevance and quality of data, equal importance to all the data samples may not be the optimal choice.While data reweighting has been explored in the context of task-specific supervised learning and LM fine-tuning, model-driven reweighting for pretraining data has not been explored.We fill this important gap and propose PRESENCE, a method for jointly reweighting samples by leveraging self-influence (SI) scores as an indicator of sample importance and pre-training.PRESENCE promotes novelty and stability for model pre-training.Through extensive analysis spanning multiple model sizes, datasets, and tasks, we present PRESENCE as an important first step in the research direction of sample reweighting for pre-training language models. Megh Thakkar, Tolga Bolukbasi, Sriram Ganapathy, Shikhar Vashishth, Sarath Chandar, Partha Talukdar |
EMNLP | 1 |
| 2023 | Learning Through Interpolative Augmentation of Dynamic Curvature SpacesabstractMixup is an efficient data augmentation technique, which improves generalization by interpolating random examples. While numerous approaches have been developed for Mixup in the Euclidean and in the hyperbolic space, they do not fully use the intrinsic properties of the examples, i.e., they manually set the geometry (Euclidean or hyperbolic) based on the overall dataset, which may be sub-optimal since each example may require a different geometry. We propose DynaMix, a framework that automatically selects an example-specific geometry and performs Mixup between the different geometries to improve training dynamics and generalization. Through extensive experiments in image and text modalities we show that DynaMix outperforms state-of-the-art methods over six downstream applications. We find that DynaMix is more useful in low-resource and semi-supervised settings likely because it displays a probabilistic view of the geometry. Parth Chhabra, Atula Tejaswi Neerkaje, Shivam Agarwal, Ramit Sawhney, Megh Thakkar, Preslav Nakov, Sudheer Chava |
SIGIR | 5 |
| 2022 | Chart-to-Text: A Large-Scale Benchmark for Chart SummarizationabstractShankar Kantharaj, Rixie Tiffany Leong, Xiang Lin, Ahmed Masry, Megh Thakkar, Enamul Hoque, Shafiq Joty. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Shankar Kantharaj, Rixie Tiffany Ko Leong, Ahmed Masry, Megh Thakkar, Enamul Hoque Prince, Shafiq R. Joty |
ACL (1) | 5 |
| 2022 | THINK: Temporal Hypergraph Hyperbolic NetworkabstractNetwork-based time series forecasting is a challenging task as it involves complex geometric properties, higher-order relations, and scale-free characteristics. Previous work has modeled network-based series as oversimplified graphs or has ignored the power law dynamics of real-world temporal and dynamic networks, which could yield suboptimal results. With the aim to address these issues, here we propose THINK, a novel framework based on hypergraph learning that captures the hyperbolic properties of time-evolving dynamic hypergraphs. We design an elegant hyperbolic distance-aware hypergraph attention mechanism to better capture informative internal structural features on the Poincaré ball. Through quantitative and conceptual analysis on seven tasks across temporal, and time-evolving dynamic hypergraphs, we demonstrate THINK’s practicality in comparison to a variety of benchmarks spanning finance, health, and energy networks. Shivam Agarwal, Ramit Sawhney, Megh Thakkar, Preslav Nakov, Jiawei Han 0001, Tyler Derr |
ICDM | 3 |
| 2022 | PISA: PoIncaré Saliency-Aware Interpolative Augmentation
Ramit Sawhney, Megh Thakkar, Vishwa Shah, Puneet Mathur, Vasu Sharma, Dinesh Manocha |
INTERSPEECH | 2 |
| 2022 | CIAug: Equipping Interpolative Augmentation with Curriculum LearningabstractRamit Sawhney, Ritesh Soun, Shrey Pandit, Megh Thakkar, Sarvagya Malaviya, Yuval Pinter. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Ramit Sawhney, Ritesh Soun, Shrey Pandit, Megh Thakkar, Sarvagya Malaviya, Yuval Pinter |
NAACL-HLT | 4 |
| 2021 | HypMix: Hyperbolic Interpolative Data AugmentationabstractInterpolation-based regularisation methods for data augmentation have proven to be effective for various tasks and modalities.These methods involve performing mathematical operations over the raw input samples or their latent states representations -vectors that often possess complex hierarchical geometries.However, these operations are performed in the Euclidean space, simplifying these representations, which may lead to distorted and noisy interpolations.We propose HypMix, a novel model-, data-, and modality-agnostic interpolative data augmentation technique operating in the hyperbolic space, which captures the complex geometry of input and hidden state hierarchies better than its contemporaries.We evaluate HypMix on benchmark and low resource datasets across speech, text, and vision modalities, showing that HypMix consistently outperforms state-of-the-art data augmentation techniques.In addition, we demonstrate the use of HypMix in semi-supervised settings.We further probe into the adversarial robustness and qualitative inferences we draw from HypMix that elucidate the efficacy of the Riemannian hyperbolic manifolds for interpolation-based data augmentation. Ramit Sawhney, Megh Thakkar, Shivam Agarwal, Diyi Yang, Lucie Flek |
EMNLP (1) | 2 |
| 2021 | Hyperbolic Online Time Stream ModelingabstractThe rapidly rising ubiquity and dissemination of online information such as social media text and news improve user accessibility towards financial markets, however, modeling these vast streams of irregular, temporal data poses a challenge. Such temporal streams of information show power-law dynamics, scale-free characteristics, and time irregularities that sequential models are unable to accurately model. In this work, we propose the first Hierarchical Time-Aware Hyperbolic LSTM (HTLSTM), which leverages the Riemannian manifold for encoding the scale-free nature of a sequence of text in a time-aware fashion. Through experiments on three financial tasks: stock trading, equity price movement prediction, and financial risk prediction, we demonstrate HTLSTM's applicability for modeling temporal sequences of online information. On real-world data from four global stock markets and three stock indices spanning data in English and Chinese, we make a step towards time-aware text modeling via hyperbolic geometry. Ramit Sawhney, Shivam Agarwal, Megh Thakkar, Arnav Wadhwa, Rajiv Ratn Shah |
SIGIR | 3 |