VLDB 2026 Research / reviewers in the wild / expert
Palash Goyal
dblp:183/3699
· DBLP profile ↗
22ranked-venue papers
5as first author
13since 2021 · last 2025
0000-0003-2455-2160ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PLAN-TUNING: Post-Training Language Models to Learn Step-by-Step Planning for Complex Problem SolvingabstractMihir Parmar, Palash Goyal, Xin Liu, Yiwen Song, Mingyang Ling, Chitta Baral, Hamid Palangi, Tomas Pfister. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Mihir Parmar, Palash Goyal, Yiwen Song, Chitta Baral, Hamid Palangi, Tomas Pfister |
EMNLP | 2 |
| 2025 | PlanGEN: A Multi-Agent Framework for Generating Planning and Reasoning Trajectories for Complex Problem SolvingabstractMihir Parmar, Xin Liu, Palash Goyal, Yanfei Chen, Long Le, Swaroop Mishra, Hossein Mobahi, Jindong Gu, Zifeng Wang, Hootan Nakhost, Chitta Baral, Chen-Yu Lee, Tomas Pfister, Hamid Palangi. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Mihir Parmar, Palash Goyal, Yanfei Chen, Long T. Le, Swaroop Mishra, Hossein Mobahi, Jindong Gu, Zifeng Wang 0002, Hootan Nakhost, Chitta Baral, Chen-Yu Lee, Tomas Pfister, Hamid Palangi |
EMNLP | 3 |
| 2025 | Heterogeneous Swarms: Jointly Optimizing Model Roles and Weights for Multi-LLM SystemsabstractWe propose Heterogeneous Swarms, an algorithm to design multi-LLM systems by jointly optimizing model roles and weights. We represent multi-LLM systems as directed acyclic graphs (DAGs) of LLMs with topological message passing for collaborative generation. Given a pool of LLM experts and a utility function, Heterogeneous Swarms employs two iterative steps: role-step and weight-step. For role-step, we interpret model roles as learning a DAG that specifies the flow of inputs and outputs between LLMs. Starting from a swarm of random continuous adjacency matrices, we decode them into discrete DAGs, call the LLMs in topological order, evaluate on the utility function (e.g. accuracy on a task), and optimize the adjacency matrices with particle swarm optimization based on the utility score. For weight-step, we assess the contribution of individual LLMs in the multi-LLM systems and optimize model weights with swarm intelligence. We propose JFK-score to quantify the individual contribution of each LLM in the best-found DAG of the role-step, then optimize model weights with particle swarm optimization based on the JFK-score. Experiments demonstrate that Heterogeneous Swarms outperforms 17 role- and/or weight-based baselines by 18.5% on average across 12 tasks. Further analysis reveals that Heterogeneous Swarms discovers multi-LLM systems with heterogeneous model roles and substantial collaborative gains, and benefits from the diversity of language models. Shangbin Feng, Zifeng Wang 0002, Palash Goyal, Yike Wang 0002, Huang Xia, Hamid Palangi, Luke Zettlemoyer, Yulia Tsvetkov, Chen-Yu Lee, Tomas Pfister |
NeurIPS | 3 |
| 2024 | FLIRT: Feedback Loop In-context Red TeamingabstractNinareh Mehrabi, Palash Goyal, Christophe Dupuy, Qian Hu, Shalini Ghosh, Richard Zemel, Kai-Wei Chang, Aram Galstyan, Rahul Gupta. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Ninareh Mehrabi, Palash Goyal, Christophe Dupuy, Shalini Ghosh, Richard S. Zemel, Kai-Wei Chang 0001, Aram Galstyan, Rahul Gupta 0001 |
EMNLP | 2 |
| 2024 | Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language ModelsabstractData is a crucial element in large language model (LLM) alignment.Recent studies have explored using LLMs for efficient data collection.However, LLM-generated data often suffers from quality issues, with underrepresented or absent aspects and low-quality datapoints.To address these problems, we propose DATA ADVISOR, an enhanced LLMbased method for generating data that takes into account the characteristics of the desired dataset.Starting from a set of pre-defined principles in hand, DATA ADVISOR monitors the status of the generated data, identifies weaknesses in the current dataset, and advises the next iteration of data generation accordingly.DATA ADVISOR can be easily integrated into existing data generation methods to enhance data quality and coverage.Experiments on safety alignment of three representative LLMs (i.e., Mistral, Llama2, and Falcon) demonstrate the effectiveness of DATA ADVISOR in enhancing model safety against various fine-grained safety issues without sacrificing model utility.Warning: this paper contains example data that may be offensive or harmful. Fei Wang 0060, Ninareh Mehrabi, Palash Goyal, Rahul Gupta 0001, Kai-Wei Chang 0001, Aram Galstyan |
EMNLP | 3 |
| 2024 | The steerability of large language models toward data-driven personasabstractJunyi Li, Charith Peris, Ninareh Mehrabi, Palash Goyal, Kai-Wei Chang, Aram Galstyan, Richard Zemel, Rahul Gupta. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Junyi Li 0002, Charith Peris, Ninareh Mehrabi, Palash Goyal, Kai-Wei Chang 0001, Aram Galstyan, Richard S. Zemel, Rahul Gupta 0001 |
NAACL-HLT | 4 |
| 2023 | Resolving Ambiguities in Text-to-Image Generative ModelsabstractNinareh Mehrabi, Palash Goyal, Apurv Verma, Jwala Dhamala, Varun Kumar, Qian Hu, Kai-Wei Chang, Richard Zemel, Aram Galstyan, Rahul Gupta. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Ninareh Mehrabi, Palash Goyal, Apurv Verma, Jwala Dhamala, Kai-Wei Chang 0001, Richard S. Zemel, Aram Galstyan, Rahul Gupta 0001 |
ACL (1) | 2 |
| 2023 | Faithful Model Evaluation for Model-Based MetricsabstractStatistical significance testing is used in natural language processing (NLP) to determine whether the results of a study or experiment are likely to be due to chance or if they reflect a genuine relationship.A key step in significance testing is the estimation of confidence interval which is a function of sample variance.Sample variance calculation is straightforward when evaluating against ground truth.However, in many cases, a metric model is often used for evaluation.For example, to compare toxicity of two large language models, a toxicity classifier is used for evaluation.Existing works usually do not consider the variance change due to metric model errors, which can lead to wrong conclusions.In this work, we establish the mathematical foundation of significance testing for model-based metrics.With experiments on public benchmark datasets and a production system, we show that considering metric model errors to calculate sample variances for model-based metrics changes the conclusions in certain experiments. Palash Goyal, Rahul Gupta 0001 |
EMNLP | 2 |
| 2023 | Incorporating Fairness in Large Scale NLU SystemsabstractNLU models power several user facing experiences such as conversations agents and chat bots. Building NLU models typically consist of 3 stages: a) building or finetuning a pre-trained model b) distilling or fine-tuning the pre-trained model to build task specific models and, c) deploying the task-specific model to production. In this presentation, we will identify fairness considerations that can be incorporated in the aforementioned three stages in the life-cycle of NLU model building: (i) selection/building of a large scale language model, (ii) distillation/fine-tuning the large model into task specific model and, (iii) deployment of the task specific model. We will present select metrics that can be used to quantify fairness in NLU models and fairness enhancement techniques that can be deployed in each of these stages. Finally, we will share some recommendations to successfully implement fairness considerations when building an industrial scale NLU system. Rahul Gupta 0001, Lisa Bauer, Kai-Wei Chang 0001, Jwala Dhamala, Aram Galstyan, Palash Goyal, Avni Khatri, Rohit Parimi, Charith Peris, Apurv Verma, Richard S. Zemel, Premkumar Natarajan |
WSDM | 6 |
| 2022 | Leveraging Local Temporal Information for Multimodal Scene ClassificationabstractRobust video scene classification models should capture the spatial (pixel-wise) and temporal (frame-wise) characteristics of a video effectively. Transformer models with self-attention which are designed to get contextualized representations for individual tokens given a sequence of tokens, are becoming increasingly popular in many computer vision tasks. However, the use of Transformer based models for video under-standing is still relatively unexplored. Moreover, these models fail to exploit the strong temporal relationships between the neighboring video frames to get potent frame-level representations. In this paper, we propose a novel self-attention block that leverages both local and global temporal relation-ships between the video frames to obtain better contextualized representations for the individual frames. This enables the model to understand the video at various granularities. We illustrate the performance of our models on the large-scale YoutTube-8M data set on the task of video categorization and further analyze the results to showcase improvement. Saurabh Sahu, Palash Goyal |
ICASSP | 2 |
| 2021 | Hierarchical Class-Based Curriculum LossabstractClassification algorithms in machine learning often assume a flat label space. However, most real world data have dependencies between the labels, which can often be captured by using a hierarchy. Utilizing this relation can help develop a model capable of satisfying the dependencies and improving model accuracy and interpretability. Further, as different levels in the hierarchy correspond to different granularities, penalizing each label equally can be detrimental to model learning. In this paper, we propose a loss function, hierarchical curriculum loss, with two properties: (i) satisfy hierarchical constraints present in the label space, and (ii) provide non-uniform weights to labels based on their levels in the hierarchy, learned implicitly by the training paradigm. We theoretically show that the proposed hierarchical class-based curriculum loss is a tight bound of 0-1 loss among all losses satisfying the hierarchical constraints. We test our loss function on real world image data sets, and show that it significantly outperforms state-of-the-art baselines. Palash Goyal, Divya Choudhary, Shalini Ghosh |
IJCAI | 1 |
| 2021 | Pykg2vec: A Python Library for Knowledge Graph EmbeddingabstractPykg2vec is a Python library for learning the representations of the entities and relations in knowledge graphs. Pykg2vec's flexible and modular software architecture currently implements 25 state-of-the-art knowledge graph embedding algorithms, and is designed to easily incorporate new algorithms.The goal of pykg2vec is to provide a practical and educational platform to accelerate research in knowledge graph representation learning. Pykg2vec is built on top of PyTorch and Python's multiprocessing framework and provides modules for batch generation, Bayesian hyperparameter optimization, evaluation of KGE tasks, embedding, and result visualization. Pykg2vec is released under the MIT License and is also available in the Python Package Index (PyPI). The source code of pykg2vec is available at https://github.com/Sujit-O/pykg2vec. Shih-Yuan Yu, Sujit Rokka Chhetri, Arquimedes Canedo, Palash Goyal, Mohammad Abdullah Al Faruque |
J. Mach. Learn. Res. | 4 |
| 2021 | ArduCode: Predictive Framework for Automation EngineeringabstractAutomation engineering is the task of integrating, via software, various sensors, actuators, and controls to automate a real-world process. Today, automation engineering is supported by a suite of software tools, including integrated development environments (IDEs), hardware configurators, compilers, and runtimes. These tools focus on the automation code itself but leave the automation engineer unassisted in their decision-making. This can lead to longer software development cycles due to the imperfections in the decision-making, which arise when integrating software and hardware. To address this problem, this article addresses multiple challenges often faced in automation engineering and proposes machine learning-based solutions to assist engineers tackle these challenges. We show that machine learning can be leveraged to assist the automation engineer in classifying automation code, finding similar code snippets, and reasoning about the hardware selection of sensors and actuators. We validate our architecture on two real data sets consisting of 2927 Arduino projects and 683 programmable logic controller (PLC) projects. Our results show that paragraph embedding techniques can be utilized to classify automation using code snippets with precision close to human annotation, giving an$F_{1}$-score of 72%. Furthermore, we show that such embedding techniques can help us find similar code snippets with high accuracy. Finally, we use autoencoder models for hardware recommendation and achieve a$p\text{@}3$of 0.79 and$p\text{@}5$of 0.95. We also present the implementation of ArduCode in a proof-of-concept user interface integrated into an existing automation engineering system platform.Note to Practitioners—This article is motivated by the use of artificial intelligence methods to improve the efficiency and quality of the automation engineering software development process. Our goal is to develop and integrate intelligent assistants in existing automation engineering development tools to minimally disrupt existing workflows. Practitioners should be able to adapt our framework to other tools and data. Our contributions address important practical problems: 1) we address the lack of realistic data sets in automation engineering with two publicly available data sources; 2) we make the reference implementation of our algorithms publicly available on GitHub for other practitioners to have a starting point for future research; and 3) we demonstrate the integration of our framework as an add-on to an existing automation engineering toolchain. Arquimedes Canedo, Palash Goyal, Gustavo Quiros Araya |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2020 | Modeling Dialogues with Hashcode Representations: A Nonparametric ApproachabstractWe propose a novel dialogue modeling framework, the first-ever nonparametric kernel functions based approach for dialogue modeling, which learns hashcodes as text representations; unlike traditional deep learning models, it handles well relatively small datasets, while also scaling to large ones. We also derive a novel lower bound on mutual information, used as a model-selection criterion favoring representations with better alignment between the utterances of participants in a collaborative dialogue setting, as well as higher predictability of the generated responses. As demonstrated on three real-life datasets, including prominently psychotherapy sessions, the proposed approach significantly outperforms several state-of-art neural network based dialogue systems, both in terms of computational efficiency, reducing training time from days or weeks to hours, and the response quality, achieving an order of magnitude improvement over competitors in frequency of being chosen as the best model by human evaluators. Sahil Garg, Irina Rish, Guillermo A. Cecchi, Palash Goyal, Sarik Ghazarian, Shuyang Gao, Greg Ver Steeg, Aram Galstyan |
AAAI | 4 |
| 2020 | Graph Representation Ensemble LearningabstractRepresentation learning on graphs has been gaining attention due to its wide applicability in predicting missing links and classifying and recommending nodes. Most embedding methods aim to preserve specific properties of the original graph in the low dimensional space. However, real-world graphs have a combination of several features that are difficult to characterize and capture by a single approach. In this work, we introduce the problem of graph representation ensemble learning and provide a first of its kind framework to aggregate multiple graph embedding methods efficiently. We provide analysis of our framework and analyze - theoretically and empirically - the dependence between state-of-the-art embedding methods. We test our models on the node classification task on four realworld graphs and show that proposed ensemble approaches can outperform the state-of-the-art methods by up to 20% on macro-F1. We further show that the strategy is even more beneficial for underrepresented classes with an improvement of up to 40%. Palash Goyal, Sachin Raja, Sujit Rokka Chhetri, Arquimedes Canedo, Ajoy Mondal, Jaya Shree, C. V. Jawahar |
ASONAM | 1 |
| 2020 | Cross-modal Non-linear Guided Attention and Temporal Coherence in Multi-modal Deep Video ModelsabstractVideos have data in multiple modalities, e.g., audio, video, text (captions). Understanding and modeling the interaction between different modalities is key for video analysis tasks like categorization, object detection, activity recognition, etc. However, data modalities are not always correlated --- so, learning when modalities are correlated and using that to guide the influence of one modality on the other is crucial. Another salient feature of videos is the coherence between successive frames due to continuity of video and audio, a property that we refer to as temporal coherence. We show how using non-linear guided cross-modal signals and temporal coherence can improve the performance of multi-modal machine learning (ML) models for video analysis tasks like categorization. Our experiments on the large-scale YouTube-8M dataset show how our approach significantly outperforms state-of-the-art multi-modal ML models for video categorization. The model trained on the YouTube-8M dataset also showed good performance on an internal dataset of video segments from actual Samsung TV Plus channels without retraining or fine-tuning, showing the generalization capabilities of our model. Saurabh Sahu, Palash Goyal, Shalini Ghosh, Chul Lee |
ACM Multimedia | 2 |
| 2020 | dyngraph2vec: Capturing network dynamics using dynamic graph representation learning
Palash Goyal, Sujit Rokka Chhetri, Arquimedes Canedo |
Knowl. Based Syst. | 1 |
| 2018 | DarkEmbed: Exploit Prediction With Neural Language ModelsabstractSoftware vulnerabilities can expose computer systems to attacks by malicious actors. With the number of vulnerabilities discovered in the recent years surging, creating timely patches for every vulnerability is not always feasible. At the same time, not every vulnerability will be exploited by attackers; hence, prioritizing vulnerabilities by assessing the likelihood they will be exploited has become an important research problem. Recent works used machine learning techniques to predict exploited vulnerabilities by analyzing discussions about vulnerabilities on social media. These methods relied on traditional text processing techniques, which represent statistical features of words, but fail to capture their context. To address this challenge, we propose DarkEmbed, a neural language modeling approach that learns low dimensional distributed representations, i.e., embeddings, of darkweb/deepweb discussions to predict whether vulnerabilities will be exploited. By capturing linguistic regularities of human language, such as syntactic, semantic similarity and logic analogy, the learned embeddings are better able to classify discussions about exploited vulnerabilities than traditional text analysis methods. Evaluations demonstrate the efficacy of learned embeddings on both structured text (such as security blog posts) and unstructured text (darkweb/deepweb posts). DarkEmbed outperforms state-of-the-art approaches on the exploit prediction task with an F1-score of 0.74. Nazgol Tavabi, Palash Goyal, Mohammed Almukaynizi, Paulo Shakarian, Kristina Lerman |
AAAI | 2 |
| 2018 | Modeling Evolution of Topics in Large-Scale Temporal Text Corpora
Elaheh Momeni, Shanika Karunasekera, Palash Goyal, Kristina Lerman |
ICWSM | 3 |
| 2018 | Graph embedding techniques, applications, and performance: A survey
Palash Goyal, Emilio Ferrara |
Knowl. Based Syst. | 1 |
| 2018 | Capturing Edge Attributes via Network EmbeddingabstractNetwork embedding, which aims to learn low-dimensional representations of nodes, has been used for various graph related tasks including visualization, link prediction, and node classification. Most existing embedding methods rely solely on network structure. However, in practice, we often have auxiliary information about the nodes and/or their interactions, e.g., the content of scientific papers in coauthorship networks, or topics of communication in Twitter mention networks. Here, we propose a novel embedding method that uses both network structure and edge attributes to learn better network representations. Our method jointly minimizes the reconstruction error for higher order node neighborhood, social roles, and edge attributes using a deep architecture that can adequately capture highly nonlinear interactions. We demonstrate the efficacy of our model over existing state-of-the-art methods on a variety of real-world networks including collaboration networks and social networks. We also observe that using edge attributes to inform network embedding yields better performance in downstream tasks such as link prediction and node classification. Palash Goyal, Homa Hosseinmardi, Emilio Ferrara, Aram Galstyan |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2017 | OReONet: Deep convolutional network for oil reservoir optimizationabstractIn recent years, deep convolutional networks have been successfully used for the tasks of image classification and speech recognition. The highly non-linear modeling combined with its emphasis on local connectivity makes them highly suitable for such tasks. However, their performance in other domains is not well explored. Specifically, in the oil industry, researchers use manual features from time series data as input to various machine learning models. In this paper, we employ deep convolutional autoencoders to extract non linear latent features from time series data. We propose a novel deep network architecture and show its efficacy in two oil field tasks related to reservoir optimization - steam job prediction and slippage detection. We show that our architecture outperforms state-of-the-art methods significantly on steam job prediction. We demonstrate the success of our model on an oil field dataset which consists of production and failure data of over two years. Our architecture achieves a precision of 98% for precision@50 in steam job prediction, and 25% improvement over the methods used in the industry. To the best of our knowledge, we are the first to attempt to automatically detect slippage failures in well pumps. We are able to classify slippage events with 70.3% accuracy, a 10.6% improvement over using manually defined input features. Chung Ming Cheung, Palash Goyal, Viktor Prasanna 0001, Arash Saber Tehrani |
IEEE BigData | 2 |