Pranjul Yadav

dblp:16/7528 · DBLP profile ↗
← Back
13ranked-venue papers in the field
2as first author
6since 2021 · last 2025
0000-0002-7860-5830ORCID · corroborated

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 10 (1 first)Information Retrieval & Web Search · 2Data Mining & Knowledge Discovery · 1 (1 first)
YearPublicationVenuePosition
2025 Survey of Agents Methodologies in the Financial Domain
Mohammad Luqman, Himaanshu Gauba, Akshar Prabhu Desai, Ritu Prajapati, Pranjul Yadav
IEEE Big Data7
2025 Multi Modal and Self Learning Agents in Finance
Mohammad Luqman, Akshar Prabhu Desai, Himaanshu Gauba, Pranjul Yadav
IEEE Big Data5
2024 Emerging Trends in LLM Benchmarking
abstract
Traditionally, machine learning models that focused on specialized tasks facilitated straightforward evaluation. However, the evolution of Large Language Models, has increased the complexity w.r.t. performance measurement. Evaluation and benchmarking of large language model is a significant challenge due to their versatility and improved capability to perform a wide range of tasks. This manuscript examines existing literature for various benchmarks and identifies a comprehensive overview of the emerging trends in benchmarking methodology.
Akshar Prabhu Desai, Ritu Prajapati, Tejasvi Ravi, Mohammad Luqman, Pranjul Yadav
IEEE Big Data5
2024 Opportunities and Challenges of Generative-AI in Finance
abstract
Gen-AI techniques are able to improve understanding of context and nuances in language modeling, translation between languages, handle large volumes of data, provide fast, low-latency responses and can be fine-tuned for various tasks and domainsIn this manuscript, we present a comprehensive overview of the applications of Gen-AI techniques in the finance domain. In particular, we present the opportunities and challenges associated with the usage of Gen-AI techniques. We also illustrate the various methodologies which can be used to train Gen-AI techniques and present the various application areas of Gen-AI technologies in the finance ecosystemTo the best of our knowledge, this work represents the most comprehensive summarization of Gen-AI techniques within the financial domain. The analysis is designed for a deep overview of areas marked for substantial advancement while simultaneously pin-point those warranting future prioritization. We also hope that this work would serve as a conduit between finance and other domains, thus fostering the cross-pollination of innovative concepts and practices.
Akshar Prabhu Desai, Tejasvi Ravi, Mohammad Luqman, Ganesh Satish Mallya, Nithya Kota, Pranjul Yadav
IEEE Big Data6
2024 Gen-AI for User Safety: A Survey
abstract
In this manuscript, we provide a comprehensive overview of the various work done while using Gen-AI techniques w.r.t user safety. In particular, we first provide the various domains (e.g. phishing, malware, content moderation, counterfeit, physical safety) across which Gen-AI techniques have been applied. Next, we provide how Gen-AI techniques can be used in conjunction with various data modalities i.e. text, images, videos, audio, executable binaries to detect violations of user-safety. Further, also provide an overview of how Gen-AI techniques can be used in an adversarial setting. We believe that this work represents the first summarization of Gen-AI techniques for user-safety.
Akshar Prabhu Desai, Tejasvi Ravi, Mohammad Luqman, Nithya Kota, Pranjul Yadav
IEEE Big Data6
2021 Addressing Stability in Classifier Explanations
abstract
Machine learning based classifiers are often a black box when considering the contribution of inputs to the output probability of a label, especially with complex non-linear models such as neural networks. A popular way to explain machine learning model outputs in a model agnostic manner is through the use of Shapley values. For our use case of abuse fighting in digital advertisements, one primary impediment of using Shapley values in explanations was a problem of instability. Specifically, the instability problem manifests as explanations for the same example varying greatly due to random sampling in the algorithm. We found it useful to view this problem explicitly as Monte Carlo integration in the form of averaging the model output while varying only a subset of features in the example to be explained. In turn, this guides the number of samples needed to achieve a stable estimate of individual Shapley values and unlocked the use of Shapley value based explainers for our models as well as classifiers in general, including neural networks.
Siavash Samiei, Nasrin Baratalipour, Pranjul Yadav, Amitabha Roy 0001, Dake He
IEEE BigData3
2019 Combining Text and Image data for Product Recommendability Modeling
abstract
Digital advertising aims to display relevant product based advertisements which matches user's intent. User's intent is usually captured via various metrics like clicks and sale. Research in the recent past have demonstrated that image of the product plays an important role in order to achieve the outcome of interest. However, product advertisements with nonrecommendable images, i.e., product depicting adult, racial or political content often leads to sub-optimal performances, along with bad reputation for the partner and the advertiser. To overcome this challenge, we propose a model for product recommendability, i.e., to identify whether the product is recommendable or not. Our proposed approach is a hybrid model build using image and text data. We demonstrate the superior performance of our proposed approach on a dataset, obtained from a major e-commerce advertiser.
Mark Capelo, Karan Aggarwal, Pranjul Yadav
IEEE BigData3
2019 Targeted display advertising: the case of preferential attachment
abstract
An average adult is exposed to hundreds of digital advertisements daily1, making the digital advertisement industry a classic example of a big-data-driven platform. As such, the ad-tech industry relies on historical engagement logs (clicks or purchases) to identify potentially interested users for the advertisement campaign of a partner (a seller who wants to target users for its products). The number of advertisements that are shown for a partner, and hence the historical campaign data available for a partner depends upon the budget constraints of the partner. Thus, enough data can be collected for the high-budget partners to make accurate predictions, while this is not the case with the low-budget partners. This skewed distribution of the data leads to preferential attachment of the targeted display advertising platforms towards the high-budget partners. In this paper, we develop domain-adaptation approaches to address the challenge of predicting interested users for the partners with insufficient data, i.e., the tail partners. Specifically, we develop simple yet effective approaches that leverage the similarity among the partners to transfer information from the partners with sufficient data to cold-start partners, i.e., partners without any campaign data. Our approaches readily adapt to the new campaign data by incremental fine-tuning, and hence work at varying points of a campaign, and not just the cold-start. We present an experimental analysis on the historical logs of a major display advertising platform2. Specifically, we evaluate our approaches across 149 partners, at varying points of their campaigns. Experimental results show that the proposed approaches outperform the other domain-adaptation approaches at different time points of the campaigns.
Saurav Manchanda, Pranjul Yadav, Khoa D. Doan, S. Sathiya Keerthi
IEEE BigData2
2019 Frequent Causal Pattern Mining: A Computationally Efficient Framework For Estimating Bias-Corrected Effects
abstract
Our aging population increasingly suffers from multiple chronic diseases simultaneously, necessitating the comprehensive treatment of these conditions. Finding the optimal set of drugs for a combinatorial set of diseases is a combinatorial pattern exploration problem. Association rule mining is a popular tool for such problems, but the requirement of health care for finding causal, rather than associative, patterns renders association rule mining unsuitable. To address this issue, we propose a novel framework based on the Rubin-Neyman causal model for extracting causal rules from observational data, correcting for a number of common biases. Specifically, given a set of interventions and a set of items that define subpopulations (e.g., diseases), we wish to find all subpopulations in which effective intervention combinations exist and in each such subpopulation, we wish to find all intervention combinations such that dropping any intervention from this combination will reduce the efficacy of the treatment. A key aspect of our framework is the concept of closed intervention sets which extend the concept of quantifying the effect of a single intervention to a set of concurrent interventions. Closed intervention sets also allow for a pruning strategy that is strictly more efficient than the traditional pruning strategy used by the Apriori algorithm. To implement our ideas, we introduce and compare five methods of estimating causal effect from observational data and rigorously evaluate them on synthetic data to mathematically prove (when possible) why they work. We also evaluated our causal rule mining framework on the Electronic Health Records (EHR) data of a large cohort of 152000 patients from Mayo Clinic and showed that the patterns we extracted are sufficiently rich to explain the controversial findings in the medical literature regarding the effect of a class of cholesterol drugs on Type-II Diabetes Mellitus (T2DM).
Pranjul Yadav, Michael S. Steinbach, Regina Castro, Pedro J. Caraballo, Vipin Kumar 0001, György J. Simon
IEEE BigData1
2019 Adversarial Factorization Autoencoder for Look-alike Modeling
abstract
Digital advertising is performed in multiple ways, for e.g., contextual, display-based and search-based advertising. Across these avenues, the primary goal of the advertiser is to maximize the return on investment. To realize this, the advertiser often aims to target the advertisements towards a targeted set of audience as this set has a high likelihood to respond positively towards the advertisements. One such form of tailored and personalized, targeted advertising is known as look-alike modeling, where the advertiser provides a set of seed users and expects the machine learning model to identify a new set of users such that the newly identified set is similar to the seed-set with respect to the online purchasing activity. Existing look-alike modeling techniques (i.e., similarity-based and regression-based) suffer from serious limitations due to the implicit constraints induced during modeling. In addition, the high-dimensional and sparse nature of the advertising data increases the complexity. To overcome these limitations, in this paper, we propose a novel Adversarial Factorization Autoencoder that can efficiently learn a binary mapping from sparse, high-dimensional data to a binary address space through the use of an adversarial training procedure. We demonstrate the effectiveness of our proposed approach on a dataset obtained from a real-world setting and also systematically compare the performance of our proposed approach with existing look-alike modeling baselines.
Khoa D. Doan, Pranjul Yadav, Chandan K. Reddy
CIKM2
2019 Domain adaptation in display advertising: an application for partner cold-start
abstract
Digital advertisements connects partners (sellers) to potentially interested online users. Within the digital advertisement domain, there are multiple platforms, e.g., user re-targeting and prospecting. Partners usually start with re-targeting campaigns and later employ prospecting campaigns to reach out to untapped customer base. There are two major challenges involved with prospecting. The first challenge is successful on-boarding of a new partner on the prospecting platform, referred to as partner cold-start problem. The second challenge revolves around the ability to leverage large amounts of re-targeting data for partner cold-start problem.
Karan Aggarwal, Pranjul Yadav, S. Sathiya Keerthi
RecSys2
2018 Reacting to Variations in Product Demand: An Application for Conversion Rate (CR) Prediction in Sponsored Search
abstract
In online internet advertising, machine learning models are widely used to compute the likelihood of a user engaging with product related advertisements. However, the performance of traditional machine learning models is often impacted due to variations in user and advertiser behavior. For example, search engine traffic for florists usually tends to peak around Valentine's day, Mother's day, etc. To overcome, this challenge, in this manuscript we propose three models which are able to incorporate the effects arising due to variations in product demand. The proposed models are a combination of product demand features, specialized data sampling methodologies and ensemble techniques. We demonstrate the performance of our proposed models on datasets obtained from a real-world setting. Our results show that the proposed models more accurately predict the outcome of users interactions with product related advertisements while simultaneously being robust to fluctuations in user and advertiser behaviors.
Marcelo Tallis, Pranjul Yadav
IEEE BigData2
2015 Forensic Style Analysis with Survival Trajectories
abstract
Electronic Health Records (EHRs) consists of patient information such as demographics, medications, laboratory test results, diagnosis codes and procedures. Mining EHRs could lead to improvement in patient healthcare management as EHRs contain detailed information related to disease prognosis for large patient populations. We hypothesize that a patient's condition does not deteriorate at random, the trajectories, sequences in which diseases appear in a patient, are determined by a finite number of underlying disease mechanisms. In this work, we exploit this idea by predicting a patient's risk of mortality in the context of the metabolic syndrome by assessing which of many available trajectories a patient is following and progression along this trajectory. Implementing this idea required innovative enhancements both for the study design and also for the fitting algorithm. We propose a forensic-style study design, which aligns patients on last follow-up and measures time backwards. We modify the time-dependent covariate Cox proportional hazards model to better capture coefficients of covariate that follow a particular temporal sequence, such as trajectories. Knowledge extracted from such analysis can lead to personalized treatments, thereby forming the basis for future trajectory-centered guidelines.
Pranjul Yadav, Michael S. Steinbach, Lisiane Pruinelli, Bonnie L. Westra, Connie White-Delaney, Vipin Kumar 0001, György J. Simon
ICDM1