VLDB 2026 Research / reviewers in the wild / expert
Rupesh Gupta
dblp:98/9811
· DBLP profile ↗
10ranked-venue papers
6as first author
5since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 4 first-author · 3 since 2021Databases, data management, data science and information retrieval · 7 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Retrieval for Semantic People SearchabstractThese days we have a large number of open-source embedding LLMs that can be leveraged in a bi-encoder architecture for retrieval. However, they do not perform very well when leveraged as is for people search because the (query, document) pairs encountered in people search are much more complex than the text pairs that these open-source embedding LLMs are trained on. In this paper we present our experiments, learnings and solution to the problem of retrieval for people search. Our solution involves query simplification, document simplification, fine-tuning of a 7B parameter embedding LLM, and compression of embeddings through Matryoshka learning. Rupesh Gupta, Chujie Zheng |
SIGIR | 1 |
| 2025 | Defining & Optimizing Quality of LinkedIn's Content SearchabstractMost search engines optimize for quality metrics in addition to engagement metrics. Quality metrics require an explicit definition of the relevance of a document for a query. Defining relevance is a non-trivial problem. Simple definitions such as on-topicness of a document to a query may not adequately capture all aspects of quality, such as document originality, depth and recency. These aspects can have an impact on the value provided by a search engine to searchers, which in-turn drives engagement metrics in the long-term. So, a more thoughtful definition can be very beneficial. However, a complex definition necessitates the need for a search system that has a deep understanding of language. In this paper we present our definition of relevance for content search, and the design of our content search system that enables optimization of Precision based on that definition. Ali Hooshmand, Rupesh Gupta |
SIGIR | 2 |
| 2023 | Practical Design of Performant Recommender Systems using Large-scale Linear Programming-based Global InferenceabstractSeveral key problems in web-scale recommender systems, such as optimal matching and allocation, can be formulated as large-scale linear programs (LPs) [4, 1]. These LPs take predictions from ML models such as probabilities of click, like, etc. as inputs and optimize recommendations made to users. In recent years, there has been an explosion in the research and development of large-scale recommender systems, but effective optimization of business objectives using the output of those systems remains a challenge. Although LPs can help optimize such business objectives, and algorithms for solving LPs have existed since the 1950s [5, 8], generic LP solvers cannot handle the scale of these problems. At LinkedIn, we have developed algorithms that can solve LPs of various forms with trillions of variables in a Spark-based library called "DuaLip" [7], a novel distributed solver that solves a perturbation of the LP problem at scale via gradient-based algorithms on the smooth dual of the perturbed LP. DuaLip has been deployed in production at LinkedIn and powers several very large-scale recommender systems. DuaLip is open-sourced and extensible in terms of features and algorithms. S. Sathiya Keerthi, Ayan Acharya, Borja Ocejo Elizondo, Rohan Ramanath, Rahul Mazumder, Kinjal Basu 0001, J. Kenneth Tay, Rupesh Gupta |
KDD | 10 |
| 2023 | Deep learning model for defect analysis in industry using casting images
Rupesh Gupta, Vatsala Anand, Sheifali Gupta, Deepika Koundal |
Expert Syst. Appl. | 1 |
| 2022 | Automated COVID-19 detection in chest X-ray images using fine-tuned deep learning architecturesabstractAbstract The COVID‐19 pandemic has a significant impact on human health globally. The illness is due to the presence of a virus manifesting itself in a widespread disease resulting in a high mortality rate in the whole world. According to the study, infected patients have distinct radiographic visual characteristics as well as dry cough, breathlessness, fever, and other symptoms. Although, the reverse transcription polymerase‐chain reaction (RT‐PCR) test has been used for COVID‐19 testing its reliability is very low. Therefore, computed tomography and X‐ray images have been widely used. Artificial intelligence coupled with X‐ray technologies has recently shown to be more effective in the diagnosis of this disease. With this motivation, a comparative analysis of fine‐tuned deep learning architectures has been made to speed up the detection and classification of COVID‐19 patients from other pneumonia groups. The models used for this analysis are MobileNetV2, ResNet50, InceptionV3, NASNetMobile, VGG16, Xception, InceptionResNetV2 DenseNet121, which have been fine‐tuned using a new set of layers replaced with the head of the network. This research work has carried out an analysis on two datasets. Dataset‐1 includes the images of three classes: Normal, COVID, and Pneumonia. Dataset‐2, in contrast, contains the same classes with more focus on two prominent pneumonia categories: bacterial pneumonia and viral pneumonia. The research was conducted on 959 X‐ray images (250 of Bacterial Pneumonia, 250 of Viral Pneumonia, 209 of COVID, and 250 of Normal cases). Using the confusion matrix, the required results of different models have been computed. For the first dataset, DenseNet121 has obtained a 97% accuracy, while for the second dataset, MobileNetV2 has performed best with an accuracy of 81%. Sonam Aggarwal, Sheifali Gupta, Adi Alhudhaif, Deepika Koundal, Rupesh Gupta, Kemal Polat |
Expert Syst. J. Knowl. Eng. | 5 |
| 2019 | Internal Promotion OptimizationabstractMost large Internet companies run internal promotions to cross-promote their different products and/or to educate members on how to obtain additional value from the products that they already use. This in turn drives engagement and/or revenue for the company. However, since these internal promotions can distract a member away from the product or page where these are shown, there is a non-zero cannibalization loss incurred for showing these internal promotions. This loss has to be carefully weighed against the gain from showing internal promotions. This can be a complex problem if different internal promotions optimize for different objectives. In that case, it is difficult to compare not just the gain from a conversion through an internal promotion against the loss incurred for showing that internal promotion, but also the gains from conversions through different internal promotions. Hence, we need a principled approach for deciding which internal promotion (if any) to serve to a member in each opportunity to serve an internal promotion. This approach should optimize not just for the net gain to the company, but also for the member's experience. In this paper, we discuss our approach for optimization of internal promotions at LinkedIn. In particular, we present a cost-benefit analysis of showing internal promotions, our formulation of internal promotion optimization as a constrained optimization problem, the architecture of the system for solving the optimization problem and serving internal promotions in real-time, and experimental results from online A/B tests. Rupesh Gupta, Guangde Chen, Shipeng Yu |
KDD | 1 |
| 2017 | Optimizing Email Volume For Sitewide EngagementabstractIn this paper we focus on the problem of optimizing email volume for maximizing sitewide engagement of an online social networking service. Email volume optimization approaches published in the past have proposed optimization of email volume for maximization of engagement metrics which are impacted exclusively by email; for example, the number of sessions that begin with clicks on links within emails. The impact of email on such downstream engagement metrics can be estimated easily because of the ease of attribution of such an engagement event to an email. However, this framework is limited in its view of the ecosystem of the networking service which comprises of several tools and utilities that contribute towards delivering value to members; with email being just one such utility. Thus, in this paper we depart from previous approaches by exploring and optimizing the contribution of email to this ecosystem. In particular, we present and contrast the differential impact of email on sitewide engagement metrics for various types of users. We propose a new email volume optimization approach which maximizes sitewide engagement metrics, such as the total number of active users. This is in sharp contrast to the previous approaches whose objective has been maximization of downstream engagement metrics. We present details of our prediction function for predicting the impact of emails on a user's activeness on the mobile or web application. We describe how certain approximations to this prediction function can be made for solving the volume optimization problem, and present results from online A/B tests. Rupesh Gupta, Guanfeng Liang, Rómer Rosales |
CIKM | 1 |
| 2016 | Email Volume Optimization at LinkedInabstractOnline social networking services distribute various types of messages to their members. Common types of messages include news, connection requests, membership notifications, promotions and event notifications. Such communication, if used judiciously, can provide an enormous value to members thereby keeping them engaged. However sending a message for every instance of news, connection request, or the like can result in an overwhelming number of messages in a member's mailbox. This may result in reduced effectiveness of communication if the messages are not sufficiently relevant to the member's interests. It may also result in a poor brand perception of the networking service. In this paper we discuss our strategy and experience with regard to the problem of email volume optimization at LinkedIn. In particular, we present a cost-benefit analysis of sending emails, the key factors to administer an effective volume optimization, our algorithm for volume optimization, the architecture of the supporting system and experimental results from online A/B tests. Rupesh Gupta, Guanfeng Liang, Hsiao-Ping Tseng, Ravi Kiran Holur Vijay, Rómer Rosales |
KDD | 1 |
| 2014 | Activity ranking in LinkedIn feedabstractUsers on an online social network site generate a large number of heterogeneous activities, ranging from connecting with other users, to sharing content, to updating their profiles. The set of activities within a user's network neighborhood forms a stream of updates for the user's consumption. In this paper, we report our experience with the problem of ranking activities in the LinkedIn homepage feed. In particular, we provide a taxonomy of social network activities, describe a system architecture (with a number of key components open-sourced) that supports fast iteration in model development, demonstrate a number of key factors for effective ranking, and report experimental results from extensive online bucket tests. Deepak Agarwal, Bee-Chung Chen, Rupesh Gupta, Joshua Hartman, Qi He 0002, Anand Iyer, Sumanth Kolar, Pannagadatta Shivaswamy, Ajit Singh, Liang Zhang 0021 |
KDD | 3 |
| 2013 | Visual saliency guided video compression algorithm
Rupesh Gupta, Meera Thapar Khanna, Santanu Chaudhury |
Signal Process. Image Commun. | 1 |