EDBT 2026 Demo / reviewers in the wild / expert
Anindya Datta
dblp:d/AnindyaDatta
· DBLP profile ↗
23ranked-venue papers in the field
13as first author
1since 2021 · last 2022
0000-0001-7966-2944ORCID · corroborated
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 15 (12 first)Big Data, Cloud & Distributed Data Systems · 5 (1 first)Information Retrieval & Web Search · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Anovos: A Scalable Feature Engineering LibraryabstractIn the current era of big data, the amount of data a company can acquire is growing exponentially. However, the data are only meaningful if they are used wisely. This paper introduces Anovos, an open-source library built on top of Apache Spark. It is designed to perform efficient, end-to-end feature engineering at scale (with TBs of Data), and helps implement a systematic and procedural data pipeline with enterprise data at one end and model-ready features at the other. Besides improving the current exploratory data analysis process, we have also introduced a few key innovations in Anovos: the concept of data stability index, a single-metric indication of the stability of an independent variable in a longitudinal way, as well as Feature Explorer and Feature Mapper, powered by semantic similarity-based AI models, in order to solve the cold-start problem of building high-quality predictive features for the model training process. Anindya Datta, Sangaralingam Kajanan, Sinuo Chen, Sourjya Sen, Ravish Ranjan |
IEEE Big Data | 1 |
| 2019 | Suspicious Location Detection Using Trajectory Analysis & Location Backfilling - A Scalable ApproachabstractThe increasing availability of GPS-embedded devices has introduced a new dimension in digital market especially location-based services. In practice, the location data is used to understand and predict consumer mobility behavior and trend for various purposes. In this paper, we propose two methodologies to first identify suspicious location from consumer location data and to infer location at both individual device and device to device level based on systematic solution. Using stay-point clustering and suspicious patterns we identified from extensive analysis, 20-30% of records with location were observed to be suspicious. After removing inaccurate location data, we have employed scalable heuristic approach to backfill records with location even for devices that originally had no available location. Our model showed the accuracy within 50 meters at 95thpercentile across different countries, including Japan, Indonesia, India, and the United States with 10-15% increase in the number of records with location and 5-10% increase in new number of devices with location. Su Won Bae, Aravind Ravi, Sangaralingam Kajanan, Nisha Verma, Anindya Datta, Varun Chugh |
IEEE BigData | 5 |
| 2019 | High Value Customer Acquisition & Retention Modelling - A Scalable Data Mashup ApproachabstractIdentifying valuable customers as well as retaining them has become key component for any business to succeed in this competitive market. Businesses have also realized that relying solely on its own transactional data, might not be sufficient any longer, to meet the required objectives. There is a need to partner and leverage the power of big data available from the external data sources to add more value. In this paper, we are detailing the methodology of mashing up Mobilewalla's high scale mobile consumer data with one of the world's largest online food delivery company in order to revamp their retention and acquisition strategy. In this deployment, Mobilewalla has helped the client, a) to identify the new potential high impact customers from Mobilewalla ecosystem, and b) to predict the unfavorable transitions such as high impact customers getting churned or falling into low impact category. We observed that correctly identified high impact customers by Mobilewalla' customer acquisition model had 21.41% higher average revenue per user (ARPU) than the expected ARPU from high impact customers. Further, the customer retention model can help the client to spend 80% of their retention budget dollars optimally. Sangaralingam Kajanan, Nisha Verma, Aravind Ravi, Su Won Bae, Anindya Datta |
IEEE BigData | 5 |
| 2018 | Predicting Age & Gender of Mobile Users at Scale - A Distributed Machine Learning ApproachabstractDemocratization of information access brought about by digital distribution has resulted in two contradictory phenomena: the ability to personalize consumer experience, and greater anonymity of users. These intensify when information is consumed on mobile devices, particularly because techniques to profile users on desktop web do not work on mobile smart-devices. Yet, the already large and still fast-growing field of mobile advertising require activation of audience segments against mobile advertising campaigns. Of particular importance are age and gender segments of mobile users, as these user characteristics are required for targeting a large number of ad campaigns. To date, there are no practical methodologies available in the literature that allow for accurate identification of age and gender of mobile users, at scale.In this paper, we propose a scalable machine learning approach to infer the age and gender of mobile users. We have successfully tested and implemented the gender prediction model for 8 countries, additionally the groundwork has been laid out for implementation of age inference. The output is integrated with our commercial products, furthermore, it is used as a part of custom client deliveries. We inferred gender for more than 500 million devices and there is a notable increase in number of devices with gender label (post prediction) within our current dataset. We have also inferred age for 17 million devices in Australia. Sangaralingam Kajanan, Nisha Verma, Aravind Ravi, Anindya Datta, Varun Chugh |
IEEE BigData | 4 |
| 2018 | Predicting Consumer Level Brand Preferences Using Persistent Mobility PatternsabstractIn the era of digital marketing, it is imperative for brands to reach target segment precisely to maximize their marketing ROI. In most cases, the target segments are created based on broad demographic and psychographic traits without any knowledge of the consumer brand preferences. In this paper, we propose a methodology to predict the individual level consumer brand preference based on the historical brand visitation patterns. We believe that this is the first attempt and a successful deployment to predict individual consumer level brand preference at large scale. In general, it is very hard to accumulate longitudinal location data of consumers. As Mobile Ad ecosystem is one of the major source in generating geospatial and temporal data which provides rich information about the mobility patterns of the mobile devices (addressed as consumers hereafter). We harnessed the power of spatio-temporal data received in our Demand Side Platform which is accumulated over a period of more than 2 years to predict brand preferences. Further, we carefully curated the brands' POI data across different countries to derive the historical brand visitation patterns of consumers. Then, this visitation pattern is used to predict propensity score for a given brand for a given consumer. We have employed a recommender system approach using distributed Alternating Least Squares - Weighted Regularization (ALS-WR) based matrix factorization to predict the brand propensities at scale. Rigorous experiments are conducted to validate the model's performance against the benchmark. Our model showed twice as much lift compared to the benchmark with 40% average recall across the brands for the Indonesian market. The full pipeline is developed and deployed in production for the countries of our business's interest such as Indonesia, Thailand, Philippines, Singapore and Malaysia in our AWS EMR1cloud environment. Aravind Ravi, Sangaralingam Kajanan, Anindya Datta |
IEEE BigData | 3 |
| 2018 | Identifying functional aspects from user reviews for functionality-based mobile app recommendationabstractThe explosive growth of mobile apps makes it difficult for users to find their needed apps in a crowded market. An effective mechanism that provides high quality app recommendations becomes necessary. However, existing recommendation techniques tend to recommend similar items but fail to consider users’ functional requirements, making them not effective in the app domain. In this article, we propose a recommendation architecture that can generate app recommendations at the functionality level. We address the redundant recommendation problem in the app domain by highlighting users’ functional requirements, an element that has received scant attention from existing recommendation research. Another main feature of our work is extracting app functionalities from textural user reviews for recommendation. We also propose an effective approach for functionality extraction. Experiments conducted on a real‐world dataset show that our proposed AppRank method outperforms other commonly used recommendation methods. In particular, it doubles the recall value of the second best method under an extremely sparse setting, increases the overall ranking accuracy of the second best method by 14.27%, and retains a high diversity of 0.99. Xiaoying Xu, Kaushik Dutta, Anindya Datta, Chunmian Ge |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2014 | Efficient automatic search query formulation using phrase-level analysisabstractOver the past decade, the volume of information available digitally over the Internet has grown enormously. Technical developments in the area of search, such as Google's Page Rank algorithm, have proved so good at serving relevant results that Internet search has become integrated into daily human activity. One can endlessly explore topics of interest simply by querying and reading through the resulting links. Yet, although search engines are well known for providing relevant results based on users' queries, users do not always receive the results they are looking for. Google's Director of Research describes clickstream evidence of frustrated users repeatedly reformulating queries and searching through page after page of results. Given the general quality of search engine results, one must consider the possibility that the frustrated user's query is not effective; that is, it does not describe the essence of the user's interest. Indeed, extensive research into human search behavior has found that humans are not very effective at formulating good search queries that describe what they are interested in. Ideally, the user should simply point to a portion of text that sparked the user's interest, and a system should automatically formulate a search query that captures the essence of the text. In this paper, we describe an implemented system that provides this capability. We first describe how our work differs from existing work in automatic query formulation, and propose a new method for improved quantification of the relevance of candidate search terms drawn from input text using phrase‐level analysis. We then propose an implementable method designed to provide relevant queries based on a user's text input. We demonstrate the quality of our results and performance of our system through experimental studies. Our results demonstrate that our system produces relevant search terms with roughly two‐thirds precision and recall compared to search terms selected by experts, and that typical users find significantly more relevant results (31% more relevant) more quickly (64% faster) using our system than self‐formulated search queries. Further, we show that our implementation can scale to request loads of up to 10 requests per second within current online responsiveness expectations (<2‐second response times at the highest loads tested). Sangaralingam Kajanan, Yang Bao 0001, Anindya Datta, Debra E. VanderMeer, Kaushik Dutta |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2013 | A partially supervised cross-collection topic model for cross-domain text classificationabstractCross-domain text classification aims to automatically train a precise text classifier for a target domain by using labelled text data from a related source domain. To this end, one of the most promising ideas is to induce a new feature representation so that the distributional difference between domains can be reduced and a more accurate classifier can be learned in this new feature space. However, most existing methods do not explore the duality of the marginal distribution of examples and the conditional distribution of class labels given labeled training examples in the source domain. Besides, few previous works attempt to explicitly distinguish the domain-independent and domain-specific latent features and align the domain-specific features to further improve the cross-domain learning. In this paper, we propose a model called Partially Supervised Cross-Collection LDA topic model (PSCCLDA) for cross-domain learning with the purpose of addressing these two issues in a unified way. Experimental results on nine datasets show that our model outperforms two standard classifiers and four state-of-the-art methods, which demonstrates the effectiveness of our proposed model. Yang Bao 0001, Nigel Collier, Anindya Datta |
CIKM | 3 |
| 2013 | Building a Scalable Database-Driven Reverse DictionaryabstractIn this paper, we describe the design and implementation of a reverse dictionary. Unlike a traditional forward dictionary, which maps from words to their definitions, a reverse dictionary takes a user input phrase describing the desired concept, and returns a set of candidate words that satisfy the input phrase. This work has significant application not only for the general public, particularly those who work closely with words, but also in the general field of conceptual search. We present a set of algorithms and the results of a set of experiments showing the retrieval accuracy of our methods and the runtime response time performance of our implementation. Our experimental results show that our approach can provide significant improvements in performance scale without sacrificing the quality of the result. Our experiments comparing the quality of our approach to that of currently available reverse dictionaries show that of our approach can provide significantly higher quality over either of the other currently available implementations. Anindya Datta, Debra E. VanderMeer, Kaushik Dutta |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2004 | Proxy-based acceleration of dynamically generated content on the world wide web: An approach and implementationabstractAs Internet traffic continues to grow and websites become increasingly complex, performance and scalability are major issues for websites. Websites are increasingly relying on dynamic content generation applications to provide website visitors with dynamic, interactive, and personalized experiences. However, dynamic content generation comes at a cost---each request requires computation as well as communication across multiple components.To address these issues, various dynamic content caching approaches have been proposed. Proxy-based caching approaches store content at various locations outside the site infrastructure and can improve website performance by reducing content generation delays, firewall processing delays, and bandwidth requirements. However, existing proxy-based caching approaches either (a) cache at the page level, which does not guarantee that correct pages are served and provides very limited reusability, or (b) cache at the fragment level, which is associated with several design-level and runtime scalability issues. To address these issues, several back-end caching approaches have been proposed, including query result caching and fragment level caching. While back-end approaches guarantee the correctness of results and offer the advantages of fine-grained caching, they neither address firewall delays nor reduce bandwidth requirements.In this article, we present an approach and an implementation of a dynamic proxy caching technique which combines the benefits of both proxy-based and back-end caching approaches, yet does not suffer from their above-mentioned limitations. Our dynamic proxy caching technique allows granular, proxy-based caching in highly dynamic scenarios, accessible outside the site infrastructure. We present two possible configurations for our dynamic proxy caching technique: (1) a reverse proxy configuration, and (2) a forward proxy configuration. Analysis of the performance of our approach indicates that it is capable of providing significant reductions in bandwidth. We have deployed our proposed dynamic proxy caching technique at a major financial institution. The results of this implementation indicate that our technique is capable of providing up to 3x reductions in bandwidth and response times in real-world dynamic Web applications when compared to existing caching solutions. Anindya Datta, Kaushik Dutta, Helen M. Thomas, Debra E. VanderMeer, Krithi Ramamritham |
ACM Trans. Database Syst. | 1 |
| 2002 | Proxy-based acceleration of dynamically generated content on the world wide web: an approach and implementationabstractAs Internet traffic continues to grow and web sites become increasingly complex, performance and scalability are major issues for web sites. Web sites are increasingly relying on dynamic content generation applications to provide web site visitors with dynamic, interactive, and personalized experiences. However, dynamic content generation comes at a cost --- each request requires computation as well as communication across multiple components.To address these issues, various dynamic content caching approaches have been proposed. Proxy-based caching approaches store content at various locations outside the site infrastructure and can improve Web site performance by reducing content generation delays, firewall processing delays, and bandwidth requirements. However, existing proxy-based caching approaches either (a) cache at the page level, which does not guarantee that correct pages are served and provides very limited reusability, or (b) cache at the fragment level, which requires the use of pre-defined page layouts. To address these issues, several back end caching approaches have been proposed, including query result caching and fragment level caching. While back end approaches guarantee the correctness of results and offer the advantages of fine-grained caching, they neither address firewall delays nor reduce bandwidth requirements.In this paper, we present an approach and an implementation of a dynamic proxy caching technique which combines the benefits of both proxy-based and back end caching approaches, yet does not suffer from their above-mentioned limitations. Our dynamic proxy caching technique allows granular, proxy-based caching where both the content and layout can be dynamic. Our analysis of the performance of our approach indicates that it is capable of providing significant reductions in bandwidth. We have also deployed our proposed dynamic proxy caching technique at a major financial institution. The results of this implementation indicate that our technique is capable of providing order-of-magnitude reductions in bandwidth and response times in real-world dynamic Web applications. Anindya Datta, Kaushik Dutta, Helen M. Thomas, Debra E. VanderMeer, Suresha, Krithi Ramamritham |
SIGMOD Conference | 1 |
| 2002 | A Study of Concurrency Control in Real-Time, Active Database SystemsabstractReal-time active database systems (RTADBSs) have attracted a considerable amount of research attention in the past and a number of important applications have been identified for such systems, such as telecommunications network management, automated air traffic control, automated financial trading, process control and military command and control systems. In spite of the recognized importance of this area, very little research has been devoted to exploring the dynamics of transaction processing in RTADBSs. Concurrency control (CC) constitutes an integral part of any transaction processing strategy and, thus, deserves special attention. We study CC strategies in RTADBSs and postulate a number of CC algorithms. These algorithms exploit the special needs and features of RTADBSs and are shown to deliver substantially superior performance to conventional real-time CC algorithms. Anindya Datta, Sang Hyuk Son |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2002 | Parallel Star Join + DataIndexes: Efficient Query Processing in Data Warehouses and OLAPabstractOn-line analytical processing (OLAP) refers to the technologies that allow users to efficiently retrieve data from the data warehouse for decision-support purposes. Data warehouses tend to be extremely large, it is quite possible for a data warehouse to be hundreds of gigabytes to terabytes in size (Chauduri and Dayal, 1997). Queries tend to be complex and ad hoc, often requiring computationally expensive operations such as joins and aggregation. Given this, we are interested in developing strategies for improving query processing in data warehouses by exploring the applicability of parallel processing techniques. In particular, we exploit the natural partitionability of a star schema and render it even more efficient by applying DataIndexes-a storage structure that serves both as an index as well as data and lends itself naturally to vertical partitioning of the data. DataIndexes are derived from the various special purpose access mechanisms currently supported in commercial OLAP products. Specifically, we propose a declustering strategy which incorporates both task and data partitioning and present the Parallel Star Join (PSJ) Algorithm, which provides a means to perform a star join in parallel using efficient operations involving only rowsets and projection columns. We compare the performance of the PSJ Algorithm with two parallel query processing strategies. The first is a parallel join strategy utilizing the Bitmap Join Index (BJI), arguably the state-of-the-art OLAP join structure in use today. For the second strategy we choose a well-known parallel join algorithm, namely the pipelined hash algorithm. To assist in the performance comparison, we first develop a cost model of the disk access and transmission costs for all three approaches. Anindya Datta, Debra E. VanderMeer, Krithi Ramamritham |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2001 | Dynamic Content Acceleration: A Caching Solution to Enable Scalable Dynamic Web Page GenerationabstractNo abstract available. Anindya Datta, Kaushik Dutta, Krithi Ramamritham, Helen M. Thomas, Debra E. VanderMeer |
SIGMOD Conference | 1 |
| 2001 | A Comparative Study of Alternative Middle Tier Caching Solutions to Support Dynamic Web Content Acceleration
Anindya Datta, Kaushik Dutta, Helen M. Thomas, Debra E. VanderMeer, Krithi Ramamritham, Dan Fishman |
VLDB | 1 |
| 2001 | An architecture to support scalable online personalization on the Web
Anindya Datta, Kaushik Dutta, Debra E. VanderMeer, Krithi Ramamritham, Shamkant B. Navathe |
VLDB J. | 1 |
| 2000 | Demonstration: Enabling Scalable Online Personalization on the Web
Kaushik Dutta, Anindya Datta, Debra E. VanderMeer, Krithi Ramamritham, Helen M. Thomas |
VLDB | 2 |
| 1999 | Curio: A Novel Solution for Efficient Storage and Indexing in Data Warehouses
Anindya Datta, Krithi Ramamritham, Helen M. Thomas |
VLDB | 1 |
| 1999 | A Novel Index Supporting High Volume Data Warehouse Insertion
Chris Jermaine, Anindya Datta, Edward Omiecinski |
VLDB | 2 |
| 1999 | Broadcast Protocols to Support Efficient Retrieval from Databases by Mobile UsersabstractMobile computing has the potential for managing information globally. Data management issues in mobile computing have received some attention in recent times, and the design of adaptive braodcast protocols has been posed as an important probllem. Such protocols are employed by database servers to decide on the content of bbroadcasts dynamically, in response to client mobility and demand patterns. In this paper we design such protocols and also propose efficient retrieval strategies that may be employed by clients to download information from broadcasts. The goal is to design cooperative strategies between server and client to provide access to information in such a way as to minimize energy expenditure by clients. We evaluate the performance of our protocols both analytically and through simulation. Anindya Datta, Debra E. VanderMeer, Aslihan Celik, Vijay Kumar 0002 |
ACM Trans. Database Syst. | 1 |
| 1997 | Adaptive Broadcast Protocols to Support Power Conservant Retrieval by Mobile UsersabstractMobile computing has the potential for managing information globally. Data management issues in mobile computing have received some attention in recent times, and the design of adaptive broadcast protocols has been posed as an important problem. Such protocols are employed by database servers to decide on the content of broadcasts dynamically, in response to client mobility and demand patterns. In this paper we design such protocols and also propose efficient retrieval strategies that may be employed by clients to download information from broadcasts. The goal is to design cooperative strategies between server and client to provide access to information in such a way as to minimize energy expenditure by clients. We evaluate the performance of our protocols analytically. Anindya Datta, Aslihan Celik, Jeong Geun Kim, Debra E. VanderMeer, Vijay Kumar 0002 |
ICDE | 1 |
| 1997 | Providing Real-Time Response, State Recency and Temporal Consistency in Databases for Rapidly Changing Environments
Anindya Datta, Igor R. Viguier |
Inf. Syst. | 1 |
| 1996 | Multiclass Transaction Scheduling and Overload Management in Firm Real-Time Database Systems
Anindya Datta, Sarit Mukherjee, Prabhudev Konana, Igor R. Viguier, Akhilesh Bajaj |
Inf. Syst. | 1 |