Sangaralingam Kajanan

dblp:122/7595 · DBLP profile ↗
← Back
9ranked-venue papers in the field
3as first author
2since 2021 · last 2022
—ORCID · none

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 8 (2 first)Information Retrieval & Web Search · 1 (1 first)
YearPublicationVenuePosition
2022 Anovos: A Scalable Feature Engineering Library
abstract
In the current era of big data, the amount of data a company can acquire is growing exponentially. However, the data are only meaningful if they are used wisely. This paper introduces Anovos, an open-source library built on top of Apache Spark. It is designed to perform efficient, end-to-end feature engineering at scale (with TBs of Data), and helps implement a systematic and procedural data pipeline with enterprise data at one end and model-ready features at the other. Besides improving the current exploratory data analysis process, we have also introduced a few key innovations in Anovos: the concept of data stability index, a single-metric indication of the stability of an independent variable in a longitudinal way, as well as Feature Explorer and Feature Mapper, powered by semantic similarity-based AI models, in order to solve the cold-start problem of building high-quality predictive features for the model training process.
Anindya Datta, Sangaralingam Kajanan, Sinuo Chen, Sourjya Sen, Ravish Ranjan
IEEE Big Data2
2022 Scalable Household Identification using Mobile Engagement Data - A Weighted Two-Mode Network Approach
abstract
It has become increasingly important for marketers to understand the customers at the household level for more meaningful and contextual targeting. In this paper, we propose a scalable graph-based approach for identifying households among mobile users, by constructing a two-mode network with mobile apps’ engagement data. The proposed solution was tested rigorously for multiple countries by benchmarking against census bureau and/or third-party data. The results in form of mean household size were found remarkably close to the public data, differing only by 3% in some countries. Additionally, we demonstrated two real-industry use-cases - first where the features were derived from the household data to predict the creditworthiness of new customers (credit risk modelling), and second where the household data was consumed directly by a telecom company to acquire and/or retain customers.
Vishnu Gowthem Thangaraj, Sangaralingam Kajanan, Nisha Verma, Sinuo Chen, Anindya Dutta
IEEE Big Data2
2019 Suspicious Location Detection Using Trajectory Analysis & Location Backfilling - A Scalable Approach
abstract
The increasing availability of GPS-embedded devices has introduced a new dimension in digital market especially location-based services. In practice, the location data is used to understand and predict consumer mobility behavior and trend for various purposes. In this paper, we propose two methodologies to first identify suspicious location from consumer location data and to infer location at both individual device and device to device level based on systematic solution. Using stay-point clustering and suspicious patterns we identified from extensive analysis, 20-30% of records with location were observed to be suspicious. After removing inaccurate location data, we have employed scalable heuristic approach to backfill records with location even for devices that originally had no available location. Our model showed the accuracy within 50 meters at 95thpercentile across different countries, including Japan, Indonesia, India, and the United States with 10-15% increase in the number of records with location and 5-10% increase in new number of devices with location.
Su Won Bae, Aravind Ravi, Sangaralingam Kajanan, Nisha Verma, Anindya Datta, Varun Chugh
IEEE BigData3
2019 High Value Customer Acquisition & Retention Modelling - A Scalable Data Mashup Approach
abstract
Identifying valuable customers as well as retaining them has become key component for any business to succeed in this competitive market. Businesses have also realized that relying solely on its own transactional data, might not be sufficient any longer, to meet the required objectives. There is a need to partner and leverage the power of big data available from the external data sources to add more value. In this paper, we are detailing the methodology of mashing up Mobilewalla's high scale mobile consumer data with one of the world's largest online food delivery company in order to revamp their retention and acquisition strategy. In this deployment, Mobilewalla has helped the client, a) to identify the new potential high impact customers from Mobilewalla ecosystem, and b) to predict the unfavorable transitions such as high impact customers getting churned or falling into low impact category. We observed that correctly identified high impact customers by Mobilewalla' customer acquisition model had 21.41% higher average revenue per user (ARPU) than the expected ARPU from high impact customers. Further, the customer retention model can help the client to spend 80% of their retention budget dollars optimally.
Sangaralingam Kajanan, Nisha Verma, Aravind Ravi, Su Won Bae, Anindya Datta
IEEE BigData1
2018 Predicting Age & Gender of Mobile Users at Scale - A Distributed Machine Learning Approach
abstract
Democratization of information access brought about by digital distribution has resulted in two contradictory phenomena: the ability to personalize consumer experience, and greater anonymity of users. These intensify when information is consumed on mobile devices, particularly because techniques to profile users on desktop web do not work on mobile smart-devices. Yet, the already large and still fast-growing field of mobile advertising require activation of audience segments against mobile advertising campaigns. Of particular importance are age and gender segments of mobile users, as these user characteristics are required for targeting a large number of ad campaigns. To date, there are no practical methodologies available in the literature that allow for accurate identification of age and gender of mobile users, at scale.In this paper, we propose a scalable machine learning approach to infer the age and gender of mobile users. We have successfully tested and implemented the gender prediction model for 8 countries, additionally the groundwork has been laid out for implementation of age inference. The output is integrated with our commercial products, furthermore, it is used as a part of custom client deliveries. We inferred gender for more than 500 million devices and there is a notable increase in number of devices with gender label (post prediction) within our current dataset. We have also inferred age for 17 million devices in Australia.
Sangaralingam Kajanan, Nisha Verma, Aravind Ravi, Anindya Datta, Varun Chugh
IEEE BigData1
2018 Predicting Consumer Level Brand Preferences Using Persistent Mobility Patterns
abstract
In the era of digital marketing, it is imperative for brands to reach target segment precisely to maximize their marketing ROI. In most cases, the target segments are created based on broad demographic and psychographic traits without any knowledge of the consumer brand preferences. In this paper, we propose a methodology to predict the individual level consumer brand preference based on the historical brand visitation patterns. We believe that this is the first attempt and a successful deployment to predict individual consumer level brand preference at large scale. In general, it is very hard to accumulate longitudinal location data of consumers. As Mobile Ad ecosystem is one of the major source in generating geospatial and temporal data which provides rich information about the mobility patterns of the mobile devices (addressed as consumers hereafter). We harnessed the power of spatio-temporal data received in our Demand Side Platform which is accumulated over a period of more than 2 years to predict brand preferences. Further, we carefully curated the brands' POI data across different countries to derive the historical brand visitation patterns of consumers. Then, this visitation pattern is used to predict propensity score for a given brand for a given consumer. We have employed a recommender system approach using distributed Alternating Least Squares - Weighted Regularization (ALS-WR) based matrix factorization to predict the brand propensities at scale. Rigorous experiments are conducted to validate the model's performance against the benchmark. Our model showed twice as much lift compared to the benchmark with 40% average recall across the brands for the Indonesian market. The full pipeline is developed and deployed in production for the countries of our business's interest such as Indonesia, Thailand, Philippines, Singapore and Malaysia in our AWS EMR1cloud environment.
Aravind Ravi, Sangaralingam Kajanan, Anindya Datta
IEEE BigData2
2016 Classification of massive mobile web log URLs for customer profiling & analytics
abstract
Many telecommunication companies today have actively started to transform the way they do business, going beyond communication infrastructure providers are repositioning themselves as data-driven service providers to create new revenue streams. In this paper, we present a novel industrial application where a scalable Big data approach combined with deep learning is used successfully to classify massive mobile web log data, to get new aggregated insights on customer web behaviors that could be applied to various industry verticals.
Kanagasabai Rajaraman, Anitha Veeramani, Shangfeng Hu, Sangaralingam Kajanan, Giuseppe Manai
IEEE BigData4
2015 Clairvoyant-push: A real-time news personalized push notifier using topic modeling and social scoring for enhanced reader engagement
abstract
Push Notification (PN) and Personalized Push Notifications (PPN) are key contemporary topics in mobile app industry today. Push notifications provide a viable content recommendation channel which complements in-app recommendation in mobile apps. There are existing algorithms for in-app content recommendation, however, the PN based recommendation systems are still under research. In this paper, we present "Clairvoyant-Push" - a novel Personalized Push Notification system based on user segmentation and social scoring. User segmentation is done by using the Latent Dirichlet Allocation (LDA) based topic modeling. Moreover, social scoring is used to assign score to each articles to filter out the quality news content for each segments. We have deployed and tested our proposed system using A/B testing framework. The results show an average of 89% lift in opening rate compared to the control group. Further, the results indicate that our system is outperforming with an opening rate of 1012% compared to the industry standard personalised push opening rate of 6-8%.
Biying Tan, Sangaralingam Kajanan, Chandra Sekhar Saripaka, Giuseppe Manai
IEEE BigData2
2014 Efficient automatic search query formulation using phrase-level analysis
abstract
Over the past decade, the volume of information available digitally over the Internet has grown enormously. Technical developments in the area of search, such as Google's Page Rank algorithm, have proved so good at serving relevant results that Internet search has become integrated into daily human activity. One can endlessly explore topics of interest simply by querying and reading through the resulting links. Yet, although search engines are well known for providing relevant results based on users' queries, users do not always receive the results they are looking for. Google's Director of Research describes clickstream evidence of frustrated users repeatedly reformulating queries and searching through page after page of results. Given the general quality of search engine results, one must consider the possibility that the frustrated user's query is not effective; that is, it does not describe the essence of the user's interest. Indeed, extensive research into human search behavior has found that humans are not very effective at formulating good search queries that describe what they are interested in. Ideally, the user should simply point to a portion of text that sparked the user's interest, and a system should automatically formulate a search query that captures the essence of the text. In this paper, we describe an implemented system that provides this capability. We first describe how our work differs from existing work in automatic query formulation, and propose a new method for improved quantification of the relevance of candidate search terms drawn from input text using phrase‐level analysis. We then propose an implementable method designed to provide relevant queries based on a user's text input. We demonstrate the quality of our results and performance of our system through experimental studies. Our results demonstrate that our system produces relevant search terms with roughly two‐thirds precision and recall compared to search terms selected by experts, and that typical users find significantly more relevant results (31% more relevant) more quickly (64% faster) using our system than self‐formulated search queries. Further, we show that our implementation can scale to request loads of up to 10 requests per second within current online responsiveness expectations (<2‐second response times at the highest loads tested).
Sangaralingam Kajanan, Yang Bao 0001, Anindya Datta, Debra E. VanderMeer, Kaushik Dutta
J. Assoc. Inf. Sci. Technol.1