Sinuo Chen

dblp:339/7788 · DBLP profile ↗
← Back
2ranked-venue papers in the field
0as first author
2since 2021 · last 2022
—ORCID · none

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 2
YearPublicationVenuePosition
2022 Anovos: A Scalable Feature Engineering Library
abstract
In the current era of big data, the amount of data a company can acquire is growing exponentially. However, the data are only meaningful if they are used wisely. This paper introduces Anovos, an open-source library built on top of Apache Spark. It is designed to perform efficient, end-to-end feature engineering at scale (with TBs of Data), and helps implement a systematic and procedural data pipeline with enterprise data at one end and model-ready features at the other. Besides improving the current exploratory data analysis process, we have also introduced a few key innovations in Anovos: the concept of data stability index, a single-metric indication of the stability of an independent variable in a longitudinal way, as well as Feature Explorer and Feature Mapper, powered by semantic similarity-based AI models, in order to solve the cold-start problem of building high-quality predictive features for the model training process.
Anindya Datta, Sangaralingam Kajanan, Sinuo Chen, Sourjya Sen, Ravish Ranjan
IEEE Big Data3
2022 Scalable Household Identification using Mobile Engagement Data - A Weighted Two-Mode Network Approach
abstract
It has become increasingly important for marketers to understand the customers at the household level for more meaningful and contextual targeting. In this paper, we propose a scalable graph-based approach for identifying households among mobile users, by constructing a two-mode network with mobile apps’ engagement data. The proposed solution was tested rigorously for multiple countries by benchmarking against census bureau and/or third-party data. The results in form of mean household size were found remarkably close to the public data, differing only by 3% in some countries. Additionally, we demonstrated two real-industry use-cases - first where the features were derived from the household data to predict the creditworthiness of new customers (credit risk modelling), and second where the household data was consumed directly by a telecom company to acquire and/or retain customers.
Vishnu Gowthem Thangaraj, Sangaralingam Kajanan, Nisha Verma, Sinuo Chen, Anindya Dutta
IEEE Big Data4