EDBT 2026 Demo / reviewers in the wild / expert
Bin Liu 0045
dblp:35/837-45
· DBLP profile ↗
30ranked-venue papers
9as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 21 · 8 first-author · 7 since 2021Artificial intelligence and machine learning · 12 · 5 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 1 since 2021Security and privacy · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorComputer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multimodal chain-of-thought reasoning with large language models to protect children from age-inappropriate apps
Chuanbo Hu, Bin Liu 0045, Minglei Yin, Yilu Zhou, Xin Li 0005 |
Inf. Manag. | 2 |
| 2025 | Discovering Time-aware Hidden Dependencies with Personalized Graphical Structure in Electronic Health RecordsabstractOver the past decade, significant advancements in mining electronic health records (EHRs) have enabled a broad range of decision-support applications and offered an unprecedented capacity for predicting critical events such as disease prognosis and mortality in healthcare. Despite the availability of comprehensive coding systems in EHRs (e.g., ICD-9), which are designed to record diverse information on diseases, procedures, and medications over time, the complex and dynamic dependencies among the recorded data are usually not captured. This limitation often hinders the contextual understanding of medical observations for effective EHR representation learning. Therefore, there is a compelling need to discover a hidden “EHR graph” that represents the medical relationship between the observed features according to a patient’s history. These hidden graphs consisting of the medical codes from the same visits can offer a comprehensive insight derived from disease-to-disease, disease-to-drug, and drug-to-drug dependencies. However, it is still unclear how to address the challenge that the dependencies may vary from patient to patient, and they can dynamically evolve from one visit to another. To this end, we propose Time-aware Personalized Graph Transformer (TPGT), a novel attention-based time-aware hidden graph model, that captures the personalized graphical structures among observed medical codes and summarizes the temporal code dependencies over time to improve patient representation for outcome prediction. Built upon an intra-visit and an inter-visit dual-attention mechanism to model patients’ EHR graphs, our model offers an interpretability of what diagnosis or medication in a patient’s history can interact, and how those interactions may change over time. We conduct extensive experiments on two real-world EHR datasets for different healthcare predictive tasks: acute kidney injury (AKI) prediction and ICU mortality prediction. The experimental results demonstrate a significant performance improvement of the proposed model over baselines through multi-aspect quantitative evaluation. Furthermore, we perform various qualitative studies to validate the interpretability of the model which highlights the application of the proposed method in the context of personalized medicine. Arya Hadizadeh Moghaddam, Mohsen Nayebi Kerdabadi, Bin Liu 0045, Zijun Yao 0001 |
ACM Trans. Knowl. Discov. Data | 3 |
| 2024 | Knowledge-prompted ChatGPT: Enhancing drug trafficking detection on social media
Chuanbo Hu, Bin Liu 0045, Xin Li 0005, Yanfang Ye 0001, Minglei Yin |
Inf. Manag. | 2 |
| 2023 | Contrastive Learning of Temporal Distinctiveness for Survival Analysis in Electronic Health RecordsabstractSurvival analysis plays a crucial role in many healthcare decisions, where the risk prediction for the events of interest can support an informative outlook for a patient's medical journey. Given the existence of data censoring, an effective way of survival analysis is to enforce the pairwise temporal concordance between censored and observed data, aiming to utilize the time interval before censoring as partially observed time-to-event labels for supervised learning. Although existing studies mostly employed ranking methods to pursue an ordering objective, contrastive methods which learn a discriminative embedding by having data contrast against each other, have not been explored thoroughly for survival analysis. Therefore, in this paper, we propose a novel Ontology-aware Temporality-based Contrastive Survival (OTCSurv) analysis framework that utilizes survival durations from both censored and observed data to define temporal distinctiveness and construct negative sample pairs with adjustable hardness for contrastive learning. Specifically, we first use an ontological encoder and a sequential self-attention encoder to represent the longitudinal EHR data with rich contexts. Second, we design a temporal contrastive loss to capture varying survival durations in a supervised setting through a hardness-aware negative sampling mechanism. Last, we incorporate the contrastive task into the time-to-event predictive task with multiple loss components. We conduct extensive experiments using a large EHR dataset to forecast the risk of hospitalized patients who are in danger of developing acute kidney injury (AKI), a critical and urgent medical condition. The effectiveness and explainability of the proposed model are validated through comprehensive quantitative and qualitative studies. Mohsen Nayebi Kerdabadi, Arya Hadizadeh Moghaddam, Bin Liu 0045, Zijun Yao 0001 |
CIKM | 3 |
| 2023 | Fine-grained classification of drug trafficking based on Instagram hashtags
Chuanbo Hu, Bin Liu 0045, Yanfang Ye 0001, Xin Li 0005 |
Decis. Support Syst. | 2 |
| 2023 | Towards Human-Machine Recognition Alignment: An Adversarilly Robust Multimodal Retrieval Hashing FrameworkabstractThe multimodality nature of web data has necessitated complex multimodal information retrieval for a wide range of web applications. Deep neural networks (DNNs) have been widely employed to extract semantic features from raw samples to improve retrieval accuracy. In addition, hashing is widely used to improve computational and storage efficiency. As such, deep hashing frameworks have been applied for multimodal retrieval tasks. However, there is still a great recognitive gap between primate brain structure-inspired DNNs and humans. On computer vision tasks, well-crafted DNN models can be easily defeated by invisible small attacks, and this phenomenon indicates a large recognition gap between DNN models and humans. Recently, adversarial defense methods have been shown to improve the human–machine recognition alignment in several classification tasks. However, the robustness problem on the retrieval tasks, especially on the deep hashing-based multimodal retrieval models, is still not well studied. Therefore, in this article, we present an adversarially robust training mechanism to improve model robustness for the purpose of human–machine recognition alignment on retrieval tasks. Through extensive experimental results on several social multimodal retrieval benchmarks, we show that the robust training hashing framework proposed can mitigate the recognition gap on retrieval tasks. Our study highlights the necessity of robustness enhancement on deep hashing models. Xingwei Zhang, Xiaolong Zheng 0001, Bin Liu 0045, Xiao Wang 0002, Wenji Mao, Daniel Dajun Zeng, Fei-Yue Wang 0001 |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2023 | Ontology-aware Prescription Recommendation in Treatment Pathways Using Multi-evidence Healthcare DataabstractFor care of chronic diseases (e.g., depression, diabetes, hypertension), it is critical to identify effective treatment pathways that aim to promptly update the medication following the change of patient state and disease progression. This task is challenging because the optimal treatment pathway for each patient needs to be personalized due to the significant heterogeneity among individuals. Therefore, it is naturally promising to investigate how to use the abundant electronic health records to recommend effective and safe prescriptions. However, prescription recommendation needs to consider multiple aspects of life-critical evidence, such as the information relevance in terms of medical concepts, the health condition in terms of diagnosis history, and the further constraint in terms of side information (e.g., patient demographics and drug side effects). To this end, in this article, we propose a novel prescription recommendation framework named OntoPath to predict the next drug in disease treatment pathways, by building an ontology-aware hierarchical-attention model that integrates multiple medical evidence from domain knowledge guidance, medical history profiling, and side information utilization. Specifically, our method can be characterized from three aspects: (1) by incorporating the longitudinal diagnosis history, we enrich the profiling of patients in terms of comprehensive health conditions, which can largely influence a drug’s outcome on individual patients; (2) using the hierarchical disease and drug ontology structures, we are able to model the domain-specific relevance between patients and drugs at multiple levels of granularity and achieve in-depth collaborative filtering; (3) we introduce a pre-training stage to enhance the discriminativeness of network representations, which helps us obtain a premium model initialization to further boost the final recommendation training. We perform extensive experiments on a large-scale depression cohort with over 37,000 patients from a real-world medical claims database. The quantitative and qualitative results demonstrate the effectiveness of OntoPath through the consistent outperformance over state-of-the-art prescription recommendation baselines and the interpretation of model mechanism in case studies. Zijun Yao 0001, Bin Liu 0045, Fei Wang 0001, Daby M. Sow, Ying Li 0053 |
ACM Trans. Inf. Syst. | 2 |
| 2021 | Detection of Illicit Drug Trafficking Events on Instagram: A Deep Multimodal Multilabel Learning ApproachabstractSocial media such as Instagram and Twitter have become important platforms for marketing and selling illicit drugs. Detection of online illicit drug trafficking has become critical to combat the online trade of illicit drugs. However, the legal status often varies spatially and temporally; even for the same drug, federal and state legislation can have different regulations about its legality. Meanwhile, more drug trafficking events are disguised as a novel form of advertising - commenting leading to information heterogeneity. Accordingly, accurate detection of illicit drug trafficking events (IDTEs) from social media has become even more challenging. In this work, we conduct the first systematic study on fine-grained detection of IDTEs on Instagram. We propose to take a deep multimodal multilabel learning (DMML) approach to detect IDTEs and demonstrate its effectiveness on a newly constructed dataset called multimodal IDTE (MM-IDTE). Specifically, our model takes text and image data as the input and combines multimodal information to predict multiple labels of illicit drugs. Inspired by the success of BERT, we have developed a self-supervised multimodal bidirectional transformer by jointly fine-tuning pretrained text and image encoders. We have constructed a large-scale dataset MM-IDTE with manually annotated multiple drug labels to support fine-grained detection of illicit drugs. Extensive experimental results on the MM-IDTE dataset show that the proposed DMML methodology can accurately detect IDTEs even in the presence of special characters and style changes attempting to evade detection. Chuanbo Hu, Minglei Yin, Bin Liu 0045, Xin Li 0005, Yanfang Ye 0001 |
CIKM | 3 |
| 2021 | Data Poisoning Attacks to Deep Learning Based Recommender Systems
Hai Huang 0014, Jiaming Mu, Neil Zhenqiang Gong, Qi Li 0002, Bin Liu 0045 |
NDSS | 5 |
| 2021 | Identifying Illicit Drug Dealers on Instagram with Large-scale Multimodal Data FusionabstractIllicit drug trafficking via social media sites such as Instagram have become a severe problem, thus drawing a great deal of attention from law enforcement and public health agencies. How to identify illicit drug dealers from social media data has remained a technical challenge for the following reasons. On the one hand, the available data are limited because of privacy concerns with crawling social media sites; on the other hand, the diversity of drug dealing patterns makes it difficult to reliably distinguish drug dealers from common drug users. Unlike existing methods that focus on posting-based detection, we propose to tackle the problem of illicit drug dealer identification by constructing a large-scale multimodal dataset named Identifying Drug Dealers on Instagram (IDDIG). Nearly 4,000 user accounts, of which more than 1,400 are drug dealers, have been collected from Instagram with multiple data sources including post comments, post images, homepage bio, and homepage images. We then design a quadruple-based multimodal fusion method to combine the multiple data sources associated with each user account for drug dealer identification. Experimental results on the constructed IDDIG dataset demonstrate the effectiveness of the proposed method in identifying drug dealers (almost 95% accuracy). Moreover, we have developed a hashtag-based community detection technique for discovering evolving patterns, especially those related to geography and drug types. Chuanbo Hu, Minglei Yin, Bin Liu 0045, Xin Li 0005, Yanfang Ye 0001 |
ACM Trans. Intell. Syst. Technol. | 3 |
| 2020 | Predicting Type 1 Diabetes Onset using Novel Survival Analysis with Biomarker Ontology
Ying Li 0053, Bin Liu 0045, Vibha Anand, Markus Lundgren, Kenney Ng, Marian Rewers, Riitta Veijola, Mohamed F. Ghalwash |
AMIA | 2 |
| 2020 | Prognostication and Outcome-specific Risk Factor Identification for Diabetes Care via Private-shared Multi-task LearningabstractDiabetes is a chronic diseases that affects nearly half a billion people around the globe, and is almost always associated with a number of complications, including kidney failure, blindness, stroke, and heart attack. An important step towards improved diabetes care is to accurately predict the risk of diabetes complications and to identify the corresponding risk factors associated with the onset of each complication. In this paper, we study the problem of risk prediction and outcome-specific risk factor identification from readily available patient medical record data. We adopt a private-shared multi-task learning (MTL) model, which jointly models multiple complications with each task corresponding to the risk modeling of one complication. The MTL formulation not only boosts prediction performance but also enables identification of outcome-specific risk factors. Specifically, we decompose the coefficient matrix, in which each column (vector) corresponds to the coefficient of one complication risk model, into a shared component and an outcome-specific private component. The shared component is assumed to be low-rank to capture the relationships among complications in terms of overall diabetes health condition. The private component is assumed to be non-overlapping and sparse so that they are discriminative among the different complication outcomes. Further, the shared component and the private component for the same complication are assumed to be orthogonal. Extensive experimental results on a type 2 diabetes cohort extracted from a large electronic medical claims database show that the proposed method outperforms baseline models by a significant margin. Also the identified outcome-specific risk factors provide meaningful clinical insights. The results demonstrate that simultaneously modeling multiple risks through MTL not only improves prediction performance but also enables identification of outcome-specific risk factors. Bin Liu 0045, Ying Li 0053, Kenney Ng |
IEEE BigData | 1 |
| 2020 | Complication Risk Profiling in Diabetes Care: A Bayesian Multi-Task and Feature Relationship Learning ApproachabstractDiabetes mellitus, commonly known as diabetes, is a chronic disease that often results in multiple complications. Risk prediction of diabetes complications is critical for healthcare professionals to design personalized treatment plans for patients in diabetes care for improved outcomes. In this paper, focusing on Type 2 diabetes mellitus (T2DM), we study the risk of developing complications after the initial T2DM diagnosis from longitudinal patient records. We propose a novel multi-task learning approach to simultaneously model multiple complications where each task corresponds to the risk modeling of one complication. Specifically, the proposed method strategically captures the relationships (1) between the risks of multiple T2DM complications, (2) between different risk factors, and (3) between the risk factor selection patterns, which assumes similar complications have similar contributing risk factors. The method uses coefficient shrinkage to identify an informative subset of risk factors from high-dimensional data, and uses a hierarchical Bayesian framework to allow domain knowledge to be incorporated as priors. The proposed method is favorable for healthcare applications because in addition to improved prediction performance, relationships among the different risks and among risk factors are also identified. Extensive experimental results on a large electronic medical claims database show that the proposed method outperforms state-of-the-art models by a significant margin. Furthermore, we show that the risk associations learned and the risk factors identified lead to meaningful clinical insights. Bin Liu 0045, Ying Li 0053, Soumya Ghosh, Zhaonan Sun, Kenney Ng, Jianying Hu |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2018 | Early Prediction of Diabetes Complications from Electronic Health Records: A Multi-Task Survival Analysis ApproachabstractType 2 diabetes mellitus (T2DM) is a chronic disease that usually results in multiple complications. Early identification of individuals at risk for complications after being diagnosed with T2DM is of significant clinical value. In this paper, we present a new data-driven predictive approach to predict when a patient will develop complications after the initial T2DM diagnosis. We propose a novel survival analysis method to model the time-to-event of T2DM complications designed to simultaneously achieve two important metrics: 1) accurate prediction of event times, and 2) good ranking of the relative risks of two patients. Moreover, to better capture the correlations of time-to-events of the multiple complications, we further develop a multi-task version of the survival model. To assess the performance of these approaches, we perform extensive experiments on patient level data extracted from a large electronic health record claims database. The results show that our new proposed survival analysis approach consistently outperforms traditional survival models and demonstrate the effectiveness of the multi-task framework over modeling each complication independently. Bin Liu 0045, Ying Li 0053, Zhaonan Sun, Soumya Ghosh, Kenney Ng |
AAAI | 1 |
| 2018 | Representing Urban Functions through Zone Embedding with Human Mobility PatternsabstractUrban functions refer to the purposes of land use in cities where each zone plays a distinct role and cooperates with each other to serve people’s various life needs. Understanding zone functions helps to solve a variety of urban related problems, such as increasing traffic capacity and enhancing location-based service. Therefore, it is beneficial to investigate how to learn the representations of city zones in terms of urban functions, for better supporting urban analytic applications. To this end, in this paper, we propose a framework to learn the vector representation (embedding) of city zones by exploiting large-scale taxi trajectories. Specifically, we extract human mobility patterns from taxi trajectories, and use the co-occurrence of origin-destination zones to learn zone embeddings. To utilize the spatio-temporal characteristics of human mobility patterns, we incorporate mobility direction, departure/arrival time, destination attraction, and travel distance into the modeling of zone embeddings. We conduct extensive experiments with real-world urban datasets of New York City. Experimental results demonstrate the effectiveness of the proposed embedding model to represent urban functions of zones with human mobility data. Zijun Yao 0001, Yanjie Fu, Bin Liu 0045, Wangsu Hu, Hui Xiong 0001 |
IJCAI | 3 |
| 2018 | Attribute Inference Attacks in Online Social NetworksabstractWe propose new privacy attacks to infer attributes (e.g., locations, occupations, and interests) of online social network users. Our attacks leverage seemingly innocent user information that is publicly available in online social networks to infer missing attributes of targeted users. Given the increasing availability of (seemingly innocent) user information online, our results have serious implications for Internet privacy—private attributes can be inferred from users’ publicly available data unless we take steps to protect users from such inference attacks. To infer attributes of a targeted user, existing inference attacks leverage either the user’s publicly available social friends or the user’s behavioral records (e.g., the web pages that the user has liked on Facebook, the apps that the user has reviewed on Google Play), but not both. As we will show, such inference attacks achieve limited success rates. However, the problem becomes qualitatively different if we consider both social friends and behavioral records. To address this challenge, we develop a novel model to integrate social friends and behavioral records, and design new attacks based on our model. We theoretically and experimentally demonstrate the effectiveness of our attacks. For instance, we observe that, in a real-world large-scale dataset with 1.1 million users, our attack can correctly infer the cities a user lived in for 57% of the users; via confidence estimation , we are able to increase the attack success rate to over 90% if the attacker selectively attacks half of the users. Moreover, we show that our attack can correctly infer attributes for significantly more users than previous attacks. Neil Zhenqiang Gong, Bin Liu 0045 |
ACM Trans. Priv. Secur. | 2 |
| 2018 | Personalized Air Travel Prediction: A Multi-factor PerspectiveabstractHuman mobility analysis is one of the most important research problems in the field of urban computing. Existing research mainly focuses on the intra-city ground travel behavior modeling, while the inter-city air travel behavior modeling has been largely ignored. Actually, the inter-city travel analysis can be of equivalent importance and complementary to the intra-city travel analysis. Understanding massive passenger-air-travel behavior delivers intelligence for airlines’ precision marketing and related socioeconomic activities, such as airport planning, emergency management, local transportation planning, and tourism-related businesses. Moreover, it provides opportunities to study the characteristics of cities and the mutual relationships between them. However, modeling and predicting air traveler behavior is challenging due to the complex factors of the market situation and individual characteristics of customers (e.g., airlines’ market share, customer membership, and travelers’ intrinsic interests on destinations). To this end, in this article, we present a systematic study on the personalized air travel prediction problem, namely where a customer will fly to and which airline carrier to fly with, by leveraging real-world anonymized Passenger Name Record (PNR) data. Specifically, we first propose a relational travel topic model, which combines the merits of latent factor model with a neighborhood-based method, to uncover the personal travel preferences of aviation customers and the latent travel topics of air routes and airline carriers simultaneously. Then we present a multi-factor travel prediction framework, which fuses complex factors of the market situation and individual characteristics of customers, to predict airline customers’ personalized travel demands. Experimental results on two real-world PNR datasets demonstrate the effectiveness of our approach on both travel topic discovery and customer travel prediction. Jie Liu 0007, Bin Liu 0045, Yanchi Liu, Huipeng Chen, Lina Feng, Hui Xiong 0001, Yalou Huang |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2016 | POI Recommendation: A Temporal Matching between POI Popularity and User RegularityabstractPoint of interest (POI) recommendation, which provides personalized recommendation of places to mobile users, is an important task in location-based social networks (LBSNs). However, quite different from traditional interest-oriented merchandise recommendation, POI recommendation is more complex due to the timing effects: we need to examine whether the POI fits a user's availability. While there are some prior studies which included the temporal effect into POI recommendations, they overlooked the compatibility between time-varying popularity of POIs and regular availability of users, which we believe has a non-negligible impact on user decision-making. To this end, in this paper, we present a novel method which incorporates the degree of temporal matching between users and POIs into personalized POI recommendations. Specifically, we first profile the temporal popularity of POIs to show when a POI is popular for visit by mining the spatio-temporal human mobility and POI category data. Secondly, we propose latent user regularities to characterize when a user is regularly available for exploring POIs, which is learned with a user-POI temporal matching function. Finally, results of extensive experiments with real-world POI check-in and human mobility data demonstrate that our proposed user-POI temporal matching method delivers substantial advantages over baseline models for POI recommendation tasks. Zijun Yao 0001, Yanjie Fu, Bin Liu 0045, Yanchi Liu, Hui Xiong 0001 |
ICDM | 3 |
| 2016 | Unified Point-of-Interest Recommendation with Temporal Interval AssessmentabstractPoint-of-interest (POI) recommendation, which helps mobile users explore new places, has become an important location-based service. Existing approaches for POI recommendation have been mainly focused on exploiting the information about user preferences, social influence, and geographical influence. However, these approaches cannot handle the scenario where users are expecting to have POI recommendation for a specific time period. To this end, in this paper, we propose a unified recommender system, named the 'Where and When to gO' (WWO) recommender system, to integrate the user interests and their evolving sequential preferences with temporal interval assessment. As a result, the WWO system can make recommendations dynamically for a specific time period and the traditional POI recommender system can be treated as the special case of the WWO system by setting this time period long enough. Specifically, to quantify users' sequential preferences, we consider the distributions of the temporal intervals between dependent POIs in the historical check-in sequences. Then, to estimate the distributions with only sparse observations, we develop the low-rank graph construction model, which identifies a set of bi-weighted graph bases so as to learn the static user preferences and the dynamic sequential preferences in a coherent way. Finally, we evaluate the proposed approach using real-world data sets from several location-based social networks (LBSNs). The experimental results show that our method outperforms the state-of-the-art approaches for POI recommendation in terms of various metrics, such as F-measure and NDCG, with a significant margin. Yanchi Liu, Chuanren Liu, Bin Liu 0045, Meng Qu, Hui Xiong 0001 |
KDD | 3 |
| 2016 | The Impact of Community Safety on House RankingabstractIt is well recognized that community safety which affects people's right to live without fear of crime has considerable impacts on housing investments. Housing investors can make more informed decisions if they are fully aware of safety related factors. To this end, we develop a safety-aware house ranking method by incorporating community safety into house assessment. Specifically, we first propose a novel framework to infer community safety level by mining community crime evidences from rich spatio-temporal historical crime data. Then we develop a ranking model which fuses multiply community safety features to rank house value based on the degree of community safety. Finally, we conduct a comprehensive evaluation of the proposed method with real-world crime and house data. The experimental results show that the proposed method substantially outperforms the baseline methods for house ranking. Zijun Yao 0001, Yanjie Fu, Bin Liu 0045, Hui Xiong 0001 |
SDM | 3 |
| 2016 | You Are Who You Know and How You Behave: Attribute Inference Attacks via Users' Social Friends and Behaviors
Neil Zhenqiang Gong, Bin Liu 0045 |
USENIX Security Symposium | 2 |
| 2016 | Structural Analysis of User Choices for Mobile App RecommendationabstractAdvances in smartphone technology have promoted the rapid development of mobile apps. However, the availability of a huge number of mobile apps in application stores has imposed the challenge of finding the right apps to meet the user needs. Indeed, there is a critical demand for personalized app recommendations. Along this line, there are opportunities and challenges posed by two unique characteristics of mobile apps. First, app markets have organized apps in a hierarchical taxonomy. Second, apps with similar functionalities are competing with each other. Although there are a variety of approaches for mobile app recommendations, these approaches do not have a focus on dealing with these opportunities and challenges. To this end, in this article, we provide a systematic study for addressing these challenges. Specifically, we develop a structural user choice model (SUCM) to learn fine-grained user preferences by exploiting the hierarchical taxonomy of apps as well as the competitive relationships among apps. Moreover, we design an efficient learning algorithm to estimate the parameters for the SUCM model. Finally, we perform extensive experiments on a large app adoption dataset collected from Google Play. The results show that SUCM consistently outperforms state-of-the-art Top-N recommendation methods by a significant margin. Bin Liu 0045, Neil Zhenqiang Gong, Junjie Wu 0002, Hui Xiong 0001, Martin Ester |
ACM Trans. Knowl. Discov. Data | 1 |
| 2015 | Protecting Your Children from Inappropriate Content in Mobile Apps: An Automatic Maturity Rating FrameworkabstractMobile applications (Apps) could expose children or adolescents to mature themes such as sexual content, violence and drug use, which results in an inappropriate security and privacy risk for them. Therefore, mobile platforms provide rating policies to label the maturity levels of Apps and the reasons why an App has a given maturity level, which enables parents to select maturity-appropriate Apps for their children. However, existing approaches to implement these maturity rating policies are either costly (because of expensive manually labeling) or inaccurate (because of no centralized controls). In this work, we aim to design and build a machine learning framework to automatically predict maturity levels for mobile Apps and the associated reasons with a high accuracy and a low cost. Bing Hu 0001, Bin Liu 0045, Neil Zhenqiang Gong, Deguang Kong, Hongxia Jin |
CIKM | 2 |
| 2015 | Personalized Mobile App Recommendation: Reconciling App Functionality and User Privacy PreferenceabstractRecent years have witnessed a rapid adoption of mobile devices and a dramatic proliferation of mobile applications (Apps for brevity). However, the large number of mobile Apps makes it difficult for users to locate relevant Apps. Therefore, recommending Apps becomes an urgent task. Traditional recommendation approaches focus on learning the interest of a user and the functionality of an item (e.g., an App) from a set of user-item ratings, and they recommend an item to a user if the item's functionality well matches the user's interest. However, Apps could have privileges to access a user's sensitive resources ( e.g., contact, message, and location). As a result, a user chooses an App not only because of its functionality, but also because it respects the user's privacy preference. To the best of our knowledge, this paper presents the first systematic study on incorporating both interest-functionality interactions and users' privacy preferences to perform personalized App recommendations. Specifically, we first construct a new model to capture the trade-off between functionality and user privacy preference. Then we crawled a real-world dataset (16,344 users, 6,157 Apps, and 263,054 ratings) from Google Play and use it to comprehensively evaluate our model and previous methods. We find that our method consistently and substantially outperforms the state-of-the-art approaches, which implies the importance of user privacy preference on personalized App recommendations. Moreover, we explore the impact of different levels of privacy information on the performances of our method, which gives us insights on what resources are more likely to be treated as private by users and influence users' behaviors at selecting Apps. Bin Liu 0045, Deguang Kong, Lei Cen, Neil Zhenqiang Gong, Hongxia Jin, Hui Xiong 0001 |
WSDM | 1 |
| 2015 | A General Geographical Probabilistic Factor Model for Point of Interest RecommendationabstractThe problem of point of interest (POI) recommendation is to provide personalized recommendations of places, such as restaurants and movie theaters. The increasing prevalence of mobile devices and of location based social networks (LBSNs) poses significant new opportunities as well as challenges, which we address. The decision process for a user to choose a POI is complex and can be influenced by numerous factors, such as personal preferences, geographical considerations, and user mobility behaviors. This is further complicated by the connection LBSNs and mobile devices. While there are some studies on POI recommendations, they lack an integrated analysis of the joint effect of multiple factors. Meanwhile, although latent factor models have been proved effective and are thus widely used for recommendations, adopting them to POI recommendations requires delicate consideration of the unique characteristics of LBSNs. To this end, in this paper, we propose a general geographical probabilistic factor model ($\sf{Geo}$-PFM) framework which strategically takes various factors into consideration. Specifically, this framework allows to capture the geographical influences on a user’s check-in behavior. Also, user mobility behaviors can be effectively leveraged in the recommendation model. Moreover, based our$\sf{Geo}$-PFM framework, we further develop a Poisson$\sf{Geo}$-PFM which provides a more rigorous probabilistic generative process for the entire model and is effective in modeling the skewed user check-in count data as implicit feedback for better POI recommendations. Finally, extensive experimental results on three real-world LBSN datasets (which differ in terms of user mobility, POI geographical distribution, implicit response data skewness, and user-POI observation sparsity), show that the proposed recommendation methods outperform state-of-the-art latent factor models by a significant margin. Bin Liu 0045, Hui Xiong 0001, Spiros Papadimitriou, Yanjie Fu, Zijun Yao 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2014 | User Preference Learning with Multiple Information Fusion for Restaurant RecommendationabstractIf properly analyzed, the multi-aspect rating data could be a source of rich intelligence for providing personalized restaurant recommendations. Indeed, while recommender systems have been studied for various applications and many recommendation techniques have been developed for general or specific recommendation tasks, there are few studies for restaurant recommendation by addressing the unique challenges of the multi-aspect restaurant reviews. As we know, traditional collaborative filtering methods are typically developed for single aspect ratings. However, multi-aspect ratings are often collected from the restaurant customers. These ratings can reflect multiple aspects of the service quality of the restaurant. Also, geographic factors play an important role in restaurant recommendation. To this end, in this paper, we develop a generative probabilistic model to exploit the multi-aspect ratings of restaurants for restaurant recommendation. Also, the geographic proximity is integrated into the probabilistic model to capture the geographic influence. Moreover, the profile information, which contains customer/restaurant-independent features and the shared features, is also integrated into the model. Finally, we conduct a comprehensive experimental study on a real-world data set. The experimental results clearly demonstrate the benefit of exploiting multi-aspect ratings and the improvement of the developed generative probabilistic model. Yanjie Fu, Bin Liu 0045, Yong Ge 0001, Zijun Yao 0001, Hui Xiong 0001 |
SDM | 2 |
| 2013 | Modeling heterogeneous time series dynamics to profile big sensor data in complex physical systemsabstractWhile a massive amount of time series can now be collected in many physical systems, it is a challenge to build an analytic model that can correctly profile the data because those time series usually exhibit various behaviors. In this paper we propose an integrated method to address the heterogeneity issue in modeling big time series data. We first extracts relevant features to summarize the underlying dynamics of those series. We present both linear and nonlinear feature extraction techniques, as well as a procedure to determine the right extraction method for individual time series. Given extracted features, our method further models the trajectory pattern of time series in the feature space. Both a regression based and a density based method are presented to profile different types of feature trajectories. Experimental results in a real power plant illustrate that our feature extraction and trajectory model are effective to profile various time series. Our method has been used to successfully detect anomalies in the system. Bin Liu 0045, Abhishek B. Sharma, Guofei Jiang, Hui Xiong 0001 |
IEEE BigData | 1 |
| 2013 | Learning geographical preferences for point-of-interest recommendationabstractThe problem of point of interest (POI) recommendation is to provide personalized recommendations of places of interests, such as restaurants, for mobile users. Due to its complexity and its connection to location based social networks (LBSNs), the decision process of a user choose a POI is complex and can be influenced by various factors, such as user preferences, geographical influences, and user mobility behaviors. While there are some studies on POI recommendations, it lacks of integrated analysis of the joint effect of multiple factors. To this end, in this paper, we propose a novel geographical probabilistic factor analysis framework which strategically takes various factors into consideration. Specifically, this framework allows to capture the geographical influences on a user's check-in behavior. Also, the user mobility behaviors can be effectively exploited in the recommendation model. Moreover, the recommendation model can effectively make use of user check-in count data as implicity user feedback for modeling user preferences. Finally, experimental results on real-world LBSNs data show that the proposed recommendation method outperforms state-of-the-art latent factor models with a significant margin. Bin Liu 0045, Yanjie Fu, Zijun Yao 0001, Hui Xiong 0001 |
KDD | 1 |
| 2013 | Point-of-Interest Recommendation in Location Based Social Networks with Topic and Location AwarenessabstractThe wide spread use of location based social networks (LBSNs) has enabled the opportunities for better location based services through Point-of-Interest (POI) recommendation. Indeed, the problem of POI recommendation is to provide personalized recommendations of places of interest. Unlike traditional recommendation tasks, POI recommendation is personalized, location-aware, and context depended. In light of this difference, this paper proposes a topic and location aware POI recommender system by exploiting associated textual and context information. Specifically, we first exploit an aggregated latent Dirichlet allocation (LDA) model to learn the interest topics of users and to infer the interest POIs by mining textual information associated with POIs. Then, a Topic and Location-aware probabilistic matrix factorization (TL-PMF) method is proposed for POI recommendation. A unique perspective of TL-PMF is to consider both the extent to which a user interest matches the POI in terms of topic distribution and the word-of-mouth opinions of the POIs. Finally, experiments on real-world LBSNs data show that the proposed recommendation method outperforms state-of-the-art probabilistic latent factor models with a significant margin. Also, we have studied the impact of personalized interest topics and word-of-mouth opinions on POI recommendations. Bin Liu 0045, Hui Xiong 0001 |
SDM | 1 |
| 2008 | MODEL: moving object detection and localization in wireless networks based on small-scale fadingabstractThis paper presents a new Moving Object Detection and Localization (MODEL) system, which is based on the smallscale fading of RF signal strength and independent from the salient characteristics of both the device and the sensor. We first validated the feasibility of applying small-scale fading effects to moving object detection and localization through experimental analysis. Then, we introduced MODEL: an embedded network system which adopts an easily-realized Rolling-Window algorithm. We applied the Region-Partition method to determine the position of the moving object, and concluded that the precision of the object position is dependant upon the density of participating nodes. MODEL is also scalable to other wireless network infrastructures and adaptable to various environments without the need for complex and time consuming training. Qingming Yao, Bin Liu 0045, Fei-Yue Wang 0001 |
SenSys | 3 |