VLDB 2026 Research / reviewers in the wild / expert
Xiao Han 0001
dblp:01/2095-1
· DBLP profile ↗
29ranked-venue papers
10as first author
14since 2021 · last 2026
0000-0003-1331-0860ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 8 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 2 since 2021Security and privacy · 6 · 2 first-author · 4 since 2021Computer networks · 5 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 first-authorTheory of computation · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Membership Inference Attack Against Time-Series Prediction ModelsabstractRecent advances in machine learning (ML) have raised growing privacy concerns. One notable concern of ML models is the membership leakage risk, which is defined as the extent to which an adversary can accurately determine whether a record was part of a model's training set through membership inference attacks (MIAs). While extensively studied in non-time-series models, this risk remains largely unexplored for time-series prediction models. To address this gap, this work aims to provide a precise evaluation of membership leakage risks for time-series prediction models. We first conduct an empirical analysis to identify distinct temporal discrepancies within sequential outputs between member and non-member samples in time-series models. Beyond global temporal discrepancies, our findings show that member records exhibit notable local discrepancies, including smaller fluctuations in their output sequences when actual values remain constant and more rapid adjustments when actual values change, compared with non-member records. Building on these insights, we propose TSP-MIA, a framework that leverages a patch-based attack model to capture both global and local discrepancies for more effective MIAs. Extensive evaluations demonstrate the effectiveness and robustness of TSP-MIA, revealing substantial membership leakage risks for time-series prediction models. We further assess the validity of existing defense mechanisms against TSP-MIA. Xiao Han 0001, Ruiyan Wang, Junjie Wu 0002, Lanjuan Liu |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2025 | Learning-Based Privacy-Preserving Graph Publishing Against Sensitive Link Inference AttacksabstractPublishing graph data is widely desired to enable a variety of structural analyses and downstream tasks. However, it also potentially poses severe privacy leakage, as attackers may leverage the released graph data to launch attacks and precisely infer private information such as the existence of hidden sensitive links in the graph. Prior studies on privacy-preserving graph data publishing relied on heuristic graph modification strategies and it is difficult to determine the graph with the optimal privacy–utility trade-off for publishing. In contrast, we propose the first privacy-preserving graph structure learning framework against sensitive link inference attacks, named PPGSL, which can automatically learn a graph with the optimal privacy–utility trade-off. The PPGSL operates by first simulating a powerful surrogate attacker conducting sensitive link attacks on a given graph. It then trains a parameterized graph to defend against the simulated adversarial attacks while maintaining the favorable utility of the original graph. To learn the parameters of both parts of the PPGSL, we introduce a secure iterative training protocol. It can enhance privacy preservation and ensure stable convergence during the training process, as supported by the theoretical proof. Additionally, we incorporate multiple acceleration techniques to improve the efficiency of the PPGSL in handling large-scale graphs. The experimental results confirm that the PPGSL achieves state-of-the-art privacy–utility trade-off performance and effectively thwarts various sensitive link inference attacks. Yucheng Wu 0002, Yuncong Yang, Xiao Han 0001, Leye Wang, Junjie Wu 0002 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | UniTrans: A Unified Vertical Federated Knowledge Transfer Framework for Enhancing Edge Healthcare CollaborationabstractCross-hospital collaboration has the potential to mitigate disparities in medical resources across different regions. However, strict privacy regulations prohibit the direct sharing of sensitive patient information between hospitals. Vertical Federated Learning (VFL) provides a novel privacy-preserving machine learning paradigm designed to maximizes data utility across multiple hospitals. Nevertheless, traditional VFL methods primarily benefit patients with overlapping data, leaving non-overlapping patients without guaranteed improvements in distributed healthcare prediction services. While some existing knowledge transfer techniques attempt to improve prediction performance for non-overlapping patients, they fail to adequately address scenarios where overlapping and non-overlapping patients originate from different domains, resulting in challenges such as feature and label heterogeneity. To address these issues, we propose UniTrans, a unified vertical federated knowledge transfer framework for edge healthcare collaboration. Our framework consists of three key steps. First, we extract the federated representation of overlapping patients by employing an effective vertical federated representation learning method to model multi-party joint features online. Next, each hospital learns a local knowledge transfer module offline, enabling the domain-adaptive transfer of knowledge from the federated representation of overlapping patients to the enriched representation of local non-overlapping patients. Finally, hospitals utilize these enriched local representations to enhance performance across various downstream medical prediction tasks. Extensive experiments on real-world medical datasets demonstrate the effectiveness and scalability of UniTrans in both intra-domain and cross-domain knowledge transfer. The code of UniTrans is available athttps://github.com/Chung-ju/Unitrans. Chung-ju Huang, Yuanpeng He, Xiao Han 0001, Wenpin Jiao, Zhi Jin 0001, Leye Wang |
IEEE Trans. Mob. Comput. | 3 |
| 2024 | Graph Contrastive Learning with Cohesive Subgraph AwarenessabstractGraph contrastive learning (GCL) has emerged as a state-of-the-art strategy for learning representations of diverse graphs including social and biomedical networks. GCL widely uses stochastic graph topology augmentation, such as uniform node dropping, to generate augmented graphs. However, such stochastic augmentations may severely damage the intrinsic properties of a graph and deteriorate the following representation learning process. We argue that incorporating an awareness of cohesive subgraphs during the graph augmentation and learning processes has the potential to enhance GCL performance. To this end, we propose a novel unified framework called CTAug, to seamlessly integrate cohesion awareness into various existing GCL mechanisms. In particular, CTAug comprises two specialized modules: topology augmentation enhancement and graph learning enhancement. The former module generates augmented graphs that carefully preserve cohesion properties, while the latter module bolsters the graph encoder's ability to discern subgraph patterns. Theoretical analysis shows that CTAug can strictly improve existing GCL mechanisms. Empirical experiments verify that CTAug can achieve state-of-the-art performance for graph representation learning, especially for graphs with high degrees. The code is available at https://doi.org/10.5281/zenodo.10594093, or https://github.com/wuyucheng2002/CTAug. Yucheng Wu 0002, Leye Wang, Xiao Han 0001, Han-Jia Ye |
WWW | 3 |
| 2024 | Privacy-Preserving Network Embedding Against Private Link Inference AttacksabstractNetwork embedding represents network nodes by a low-dimensional informative vector. While it is generally effective for various downstream tasks, it may leak some private information of networks, such as hidden private links. In this work, we address a novel problem ofprivacy-preserving network embedding against private link inference attacks. Basically, we propose to perturb the original network by adding or removing links, and expect the embedding generated on the perturbed network can leak little information about private links but hold high utility for various downstream tasks. Towards this goal, we first propose general measurements to quantify privacy gain and utility loss incurred by candidate network perturbations; we then design aPrivacy-PreservingNetworkEmbedding (i.e., PPNE) framework to identify the optimal perturbation solution with the best privacy-utility trade-off in an iterative way. Furthermore, we propose many techniques to accelerate PPNE and ensure its scalability. For instance, as the skip-gram embedding methods including DeepWalk and LINE can be seen as matrix factorization with closed-form embedding results, we devise efficient privacy gain and utility loss approximation methods to avoid the repetitive time-consuming embedding training for every candidate network perturbation in each iteration. Experiments on real-life network datasets (with up to millions of nodes) verify that PPNE outperforms baselines by sacrificing less utility and obtaining higher privacy protection. Xiao Han 0001, Yuncong Yang, Leye Wang, Junjie Wu 0002 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2024 | HyObscure: Hybrid Obscuring for Privacy-Preserving Data PublishingabstractMinimizing privacy leakage while ensuring data utility is a critical problem in a privacy-preserving data publishing task, from which data holders can boost platform engagements or enlarge data values. Most prior research concerned only with either privacy-insensitive or exact private data and resorts to a single obscuring method to achieve a privacy-utility tradeoff, which is inadequate for real-life hybrid data especially when facing machine learning-based inference attacks. This work takes a pilot study on privacy-preserving data publishing when both widely adopted generalization and obfuscation operations are employed for privacy-heterogeneous data protection. Specifically, we first propose novel measures for privacy and utility values quantification and formulate the hybrid privacy-preserving data obscuring problem to account for the joint effect of generalization and obfuscation. We then design a novel protection mechanism called HyObscure, which decomposes the original problem into three sub-problems to cross-iteratively optimize the hybrid operations for maximum privacy protection under a certain data utility guarantee. The convergence of the iterative process and the privacy leakage bound of HyObscure are also provided in theory. Extensive experiments demonstrate that HyObscure significantly outperforms a variety of state-of-the-art baseline methods when facing various inference attacks in different scenarios. Xiao Han 0001, Yuncong Yang, Junjie Wu 0002, Hui Xiong 0001 |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2024 | Safety and Performance, Why Not Both? Bi-Objective Optimized Model Compression Against Heterogeneous Attacks Toward AI Software DeploymentabstractThe size of deep learning models in artificial intelligence (AI) software is increasing rapidly, which hinders the large-scale deployment on resource-restricted devices (e.g., smartphones). To mitigate this issue, AI software compression plays a crucial role, which aims to compress model size while keeping high performance. However, the intrinsic defects in the big model may be inherited by the compressed one. Such defects may be easily leveraged by attackers, since the compressed models are usually deployed in a large number of devices without adequate protection. In this paper, we try to address the safe model compression problem from a safety-performance co-optimization perspective. Specifically, inspired by the test-driven development (TDD) paradigm in software engineering, we propose a test-driven sparse training framework calledSafeCompress. By simulating the attack mechanism as the safety test, SafeCompress can automatically compress a big model to a small one following the dynamic sparse training paradigm. Then, considering two kinds of representative and heterogeneous attack mechanisms i.e., black-box membership inference attack and white-box membership inference attack, we develop two concrete instances called BMIA-SafeCompress and WMIA-SafeCompress. Further, we implement another instance called MMIA-SafeCompress by extending SafeCompress to defend against the occasion when attackers conduct black-box and white-box membership inference attacks simultaneously. Extensive experiments are conducted on five datasets for both computer vision and natural language processing tasks. The results verify the effectiveness and generalizability of our method. We also discuss how to adapt SafeCompress to other attacks besides MIA, demonstrating the flexibility of SafeCompress. Leye Wang, Xiao Han 0001, Anmin Liu, Tao Xie 0001 |
IEEE Trans. Software Eng. | 3 |
| 2023 | Cross-center Early Sepsis Recognition by Medical Knowledge Guided Collaborative Learning for Data-scarce HospitalsabstractThere are significant regional inequities in health resources around the world. It has become one of the most focused topics to improve health services for data-scarce hospitals and promote health equity through knowledge sharing among medical institutions. Because electronic medical records (EMRs) contain sensitive personal information, privacy protection is unavoidable and essential for multi-hospital collaboration. In this paper, for a common disease in ICU patients, sepsis, we propose a novel cross-center collaborative learning framework guided by medical knowledge, SofaNet, to achieve early recognition of this disease. The Sepsis-3 guideline, published in 2016, defines that sepsis can be diagnosed by satisfying both suspicion of infection and Sequential Organ Failure Assessment (SOFA) greater than or equal to 2. Based on this knowledge, SofaNet adopts a multi-channel GRU structure to predict SOFA values of different systems, which can be seen as an auxiliary task to generate better health status representations for sepsis recognition. Moreover, we only achieve feature distribution alignment in the hidden space during cross-center collaborative learning, which ensures secure and compliant knowledge transfer without raw data exchange. Extensive experiments on two open clinical datasets, MIMIC-III and Challenge, demonstrate that SofaNet can benefit early sepsis recognition when hospitals only have limited EMRs. Ruiqing Ding, Fangjie Rong, Xiao Han 0001, Leye Wang |
WWW | 3 |
| 2023 | Vertical Federated Knowledge Transfer via Representation Distillation for Healthcare Collaboration NetworksabstractCollaboration between healthcare institutions can significantly lessen the imbalance in medical resources across various geographic areas. However, directly sharing diagnostic information between institutions is typically not permitted due to the protection of patients’ highly sensitive privacy. As a novel privacy-preserving machine learning paradigm, federated learning (FL) makes it possible to maximize the data utility among multiple medical institutions. These feature-enrichment FL techniques are referred to as vertical FL (VFL). Traditional VFL can only benefit multi-parties’ shared samples, which strongly restricts its application scope. In order to improve the information-sharing capability and innovation of various healthcare-related institutions, and then to establish a next-generation open medical collaboration network, we propose a unified framework for vertical federated knowledge transfer mechanism (VFedTrans) based on a novel cross-hospital representation distillation component. Specifically, our framework includes three steps. First, shared samples’ federated representations are extracted by collaboratively modeling multi-parties’ joint features with current efficient vertical federated representation learning methods. Second, for each hospital, we learn a local-representation-distilled module, which can transfer the knowledge from shared samples’ federated representations to enrich local samples’ representations. Finally, each hospital can leverage local samples’ representations enriched by the distillation module to boost arbitrary downstream machine learning tasks. The experiments on real-life medical datasets verify the knowledge transfer effectiveness of our framework. Chung-ju Huang, Leye Wang, Xiao Han 0001 |
WWW | 3 |
| 2023 | Cost-Effective Social Media Influencer MarketingabstractIt is becoming more and more promising that marketers hire influencers to launch campaigns for spreading items (e.g., articles or videos about products) over social media platforms. Such social media influencer marketing may generate tremendous utility if the influencers persuade their followers to adopt the recommended items. This could further spur extensive spontaneous item propagation on social media. Although prior studies mainly focus on influencer-selection strategies by the influencers’ traits, marketers with a number of items are often requested to determine both influencers and marketing items. The appropriateness between influencers and items is critical, but rarely considered in prior influencer-identification methods. We thus formulate and solve a novel cost-effective social media influencer marketing problem to maximize marketers’ utility by selecting appropriate pairwise combinations of influencers and items (i.e., item-influencer pairs). In particular, we first model utility functions and propose a simulation-based method to estimate the appropriateness of arbitrarily given item-influencer pairs by their potential utility. With the estimated utility, we devise an algorithm to iteratively select appropriate item-influencer pairs under various realistic conditions, including marketers’ budget, influencers’ payments, item-user fitness, social propagation, and influencers’ marketing slots. We theoretically prove that the marketing utility achieved by our method is near-optimal. We also conduct extensive empirical experiments with three real-world data sets to verify the superiority of our method in terms of cost-effectiveness and computational efficiency. Lastly, we discuss insightful theoretical and practical implications. History: Accepted by Ram Ramesh, Area Editor for Data Science and Machine Learning. Funding: This study was partially funded by the National Natural Science Foundation of China [Grants 72071125, 72031001, and 61972008]. Supplemental Material: The online appendix is available at https://doi.org/10.1287/ijoc.2022.1246 . Xiao Han 0001, Leye Wang, Weiguo Fan |
INFORMS J. Comput. | 1 |
| 2023 | Video summarization for event-centric videos
Jianni Chen, Qiqin Xie, Xiao Han 0001 |
Neural Networks | 4 |
| 2022 | Safety and Performance, Why not Both? Bi-Objective Optimized Model Compression toward AI Software DeploymentabstractThe size of deep learning models in artificial intelligence (AI) software is increasing rapidly, which hinders the large-scale deployment on resource-restricted devices (e.g., smartphones). To mitigate this issue, AI software compression plays a crucial role, which aims to compress model size while keeping high performance. However, the intrinsic defects in the big model may be inherited by the compressed one. Such defects may be easily leveraged by attackers, since the compressed models are usually deployed in a large number of devices without adequate protection. In this paper, we try to address the safe model compression problem from a safety-performance co-optimization perspective. Specifically, inspired by the test-driven development (TDD) paradigm in software engineering, we propose a test-driven sparse training framework called SafeCompress. By simulating the attack mechanism as the safety test, SafeCompress can automatically compress a big model to a small one following the dynamic sparse training paradigm. Further, considering a representative attack, i.e., membership inference attack (MIA), we develop a concrete safe model compression mechanism, called MIA-SafeCompress. Extensive experiments are conducted to evaluate MIA-SafeCompress on five datasets for both computer vision and natural language processing tasks. The results verify the effectiveness and generalization of our method. We also discuss how to adapt SafeCompress to other attacks besides MIA, demonstrating the flexibility of SafeCompress. Leye Wang, Xiao Han 0001 |
ASE | 3 |
| 2021 | Label Confusion Learning to Enhance Text Classification ModelsabstractRepresenting the true label as one-hot vector is the common practice in training text classification models. However, the one-hot representation may not adequately reflect the relation between the instance and labels, as labels are often not completely independent and instances may relate to multiple labels in practice. The inadequate one-hot representations tend to train the model to be over-confident, which may result in arbitrary prediction and model overfitting, especially for confused datasets (datasets with very similar labels) or noisy datasets (datasets with labeling errors). While training models with label smoothing can ease this problem in some degree, it still fails to capture the realistic relation among labels. In this paper, we propose a novel Label Confusion Model (LCM) as an enhancement component to current popular text classification models. LCM can learn label confusion to capture semantic overlap among labels by calculating the similarity between instance and labels during training and generate a better label distribution to replace the original one-hot label vector, thus improving the final classification performance. Extensive experiments on five text classification benchmark datasets reveal the effectiveness of LCM for several widely used deep learning classification models. Further experiments also verify that LCM is especially helpful for confused or noisy datasets and superior to the label smoothing method. Biyang Guo, Songqiao Han, Xiao Han 0001, Hailiang Huang 0003 |
AAAI | 3 |
| 2021 | Mobile Crowdsourcing Task Allocation with Differential-and-Distortion Geo-ObfuscationabstractIn mobile crowdsourcing, organizers usually need participants' precise locations for optimal task allocation, e.g., minimizing selected workers' travel distance to task locations. However, the exposure of users' locations raises privacy concerns. In this paper, we propose a location privacy-preserving task allocation framework with geo-obfuscation to protect users' locations during task assignments. More specifically, we make participants obfuscate their reported locations under the guarantee of two rigorous privacy-preserving schemes, differential and distortion privacy, without the need to involve any third-party trusted entity. In order to achieve optimal task allocation with the differential-and-distortion geo-obfuscation, we formulate a mixed-integer non-linear programming problem to minimize the expected travel distance of the selected workers under the constraints of differential and distortion privacy. Moreover, a worker may be willing to accept multiple tasks, and a task organizer may be concerned with multiple utility objectives such as task acceptance ratio in addition to travel distance. Against this background, we also extend our solution to the multi-task allocation and multi-objective optimization cases. Evaluation results on both simulation and real-world user mobility traces verify the effectiveness of our framework. Particularly, our framework outperforms Laplace obfuscation, a state-of-the-art geo-obfuscation mechanism, by achieving up to 47 percent shorter average travel distance on real-world data under the same level of privacy protection. Leye Wang, Dingqi Yang, Xiao Han 0001, Daqing Zhang 0001, Xiaojuan Ma |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2020 | Sparse Mobile Crowdsensing With Differential and Distortion Location PrivacyabstractSparse Mobile Crowdsensing (MCS) has become a compelling approach to acquire and infer urban-scale sensing data. However, participants risk their location privacy when reporting data with their actual sensing positions. To address this issue, we propose a novel location obfuscation mechanism combining E-differential-privacy and δ-distortion-privacy in Sparse MCS. More specifically, differential privacy bounds adversaries' relative information gain regardless of their prior knowledge, while distortion privacy ensures that the expected inference error is larger than a threshold under an assumption of adversaries' prior knowledge. To reduce the data quality loss incurred by location obfuscation, we design a differential-and-distortion privacy-preserving framework with three components. First, we learn a data adjustment function to fit the original sensing data to the obfuscated location. Second, we apply a linear program to select an optimal location obfuscation function. The linear program aims to minimize the uncertainty in data adjustment under the constraints of E-differential-privacy, δ-distortion-privacy, and evenly-distributed obfuscation. We also design an approximated method to reduce the required computation resources. Third, we propose an uncertainty-aware inference algorithm to improve the inference accuracy for the obfuscated data. Evaluations with real environment and traffic datasets show that our optimal method reduces the data quality loss by up to 42% compared to the state-of-the-art methods with the same level of privacy protection; the approximated method incurs <; 3% additional quality loss than the optimal method, but only needs <; 1% of the computation time. Leye Wang, Daqing Zhang 0001, Dingqi Yang, Brian Y. Lim, Xiao Han 0001, Xiaojuan Ma |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2019 | F-PAD: Private Attribute Disclosure Risk Estimation in Online Social NetworksabstractIn online social networks, users always expect to share some information for benefits (e.g., personalized services) while hiding the others for privacy. Unfortunately, the hidden information is likely to be predicted by various powerful inference attacks with the rapid advances in machine learning. Then, what is the risk that a user's private information could be disclosed? What countermeasures can be taken to fight against the privacy violation for the user? To tackle these issues, this article proposes a general Framework for Private Attribute Disclosure estimation (F-PAD) including three steps: 1) private attribute prediction; 2) disclosure model training; 3) disclosure risk estimation. Not like most prior risk estimation studies focusing on one specific attack model and private attribute, F-PAD can estimate disclosure risk for individual users in terms of disclosure probability and risk level within a high confidence given a basket of potential inference attack models; furthermore, F-PAD can adapt to various attributes (e.g., gender, age) and offer countermeasures to help users lower the risk. Extensive experiment studies on two real social network datasets, Facebook and Book-Crossing, have verified the effectiveness of F-PAD in `current city', `gender' and `age' disclosure risk estimation. Xiao Han 0001, Hailiang Huang 0003, Leye Wang |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2018 | Geographic Differential Privacy for Mobile Crowd Coverage MaximizationabstractFor real-world mobile applications such as location-based advertising and spatial crowdsourcing, a key to success is targeting mobile users that can maximally cover certain locations in a future period. To find an optimal group of users, existing methods often require information about users' mobility history, which may cause privacy breaches. In this paper, we propose a method to maximize mobile crowd's future location coverage under a guaranteed location privacy protection scheme. In our approach, users only need to upload one of their frequently visited locations, and more importantly, the uploaded location is obfuscated using a geographic differential privacy policy. We propose both analytic and practical solutions to this problem. Experiments on real user mobility datasets show that our method significantly outperforms the state-of-the-art geographic differential privacy methods by achieving a higher coverage under the same level of privacy protection. Leye Wang, Gehua Qin, Dingqi Yang, Xiao Han 0001, Xiaojuan Ma |
AAAI | 4 |
| 2018 | SPACE-TA: Cost-Effective Task Allocation Exploiting Intradata and Interdata Correlations in Sparse CrowdsensingabstractData quality and budget are two primary concerns in urban-scale mobile crowdsensing. Traditional research on mobile crowdsensing mainly takes sensing coverage ratio as the data quality metric rather than the overall sensed data error in the target-sensing area. In this article, we propose to leverage spatiotemporal correlations among the sensed data in the target-sensing area to significantly reduce the number of sensing task assignments. In particular, we exploit both intradata correlations within the same type of sensed data and interdata correlations among different types of sensed data in the sensing task. We propose a novel crowdsensing task allocation framework called SPACE-TA (SPArse Cost-Effective Task Allocation) , combining compressive sensing, statistical analysis, active learning, and transfer learning, to dynamically select a small set of subareas for sensing in each timeslot (cycle), while inferring the data of unsensed subareas under a probabilistic data quality guarantee. Evaluations on real-life temperature, humidity, air quality, and traffic monitoring datasets verify the effectiveness of SPACE-TA. In the temperature-monitoring task leveraging intradata correlations, SPACE-TA requires data from only 15.5% of the subareas while keeping the inference error below 0.25°C in 95% of the cycles, reducing the number of sensed subareas by 18.0% to 26.5% compared to baselines. When multiple tasks run simultaneously, for example, for temperature and humidity monitoring, SPACE-TA can further reduce ∼10% of the sensed subareas by exploiting interdata correlations. Leye Wang, Daqing Zhang 0001, Dingqi Yang, Animesh Pathak, Chao Chen 0004, Xiao Han 0001, Haoyi Xiong, Yasha Wang |
ACM Trans. Intell. Syst. Technol. | 6 |
| 2017 | Location Privacy-Preserving Task Allocation for Mobile Crowdsensing with Differential Geo-ObfuscationabstractIn traditional mobile crowdsensing applications, organizers need participants' precise locations for optimal task allocation, e.g., minimizing selected workers' travel distance to task locations. However, the exposure of their locations raises privacy concerns. Especially for those who are not eventually selected for any task, their location privacy is sacrificed in vain. Hence, in this paper, we propose a location privacy-preserving task allocation framework with geo-obfuscation to protect users' locations during task assignments. Specifically, we make participants obfuscate their reported locations under the guarantee of differential privacy, which can provide privacy protection regardless of adversaries' prior knowledge and without the involvement of any third-part entity. In order to achieve optimal task allocation with such differential geo-obfuscation, we formulate a mixed-integer non-linear programming problem to minimize the expected travel distance of the selected workers under the constraint of differential privacy. Evaluation results on both simulation and real-world user mobility traces show the effectiveness of our proposed framework. Particularly, our framework outperforms Laplace obfuscation, a state-of-the-art differential geo-obfuscation mechanism, by achieving 45% less average travel distance on the real-world data. Leye Wang, Dingqi Yang, Xiao Han 0001, Tianben Wang, Daqing Zhang 0001, Xiaojuan Ma |
WWW | 3 |
| 2016 | CSD: A multi-user similarity metric for community recommendation in online social networks
Xiao Han 0001, Leye Wang, Reza Farahbakhsh, Ángel Cuevas, Rubén Cuevas Rumín, Noël Crespi |
Expert Syst. Appl. | 1 |
| 2015 | Link prediction for new users in Social NetworksabstractLink prediction for new users who have not created any link is a fundamental problem in Online Social Networks (OSNs). It can be used to recommend friends for new users to start building their social networks. The existing studies use cross-platform approaches to predict a new user's links on a certain OSN by porting his existing links from other OSNs. However, it cannot work when OSNs are not willing to share their data or users do not want to connect different OSN accounts. In this paper, we use a single-platform approach to carry out the link prediction. We explore the users' profile attributes (e.g., workplace, high school and hometown) which can be easily obtained during the new users' sign up procedure. Based on the limited available information from the new user, along with the attributes and links from existing users, we extract three types of social features: basic feature, derived feature and latent relation feature. We propose a link prediction model using these social features based on Support Vector Machines. Eventually, we rely on a large Facebook data set consisting of 479,000 users to evaluate our proposed model. The result reveals that our model outperforms the baselines by achieving the AUC value of 0.83; it also demonstrates that each of the proposed social features contribute significantly to the prediction model. Xiao Han 0001, Leye Wang, Son N. Han, Chao Chen 0004, Noël Crespi, Reza Farahbakhsh |
ICC | 1 |
| 2015 | Alike people, alike interests? Inferring interest similarity in online social networks
Xiao Han 0001, Leye Wang, Noël Crespi, Soochang Park, Ángel Cuevas |
Decis. Support Syst. | 1 |
| 2014 | Alike people, alike interests? A large-scale study on interest similarity in social networksabstractThis paper presents a comprehensive empirical study on the correlations between users' interest similarity and various social features across three interest domains (i.e., movie, music and TV). This study relies on a large dataset, containing 479, 048 users and 5, 263, 351 user-generated interests, captured from Facebook. We identify the social features from three types of the users' information - demographic information (e.g., age, gender, location), social relations (i.e., friendship), and users' interests. The results reveal that the interest similarity follows the homophily principle. Particularly, the results show that two users are more likely to be alike in their interests 1) if they exhibit more similarity in their demographic characteristics (e.g., similar age, same gender, or close to each other geographically), or 2) if they are more intimate in their friendship, or 3) if they present a higher average interest individuality (i.e., a measurement for estimating the personalized characteristics of a user's interests). The empirical observations could be exploited to infer how two users are alike in their interests according to the social features, which could be further harnessed by various practical applications and services, such as recommendation system and advertisement service. Xiao Han 0001, Leye Wang, Soochang Park, Ángel Cuevas, Noël Crespi |
ASONAM | 1 |
| 2014 | On exploiting social relationship and personal background for content discovery in P2P networksabstractContent discovery is a critical issue in unstructured Peer-to-Peer (P2P) networks as nodes maintain only local network information. However, similarly without global information about human networks, one still can find specific persons via his/her friends by using social information. Therefore, in this paper, we investigate the problem of how social information (i.e., friends and background information) could benefit content discovery in P2P networks. We collect social information of 384,494 user profiles from Facebook, and build a social P2P network model based on the empirical analysis. In this model, we enrich nodes in P2P networks with social information and link nodes via their friendships. Each node extracts two types of social features–Knowledge and Similarity–and assigns more weight to the friends that have higher similarity and more knowledge. Furthermore, we present a novel content discovery algorithm which can explore the latent relationships among a node’s friends. A node computes stable scores for all its friends regarding their weight and the latent relationships. It then selects the top friends with higher scores to query content. Extensive experiments validate performance of the proposed mechanism. In particular, for personal interests searching, the proposed mechanism can achieve 100% of Search Success Rate by selecting the top 20 friends within two-hop. It also achieves 6.5 Hits on average, which improves 8x the performance of the compared methods. Xiao Han 0001, Ángel Cuevas, Noël Crespi, Rubén Cuevas Rumín, Xiaodi Huang 0001 |
Future Gener. Comput. Syst. | 1 |
| 2013 | Analysis of publicly disclosed information in Facebook profilesabstractFacebook, the most popular Online social network is a virtual environment where users share information and are in contact with friends. Apart from many useful aspects, there is a large amount of personal and sensitive information publicly available that is accessible to external entities/users. In this paper we study the public exposure of Facebook profile attributes to understand what type of attributes are considered more sensitive by Facebook users in terms of privacy, and thus are rarely disclosed, and which attributes are available in most Facebook profiles. Furthermore, we also analyze the public exposure of Facebook users by accounting the number of attributes that users make publicly available on average. To complete our analysis we have crawled the profile information of 479K randomly selected Facebook users. Finally, in order to demonstrate the utility of the publicly available information in Facebook profiles we show in this paper three case studies. The first one carries out a gender-based analysis to understand whether men or women share more or less information. The second case study depicts the age distribution of Facebook users. The last case study uses data inferred from Facebook profiles to map the distribution of worldwide population across cities according to its size. Reza Farahbakhsh, Xiao Han 0001, Ángel Cuevas, Noël Crespi |
ASONAM | 2 |
| 2013 | A Framework for Social Device NetworkingabstractThe concept of connectedness, as inspired by Social Networking Service, is a key factor which participates in changing the way people interact with each other over the Internet. On the other hand, connected world as envisioned by the Internet of Things aims to expand the idea of connectivity to include everything in the physical world to a big network called the Internet. In order to realize the integration between the world of connected people and the world of connected devices, intelligence including semantics and recommendation acts as a key factor to expand the basic communication functionalities to include search, discovery, mashup of new services and filtering. We propose a framework to facilitate the next generation of communication between people and devices, and a preliminary prototype including three modules DPWSim, ThingsGate, ThingsChat along with a use case discussion. Dina Hussein, Son N. Han, Xiao Han 0001, Gyu Myoung Lee, Noël Crespi |
DCOSS | 3 |
| 2013 | "Current City" prediction for coarse location based applications on FacebookabstractLocation-Based services with social networks improve users' experience and enrich people's social live. However, location information is often inadequate due to privacy and security concerns. We seek to infer users' ‘Current City’ on Facebook for coarse location based applications. We first extract users' multiple explicit and implicit location attributes, and analyze correlations of these attributes from two perspective: user-centric and user-friends. We observe that both user-centric and user-friends location attributes tightly correlate to a user's Current City (e.g., 60% of users stay in their hometown, 60% of users live in the same city as 50% of their friends). Based on extensive analysis and observations on location attributes correlations, we have constructed a Current City Prediction model (CCP) using artificial neural network (ANN) learning frameworks. The experimental results indicate that we achieve accuracy levels of 84% for city-level prediction and 98% for country-level which are increases of 9% and 18%, respectively than what is possible with Tweecalization. Wipada Chanthaweethip, Xiao Han 0001, Noël Crespi, Yuanfang Chen, Reza Farahbakhsh, Ángel Cuevas |
GLOBECOM | 2 |
| 2013 | Reality Mining: Digging the Impact of Friendship and Location on Crowd BehaviorabstractFinding basic laws that govern human crowd behavior is a subject that deserves to study. Crowd behavior is a natural instinct for human, which directly impacts how we form opinions and make decisions. Moreover, it is common that people change their behavior in a group. In pervasive computing research, substantial work has been directed towards discovering human movement patterns based on wireless networks. This research has, however, mainly been focused on movements of individuals. Mobile phones offer on-body tracking and they are already deployed on a large scale, allowing the characterization of user behavior through the information related to individual movements. In this paper, we observe and analyze the impact of friendship and location attributes on crowd behavior, using location-based wireless mobility information. These preliminary studies will be a good cornerstone for a crowd behavior prediction. Yuanfang Chen, Lei Shu 0001, Xiao Han 0001, Lin Lv, Xuemin Cheng |
MASS | 3 |
| 2013 | WX-MAC: An Energy Efficient MAC Protocol for Wireless Sensor NetworksabstractIn this poster, we propose a novel energy efficient asynchronous MAC protocol for WSNs-WX-MAC. WX-MAC supports nodes to exchange their sampling schedules by query and report mechanisms. With the aid of neighbors' information and strobed preamble approach, WX-MAC shortens senders' preamble before data transmission and reduces nodes' wakeup duration. WX-MAC not only reduces the energy consumption on the senders and targets, but also decreases the possibility of overhearing on the surrounding neighbors, leading to a further energy conservation. Preliminary simulations have shown that WX-MAC saves more energy than B-MAC, Wise MAC, and X-MAC. Xiao Han 0001, Lei Shu 0001, Yuanfang Chen, Hairui Zhou |
MASS | 1 |