EDBT 2026 Demo / reviewers in the wild / expert
Xinghua Shi
dblp:56/4047 · also Xinghua Mindy Shi
· DBLP profile ↗
31ranked-venue papers
1as first author
10since 2021 · last 2026
0000-0003-4662-3177ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 12 · 2 since 2021Computer networks · 8Artificial intelligence and machine learning · 4 · 4 since 2021Systems, architecture and hardware · 4 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Security and privacy · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GPUFast-$Q$: Novel High-Throughput Sequence Data Compression on GPUs Using Fine-Grained ParallelismabstractNext-generation sequencing (NGS) platforms generate multi-terabyte data, making compression a major bottleneck for storage, transfer, and downstream analysis. Taolue Yang, Youyuan Liu, Chong Li 0001, Xinghua Shi, Sian Jin |
DCC | 5 |
| 2026 | GFAz: State-of-the-Art Graphical Fragment Assembly Compression
Taolue Yang, Youyuan Liu, Xinghua Shi, Sian Jin |
ICS | 4 |
| 2025 | FSLearning: An Efficient Federated Split Learning Framework for Privacy-Preserving Disease Prediction
Xiaoqian Jiang, Yu-Chun Hsu, Arif Ozgun Harmanci, Hongchang Gao, Xinghua Shi |
AIME (1) | 6 |
| 2025 | Uncertainty Quantification of Conformalized Spatio-Temporal Graph Convolutional Network on ForecastingabstractAccurately forecasting spatio-temporal dynamics in high-resolution and sparsely populated datasets remains a fundamental challenge in many real-world applications such as urban mobility and environmental monitoring. Traditional deterministic machine learning models often rely on point estimates, which inherently lack the ability to capture and quantify predictive uncertainty—particularly in settings where data sparsity and topological complexity prevail. While existing approaches to spatio-temporal prediction have made progress in modeling data distributions or estimating model uncertainty, they typically rely on strong statistical assumptions and fail to provide comprehensive, distribution-free uncertainty quantification. In this work, we introduce a novel hybrid framework that integrates Conformal Prediction (CP) with Spatio-Temporal Graph Convolutional Networks (STGCNs) to address these limitations. Our method builds upon pretrained STGCN models by incorporating a conformal calibration procedure, which constructs a calibration set tailored to the graph's temporal and topological structure under the assumption of sequential exchangeability. This enables the model to produce valid, data-driven prediction intervals without retraining, offering statistically rigorous uncertainty estimates alongside point predictions. The proposed CP-STGCN framework dynamically adjusts its predictive intervals based on the variability observed in the calibration set, ensuring robust coverage across multiple prediction horizons and spatial resolutions. We empirically evaluate our approach on large-scale ride-sharing and taxi demand datasets spanning diverse urban regions and demonstrate that CP-STGCN significantly improves uncertainty quantification. The results show consistent improvements in interval validity, sharpness, and overall predictive accuracy, confirming the efficacy of our method in capturing both epistemic and aleatoric uncertainties in complex spatio-temporal environments. Xinghua Shi |
ICDM | 3 |
| 2024 | Discriminative Forests Improve Generative Diversity for Generative Adversarial NetworksabstractImproving the diversity of Artificial Intelligence Generated Content (AIGC) is one of the fundamental problems in the theory of generative models such as generative adversarial networks (GANs). Previous studies have demonstrated that the discriminator in GANs should have high capacity and robustness to achieve the diversity of generated data. However, a discriminator with high capacity tends to overfit and guide the generator toward collapsed equilibrium. In this study, we propose a novel discriminative forest GAN, named Forest-GAN, that replaces the discriminator to improve the capacity and robustness for modeling statistics in real-world data distribution. A discriminative forest is composed of multiple independent discriminators built on bootstrapped data. We prove that a discriminative forest has a generalization error bound, which is determined by the strength of individual discriminators and the correlations among them. Hence, a discriminative forest can provide very large capacity without any risk of overfitting, which subsequently improves the generative diversity. With the discriminative forest framework, we significantly improved the performance of AutoGAN with a new record FID of 19.27 from 30.71 on STL10 and improved the performance of StyleGAN2-ADA with a new record FID of 6.87 from 9.22 on LSUN-cat. Junjie Chen 0004, Qingcai Chen, Hongchang Gao, Wendy Hui Wang, Zenglin Xu, Xinghua Shi |
AAAI | 9 |
| 2024 | Joint Participant and Learning Topology Selection for Federated Learning in Edge CloudsabstractDeploying federated learning (FL) in edge clouds poses challenges, especially when multiple models are concurrently trained in resource-constrained edge environments. Existing research on federated edge learning has predominantly focused on client selection for training a single FL model, typically with a fixed learning topology. Preliminary experiments indicate that FL models with adaptable topologies exhibit lower learning costs compared to those with fixed topologies. This paper delves into the intricacies of jointly selecting participants and learning topologies for multiple FL models simultaneously trained in the edge cloud. The problem is formulated as an integer non-linear programming problem, aiming to minimize total learning costs associated with all FL models while adhering to edge resource constraints. To tackle this challenging optimization problem, we introduce a two-stage algorithm that decouples the original problem into two sub-problems and iteratively addresses them separately with efficient heuristics. Our method enhances resource competition and load balancing in edge clouds by allowing FL models to choose participants and learning topologies independently. Extensive experiments conducted with real-world networks and FL datasets affirm the better performance of our algorithm, demonstrating lower average total costs with up to 33.5% and 39.6% compared to previous methods designed for multi-model FL. Xinliang Wei, Kejiang Ye, Xinghua Shi, Cheng-Zhong Xu 0001, Yu Wang 0003 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2022 | Participant Selection for Hierarchical Federated Learning in Edge CloudsabstractFederated learning (FL) has been emerging as a new distributed machine learning paradigm recently. Although FL can protect the data privacy of participants by keeping their training data on local devices, there are recent works raising new privacy concerns especially when workers or the parameter server of FL are untrustworthy or malicious. One effective way to solve the problem is using hierarchical federated learning (HFL) where a few middle-layer aggregators (or called group leaders) are used to aggregate local model updates from workers and send group model updates to the parameter server. In this paper, we consider the participant selection problem of HFL in an edge cloud with multiple FL models, where each model needs to select one parameter server, a few group leaders and a certain amount of workers from edge servers to jointly perform HFL. We first formulate this problem as a non-linear integer programming, aiming to minimize the total learning cost of all models while satisfying the constrained edge resources. We then design a three-stage algorithm by decoupling the original problem into three sub-problems and solving them iteratively. Simulations with real-world datasets and FL models confirm that our proposed algorithm can efficiently reduce the average total learning cost in edge cloud compared with existing methods. Xinliang Wei, Jiyao Liu, Xinghua Shi, Yu Wang 0003 |
NAS | 3 |
| 2021 | On the Convergence of Stochastic Compositional Gradient Descent Ascent MethodabstractThe compositional minimax problem covers plenty of machine learning models such as the distributionally robust compositional optimization problem. However, it is yet another understudied problem to optimize the compositional minimax problem. In this paper, we develop a novel efficient stochastic compositional gradient descent ascent method for optimizing the compositional minimax problem. Moreover, we establish the theoretical convergence rate of our proposed method. To the best of our knowledge, this is the first work achieving such a convergence rate for the compositional minimax problem. Finally, we conduct extensive experiments to demonstrate the effectiveness of our proposed method. Hongchang Gao, Xiaoqian Wang 0001, Xinghua Shi |
IJCAI | 4 |
| 2021 | PAR-GAN: Improving the Generalization of Generative Adversarial Networks Against Membership Inference AttacksabstractRecent works have shown that Generative Adversarial Networks (GANs) may generalize poorly and thus are vulnerable to privacy attacks. In this paper, we seek to improve the generalization of GANs from a perspective of privacy protection, specifically in terms of defending against the membership inference attack (MIA) which aims to infer whether a particular sample was used for model training. We design a GAN framework, partition GAN (PAR-GAN), which consists of one generator and multiple discriminators trained over disjoint partitions of the training data. The key idea of PAR-GAN is to reduce the generalization gap by approximating a mixture distribution of all partitions of the training data. Our theoretical analysis shows that PAR-GAN can achieve global optimality just like the original GAN. Our experimental results on simulated data and multiple popular datasets demonstrate that PAR-GAN can improve the generalization of GANs while mitigating information leakage induced by MIA. Junjie Chen 0004, Wendy Hui Wang, Hongchang Gao, Xinghua Shi |
KDD | 4 |
| 2021 | DeepDSC: A Deep Learning Method to Predict Drug Sensitivity of Cancer Cell LinesabstractHigh-throughput screening technologies have provided a large amount of drug sensitivity data for a panel of cancer cell lines and hundreds of compounds. Computational approaches to analyzing these data can benefit anticancer therapeutics by identifying molecular genomic determinants of drug sensitivity and developing new anticancer drugs. In this study, we have developed a deep learning architecture to improve the performance of drug sensitivity prediction based on these data. We integrated both genomic features of cell lines and chemical information of compounds to predict the half maximal inhibitory concentrations [Formula: see text] on the Cancer Cell Line Encyclopedia (CCLE) and the Genomics of Drug Sensitivity in Cancer (GDSC) datasets using a deep neural network, which we called DeepDSC. Specifically, we first applied a stacked deep autoencoder to extract genomic features of cell lines from gene expression data, and then combined the compounds' chemical features to these genomic features to produce final response data. We conducted 10-fold cross-validation to demonstrate the performance of our deep model in terms of root-mean-square error (RMSE) and coefficient of determination [Formula: see text]. We show that our model outperforms the previous approaches with RMSE of 0.23 and [Formula: see text] of 0.78 on CCLE dataset, and RMSE of 0.52 and [Formula: see text] of 0.78 on GDSC dataset, respectively. Moreover, to demonstrate the prediction ability of our models on novel cell lines or novel compounds, we left cell lines originating from the same tissue and each compound out as the test sets, respectively, and the rest as training sets. The performance was comparable to other methods. Min Li 0007, Yake Wang, Ruiqing Zheng, Xinghua Shi, Yaohang Li, Fang-Xiang Wu, Jianxin Wang 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2020 | A parallelized strategy for epistasis analysis based on Empirical Bayesian Elastic Net modelsabstractMOTIVATION: Epistasis reflects the distortion on a particular trait or phenotype resulting from the combinatorial effect of two or more genes or genetic variants. Epistasis is an important genetic foundation underlying quantitative traits in many organisms as well as in complex human diseases. However, there are two major barriers in identifying epistasis using large genomic datasets. One is that epistasis analysis will induce over-fitting of an over-saturated model with the high-dimensionality of a genomic dataset. Therefore, the problem of identifying epistasis demands efficient statistical methods. The second barrier comes from the intensive computing time for epistasis analysis, even when the appropriate model and data are specified. RESULTS: In this study, we combine statistical techniques and computational techniques to scale up epistasis analysis using Empirical Bayesian Elastic Net (EBEN) models. Specifically, we first apply a matrix manipulation strategy for pre-computing the correlation matrix and pre-filter to narrow down the search space for epistasis analysis. We then develop a parallelized approach to further accelerate the modeling process. Our experiments on synthetic and empirical genomic data demonstrate that our parallelized methods offer tens of fold speed up in comparison with the classical EBEN method which runs in a sequential manner. We applied our parallelized approach to a yeast dataset, and we were able to identify both main and epistatic effects of genetic variants associated with traits such as fitness. AVAILABILITY AND IMPLEMENTATION: The software is available at github.com/shilab/parEBEN. Colby T. Ford, Daniel Janies, Xinghua Shi |
Bioinform. | 4 |
| 2020 | Accelerating bioinformatics research with International Conference on Intelligent Biology and Medicine 2020abstractThe International Association for Intelligent Biology and Medicine (IAIBM) is a nonprofit organization that promotes intelligent biology and medical science. It hosts an annual International Conference on Intelligent Biology and Medicine (ICIBM), which was initially established in 2012. Due to the coronavirus (COVID-19) pandemic, the ICIBM 2020 was held for the first time as a virtual online conference on August 9 to 10. The virtual conference had ~ 300 registered participants and featured 41 online real-time presentations. ICIBM 2020 received a total of 75 manuscript submissions, and 12 were selected to be published in this special issue of BMC Bioinformatics. These 12 manuscripts cover a wide range of bioinformatics topics including network analysis, imaging analysis, machine learning, gene expression analysis, and sequence analysis. Li Shen 0001, Xinghua Shi, Kai Wang 0063, Yulin Dai, Zhongming Zhao |
BMC Bioinform. | 3 |
| 2019 | A Survey on RFID Security and Privacy in Smart Medical: Threats and ProtectionsabstractIn recent years, with the rapid development of the Internet of things, smart medical has been gradually integrated into people's lives. Among them, RFID technology in the Internet of things is the most prominent application in the medical and health industry. However, these emerging technologies will bring many security and privacy problems when they are integrated into people's lives. After examining the possible security and privacy threats brought by RFID in smart medical, this paper surveys security requirements most suitable for this industry, and then compares all current RFID security and privacy protection technologies to analysis whether they are suitable for the smart medical industry. Next, the survey sets out the most far-reaching RFID standards and summarized their advantages and shortcomings in smart medical. Finally, the survey puts forward constructive suggestions on security and privacy protections for hospitals and patients involved. Xinghua Shi, Jinxuan Cao, Tianliang Lu, Victor Chang 0001 |
IoTBDS | 1 |
| 2019 | Bayesian Network Construction and Genotype-Phenotype Inference Using GWAS StatisticsabstractGenome-wide association studies (GWASs) have received increasing attention to understand how genetic variation affects different human traits. In this paper, we study whether and to what extent exploiting the GWAS statistics can be used for inferring private information about a human individual. We first provide a method to construct a three-layered Bayesian network explicitly revealing the conditional dependency between single-nucleotide polymorphisms (SNPs) and traits from public GWAS catalog. The key challenge in building a Bayesian network from GWAS statistics is the specification of the conditional probability table of a variable with multiple parent variables. We employ the models of independence of causal influences which assume that the causal mechanism of each parent variable is mutually independent. We then formulate three inference problems based on the dependency relationship captured in the Bayesian network, namely trait inference given SNP genotype, genotype inference given trait, and trait inference given known traits, and develop efficient formulas and algorithms. Different from previous work, the possible target of these inference problems we study may be any individual, not limited to GWAS participants. Empirical evaluations show the effectiveness of our proposed methods. In summary, our work implies that meaningful information can be inferred from modeling GWAS statistics, and appropriate privacy protection mechanisms need to be developed to protect genetic privacy not only of GWAS participants but also regular individuals. Lu Zhang 0021, Qiuping Pan, Yue Wang 0009, Xintao Wu, Xinghua Shi |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2018 | Using Deep Neural Network to Predict Drug Sensitivity of Cancer Cell Lines
Yake Wang, Min Li 0007, Ruiqing Zheng, Xinghua Shi, Yaohang Li, Fang-Xiang Wu, Jianxin Wang 0001 |
ICIC (2) | 4 |
| 2018 | Participant Grouping for Privacy Preservation in Mobile Crowdsensing over Hierarchical Edge CloudsabstractIn mobile crowdsensing (MCS), to select the optimal set of participants for a particular sensing task, the cloud-based MCS platform requires mobile users to submit their bids and their sensing quality data. This can cause privacy breaches. One possible solution is to leverage secure sharing or bidding schemes to protect participants' personal information during selection. However, these schemes suffer from high overheads, poor scalability and more importantly, the group formation has never been studied. To address this issue and to enhance the protection of user privacy, we propose a set of novel privacy-preserving grouping methods, which place participants into small groups over hierarchical edge clouds. By doing this, not only can the participants be hidden in groups, but also the overall privacy-preserving participant selection becomes more scalable. The design goal to minimize the communication cost during secure sharing/bidding within groups, while satisfying each participant's requirement for privacy preservation. For different scenarios and optimization functions, we propose a set of grouping schemes to fulfill this goal. Extensive simulations over both synthetic and real-life datasets illustrate the efficiency of proposed mechanisms. Ting Li 0010, Zhijin Qiu, Lijuan Cao, Hanshang Li, Zhongwen Guo, Fan Li 0001, Xinghua Shi, Yu Wang 0003 |
IPCCC | 7 |
| 2018 | Group spike-and-slab lasso generalized linear models for disease prediction and associated genes detection by incorporating pathway informationabstractMotivation: Large-scale molecular data have been increasingly used as an important resource for prognostic prediction of diseases and detection of associated genes. However, standard approaches for omics data analysis ignore the group structure among genes encoded in functional relationships or pathway information. Results: We propose new Bayesian hierarchical generalized linear models, called group spike-and-slab lasso GLMs, for predicting disease outcomes and detecting associated genes by incorporating large-scale molecular data and group structures. The proposed model employs a mixture double-exponential prior for coefficients that induces self-adaptive shrinkage amount on different coefficients. The group information is incorporated into the model by setting group-specific parameters. We have developed a fast and stable deterministic algorithm to fit the proposed hierarchal GLMs, which can perform variable selection within groups. We assess the performance of the proposed method on several simulated scenarios, by varying the overlap among groups, group size, number of non-null groups, and the correlation within group. Compared with existing methods, the proposed method provides not only more accurate estimates of the parameters but also better prediction. We further demonstrate the application of the proposed procedure on three cancer datasets by utilizing pathway structures of genes. Our results show that the proposed method generates powerful models for predicting disease outcomes and detecting associated genes. Availability and implementation: The methods have been implemented in a freely available R package BhGLM (http://www.ssg.uab.edu/bhglm/). Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Zaixiang Tang, Yueping Shen, Chen'ao Qian, Wenzhuo Zhuang, Xinghua Shi, Nengjun Yi |
Bioinform. | 8 |
| 2017 | Lightweight mutual authentication among sensors in body area networks through Physical Unclonable FunctionsabstractMedical sensors are usually attached to or implanted inside patient body. Since sensing results of Body Area Networks (BAN) can directly impact the control of medical equipment, the authenticity and integrity of sensing data is essential for safety of patients. Restricted by the limited resources available to BAN sensors, researchers have referred to Physical Unclonable Function of the nodes to achieve authentication. Existing approaches focus on the authentication between control unit and sensors. Mutual authentication among body sensors has not been carefully studied. In this paper, we propose to design a lightweight mutual authentication mechanism for BAN sensors with physical unclonable functions (PUF). Using control unit as a middle point, a pair of body sensors can establish shared secrets so that authenticity of exchanged data can be protected. The proposed approach does not require sensors to conduct any encryption operations, which suits the restricted resources available to BAN nodes. The analysis shows that the proposed approach has very low overhead and does not introduce new vulnerabilities into the system. Weichao Wang, Xinghua Shi, Tuanfa Qin |
ICC | 3 |
| 2017 | A Sparse Learning Framework for Joint Effect Analysis of Copy Number VariantsabstractCopy number variants (CNVs), including large deletions and duplications, represent an unbalanced change of DNA segments. Abundant in human genomes, CNVs contribute to a large proportion of human genetic diversity, with impact on many human phenotypes. Although recent advances in genetic studies have shed light on the impact of individual CNVs on different traits, the analysis of joint effect of multiple interactive CNVs lags behind from many perspectives. A primary reason is that the large number of CNV combinations and interactions in the human genome make it computationally challenging to perform such joint analysis. To address this challenge, we developed a novel sparse learning framework that combines sparse learning with biological networks to identify interacting CNVs with joint effect on particular traits. We showed that our approach performs well in identifying CNVs with joint phenotypic effect using simulated data. Applied to a real human genomic dataset from the 1,000 Genomes Project, our approach identified multiple CNVs that collectively contribute to population differentiation. We found a set of multiple CNVs that have joint effect in different populations, and affect gene expression differently in distinct populations. These results provided a collection of CNVs that likely have downstream biomedical implications in individuals from diverse population backgrounds. Zhiyong Wang 0005, Benika Hall, Jinbo Xu, Xinghua Shi |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2016 | A predictive model of gene expression using a deep learning frameworkabstractWith an unprecedented amount of data available, it is important to explore new methods for developing predictive models to mine this data for scientific discoveries. In this study, we propose a deep learning regression model based on MultiLayer Perceptron and Stacked Denoising Auto-encoder (MLP-SAE) to predict gene expression from genotypes of genetic variation. Specifically, we use a stacked denoising auto-encoder to train our regression model in order to extract useful features, and utilize the multilayer perceptron for backpropagation. We further improve our model by adding a dropout technique to prevent overfitting. Our results on a real genomic dataset show that our MLP-SAE model with dropout outperform Lasso, Random Forests, and MLP-SAE without dropout. Our study provides a new application of deep learning in mining genomics data, and demonstrates that deep learning has great potentials in building predictive models to help understand biological systems. Andrew Quitadamo, Jianlin Cheng, Xinghua Shi |
BIBM | 4 |
| 2016 | Building Bayesian networks from GWAS statistics based on Independence of Causal InfluenceabstractGenome-wide association studies (GWASs) have received an increasing attention to understand genotype-phenotype relationships. In this paper, we study how to build Bayesian networks from publicly released GWAS statistics to explicitly reveal the conditional dependency between single-nucleotide polymorphisms (SNPs) and traits. The key challenge in building a Bayesian network is the specification of the conditional probability table (CPT) of an variable with multiple parent variables. We employ the Independence of Causal Influences (ICI) which assumes that the causal mechanism of each parent variable is mutually independent. Specifically, we derive a formulation from the Noisy-or model, one of the ICI models, to specify the CPT using the released GWAS statistics. We prove that the specified CPT is accurate as long as the underlying individual-level genotype and phenotype profile data follows the Noisy-or model. We empirically evaluate the Noisy-or model and its derived formulation using data from openSNP. Experimental results demonstrate the effectiveness of our approach. Lu Zhang 0021, Qiuping Pan, Xintao Wu, Xinghua Shi |
BIBM | 4 |
| 2015 | An integrated network of microRNA and gene expression in ovarian cancerabstractBACKGROUND: Ovarian cancer is a deadly female reproductive cancer. Understanding the biological mechanisms underlying ovarian cancer could help lead to quicker and more accurate diagnosis and more effective treatments. Both changes in microRNA(miRNA) expression and miRNA/mRNA dysregulation have been associated with ovarian cancer. With the availability of whole-genome miRNA and mRNA sequencing we now have new potentials to study these associations. In this study, we performed a comprehensive analysis of miRNA and mRNA expression in ovarian cancer using an integrative network approach combined with association analysis. RESULTS: We developed an integrative approach to construct a network that illustrates the complex interplay among miRNA and gene expression from a systems perspective. Our method is composed of expanding networks from eQTL associations, building network associations in eQTL analysis, and then combine the networks into an integrated network. This integrated network takes account of miRNA expression quantitative trait loci (eQTL) associations, miRNAs and their targets, protein-protein interactions, co-expressions among miRNAs and genes respectively. Applied to the ovarian cancer data set from The Cancer Genome Atlas (TCGA), we created an integrated network with 167 nodes containing 108 miRNA-target interactions and 145 from protein-protein interactions, starting from 44 initial eQTLs. This integrated network encompassed 26 genes and 14 miRNAs associated with cancer. In particular, 11 genes and 12 miRNAs in the integrated network are associated with ovarian cancer. CONCLUSION: We demonstrated an integrated network approach that integrates multiple data sources at a systems level. We applied this approach to the TCGA ovarian cancer dataset, and constructed a network that provided a more inclusive view of miRNA and gene expression in ovarian cancer. This network included four separate types of interactions among miRNAs and genes. Simply analyzing each interaction component in isolation, such as the eQTL associations, the miRNA-target interactions or the protein-protein interactions, would create a much more limited network than the integrated one. Andrew Quitadamo, Benika Hall, Xinghua Shi |
BMC Bioinform. | 4 |
| 2013 | Using aggregate human genome data for individual identificationabstractData privacy in genome-wide association studies (GWAS) is a critical yet under-exploited research area. In this paper, we first provide a method to construct a two-layered bayesian network explicitly revealing the conditional dependency between SNPs and traits, from the public GWAS catalog. Then we develop efficient algorithms for two attacks: identity inference attack and trait inference attack based on reasoning with the dependency relationship captured in the constructed bayesian network. Different from previously proposed attacks, the possible target of our attacks may be any common people, not limited to GWAS participants. The empirical evaluations show that unprotected statistics released from GWAS can be exploited by attackers to identify individual or derive private information. Thus we show that mining GWAS statistics threatens the privacy of a much wider population and privacy protection mechanisms should be employed. Yue Wang 0009, Xintao Wu, Xinghua Shi |
BIBM | 3 |
| 2012 | User identification and anonymization in 802.11 wireless LANsabstractABSTRACT Privacy issues have been a serious concern for 802.11 Wireless LAN users. Prior research on privacy issues usually focuses on pseudonym techniques, where the unique, consistent explicit identifiers such as MAC addresses from the users can be frequently changed, and thus it is challenging to track users through those always changing identifiers. However, recent research done by Pang et al. (Pang et al. Proceedings of the 13th Annual ACM International Conference on Mobile Computing and Networking, 1997; 99–110) has demonstrated that pseudonyms are not adequate to protect user privacy. The key idea of Pang et al.'s method is to locate implicit identifiers (e.g., IP addresses and port numbers a user frequently visits), build user behavior patterns based on these implicit identifiers, and then apply classification techniques to identify users. In this paper, we first propose a new 802.11 user identification approach through enhanced feature selection and generation for building more accurate user behavior patterns. Our simulation results on 9.27 GB SIGCOMM 2004 wireless data sets demonstrate that our method can achieve better classification rates compared with Pang et al.'s method. Then, we further study how to provide user anonymity, even if implicit identifiers based identification is applied, by introducing a set of 802.11 user anonymization approaches based on bogus traffic injection. We propose eight different methods to artificially generate bogus data and inject them into original traffic, and thus users' behavior patterns are disturbed. Our simulation results on the same SIGCOMM 2004 data sets demonstrate that our anonymization methods can efficiently decrease user identification rates and hence improve user anonymity. Copyright © 2011 John Wiley & Sons, Ltd. Dingbang Xu, Yu Wang 0003, Xinghua Shi, Xiaohang Yin |
Secur. Commun. Networks | 3 |
| 2011 | Capacity of data collection in randomly-deployed wireless sensor networks
Siyuan Chen 0001, Yu Wang 0003, Xiang-Yang Li 0001, Xinghua Shi |
Wirel. Networks | 4 |
| 2010 | 802.11 User AnonymizationabstractPrivacy issues have been a serious concern for 802.11 Wireless LAN users. As demonstrated by Pang et al. and Xu et al., applying pseudonym techniques does not completely protect users' privacy. In particular, users' identities can be disclosed through implicit identifiers such as the IP addresses and port numbers users often access. In this paper, we study how to improve user anonymity even if implicit identifiers based identification is applied. The basic idea of our approach is to artificially generate bogus data and inject them into original traffic, and thus users' behavior patterns are disturbed. Specifically, we propose eight different methods to generate bogus data, where each of them applies different algorithms and metrics to generate bogus packets. Our simulation results with SIGCOMM 2004 wireless trace demonstrate that our anonymization methods can decrease user identification rates, and hence improve user anonymity. Dingbang Xu, Yu Wang 0003, Xinghua Shi, Xiaohang Yin |
GLOBECOM | 3 |
| 2009 | Data Collection Capacity of Random-Deployed Wireless Sensor NetworksabstractData collection is one of the most important functions provided by wireless sensor networks. Here, we study the theoretical limitations of data collection and data aggregation in terms of delay and capacity for a wireless sensor network where n sensors are randomly deployed. We consider two different communication scenarios (with or without aggregation) under physical interference model. For each scenario, we first propose a new collection method and analyze its performance in terms of delay and capacity, then theoretically prove that our method can achieve the optimal order. Particularly, the capacity of data collection is in order of Θ(W) where W is the fixed data-rate on individual links. If each sensor can aggregate its receiving packets into a single packet to send, the capacity of data collection increases to Θ((n/log n) W). Siyuan Chen 0001, Yu Wang 0003, Xiang-Yang Li 0001, Xinghua Shi |
GLOBECOM | 4 |
| 2009 | Enhanced Feature Selection and Generation for 802.11 User IdentificationabstractTo provide user privacy, several anonymization techniques (e.g., pseudonyms applied to MAC addresses) have been proposed in 802.11 networks. However, recent research done by Pang et al. has demonstrated that pseudonyms are not adequate to protect user privacy. The key idea of Pang et al.'s method is to locate implicit identifiers (e.g., IP addresses and port numbers a user frequently visits), build user profiles based on these implicit identifiers in the training data sets, and then apply classification techniques to identify unlabeled (testing) users. Our method proposed in this paper partly focuses on building user profiles. Compared with the method proposed by Pang et al. in, we propose a novel approach to selecting and generating features, which is critical to build user profiles. The feature selection and generation procedure can be dynamically controlled through setting a few important parameters. We did a series of simulations using 9.27 GB SIGCOMM 2004 wireless data sets, and our simulation results demonstrate better classification rates compared with Pang et al.'s method. Dingbang Xu, Yu Wang 0003, Xinghua Shi |
ICCCN | 3 |
| 2009 | Order-Optimal Data Collection in Wireless Sensor Networks: Delay and CapacityabstractData collection is one of the most important functions provided by wireless sensor networks. In this paper, we study theoretical limitations of data collection and data aggregation in terms of delay and capacity for a wireless sensor network where n sensors are randomly deployed. We consider different communication scenarios (single sink or multiple sinks, regularly-deployed or randomly-deployed sinks, with or without aggregation) under protocol interference model. For each scenario, we first propose a new collection/aggregation method and analyze its performance in terms of delay and capacity, then theoretically prove that our method can achieve the optimal order (i.e., its performance is within a constant factor of the optimal). Particularly, with a single sink, the capacity of data collection is in order of Theta(W) where W is the fixed data-rate on individual links. With k sinks, the capacity of data collection is increased to Theta(kW) when k=O(n/log n) or Theta(n/log n) when k=Omega(n/log n). If each sensor can aggregate its receiving packets into a single packet to send, the capacity of data collection with a single sink is also increased to Theta(n/log nW). Siyuan Chen 0001, Yu Wang 0003, Xiang-Yang Li 0001, Xinghua Shi |
SECON | 4 |
| 2009 | Self-organizing fault-tolerant topology control in large-scale three-dimensional wireless networksabstractTopology control protocol aims to efficiently adjust the network topology of wireless networks in a self-adaptive fashion to improve the performance and scalability of networks. This is especially essential to large-scale multihop wireless networks (e.g., wireless sensor networks). Fault-tolerant topology control has been studied recently. In order to achieve both sparseness (i.e., the number of links is linear with the number of nodes) and fault tolerance (i.e., can survive certain level of node/link failures), different geometric topologies were proposed and used as the underlying network topologies for wireless networks. However, most of the existing topology control algorithms can only be applied to two-dimensional (2D) networks where all nodes are distributed in a 2D plane. In practice, wireless networks may be deployed in three-dimensional (3D) space, such as under water wireless sensor networks in ocean or mobile ad hoc networks among space shuttles in space. This article seeks to investigate self-organizing fault-tolerant topology control protocols for large-scale 3D wireless networks. Our new protocols not only guarantee k -connectivity of the network, but also ensure the bounded node degree and constant power stretch factor even under k −1 node failures. All of our proposed protocols are localized algorithms, which only use one-hop neighbor information and constant messages with small time complexity. Thus, it is easy to update the topology efficiently and self-adaptively for large-scale dynamic networks. Our simulation confirms our theoretical proofs for all proposed 3D topologies. Yu Wang 0003, Lijuan Cao, Teresa A. Dahlberg, Fan Li 0001, Xinghua Shi |
ACM Trans. Auton. Adapt. Syst. | 5 |
| 2005 | Efficient on-demand topology control for wireless ad hoc networksabstractTopology control in wireless ad hoc networks has been heavily studied recently. Different geometric topologies were proposed to be used as the underlying network topologies, in order to achieve the sparseness of the communication network or to guarantee the package delivery of specific routing methods. However, most of the proposed topology control algorithms were applied for all nodes in the networks at the initial network startup stage, and the constructed topologies were kept being maintained thereafter. The overhead of topology control at each node at any time is notable, and this affects the performance of the network and wastes energy for each wireless node. This paper seeks to investigate the practices of efficient on-demand topology control protocol for wireless ad hoc networks. In our new protocol, we only apply the specific topology control technique where and when the wireless node needs it. Our simulation confirms that our scheme has better performance than several existing methods. Yu Wang 0003, Xinghua Shi |
ICCCN | 2 |