Shumin Guo

dblp:07/2263 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4Systems, architecture and hardware · 2Security and privacy · 2 · 1 first-authorArtificial intelligence and machine learning · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Query processing and optimization · 53% Spatial and temporal data management · 27% Data mining · 20%
Computer architecture, parallel and distributed computing, and storage systems
5 papers
Cloud and datacenter computing · 92% Performance modeling and evaluation · 8%
Network and information security
2 papers
Privacy and data protection · 100%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 100%

Topics — the 12 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Spatial and temporal data management › spatial query processing › nearest neighbor query
k-nearest neighbor query
0.212014
Building Confidential and Efficient Query Services in the Cloud with RASP Data Perturbation · IEEE Trans. Knowl. Data Eng. 2014
Query processing and optimization
range query
0.212014
Building Confidential and Efficient Query Services in the Cloud with RASP Data Perturbation · IEEE Trans. Knowl. Data Eng. 2014
Query processing and optimization › secure query processing
secure kNN query
0.212014
Building Confidential and Efficient Query Services in the Cloud with RASP Data Perturbation · IEEE Trans. Knowl. Data Eng. 2014
Cloud and datacenter computing › resource allocation
cloud resource allocation
0.212014
CRESP: Towards Optimal Resource Provisioning for MapReduce Computing in Public Clouds · IEEE Trans. Parallel Distributed Syst. 2014
Cloud and datacenter computing › job scheduling › economic scheduling
cost-aware scheduling
0.212014
CRESP: Towards Optimal Resource Provisioning for MapReduce Computing in Public Clouds · IEEE Trans. Parallel Distributed Syst. 2014
Privacy and data protection › privacy-preserving data analysis
privacy-preserving data mining
0.212013
PerturBoost: Practical Confidential Classifier Learning in the Cloud · ICDM 2013
Data mining
clustering
0.112012
CloudVista: Interactive and Economical Visual Cluster Analysis for Big Data in the Cloud · Proc. VLDB Endow. 2012
Visualization and visual analytics › information visualization
large-scale data visualization
0.112012
CloudVista: Interactive and Economical Visual Cluster Analysis for Big Data in the Cloud · Proc. VLDB Endow. 2012
Privacy and data protection
privacy-preserving machine learning
0.112012
Privacy preserving boosting in the cloud with secure half-space queries · CCS 2012
Cloud and datacenter computing › cloud data management
cloud query service
0.112014
Building Confidential and Efficient Query Services in the Cloud with RASP Data Perturbation · IEEE Trans. Knowl. Data Eng. 2014
Performance modeling and evaluation
cost modeling
0.112014
CRESP: Towards Optimal Resource Provisioning for MapReduce Computing in Public Clouds · IEEE Trans. Parallel Distributed Syst. 2014
Cloud and datacenter computing
cloud data analytics
0.012012
CloudVista: Interactive and Economical Visual Cluster Analysis for Big Data in the Cloud · Proc. VLDB Endow. 2012

Methods — techniques the papers use, named apart from their topics

random space perturbation · 0.7boosting · 0.6randgen · 0.4parallel visualization · 0.4random projection · 0.4order-preserving encryption · 0.4dimensionality expansion · 0.4adaboost · 0.3parameter learning from test runs · 0.2cost function modeling · 0.2perturbation · 0.1
YearPublicationVenuePosition
2018 RASP-Boost: Confidential Boosting-Model Learning with Perturbed Data in the Cloud
abstract
Mining large data requires intensive computing resources and data mining expertise, which might be unavailable for many users. With widely available cloud computing resources, data mining tasks can now be moved to the cloud or outsourced to third parties to save costs. In this new paradigm, data and model confidentiality becomes the major concern to the data owner. Data owners have to understand the potential trade-offs among client-side costs, model quality, and confidentiality to justify outsourcing solutions. In this paper, we propose the RASP-Boost framework to address these problems in confidential cloud-based learning. The RASP-Boost approach works with our previous developed RAndom Space Perturbation (RASP) method to protect data confidentiality and uses the boosting framework to overcome the difficulty of learning high-quality classifiers from RASP perturbed data. We develop several cloudclient collaborative boosting algorithms. These algorithms require low client-side computation and communication costs. The client does not need to stay online in the process of learning models. We have thoroughly studied the confidentiality of data, model, and learning process under a practical security model. Experiments on public datasets show that the RASP-Boost approach can provide high-quality classifiers, while preserving high data and model confidentiality and requiring low client-side costs.
Keke Chen, Shumin Guo
IEEE Trans. Cloud Comput.2
2014 Building Confidential and Efficient Query Services in the Cloud with RASP Data Perturbation
abstract
With the wide deployment of public cloud computing infrastructures, using clouds to host data query services has become an appealing solution for the advantages on scalability and cost-saving. However, some data might be sensitive that the data owner does not want to move to the cloud unless the data confidentiality and query privacy are guaranteed. On the other hand, a secured query service should still provide efficient query processing and significantly reduce the in-house workload to fully realize the benefits of cloud computing. We propose the random space perturbation (RASP) data perturbation method to provide secure and efficient range query and kNN query services for protected data in the cloud. The RASP data perturbation method combines order preserving encryption, dimensionality expansion, random noise injection, and random projection, to provide strong resilience to attacks on the perturbed data and queries. It also preserves multidimensional ranges, which allows existing indexing techniques to be applied to speedup range query processing. The kNN-R algorithm is designed to work with the RASP range query algorithm to process the kNN queries. We have carefully analyzed the attacks on data and queries under a precisely defined threat model and realistic security assumptions. Extensive experiments have been conducted to show the advantages of this approach on efficiency and security.
Huiqi Xu, Shumin Guo, Keke Chen
IEEE Trans. Knowl. Data Eng.2
2014 CRESP: Towards Optimal Resource Provisioning for MapReduce Computing in Public Clouds
abstract
Running MapReduce programs in the cloud introduces this unique problem: how to optimize resource provisioning to minimize the monetary cost or job finish time for a specific job? We study the whole process of MapReduce processing and build up a cost function that explicitly models the relationship among the time cost, the amount of input data, the available system resources (Map and Reduce slots), and the complexity of the Reduce function for the target MapReduce job. The model parameters can be learned from test runs. Based on this cost function, we can solve a number of decision problems, such as the optimal amount of resources that can minimize monetary cost within a job finish deadline, minimize time cost under a certain monetary budget, or find the optimal tradeoffs between time and monetary costs. Experimental results show that the proposed approach performs well on a number of sample MapReduce programs in both the in-house cluster and Amazon EC2. We also conducted a variance analysis on different components of the MapReduce workflow to show the possible sources of modeling error. Our optimization results show that with the proposed approach we can save a significant amount of time and money, compared to randomly selected settings.
Keke Chen, James Powers, Shumin Guo, Fengguang Tian
IEEE Trans. Parallel Distributed Syst.3
2013 PerturBoost: Practical Confidential Classifier Learning in the Cloud
abstract
Mining large data requires intensive computing resources and data mining expertise, which might not be available for many users. With the development of cloud computing and services computing, data mining tasks can now be moved to the cloud or outsourced to third parties to save costs. In this new paradigm, data and model confidentiality becomes the major concern to the data owner. Meanwhile, users are also concerned about the potential tradeoff among costs, model quality, and confidentiality. In this paper, we propose the PerturBoost framework to address the problems in confidential cloud or outsourced learning. PerturBoost combined with the random space perturbation (RASP) method that was also developed by us can effectively protect data confidentiality, model confidentiality, and model quality with low client-side costs. Based on the boosting framework, we develop a number of base learner algorithms that can learn linear classifiers from the RASP-perturbed data. This approach has been evaluated with public datasets. The result shows that the RASP-based PerturBoost can provide model accuracy very close to the classifiers trained with the original data and the AdaBoost method, with high confidentiality guarantee and acceptable costs.
Keke Chen, Shumin Guo
ICDM2
2012 Privacy preserving boosting in the cloud with secure half-space queries
abstract
This poster presents a preliminary study on the PerturBoost approach that aims to provide efficient and secure classifier learning in the cloud with both data and model privacy preserved.
Shumin Guo, Keke Chen
CCS1
2012 CloudVista: Interactive and Economical Visual Cluster Analysis for Big Data in the Cloud
abstract
Analysis of big data has become an important problem for many business and scientific applications, among which clustering and visualizing clusters in big data raise some unique challenges. This demonstration presents the CloudVista prototype system to address the problems with big data caused by using existing data reduction approaches. It promotes a whole-big-data visualization approach that preserves the details of clustering structure. The prototype system has several merits. (1) Its visualization model is naturally parallel, which guarantees the scalability. (2) The visual frame structure minimizes the data transferred between the cloud and the client. (3) The RandGen algorithm is used to achieve a good balance between interactivity and batch processing. (4) This approach is also designed to minimize the financial cost of interactive exploration in the cloud. The demonstration will highlight the problems with existing approaches and show the advantages of the CloudVista approach. The viewers will have the chance to play with the CloudVista prototype system and compare the visualization results generated with different approaches.
Huiqi Xu, Shumin Guo, Keke Chen
Proc. VLDB Endow.3
2011 RASP: efficient multidimensional range query on attack-resilient encrypted databases
abstract
Range query is one of the most frequently used queries for online data analytics. Providing such a query service could be expensive for the data owner. With the development of services computing and cloud computing, it has become possible to outsource large databases to database service providers and let the providers maintain the range-query service. With outsourced services, the data owner can greatly reduce the cost in maintaining computing infrastructure and data-rich applications. However, the service provider, although honestly processing queries, may be curious about the hosted data and received queries. Most existing encryption based approaches require linear scan over the entire database, which is inappropriate for online data analytics on large databases. While a few encryption solutions are more focused on efficiency side, they are vulnerable to attackers equipped with certain prior knowledge. We propose the Random Space Encryption (RASP) approach that allows efficient range search with stronger attack resilience than existing efficiency-focused approaches. We use RASP to generate indexable auxiliary data that is resilient to prior knowledge enhanced attacks. Range queries are securely transformed to the encrypted data space and then efficiently processed with a two-stage processing algorithm. We thoroughly studied the potential attacks on the encrypted data and queries at three different levels of prior knowledge available to an attacker. Experimental results on synthetic and real datasets show that this encryption approach allows efficient processing of range queries with high resilience to attacks.
Keke Chen, Ramakanth Kavuluru, Shumin Guo
CODASPY3
2011 CloudVista: Visual Cluster Exploration for Extreme Scale Data in the Cloud
Keke Chen, Huiqi Xu, Fengguang Tian, Shumin Guo
SSDBM4