VLDB 2026 Research / reviewers in the wild / expert
Chang Liu 0001
dblp:52/5716-1
· DBLP profile ↗
37ranked-venue papers
12as first author
13since 2021 · last 2026
0000-0002-3914-754XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 7 first-author · 4 since 2021Computer networks · 9 · 2 first-author · 8 since 2021Security and privacy · 4 · 2 first-author · 1 since 2021Theory of computation · 3Software engineering, systems software and programming languages · 2Human-computer interaction and ubiquitous computing · 2Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UBEP: Re-architecting Expert Parallelism Communication Library for Production SuperpodsabstractThe deployment of Mixture-of-Experts (MoE) models on production high-bandwidth superpods, such as NVIDIA's NVL72/576 and Huawei's CloudMatrix384, introduces critical challenges beyond raw interconnect bandwidth. While these systems provide unified global address spaces and high-bandwidth fabrics, their full potential for sparse MoE communication is hindered by three fundamental bottlenecks: (1) Strict execution serialization imposed by coarse-grained Bulk Synchronous Parallel (BSP) orchestration of interdependent communication phases; (2) Prohibitive synchronization overhead that fails to scale alongside high interconnect bandwidth; and (3) Severe load imbalance resulting from distance-agnostic scheduling of irregular token traffic. To eliminate these bottlenecks, we introduce UBEP (Unified-Bus Expert Parallelism), a production-ready communication library that rethinks MoE's All-to-All primitives for modern superpod architectures. Through large-scale experiments, UBEP reduces All-to-All latency by up to 52.4% and MoE inference Time Per Output Token (TPOT) by up to 11.1%. Chang Liu 0001, Si Shen, Jiaqi Zheng 0001, Mingfan Li, Yuyang Yang, Guanhua Li, Yuquan Zhang, Zhongzhe Hu, Qihang Duan, Wenkai Ling, Baochuan Yang, Xianzhi Yu, Guihai Chen |
SIGCOMM | 2 |
| 2026 | Anytest: Localizing the Root Cause of Hardware Transport Performance Anomalies
Zhaochen Zhang, Sheng Cheng 0002, Feiyang Xue, Chang Liu 0001, Boliang Liu, Rui Li 0020, Li Wang 0110, Peirui Cao, Qingkai Meng 0001, Guihai Chen, Shuguang Cheng, Yongqing Xi, Binzhang Fu, Dennis Cai, Chen Tian 0001 |
SIGCOMM | 6 |
| 2026 | Revisiting Flow Control in Node-Centric Datacenter NetworksabstractNode-centric Data Centers (NDCs) are highly flexible, cost-efficient, and failure-resilient, and have gained growing popularity in recent years. However, RDMA technology used in NDC still faces challenges, including high retransmission overhead, Head-of-Line Blocking (HoLB) and deadlock problems. Existing solutions for traditional data centers cannot simultaneously address these issues due to the unique topology and server transmission characteristics of NDC. In this paper, we propose a per-port flow control named PortFC for NDC. PortFC addresses the above problems through the designs of a Pause/Resume control signal, a per-port queue allocation method, an egress-detecting per-port flow control mechanism, and a server-aware queue scheduling method. Our evaluation shows that PortFC is free from retransmission, capable of eliminating HoLB and avoiding deadlocks. PortFC achieves 1.7-8.0 times higher throughput and reduces latency by 11.7%-87.7% compared to the state-of-the-art lossy RDMA based on IRN and the lossless RDMA method based on PFC. In particular, PortFC still demonstrates good performance in a Rail-only NDC with heterogeneous bandwidth domains. Peirui Cao, Rui Ning, Guangyu Zhao, Zhaochen Zhang, Chang Liu 0001, Yunzhuo Liu, Rui Li 0020, Chengyuan Huang, Tao Sun 0010, Guihai Chen, Baochun Li, Chen Tian 0001 |
IEEE Trans. Netw. | 5 |
| 2026 | SRViT: A Robust Online Encrypted Traffic Classification Based on Vision TransformerabstractThe dramatic rise in encrypted traffic brings huge challenges to traditional traffic classification methods. Deep learning-based traffic classification methods have been demonstrated to significantly improve performance. However, the following limitations remain: i) It is challenging to concurrently focus on both global and local information in traffic flows, resulting in the absence of important information. ii) The existing methods relying on temporal information suffer from low robustness in case of packet disordering or loss. iii) The use of multi-layer encryption and random routing in Tor technology poses more challenges for traffic identification. In this paper, we propose a novel ViT-based model for more accurate encrypted traffic classification, called SRViT to overcome the above challenges. Firstly, SRViT proposes a novel mechanism of multi-size patch division to learn comprehensive hidden knowledge and dependencies between packets. Secondly, we propose a self-attention operation with a relative position bias to learn the relative position relationship. After that, an incremental update mechanism is proposed to adapt to dynamic changes in the real traffic environment. At last, the comprehensive experiments on 5 real-world encrypted traffic datasets are carried out. The experimental results indicate that SRViT outperforms the state-of-the-art methods with an average accuracy improvement of 24.62% while keeping higher robustness and execution efficiency. Chang Liu 0001, Zulong Diao, Xin He 0010, Weibei Fan, Fu Xiao 0001 |
IEEE Trans. Netw. | 1 |
| 2026 | Analysis of Pyrrha: Congestion-Root-Based Flow Control Is Most Cost-Effective to Eliminate Head-of-Line BlockingabstractIn modern datacenters, the effectiveness of end-to-end congestion control (CC) is quickly diminishing with the rapid bandwidth evolution. Per-hop flow control (FC) can react to congestion more promptly. However, a coarse-grained FC can result in Head-Of-Line (HOL) blocking. A fine-grained, per-flow FC can eliminate HOL blocking caused by flow control, however, it does not scale well. This paper presents Pyrrha, a scalable flow control approach that provably eliminates HOL blocking while using a minimum number of queues. In Pyrrha, flow control first takes effect on the root of the congestion, i.e., the port where congestion occurs. And then flows are controlled according to their contributed congestion roots. A prototype of Pyrrha is implemented on Tofino2 switches. Compared with state-of-the-art approaches, the average FCT of uncongested flows is reduced by 42%-98%, and 99th-tail latency can be$1.6\times $-$215\times $lower, without compromising the performance of congested flows. Zhaochen Zhang, Peirui Cao, Chang Liu 0001, Yizhi Wang 0004, Vamsi Addanki, Stefan Schmid 0001, Qingyue Wang, Xiaoliang Wang 0001, Jiaqi Zheng 0001, Tao Wu 0011, Bingyang Liu, Wan-Chun Dou, Guihai Chen, Chen Tian 0001, Fu Xiao 0001 |
IEEE Trans. Netw. | 4 |
| 2025 | PortFC: Designing High-performance Deadlock-free BCube NetworksabstractBCube is a modular data center network.Compared with other topologies, BCube has natural advantages, such as lower deployment costs and stronger failure recovery capabilities.However, RDMA technology used in BCube still faces challenges, including high retransmission overhead, Head-of-Line Blocking (HoLB) and deadlock problems.Existing solutions for traditional data centers cannot simultaneously address these issues due to the unique topology and server transmission characteristics of BCube.In this paper, we propose a per-port flow control named PortFC for BCube.PortFC addresses the above problems through the designs of a Pause/Resume control signal, a per-port queue allocation method, an egress-detecting per-port flow control mechanism, and a serveraware queue scheduling method.Our evaluation shows that PortFC is free from retransmission, capable of eliminating HoLB and avoiding deadlocks.PortFC achieves 1.7-8.0times higher throughput and reduces latency by 11.7%-87.7%compared to the state-of-the-art Peirui Cao, Rui Ning, Zhaochen Zhang, Chang Liu 0001, Rui Li 0020, Yongqi Yang, Yunzhuo Liu, Chengyuan Huang, Tao Sun 0010, Xiaodong Duan, Guihai Chen, Chen Tian 0001 |
ICS | 5 |
| 2025 | Pyrrha: Congestion-Root-Based Flow Control to Eliminate Head-of-Line Blocking in Datacenter
Zhaochen Zhang, Chang Liu 0001, Yizhi Wang 0004, Vamsi Addanki, Stefan Schmid 0001, Qingyue Wang, Xiaoliang Wang 0001, Jiaqi Zheng 0001, Tao Wu 0011, Bingyang Liu, Wan-Chun Dou, Guihai Chen, Chen Tian 0001 |
NSDI | 3 |
| 2025 | An Anatomy of Token-Based Congestion ControlabstractCongestion control protocols play a vital role in enhancing the performance of various applications within datacenter networks. While reactive congestion control (RCC) protocols are widely deployed in commercial datacenters, the research community has actively explored token-based proactive congestion control (TCC) protocols to further push the boundaries of performance. However, despite the emergence of numerous TCC variants, there has been a lack of systematic exploration in the design space of TCC. This paper aims to bridge this gap by proposing a framework for understanding the design choices within the TCC approach. In this study, we systematically analyze different design choices of TCC approaches and leverage this understanding to develop a novel TCC protocol called ToCC. To implement ToCC, we address a set of challenges and deploy it in NP-based smart NICs. We compare ToCC with state-of-the-art TCC and RCC protocols through extensive large-scale simulations and testbed evaluations. The results demonstrate that ToCC exhibits robustness in achieving low latency across various scenarios. Additionally, ToCC effectively reduces buffer occupancy by 4.8 times compared to existing approaches, and under incast scenarios, it significantly shortens flow completion time by up to 90%. Congestion control protocols are crucial for optimizing the performance of datacenter network applications. Although reactive congestion control (RCC) protocols are commonly used in commercial datacenters, researchers have been exploring token-based proactive congestion control (TCC) protocols to further enhance network performance. Despite the development of numerous TCC variants, there has not been a thorough examination of the design space of TCC protocols until now. This paper aims to address this gap by introducing a framework for understanding the design choices within the TCC approach for TCC protocols. By analyzing various design aspects of TCC approaches, we create a novel TCC protocol called ToCC. At the central of ToCC design is that it leverages congestion control mechanisms over tokens. To implement ToCC, we tackle several challenges and integrate it into NP-based smart NICs. Comparing ToCC with state-of-the-art TCC and RCC protocols through extensive large-scale simulations and testbed evaluations, we find that ToCC consistently achieves low latency across different scenarios. Moreover, ToCC significantly reduces buffer occupancy by 4.8 times compared to existing methods, and during incast scenarios, it decreases flow completion time by up to 90%. Chang Liu 0001, Qingyue Wang, Lu Lu 0016, Xiaoliang Wang 0001, Fu Xiao 0001, Ying Zhang 0022, Wan-Chun Dou, Guihai Chen, Chen Tian 0001 |
IEEE Trans. Netw. | 2 |
| 2024 | Unison: A Parallel-Efficient and User-Transparent Network Simulation KernelabstractDiscrete-event simulation (DES) is a prevalent tool for evaluating network designs. Although DES offers full fidelity and generality, its slow performance limits its application. To speed up DES, many network simulators employ parallel discrete-event simulation (PDES). However, adapting existing network simulation models to PDES requires complex reconfigurations and often yields limited performance improvement. In this paper, we address this gap by proposing a parallel-efficient and user-transparent network simulation kernel, Unison, that adopts fine-grained partition and load-adaptive scheduling optimized for network scenarios. We prototype Unison based on ns-3. Existing network simulation models of ns-3 can be seamlessly transitioned to Unison. Testbed experiments on commodity servers demonstrate that Unison can achieve a 40× speedup over DES using 24 CPU cores, and a 10× speedup compared with existing PDES algorithms under the same CPU cores. Songyuan Bai, Chen Tian 0001, Xiaoliang Wang 0001, Chang Liu 0001, Xin Jin 0008, Fu Xiao 0001, Qiao Xiang, Wan-Chun Dou, Guihai Chen |
EuroSys | 5 |
| 2023 | Identifying Performance Bottleneck in Shared In-Network Aggregation during Distributed TrainingabstractAs the emergence of recently popular large language model, distributed training (DT) optimizes the performance via using different parallelization strategies, resource schedulers and advanced compression techniques. Meanwhile, a promising acceleration primitive, In-Network Aggregation (INA), offloads the gradient aggregation to programmable switches to further reduce the communication overhead using the switch memory. However, to the best of our knowledge, how to identify the performance bottleneck in real time remains challenging. In this paper, we build Argus, a performance bottleneck monitoring framework for INA. Argus implements an aggregation digest extracting mechanism for real-time monitoring of DT jobs at the multi-tenant, multi-rack clusters. Argus models aggregation to identify performance bottlenecks in aggregation, which assists the scheduler in deciding the resource allocation. Extensive evaluation and prototype implementation show that Argus provides real-time and packet-level aggregation monitoring for identifying bottlenecks in INA with minimal performance overhead. Chang Liu 0001, Jiaqi Zheng 0001, Wenfei Wu, Bohan Zhao, Guihai Chen |
ICPADS | 1 |
| 2023 | Swing: Providing Long-Range Lossless RDMA via PFC-RelayabstractRemote Direct Memory Access (RDMA) has been widely deployed in datacenters for its high performance. Large-scale high performance cloud services built on geographically distributed datacenters require long-range RDMA for performance requirements. However, existing RDMA solutions can hardly satisfy the stringent requirements of the emerging large-scale high-performance cloud services built on geo-distributed datacenters in terms of throughput and delay. On the one hand, lossless RDMA suffers from a deep buffer and potential suboptimal throughput for inter-datacenter traffic due to delayed response to Priority Flow Control (PFC) messages. On the other hand, lossy RDMA with selective retransmissions suffers from poor performance when multiple flows with different round-trip times (RTTs) coexist in cross-datacenter scenarios. This article proposesSwing, which expands the high-performance lossless RDMA to long-distance links through PFC-Relay.Swingensures the throughput of long-distance links while minimizing the buffer requirement for long-range RDMA. It enables long-range RDMA without making any modifications to existing in-datacenter networks. The evaluation shows thatSwingcan reduce the average flow completion time (FCT) by 14%-66% in a variety of traffic scenarios. Chen Tian 0001, Jiaqing Dong, Xu Zhang 0006, Chang Liu 0001, Nai Xia, Wan-Chun Dou, Guihai Chen |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2022 | FlyMon: enabling on-the-fly task reconfiguration for network measurementabstractNetwork measurement is important to data center operators. Most existing efforts focus on developing new implementation schemes for measurement tasks. Little attention is paid to on-the-fly task reconfiguration. Due to resource constraints, it is impossible to configure all needed tasks at start-up and dynamically turn on/of them. To support real-time reconfiguration of many different tasks, a key observation is that it is unnecessary to bind a task and its implementation at the compilation phase. We design FlyMon, the first sketch-based measurement system that can make on-the-fly reconfigurations on a large set of measurement tasks. FlyMon introduces the concept of Composable Measurement Units (CMUs), which are general operation units that support reconfigurable implementation for measurement tasks combined from different flow keys and flow attributes. FlyMon maps the design of CMUs to programmable switches' data planes so that the number of compacted CMUs can be maximized. FlyMon also provides dynamic memory management. We prototype FlyMon on Tofino and currently enable four frequently used flow attributes. Each CMU Group (with 3 CMUs) can concurrently perform up to 96 isolated measurement tasks with less than 8.3% hardware resources. The tasks can be deployed with configurable memory size at the millisecond level. By cross-stacking, FlyMon can deploy up to 27 CMUs in one pipeline of Tofino. Chen Tian 0001, Tong Yang 0003, Chang Liu 0001, Zhaochen Zhang, Wan-Chun Dou, Guihai Chen |
SIGCOMM | 5 |
| 2021 | Enabling Privacy-Preserving Shortest Distance Queries on Encrypted Graph DataabstractWhen coming to perform shortest distance queries on encrypted graph data outsourced in external storage infrastructure such as cloud, a significant challenge is how to compute the shortest distance in an accurate, efficient and secure way. This issue is addressed by a recent work, which makes use of somewhat homomorphic encryption (SWHE) to encrypt distance values output by a 2-hop cover labeling (2HCL) scheme. However, it may import large errors and even yield negative results. Besides, SWHE would be too inefficient for normal clients. In this paper, we propose GENOA, a novel Graph ENcryption scheme for shOrtest distAnce queries. GENOA employs only efficient symmetric-key primitives while significantly enhances the accuracy compared to the prior work. As a reasonable trade-off, it additionally reveals the order information among queried distance values in the 2HCL index. We theoretically prove the accuracy and security of GENOA under rigorous cryptographic model. Detailed experiments on eight real-world graphs demonstrate that GENOA is efficient and can produce almost exact results. Chang Liu 0001, Liehuang Zhu, Xiangjian He, Jinjun Chen |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2019 | Cross-Layer Multi-Cloud Real-Time Application QoS Monitoring and Benchmarking As-a-Service FrameworkabstractCloud computing provides on-demand access to affordable hardware (e.g., multi-core CPUs, GPUs, disks, and networking equipment) and software (e.g., databases, application servers and data processing frameworks) platforms with features such as elasticity, pay-per-use, low upfront investment and low time to market. This has led to the proliferation of business critical applications that leverage various cloud platforms. Such applications hosted on single/multiple cloud provider platforms have diverse characteristics requiring extensive monitoring and benchmarking mechanisms to ensure run-time Quality of Service (QoS) (e.g., latency and throughput). This paper proposes, develops and validates CLAMBS-Cross-Layer Multi-Cloud Application Monitoring and Benchmarking as-a-Service for efficient QoS monitoring and benchmarking of cloud applications hosted on multi-clouds environments. The major highlight of CLAMBS is its capability of monitoring and benchmarking individual application components such as databases and web servers, distributed across cloud layers (*-aaS), spread among multiple cloud providers. We validate CLAMBS using prototype implementation and extensive experimentation and show that CLAMBS efficiently monitors and benchmarks application components on multi-cloud platforms including Amazon EC2 and Microsoft Azure. Khalid Alhamazani, Rajiv Ranjan 0001, Prem Prakash Jayaraman, Karan Mitra, Chang Liu 0001, Fethi A. Rabhi, Dimitrios Georgakopoulos 0001, Lizhe Wang 0001 |
IEEE Trans. Cloud Comput. | 5 |
| 2018 | Privacy Issues in Big Data Mining Infrastructure, Platforms, and Applications
Xuyun Zhang, Julian Jang, Lianyong Qi, Md. Zakirul Alam Bhuiyan, Chang Liu 0001 |
Secur. Commun. Networks | 5 |
| 2017 | Efficient searchable symmetric encryption for storing multiple source dynamic social data on cloud
Chang Liu 0001, Liehuang Zhu, Jinjun Chen |
J. Netw. Comput. Appl. | 1 |
| 2017 | Graph Encryption for Top-K Nearest Keyword Search Queries on CloudabstractDriven by the growing security demands of data outsourcing applications in sustainable smart cities, encrypting clients' data has been widely accepted by academia and industry. Data encryptions should be done at the client side before outsourcing, because clouds and edges are not trusted. Therefore, how to properly encrypt data in a way that the encrypted and remotely stored data can still be queried has become a challenging issue. Though keyword searches over encrypted textual data have been extensively studied, approaches for encrypting graph-structured data with support for answering graph queries are still lacking in the literature. In this paper, we specially investigate graph encryption method for an important graph query type, called top-k Nearest Keyword (kNK) searches. We design several indexes to store necessary information for answering queries and guarantee that private information about the graph such as vertex identifiers, keywords and edges are encrypted or excluded. Security and efficiency of our graph encryption scheme are demonstrated by theoretical proofs and experiments on real-world datasets, respectively. Chang Liu 0001, Liehuang Zhu, Jinjun Chen |
IEEE Trans. Sustain. Comput. | 1 |
| 2016 | HKE-BC: hierarchical key exchange for secure scheduling and auditing of big data in cloud computingabstractSummary Big data is one of the most referred key words in recent information and communications technology industry. As the new‐generation distributed computing platform, cloud environments offer high efficiency and low cost for data‐intensive storage and computation for big data applications. Cloud resources and services are available in pay‐as‐you‐go mode, which brings extraordinary flexibility and cost‐effectiveness as well as minimal investments in their own computing infrastructure. However, these advantages come at a price—people no longer have direct control over their own data. Based on this view, data security becomes a major concern in the adoption of cloud computing. Authenticated key exchange is essential to a security system that is based on high‐efficiency symmetric‐key encryptions. With virtualisation technology being applied, existing key exchange schemes such as Internet key exchange become time consuming when directly deployed into cloud computing environment, especially for large‐scale tasks that involve intensive user–cloud interactions, such as scheduling and data auditing. In this paper, we propose a novel hierarchical key exchange scheme, namely hierarchical key exchange for big data in cloud, which aims at providing efficient security‐aware scheduling and auditing for cloud environments. In this novel key exchange scheme, we developed a two‐phase layer‐by‐layer iterative key exchange strategy to achieve more efficient authenticated key exchange without sacrificing the level of data security. Both theoretical analysis and experimental results demonstrate that when deployed in cloud environments with diverse server layouts, efficiency of the proposed scheme is dramatically superior to its predecessors cloud computing background key exchange and Internet key exchange schemes. Copyright © 2014 John Wiley & Sons, Ltd. Chang Liu 0001, Nick Beaugeard, Chi Yang, Xuyun Zhang, Jinjun Chen |
Concurr. Comput. Pract. Exp. | 1 |
| 2015 | External integrity verification for outsourced big data in cloud and IoT: A big picture
Chang Liu 0001, Chi Yang, Xuyun Zhang, Jinjun Chen |
Future Gener. Comput. Syst. | 1 |
| 2015 | A cloud-based framework for Home-diagnosis service over big medical data
Wenmin Lin, Wan-Chun Dou, Zuojian Zhou, Chang Liu 0001 |
J. Syst. Softw. | 4 |
| 2015 | MuR-DPA: Top-Down Levelled Multi-Replica Merkle Hash Tree Based Secure Public Auditing for Dynamic Big Data Storage on CloudabstractCloud computing that provides elastic computing and storage resource on demand has become increasingly important due to the emergence of “big data”. Cloud computing resources are a natural fit for processing big data streams as they allow big data application to run at a scale which is required for handling its complexities (data volume, variety and velocity). With the data no longer under users' direct control, data security in cloud computing is becoming one of the most concerns in the adoption of cloud computing resources. In order to improve data reliability and availability, storing multiple replicas along with original datasets is a common strategy for cloud service providers. Public data auditing schemes allow users to verify their outsourced data storage without having to retrieve the whole dataset. However, existing data auditing techniques suffers from efficiency and security problems. First, for dynamic datasets with multiple replicas, the communication overhead for update verifications is very large, because each update requires updating of all replicas, where verification for each update requires O(log n ) communication complexity. Second, existing schemes cannot provide public auditing and authentication of block indices at the same time. Without authentication of block indices, the server can build a valid proof based on data blocks other than the blocks client requested to verify. In order to address these problems, in this paper, we present a novel public auditing scheme named MuR-DPA. The new scheme incorporated a novel authenticated data structure (ADS) based on the Merkle hash tree (MHT), which we call MR-MHT. To support full dynamic data updates and authentication of block indices, we included rank and level values in computation of MHT nodes. In contrast to existing schemes, level values of nodes in MR-MHT are assigned in a top-down order, and all replica blocks for each data block are organized into a same replica sub-tree. Such a configuration allows efficient verification of updates for multiple replicas. Compared to existing integrity verification and public auditing schemes, theoretical analysis and experimental results show that the proposed MuR-DPA scheme can not only incur much less communication overhead for both update verification and integrity verification of cloud datasets with multiple replicas, but also provide enhanced security against dishonest cloud service providers. Chang Liu 0001, Rajiv Ranjan 0001, Chi Yang, Xuyun Zhang, Lizhe Wang 0001, Jinjun Chen |
IEEE Trans. Computers | 1 |
| 2015 | Proximity-Aware Local-Recoding Anonymization with MapReduce for Scalable Big Data Privacy Preservation in CloudabstractCloud computing provides promising scalable IT infrastructure to support various processing of a variety of big data applications in sectors such as healthcare and business. Data sets like electronic health records in such applications often contain privacy-sensitive information, which brings about privacy concerns potentially if the information is released or shared to third-parties in cloud. A practical and widely-adopted technique for data privacy preservation is to anonymize data via generalization to satisfy a given privacy model. However, most existing privacy preserving approaches tailored to small-scale data sets often fall short when encountering big data, due to their insufficiency or poor scalability. In this paper, we investigate the local-recoding problem for big data anonymization against proximity privacy breaches and attempt to identify a scalable solution to this problem. Specifically, we present a proximity privacy model with allowing semantic proximity of sensitive values and multiple sensitive attributes, and model the problem of local recoding as a proximity-aware clustering problem. A scalable two-phase clustering approach consisting of a t-ancestors clustering (similar to k-means) algorithm and a proximity-aware agglomerative clustering algorithm is proposed to address the above problem. We design the algorithms with MapReduce to gain high scalability by performing data-parallel computation in cloud. Extensive experiments on real-life data sets demonstrate that our approach significantly improves the capability of defending the proximity privacy breaches, the scalability and the time-efficiency of local-recoding anonymization over existing approaches. Xuyun Zhang, Wan-Chun Dou, Jian Pei 0001, Surya Nepal, Chi Yang, Chang Liu 0001, Jinjun Chen |
IEEE Trans. Computers | 6 |
| 2015 | A Time Efficient Approach for Detecting Errors in Big Sensor Data on CloudabstractBig sensor data is prevalent in both industry and scientific research applications where the data is generated with high volume and velocity it is difficult to process using on-hand database management tools or traditional data processing applications. Cloud computing provides a promising platform to support the addressing of this challenge as it provides a flexible stack of massive computing, storage, and software services in a scalable manner at low cost. Some techniques have been developed in recent years for processing sensor data on cloud, such as sensor-cloud. However, these techniques do not provide efficient support on fast detection and locating of errors in big sensor data sets. For fast data error detection in big sensor data sets, in this paper, we develop a novel data error detection approach which exploits the full computation potential of cloud platform and the network feature of WSN. Firstly, a set of sensor data error types are classified and defined. Based on that classification, the network feature of a clustered WSN is introduced and analyzed to support fast error detection and location. Specifically, in our proposed approach, the error detection is based on the scale-free network topology and most of detection operations can be conducted in limited temporal or spatial data blocks instead of a whole big data set. Hence the detection and location process can be dramatically accelerated. Furthermore, the detection and location tasks can be distributed to cloud platform to fully exploit the computation power and massive storage. Through the experiment on our cloud computing platform of U-Cloud, it is demonstrated that our proposed approach can significantly reduce the time for error detection and location in big data sets generated by large scale sensor network systems with acceptable error detecting accuracy. Chi Yang, Chang Liu 0001, Xuyun Zhang, Surya Nepal, Jinjun Chen |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2014 | Search pattern leakage in searchable encryption: Attacks and new construction
Chang Liu 0001, Liehuang Zhu, Mingzhong Wang, Yu-an Tan 0001 |
Inf. Sci. | 1 |
| 2014 | A spatiotemporal compression based approach for efficient big data processing on Cloud
Chi Yang, Xuyun Zhang, Changmin Zhong, Chang Liu 0001, Jian Pei 0001, Kotagiri Ramamohanarao, Jinjun Chen |
J. Comput. Syst. Sci. | 4 |
| 2014 | A hybrid approach for scalable sub-tree anonymization over big data using MapReduce on cloud
Xuyun Zhang, Chang Liu 0001, Surya Nepal, Chi Yang, Wan-Chun Dou, Jinjun Chen |
J. Comput. Syst. Sci. | 2 |
| 2014 | Authorized Public Auditing of Dynamic Big Data Storage on Cloud with Efficient Verifiable Fine-Grained UpdatesabstractCloud computing opens a new era in IT as it can provide various elastic and scalable IT services in a pay-as-you-go fashion, where its users can reduce the huge capital investments in their own IT infrastructure. In this philosophy, users of cloud storage services no longer physically maintain direct control over their data, which makes data security one of the major concerns of using cloud. Existing research work already allows data integrity to be verified without possession of the actual data file. When the verification is done by a trusted third party, this verification process is also called data auditing, and this third party is called an auditor. However, such schemes in existence suffer from several common drawbacks. First, a necessary authorization/authentication process is missing between the auditor and cloud service provider, i.e., anyone can challenge the cloud service provider for a proof of integrity of certain file, which potentially puts the quality of the so-called ‘auditing-as-a-service’ at risk; Second, although some of the recent work based on BLS signature can already support fully dynamic data updates over fixed-size data blocks, they only support updates with fixed-sized blocks as basic unit, which we call coarse-grained updates. As a result, every small update will cause re-computation and updating of the authenticator for an entire file block, which in turn causes higher storage and communication overheads. In this paper, we provide a formal analysis for possible types of fine-grained data updates and propose a scheme that can fully support authorized auditing and fine-grained update requests. Based on our scheme, we also propose an enhancement that can dramatically reduce communication overheads for verifying small updates. Theoretical analysis and experimental results demonstrate that our scheme can offer not only enhanced security and flexibility, but also significantly lower overhead for big data applications with a large number of frequent small updates, such as applications in social media and business transactions. Chang Liu 0001, Jinjun Chen, Laurence T. Yang, Xuyun Zhang, Chi Yang, Rajiv Ranjan 0001, Kotagiri Ramamohanarao |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2014 | A Scalable Two-Phase Top-Down Specialization Approach for Data Anonymization Using MapReduce on CloudabstractA large number of cloud services require users to share private data like electronic health records for data analysis or mining, bringing privacy concerns. Anonymizing data sets via generalization to satisfy certain privacy requirements such as k-anonymity is a widely used category of privacy preserving techniques. At present, the scale of data in many cloud applications increases tremendously in accordance with the Big Data trend, thereby making it a challenge for commonly used software tools to capture, manage, and process such large-scale data within a tolerable elapsed time. As a result, it is a challenge for existing anonymization approaches to achieve privacy preservation on privacy-sensitive large-scale data sets due to their insufficiency of scalability. In this paper, we propose a scalable two-phase top-down specialization (TDS) approach to anonymize large-scale data sets using the MapReduce framework on cloud. In both phases of our approach, we deliberately design a group of innovative MapReduce jobs to concretely accomplish the specialization computation in a highly scalable way. Experimental evaluation results demonstrate that with our approach, the scalability and efficiency of TDS can be significantly improved over existing approaches. Xuyun Zhang, Laurence T. Yang, Chang Liu 0001, Jinjun Chen |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2013 | SaC-FRAPP: a scalable and cost-effective framework for privacy preservation over big data on cloudabstractSUMMARY Big data and cloud computing are two disruptive trends nowadays, provisioning numerous opportunities to the current information technology industry and research communities while posing significant challenges on them as well. Cloud computing provides powerful and economical infrastructural resources for cloud users to handle ever increasing data sets in big data applications. However, processing or sharing privacy‐sensitive data sets on cloud probably engenders severe privacy concerns because of multi‐tenancy. Data encryption and anonymization are two widely‐adopted ways to combat privacy breach. However, encryption is not suitable for data that are processed and shared frequently, and anonymizing big data and manage numerous anonymized data sets are still challenges for traditional anonymization approaches. As such, we propose a scalable and cost‐effective framework for privacy preservation over big data on cloud in this paper. The key idea of the framework is that it leverages cloud‐based MapReduce to conduct data anonymization and manage anonymous data sets, before releasing data to others. The framework provides a holistic conceptual foundation for privacy preservation over big data. Further, a corresponding proof‐of‐concept prototype system is implemented. Empirical evaluations demonstrate that scalable and cost‐effective framework for privacy preservation can anonymize large‐scale data sets and mange anonymous data sets in a highly flexible, scalable, efficient, and cost‐effective fashion. Copyright © 2013 John Wiley & Sons, Ltd. Xuyun Zhang, Chang Liu 0001, Surya Nepal, Chi Yang, Wan-Chun Dou, Jinjun Chen |
Concurr. Comput. Pract. Exp. | 2 |
| 2013 | CCBKE - Session key negotiation for fast and secure scheduling of scientific applications in cloud computing
Chang Liu 0001, Xuyun Zhang, Chi Yang, Jinjun Chen |
Future Gener. Comput. Syst. | 1 |
| 2013 | An efficient quasi-identifier index based approach for privacy preservation over incremental data sets on cloud
Xuyun Zhang, Chang Liu 0001, Surya Nepal, Jinjun Chen |
J. Comput. Syst. Sci. | 2 |
| 2013 | A Privacy Leakage Upper Bound Constraint-Based Approach for Cost-Effective Privacy Preserving of Intermediate Data Sets in CloudabstractCloud computing provides massive computation power and storage capacity which enable users to deploy computation and data-intensive applications without infrastructure investment. Along the processing of such applications, a large volume of intermediate data sets will be generated, and often stored to save the cost of recomputing them. However, preserving the privacy of intermediate data sets becomes a challenging problem because adversaries may recover privacy-sensitive information by analyzing multiple intermediate data sets. Encrypting ALL data sets in cloud is widely adopted in existing approaches to address this challenge. But we argue that encrypting all intermediate data sets are neither efficient nor cost-effective because it is very time consuming and costly for data-intensive applications to en/decrypt data sets frequently while performing any operation on them. In this paper, we propose a novel upper bound privacy leakage constraint-based approach to identify which intermediate data sets need to be encrypted and which do not, so that privacy-preserving cost can be saved while the privacy requirements of data holders can still be satisfied. Evaluation results demonstrate that the privacy-preserving cost of intermediate data sets can be significantly reduced with our approach over existing ones where all data sets are encrypted. Xuyun Zhang, Chang Liu 0001, Surya Nepal, Suraj Pandey, Jinjun Chen |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2012 | An Association Probability Based Noise Generation Strategy for Privacy Protection in Cloud Computing
Gaofeng Zhang, Xuyun Zhang, Yun Yang 0001, Chang Liu 0001, Jinjun Chen |
ICSOC | 4 |
| 2011 | A CSMA-based approach for detecting composite data aggregate events with collaborative sensors in WSNabstractNowadays, wireless sensor networks are widely used to monitor real time events and answer the ad hoc queries from a certain member node. However, computing and maintaining the information of aggregate queries in event monitoring wireless sensor networks incurs high spatial and temporal overhead for storage and transmission where potentially high volumes of unnecessary data may run through with changing time. Failure of processing that data can lead to unsuccessful event detection which can be very dangerous and costly in real world application. In order to reduce the overhead caused by unnecessary data for aggregate, suppression techniques such as data fusion, data sharing, data prediction, lossless data compression and base station side query rewriting are widely discussed in the WSN research community. In this paper, a technique which makes use of the spatial data relationship of local sensor nodes collaboratively is proposed to rein the detection of composite events with data aggregate. An empirical study is carried out to show the efficiency of the new technique. In addition, the new algorithm is compared to the previous event detection algorithms without spatial data suppression technique to demonstrate the significant performance gains. Chi Yang, Kaijun Ren, Zhimin Yang, Chang Liu 0001 |
CSCWD | 5 |
| 2011 | An Authenticated Key Exchange Scheme for Efficient Security-Aware Scheduling of Scientific Applications in Cloud ComputingabstractInstead of purchasing and maintaining their own computing infrastructure, scientists can now run data-intensive scientific applications in cloud computing environment by facilitating its vast storage and computation capabilities. During the scheduling of such scientific applications for execution, various computation data flows will happen between the controller and computing server instances. Amongst various quality-of-service (QoS) metrics, data security is one of the greatest concerns to scientists because their data may be intercepted or stolen by malicious parties during those data flows. An existing typical method for addressing this issue is to apply Internet Key Exchange (IKE) scheme to generate and exchange session keys, and then to apply these keys for performing symmetric-key encryption which will encrypt those data flows. However, the IKE scheme suffers from low efficiency due to its low performance of asymmetric-key crypto logical operations over a large amount of data and high-density operations which are exactly the characteristics of scientific applications. In this paper, we propose Cloud Computing Background Key Exchange (CCBKE), a novel authenticated key exchange scheme that aims at efficient security-aware scheduling of scientific applications. Our scheme is designed based on randomness-reuse strategy and Internet Key Exchange (IKE) scheme. Theoretical analyses and simulation results demonstrate that, compared with the IKE scheme, our CCBKE scheme can significantly improve the efficiency by dramatically reducing time consumption and computation load without sacrificing the level of security. Chang Liu 0001, Xuyun Zhang, Jinjun Chen, Chi Yang |
DASC | 1 |
| 2011 | An Upper-Bound Control Approach for Cost-Effective Privacy Protection of Intermediate Dataset Storage in CloudabstractAlong with more and more data intensive applications have been migrated into cloud environments, storing some valuable intermediate datasets has been accommodated in order to avoid the high cost of re-computing them. However, this poses a risk on data privacy protection because malicious parties may deduce the private information of the parent dataset or original dataset by analyzing some of those stored intermediate datasets. The traditional way for addressing this issue is to encrypt all of those stored datasets so that they can be hidden. We argue that this is neither efficient nor cost-effective because it is not necessary to encrypt ALL of those datasets and encryption of all large amounts of datasets can be very costly. In this paper, we propose a new approach to identify which stored datasets need to be encrypted and which not. Through intensive analysis of information theory, our approach designs an upper bound on privacy measure. As long as the overall mixed information amount of some stored datasets is no more than that upper bound, those datasets do not need to be encrypted while privacy can still be protected. A tree model is leveraged to analyze privacy disclosure of datasets, and privacy requirements are decomposed and satisfied layer by layer. With a heuristic implementation of this approach, evaluation results demonstrate that the cost for encrypting intermediate datasets decreases significantly compared with the traditional approach while the privacy protection of parent or original dataset is guaranteed. Xuyun Zhang, Chang Liu 0001, Jinjun Chen, Wan-Chun Dou |
DASC | 2 |
| 2010 | A collaborative trust model of firewall-through based on Cloud ComputingabstractIn this paper, the existing trust models and firewall technology are studied. Then we develop an approach based on Cloud Computing to estimate dynamic context and present the definition of risk signal. A collaborative trust model of firewall-through based on Cloud Environment is proposed. The model has three advantages: Firstly, there are different security policies for different domains. Secondly, the model considers the transaction context, the historical data of entity influences and the measurement of trust value dynamically. Thirdly, the trust model is compatible with the firewall and does not break the firewall's local control policies. To verify the reliability and effectiveness of the coordinated trust model, a simulation is carried out. The simulation result shows that this trust model is robust and out-performances traditional trust models based on domains because it successfully blocks malicious access under Cloud Computing environment. Zhimin Yang, Lixiang Qiao, Chang Liu 0001, Chi Yang, Guangming Wan |
CSCWD | 3 |