Zhiping Cai

dblp:64/2739 · DBLP profile ↗
← Back
19ranked-venue papers in the field
0as first author
18since 2021 · last 2026
0000-0001-5726-833XORCID · verified

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 7Information Retrieval & Web Search · 6Other / Interdisciplinary · 4Data Mining & Knowledge Discovery · 2
YearPublicationVenuePosition
2026 Exploring and Exploiting Security Vulnerabilities in Self-Hosted LLM Services
Zhihuang Liu, Ling Hu 0001, Yonghao Tang, Tongqing Zhou, Fang Liu 0002, Zhiping Cai
WWW6
2025 Spatiotemporal attention-based real-time video watermarking
Quan Yan, Yuanjing Luo, Zhangdong Wang, Junhua Xi, Geming Xia, Zhiping Cai
Data Min. Knowl. Discov.6
2025 ISSD: Indicator Selection for Time Series State Detection
abstract
Time series data from monitoring applications captures the behaviours of objects, which can often be split into distinguishable segments that reflect the underlying state changes. Despite the recent advances in time series state detection, the indicator selection for state detection is rarely studied, most of state detection work assumes the input indicators have been properly or manually selected. However, this assumption is disconnected from practice, on one hand, manual selection is not scalable, there can be up to thousands of indicators for the runtime monitoring of certain objects, e.g., supercomputer systems. On the other hand, performing state detection on a large amount of raw indicators is both inefficient and redundant. We argue that indicator selection should be made an upstream task for selecting a subset of indicators to facilitate state analysis. To this end, we propose ISSD ( I ndicator S election for S tate D etection), an indicator selection method for time series state detection. At its core, ISSD attempts to find an indicator subset that has as much high-quality states, which is measured by the channel set completeness and quality we invent based on segment-level sampling statistics. Such an indicator selection process is transformed into a multi-objective optimization problem and an approximation algorithm is designed to solve the NP-hard searching for specific end point in the Pareto front. Experiments on 5 datasets and 4 downstream methods show that ISSD has significant selection superiority compared with 6 baselines. We also elaborate on two observations of selection resilience and channel sensitivity of existing state detection methods and appeal to further research on them.
Chengyu Wang 0008, Tongqing Zhou, Lin Chen 0028, Shan Zhao 0002, Zhiping Cai
Proc. ACM Manag. Data5
2024 RARE: Robust Masked Graph Autoencoder
abstract
Masked graph autoencoder (MGAE) has emerged as a promising self-supervised graph pre-training (SGP) paradigm due to its simplicity and effectiveness. However, existing efforts perform the mask-then-reconstruct operation in the raw data space as is done in computer vision (CV) and natural language processing (NLP) areas, while neglecting the important non-Euclidean property of graph data. As a result, the highly unstable local structures largely increase the uncertainty in inferring masked data and decrease the reliability of the exploited self-supervision signals, leading to inferior representations for downstream evaluations. To address this issue, we propose a novel SGP method termed Robust mAsked gRaph autoEncoder (RARE) to improve the certainty in inferring masked data and the reliability of the self-supervision mechanism by further masking and reconstructing node samples in the high-order latent feature space. Through both theoretical and empirical analyses, we have discovered that performing a joint mask-then-reconstruct strategy in both latent feature and raw data spaces could yield improved stability and performance. To this end, we elaborately design a masked latent feature completion scheme, which predicts latent features of masked nodes under the guidance of high-order sample correlations that are hard to be observed from the raw data perspective. Specifically, we first adopt a latent feature predictor to predict the masked latent features from the visible ones. Next, we encode the raw data of masked samples with a momentum graph encoder and subsequently employ the resulting representations to improve the predicted results through latent feature matching. Extensive experiments on seventeen datasets have demonstrated the effectiveness and robustness of RARE against state-of-the-art (SOTA) competitors across three downstream tasks. Our source code is available athttps://github.com/WxTu/RARE.
Wenxuan Tu, Qing Liao 0001, Sihang Zhou 0001, Xin Peng 0010, Chuan Ma 0001, Zhe Liu 0001, Xinwang Liu 0002, Zhiping Cai, Kunlun He
IEEE Trans. Knowl. Data Eng.8
2023 Consensus One-step Multi-view Subspace Clustering (Extended abstract)
abstract
Multi-view clustering has attracted increasing attention in data mining communities. Despite superior clustering performance, we observe that existing multi-view subspace clustering methods directly fuse multi-view information in the similarity level by merging noisy affinity matrices; and isolate the processes of affinity learning, multiple information fusion and clustering. Both factors may cause insufficient utilization of multi-view information, leading to unsatisfying clustering performance. This paper proposes a novel consensus one-step multi-view subspace clustering (COMVSC) method to address these issues. Instead of directly fusing affinity matrices, COMVSC optimally integrates discriminative partition-level information, which is helpful in eliminating noise among data. Moreover, the affinity matrices, consensus representation and final clustering labels are learned simultaneously in a unified framework. Extensive experiment results on benchmark datasets demonstrate the superiority of our method over other state-of-the-art approaches.
Pei Zhang 0008, Xinwang Liu 0002, Jian Xiong 0002, Sihang Zhou 0001, En Zhu, Zhiping Cai
ICDE7
2023 IRWArt: Levering Watermarking Performance for Protecting High-quality Artwork Images
abstract
Increasing artwork plagiarism incidents underscores the urgent need for reliable copyright protection for high-quality artwork images. Although watermarking is helpful to this issue, existing methods are limited in imperceptibility and robustness. To provide high-level protection for valuable artwork images, we propose a novel invisible robust watermarking framework, dubbed as IRWArt. In our architecture, the embedding and recovery of the watermark are treated as a pair of image transformations’ inverse problems, and can be implemented through the forward and backward processes of an invertible neural networks (INN), respectively. For high visual quality, we embed the watermark in high-frequency domains with minimal impact on artwork and supervise image reconstruction using a human visual system(HVS)-consistent deep perceptual loss. For strong plagiarism-resistant, we construct a quality enhancement module for the embedded image against possible distortions caused by plagiarism actions. Moreover, the two-stagecontrastive training strategy enables the simultaneous realization of the above two goals. Experimental results on 4 datasets demonstrate the superiority of our IRWArt over other state-of-the-art watermarking methods. Code: https://github.com/1024yy/IRWArt.
Yuanjing Luo, Tongqing Zhou, Fang Liu 0002, Zhiping Cai
WWW4
2023 Seeing is believing: Towards interactive visual exploration of data privacy in federated learning
Yeting Guo, Fang Liu 0002, Tongqing Zhou, Zhiping Cai, Nong Xiao 0001
Inf. Process. Manag.4
2023 Leveraging heuristic client selection for enhanced secure federated submodel learning
Panyu Liu, Tongqing Zhou, Zhiping Cai, Fang Liu 0002, Yeting Guo
Inf. Process. Manag.3
2023 Turning backdoors for efficient privacy protection against image retrieval violations
Qiang Liu 0004, Tongqing Zhou, Zhiping Cai, Yuan Yuan 0034, Ming Xu 0002, Jiaohua Qin, Wentao Ma 0003
Inf. Process. Manag.3
2023 Adaptive multi-feature fusion via cross-entropy normalization for effective image retrieval
Wentao Ma 0003, Tongqing Zhou, Jiaohua Qin, Xuyu Xiang, Yun Tan, Zhiping Cai
Inf. Process. Manag.6
2023 Time2State: An Unsupervised Framework for Inferring the Latent States in Time Series Data
abstract
Time series data from monitoring applications reflect the physical or logical states of the objects, which may produce time series of distinguishable characteristics in different states. Thus, time series data can usually be split into different segments, each reflecting a state of the objects. These states carry rich high-level semantic information, e.g., run, walk, or jump, which helps people better understand the behaviour of the monitored objects. Nevertheless, these states are latent and hard to discover, because the characteristic of time series is complicated and the computational cost is high. This paper develops an efficient and effective unsupervised approach for inferring the latent states of massive multivariate time data. To reduce the computational cost, we present Time2State, a scalable framework that utilizes a sliding window and an encoder to greatly reduce the length of raw time series. To train the encoder, we propose a novel unsupervised loss function, LSE-Loss. Extensive experiments show that compared to the state-of-the-art time series representation learning methods of the same kind, LSE-Loss brings a performance improvement of up to 15% in accuracy.
Chengyu Wang 0008, Kui Wu 0001, Tongqing Zhou, Zhiping Cai
Proc. ACM Manag. Data4
2023 In Pursuit of Beauty: Aesthetic-Aware and Context-Adaptive Photo Selection in Crowdsensing
abstract
The pervasive view of the mobile crowd bridges various real-world scenes and people's perceptions with the gathering of distributed crowdsensing photos. To elaborate informative visuals for viewers, existing techniques introduce photo selection as an essential step in crowdsensing. Yet, the aesthetic preference of viewers, at the very heart of their experiences under various crowdsensing contexts (e.g., travel planning), is seldom considered and hardly guaranteed. We propose CrowdPicker, a novel photo selection framework with adaptive aesthetic awareness for crowdsensing. With the observations on aesthetic uncertainty and bias in different crowdsensing contexts, we exploit a joint effort of mobile crowdsourcing and domain adaptation to actively learn contextual knowledge for dynamically tailoring the aesthetic predictor. Concretely, an aesthetic utility measure is invented based on the probabilistic balance formalization to quantify the benefit of photos in improving the adaptation performance. We prove the NP-hardness of sampling the best-utility photos for crowdsourcing annotation and present a (1-1/e) approximate solution. Furthermore, a two-stage distillation-based adaptation architecture is designed based on fusing contextual and common aesthetic preferences. Extensive experiments on three datasets and four raw models demonstrate the performance superiority of CrowdPicker over four photo selection baselines and four typical sampling strategies. Cross-dataset evaluation illustrates the impacts of aesthetic bias on selection.
Tongqing Zhou, Zhiping Cai, Fang Liu 0002, Jinshu Su
IEEE Trans. Knowl. Data Eng.2
2022 PARA: Performability-aware resource allocation on the edges for cloud-native services
abstract
This paper explores resource allocation strategy in the Baidu Over The Edge system to enable mobile edge computing (MEC) datacenters to effectively support cloud-native services downstream to the network edge. There are many challenges to this issue. First, MEC datacenters are resource-constrained to fully meet resource demands. Second, previous works regard the resource requirements of each service as an indivisible unit, resulting in idle MEC resources, even if the resources can meet the demands of some microservices decoupled by the service. Third, they are confined to optimize the allocation for a single slot, failing to adapt to the dynamic demands. To improve resource utilization, we propose performability-aware resource allocation (PARA), a PARA on the edges for cloud-native services. It takes microservices as the unit of resource allocation and allows services to perform with degraded services when only part of microservices' demands are met. It also considers dependency among microservices, dynamic resource requirements, and resource supply characteristics of MEC and cloud. Performability is a unified performance-reliability measure for evaluating such degradable systems. To maximize the long-term overall performability, we model the resource optimization problem and then develop an online greedy heuristic algorithm. The algorithm predicts services' resource demands and then adapts the online allocation. The experimental results show that PARA reduces the reallocation overhead by 47.7%–53.6%, and improves the long-term overall performability by 23.14%–43.25% of existing state-of-the-art works.
Yeting Guo, Fang Liu 0002, Nong Xiao 0001, Zhaogeng Li, Zhiping Cai, Guoming Tang, Ning Liu 0015
Int. J. Intell. Syst.5
2022 EviChain: A scalable blockchain for accountable intelligent surveillance systems
abstract
Smart cameras, as typical IoT devices, are widely adopted to provide surveillance on individuals, homes, and the environment. The unavoidably captured sensitive visuals via these cameras may raise significant security concerns, while the prevalent software defects and authentication misconfiguration issues aggravate the vulnerability of such devices. However, traditional cryptography techniques are inadequate to provide full protection of these devices due to the large computation overhead. In this context, realizing accountability for these surveillance systems shall be the last line of defense in the presence of fast-evolving and high-influential threats. We propose EviChain, a scalable blockchain-based solution to trace the operations on intelligent surveillance cameras and reserve the evidence for any misuse in tamper-proofing manipulation records. Building a blockchain over the distributed cameras is challenging due to the limited capacity of on-board memory. To tackle this challenge, we design a cooperative mechanism that enables cameras to adaptively join in groups and share storage for recording blocks. In addition, we present a computation efficiency and delay-aware block generation strategy to reduce the cost of the consensus process. We perform extensive simulations to validate the superior performance of EviChain over other baselines, for example, Practical Byzantine Fault Tolerance (PBFT).
Jiaping Yu, Haiwen Chen, Kui Wu 0001, Tongqing Zhou, Zhiping Cai, Fang Liu 0002
Int. J. Intell. Syst.5
2022 Multi-View Spectral Clustering With High-Order Optimal Neighborhood Laplacian Matrix
abstract
Multi-view spectral clustering can effectively reveal the intrinsic cluster structure among data by performing clustering on the learned optimal embedding across views. Though demonstrating promising performance in various applications, most of existing methods usually linearly combine a group of pre-specified first-order Laplacian matrices to construct the optimal Laplacian matrix, which may result in limited representation capability and insufficient information exploitation. Also, storing and implementing complex operations on the{$n\times n}$Laplacian matrices incurs intensive storage and computation complexity. To address these issues, this paper first proposes a multi-view spectral clustering algorithm that learns a high-order optimal neighborhood Laplacian matrix, and then extends it to the late fusion version for accurate and efficient multi-view clustering. Specifically, our proposed algorithm generates the optimal Laplacian matrix by searching the neighborhood of the linear combination of both the first-order and high-order base Laplacian matrices simultaneously. By this way, the representative capacity of the learned optimal Laplacian matrix is enhanced, which is helpful to better utilize the hidden high-order connection information among data, leading to improved clustering performance. We design an efficient algorithm with proved convergence to solve the resultant optimization problem. Extensive experimental results on nine datasets demonstrate the superiority of the proposed algorithm
Weixuan Liang, Sihang Zhou 0001, Jian Xiong 0002, Xinwang Liu 0002, Siwei Wang 0001, En Zhu, Zhiping Cai, Xin Xu 0001
IEEE Trans. Knowl. Data Eng.7
2022 Consensus One-Step Multi-View Subspace Clustering
abstract
Multi-view clustering has attracted increasing attention in multimedia, machine learning and data mining communities. As one kind of the essential multi-view clustering algorithm, multi-view subspace clustering (MVSC) becomes more and more popular due to its strong ability to reveal the intrinsic low dimensional clustering structure hidden across views. Despite superior clustering performance in various applications, we observe that existing MVSC methodsdirectly fuse multi-view information in the similarity level by merging noisy affinity matrices; andisolate the processes of affinity learning, multi-view information fusion and clustering. Both factors may cause insufficient utilization of multi-view information, leading to unsatisfying clustering performance. This paper proposes a novel consensus one-step multi-view subspace clustering (COMVSC) method to address these issues. Instead of directly fusing multiple affinity matrices, COMVSC optimally integrates discriminative partition-level information, which is helpful to eliminate noise among data. Moreover, the affinity matrices, consensus representation and final clustering labels matrix are learned simultaneously in a unified framework. By doing so, the three steps can negotiate with each other to best serve the clustering task, leading to improved performance. Accordingly, we propose an iterative algorithm to solve the resulting optimization problem. Extensive experiment results on benchmark datasets demonstrate the superiority of our method against other state-of-the-art approaches.
Pei Zhang 0008, Xinwang Liu 0002, Jian Xiong 0002, Sihang Zhou 0001, En Zhu, Zhiping Cai
IEEE Trans. Knowl. Data Eng.7
2021 SmartStore: A blockchain and clustering based intelligent edge storage system with fairness and resilience
abstract
With the development of edge computing, edge storage solutions are attracting widespread attention. When facing the requirements of lower latency and faster access speed from end devices, edge storage solutions are considered to be an alternative to the cloud. However, edges are usually owned by small organizations which have limited operations and maintenance capabilities. This makes these edge devices can be easily disabled by external attacks or internal hardware failures. Besides, the heterogeneity of the edge devices will also make it difficult to price the edge resources uniformly. To tackle these problems, we propose SmartStore: an auction mechanism based on blockchain to allocate edge resources. Considering centralized solutions have access bottlenecks and trust issues, we built SmartStore on the smart contract. With Bayesian game theory, SmartStore can analyze how data owners (DO) and edges price the resources can maximize their benefits. From an economic perspective, both DO and edges can make full use of edge heterogeneous resources with SmartStore. Besides, a two-stage submission strategy is proposed to complete the sealed auction. Furthermore, considering the reliability of edge storage, we propose a cluster-based block distribution algorithm for SmartStore's intelligent edge recommendation process. SmartStore ensures the reliability of edge storage while maximizing the benefits and resource utilization of both parties. Finally, we conduct specific experiments on the proposed auction smart contract through “Ethereum” and the experimental results of implementation show the effectiveness and efficiency of our SmartStore.
Haiwen Chen, Jiaping Yu, Huan Zhou 0006, Tongqing Zhou, Fang Liu 0002, Zhiping Cai
Int. J. Intell. Syst.6
2021 Trusted audit with untrusted auditors: A decentralized data integrity Crowdauditing approach based on blockchain
abstract
Edge computing emerges as an alternative to cloud computing in the scenarios where the end devices require lower latency and faster access speeds. Edge nodes are deployed at the proximity of the end devices to reduce response time. On the other hand, the edge nodes are usually owned by small organizations that have limited operations and maintenance capabilities. Data on the edge may be easily damaged, due to external attacks or internal hardware failures. Therefore, it is essential to verify data integrity in edge computing. However, edge environment requires a different trust model compared with other computing and storage paradigm. Besides, compared with cloud storage, edge storage is decentralized and storage service participants may pose greater internal and external threats. This paper proposes a blockchain-based intelligent crowdsourcing audit approach (Crowdauditing) to achieve on-chain and off-chain credibility of audit results. The model relies on an untrusted auditor committee from the crowd to audit data integrity and uses smart contracts as the core of the intelligent system to ensure the reliability of result submission, the accuracy of the result judgment, and reasonable punishments and rewards. Specifically, an unbiased selection algorithm is proposed to achieve fairness during the auditor committee construction. An innovative two-stage submission strategy is proposed to ensure that the auditor committee can reach a consensus on the off-chain audit results. An incentive mechanism is carefully designed to force auditors providing audit services honestly to maximize their own rewards. Moreover, we modeled that as a game of n players, which proves the reliability of the result. Finally, we implement a prototype of Crowdauditing based on smart contracts. The extensive experimental results demonstrate the effectiveness of Crowdauditing.
Haiwen Chen, Huan Zhou 0006, Jiaping Yu, Kui Wu 0001, Fang Liu 0002, Tongqing Zhou, Zhiping Cai
Int. J. Intell. Syst.7
2019 Edge-enabled Disaster Rescue: A Case Study of Searching for Missing People
abstract
In the aftermath of earthquakes, floods, and other disasters, photos are increasingly playing more significant roles, such as finding missing people and assessing disasters, in rescue and recovery efforts. These disaster photos are taken in real time by the crowd, unmanned aerial vehicles, and wireless sensors. However, communications equipment is often damaged in disasters, and the very limited communication bandwidth restricts the upload of photos to the cloud center, seriously impeding disaster rescue endeavors. Based on edge computing, we propose Echo, a highly time-efficient disaster rescue framework. By utilizing the computing, storage, and communication abilities of edge servers, disaster photos are preprocessed and analyzed in real time, and more specific visuals are immensely helpful for conducting emergency response and rescue. This article takes the search for missing people as a case study to show that Echo can be more advantageous in terms of disaster rescue. To greatly conserve valuable communication bandwidth, only significantly associated images are extracted and uploaded to the cloud center for subsequent facial recognition. Furthermore, an adaptive photo detector is designed to utilize the precious and unstable communication bandwidth effectively, as well as ensure the photo detection precision and recall rate. The effectiveness and efficiency of the proposed method are demonstrated by simulation experiments.
Fang Liu 0002, Yeting Guo, Zhiping Cai, Nong Xiao 0001, Ziming Zhao 0002
ACM Trans. Intell. Syst. Technol.3