Xinkui Zhao

dblp:135/5118 · DBLP profile ↗
← Back
54ranked-venue papers
9as first author
41since 2021 · last 2026
0000-0002-1115-5652ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 14 · 4 first-author · 7 since 2021Software engineering, systems software and programming languages · 14 · 2 first-author · 12 since 2021Artificial intelligence and machine learning · 13 · 13 since 2021Databases, data management, data science and information retrieval · 9 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 3 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DyC-STG: Dynamic Causal Spatio-Temporal Graph Network for Real-time Data Credibility Analysis in IoT
abstract
The wide spreading of Internet of Things (IoT) sensors generates vast spatio-temporal data streams, but ensuring data credibility is a critical yet unsolved challenge for applications like smart homes. While spatio-temporal graph (STG) models are a leading paradigm for such data, they often fall short in dynamic, human-centric environments due to two fundamental limitations: (1) their reliance on static graph topologies, which fail to capture physical, event-driven dynamics, and (2) their tendency to confuse spurious correlations with true causality, undermining robustness in human-centric environments. To address these gaps, we propose the Dynamic Causal Spatio-Temporal Graph Network (DyC-STG), a novel framework designed for real-time data credibility analysis in IoT. Our framework features two synergistic contributions: an event-driven dynamic graph module that adapts the graph topology in real-time to reflect physical state changes, and a causal reasoning module to distill causally-aware representations by strictly enforcing temporal precedence. To facilitate the research in this domain we release two new real-world datasets. Comprehensive experiments show that DyC-STG establishes a new state-of-the-art, outperforming the strongest baselines by 1.4 percentage points and achieving an F1-Score of up to 0.930.
Guanjie Cheng, Peihan Wu, Feiyi Chen, Xinkui Zhao, Mengying Zhu, Shuiguang Deng
AAAI5
2026 LSHFed: Robust and Communication-Efficient Federated Learning with Locally-Sensitive Hashing Gradient Mapping
abstract
Federated learning (FL) enables collaborative model training across distributed nodes without exposing raw data, but its decentralized nature makes it vulnerable in trust-deficient environments. Inference attacks may recover sensitive information from gradient updates, while poisoning attacks can degrade model performance or induce malicious behaviors. Existing defenses often suffer from high communication and computation costs, or limited detection precision. To address these issues, we propose LSHFed, a robust and communication-efficient FL framework that simultaneously enhances aggregation robustness and privacy preservation. At its core, LSHFed incorporates LSHGM, a novel gradient verification mechanism that projects high-dimensional gradients into compact binary representations via multi-hyperplane locality-sensitive hashing. This enables accurate detection and filtering of malicious gradients using only their irreversible hash forms, thus mitigating privacy leakage risks and substantially reducing transmission overhead. Extensive experiments demonstrate that LSHFed maintains high model performance even when up to 50% of participants are collusive adversaries, while achieving up to a 1000× reduction in gradient verification communication compared to full-gradient methods.
Guanjie Cheng, Mengzhen Yang, Xinkui Zhao, Shuyi Yu, Tianyu Du, Mengying Zhu, Shuiguang Deng
AAAI3
2026 Swarm: Efficient Logical Memory Disaggregation With Shared CXL Memory
abstract
Memory disaggregation has become a research trend in data centers. Existing studies fall into two paths: Network-based logical memory disaggregation (LMD) and Compute Express Link (CXL)-based physical memory disaggregation (PMD). However, LMD suffers from network overhead, while PMD incurs expensive hardware costs and lacks flexibility. This paper advocates for building LMD systems on shared CXL memory, taking advantage of its low latency and cache-coherent memory access. However, shared CXL memory has severe scalability issues due to its strict coherence model.This paper introduces Swarm, an efficient LMP system built on shared CXL memory. Swarm divides shared CXL memory into small cacheable memory and large non-cacheable memory. Hardware only needs to maintain coherence for cacheable memory, while software handles coherence for non-cacheable memory, thereby enabling all CXL memory to be shared. Then, Swarm implements cross-node RPC and dynamic global memory allocation on shared CXL memory to improve performance and memory utilization. Swarm also proposes distributed computing offloading to fully leverage compute power on memory nodes for acceleration. Our evaluation shows that Swarm not only improves the throughput (e.g., by 3.6× and 1.8×) compared with representative network-based LMD system, AIFM and CXL-based PMD system when computing offloading is enabled but also achieves an advantage in TCO.
Xinkui Zhao, Guanjie Cheng, Yueshen Xu, Shuiguang Deng, Jianwei Yin
IEEE Trans. Computers2
2026 Secure and Efficient Personalized Multi-Receiver Data Sharing With Cross-Domain Authentication for Internet of Vehicles
Taolong Su, Guanjie Cheng, Junqin Huang, Xinkui Zhao, Shuiguang Deng
IEEE Trans. Dependable Secur. Comput.4
2026 Service Pattern Fusion: Toward Self-Evolving of Service Ecosystems
abstract
A service ecosystem refers to a multilateral network composed of heterogeneous service entities, where the exchange of data, resources, and value through interactions among specific participants forms a service pattern. As service ecosystems like virtual hospital alliance (VHA) evolve towards large-scale, multi-domain integration to meet complex user needs, service pattern fusion has emerged as a fundamental approach to leverage data, resources, and value aggregation. By converging elements from multiple patterns, service pattern fusion enables the fulfillment of composite business objectives with reduced redundancy and lower costs. Existing works primarily address fusion requirements by reorganizing existing services through approaches such as service composition and business process management where only service functions and workflows are considered. However, they lack formalization of pattern fusion constraints and fail to support comprehensive integration of participants, data, resources, and value, let alone identifying optimal fusion solutions that account for participant collaboration and service integration. In this study, we formally define the Service Pattern Fusion Problem (SPFP) as an optimization task aimed at identifying the most efficient and cost-effective pattern by integrating, combining, and pruning elements from multiple patterns while preserving their objectives and meeting business constraints. We adapt traditional heuristic methods to SPFP and propose the Fusion-Oriented Confidence-Aware genetic algorithm (FoCa). FoCa dynamically adjusts the search space and transition probabilities in each iteration, achieving optimal fusion results with a 41.09% reduction in pattern loss and the fastest convergence. In addition, we designed a set of pattern features and conducted random fusion experiments on the public service pattern dataset S-SPD, to explore the correlation between those features and the optimization magnitude across various metrics. The analysis helps identify which types of service patterns benefit most from fusion, providing valuable insights for researchers and practitioners in both academic and engineering contexts.
Meng Xi 0002, Yechen Jin, Jinshan Zhang 0001, Ying Li 0001, Xinkui Zhao, Jianwei Yin
IEEE Trans. Serv. Comput.7
2026 Online Microservice Deployment in Edge Networks via Multiobjective Deep Reinforcement Learning
abstract
In recent years, edge networks have been deployed broadly at large scale, hosting a wide variety of services. Among these, microservices have emerged as one of the predominant service paradigms. Typically, microservices run on edge servers with varying configurations, while new microservice instances are usually online generated and join in edge networks due to the dynamic attributes of requests and networks. In those cases, an effective microservice online deployment solution are expected to be vital to system performance. So it becomes a critical issue to design online deployment solutions for microservices in edge. Existing research has always focused on offline deployment of microservices. However, edge networks are characterized by dynamics, real time, and concurrency, and when the environments or requests change, traditional offline deployment solutions usually cannot handle the deployment task in such cases. To address these issues, we carry out a comprehensive investigation on those potential influencing factors in edge, fully covering deployment cost, load balance, packet loss, and network delay. We further develop an innovative holistic online deployment solution that encompasses a system model, constraint analysis, multiobjective optimization, and a deep reinforcement learning algorithm. We conducted extensive experiments and evaluated our solution over a set of metrics using a real-world microservice prototype system. The results show that our online deployment solution produces superior performance, for example, reducing deployment cost by an average of 70.14% compared to all baselines. We also evaluated our solution under varying volumes of requests and gave analysis for performance stability and parameter sensitivity. We have released the code on GitHub.
Yueshen Xu, Fanhao Zeng, Qingshan Li, Xinkui Zhao, Wei Shao 0006, Shuiguang Deng, Rui Li 0047
IEEE Trans. Serv. Comput.4
2025 DAPoinTr: Domain Adaptive Point Transformer for Point Cloud Completion
abstract
Point Transformers (PoinTr) have shown great potential in point cloud completion recently. Nevertheless, effective domain adaptation that improves transferability toward target domains remains unexplored. In this paper, we delve into this topic and empirically discover that direct feature alignment on point Transformer’s CNN backbone only brings limited improvements since it cannot guarantee sequence-wise domain-invariant features in the Transformer. To this end, we propose a pioneering Domain Adaptive Point Transformer (DAPoinTr) framework for point cloud completion. DAPoinTr consists of three novel components: Domain Query-based Feature Alignment (DQFA), Point Token-wise Feature alignment (PTFA), and Voted Prediction Consistency (VPC). In particular, DQFA is presented to narrow the global domain gaps from the sequence via the presented domain proxy and domain query at the Transformer encoder and decoder, respectively. PTFA is proposed to close the local domain shifts by aligning the tokens, i.e., point proxy and dynamic query, at the Transformer encoder and decoder, respectively. VPC is designed to consider different Transformer decoders as multiple of experts (MoE) for ensembled prediction voting and pseudo-label generation. Extensive experiments with visualization on several challenging domain adaptation benchmarks demonstrate the effectiveness and superiority of our DAPoinTr compared with other state-of-the-art methods.
Qianyu Zhou 0001, Jingyu Gong, Ye Zhu 0002, Richard Dazeley, Xinkui Zhao, Xuequan Lu
AAAI6
2025 STraj: Self-training for Bridging the Cross-Geography Gap in Trajectory Prediction
abstract
Accurate trajectory prediction has prominent significance in autonomous driving scenarios. Most existing methods predict the trajectory of an agent by learning its interaction with other agents and the map within the scenario. However, the heterogeneous distribution of these elements across different geographical scenarios is always ignored. Thus, trajectory predictors might struggle to generalize well when deployed in different geographical scenarios. To bridge the cross-geography gap, in this paper, we propose a plug-and-play self-training pipeline, termed STraj, for cross-geography trajectory prediction. STraj comprises three progressive steps: pseudo label (i.e., time-series trajectory) generation, update, and utilization. First, to generate pseudo labels that generalize to the cross-geography scenarios, STraj pre-trains the predictor through the complementary agent and map augmentations. Second, to facilitate the stable training of the predictor, we design a specific pseudo label update strategy. This strategy selects high-consistency pseudo trajectories from the current and historical epochs to supervise the target domain samples. Third, with generated pseudo trajectories, we introduce trajectory-induced contrastive learning to mitigate the representation bias of cross-geography agents. Extensive experiment results on various cross-geography trajectory prediction benchmarks demonstrate the effectiveness of STraj.
Zhanwei Zhang, Minghao Chen 0001, Zhihong Gu, Xinkui Zhao, Zheng Yang 0008, Binbin Lin 0001, Deng Cai 0001, Wenxiao Wang 0001
AAAI4
2025 CADRef: Robust Out-of-Distribution Detection via Class-Aware Decoupled Relative Feature Leveraging
abstract
Deep neural networks (DNNs) have been widely criticized for their overconfidence when dealing with out-of-distribution (OOD) samples, highlighting the critical need for effective OOD detection to ensure the safe deployment of DNNs in real-world settings. Existing post-hoc OOD detection methods primarily enhance the discriminative power of logit-based approaches by reshaping sample features, yet they often neglect critical information inherent in the features themselves. In this paper, we propose the Class-Aware Relative Feature-based method (CARef), which utilizes the error between a sample’s feature and its class-aware average feature as a discriminative criterion. To further refine this approach, we introduce the Class-Aware Decoupled Relative Feature-based method (CADRef), which decouples sample features based on the alignment of signs between the relative feature and corresponding model weights, enhancing the discriminative capabilities of CARef. Extensive experimental results across multiple datasets and models demonstrate that both proposed methods exhibit effectiveness and robustness in OOD detection compared to state-of-the-art methods. Specifically, our two methods outperform the best baseline by 2.82% and 3.27% in AUROC, with improvements of 4.03% and 6.32% in FPR95, respectively.
Zhiwei Ling, Yachen Chang, Hailiang Zhao, Xinkui Zhao, Kingsum Chow, Shuiguang Deng
CVPR4
2025 DyREM: Dynamically Mitigating Quantum Readout Error with Embedded Accelerator
abstract
Quantum readout error is the most significant source of error, substantially reducing the measurement fidelity. Tensor-product-based readout error mitigation has been proposed to address this issue by approximating the mitigation matrix. However, this method inevitably encounters the dynamic generation of the mitigation matrix, leading to long latency. In this paper, we propose DyREM, a software-hardware codesign approach that mitigates readout errors with an embedded accelerator. The main innovation lies in leveraging the inherent sparsity in the nonzero probability distribution of quantum states and calculating the tensor product on an embedded accelerator. Specifically, using the output sparsity, our dataflow dynamically downsamples the original mitigation matrix, which dramatically reduces the memory requirement. Then, we design DyREM architecture that can flexibly gate the redundant computation of nonzero quantum states. Experiments demonstrate that DyREM achieves an average speedup of $9.6 \times \sim 2000 \times$ and fidelity improvements of $1.03 \times \sim 1.15 \times$ compared to state-of-the-art readout error mitigation methods.
Kaiwen Zhou 0003, Liqiang Lu, Debin Xiang, Chenning Tao, Xinkui Zhao, Size Zheng 0001, Jianwei Yin
DAC6
2025 AgentPro: Enhancing LLM Agents with Automated Process Supervision
abstract
Large language model (LLM) agents have demonstrated significant potential for addressing complex tasks through mechanisms such as chain-of-thought reasoning and tool invocation.However, current frameworks lack explicit supervision during the reasoning process, which may lead to error propagation across reasoning chains and hinder the optimization of intermediate decision-making stages.This paper introduces a novel framework, AgentPro, which enhances LLM agent performance by automated process supervision.AgentPro employs Monte Carlo Tree Search to automatically generate step-level annotations, and develops a process reward model based on these annotations to facilitate fine-grained quality assessment of reasoning.By employing a rejection sampling strategy, the LLM agent dynamically adjusts generation probability distributions to prevent the continuation of erroneous paths, thereby improving reasoning capabilities.Extensive experiments on four datasets indicate that our method significantly outperforms existing agent-based LLM methods (e.g., achieving a 6.32% increase in accuracy on the HotpotQA dataset), underscoring its proficiency in managing intricate reasoning chains.
Yuchen Deng, Shichen Fan, Naibo Wang, Xinkui Zhao, See-Kiong Ng
EMNLP4
2025 MISS: An Incomplete Tabular Data Representation System with Missing Mechanism Learning
abstract
The missing data problem widely exists in real-life scenarios. The incomplete data analysis through imputation can amplify the errors or bias, hindering the effective analysis. Ex-isting tabular data representation methods overlook the missing state of data values, and thus cannot effectively deal with the incomplete data. In this paper, we propose a novel incomplete tabular data representation system, named MISS. It is capable of enabling all Transformer-based tabular representation methods to effectively handle incomplete data. MISS consists of two modules, i.e., missing mechanism learning (MML) and incomplete data representation (IDR). MML leverages a new missingness propensity score calculation strategy to learn the observed data distribution and missing mechanisms within incomplete data. IDR introduces a novel probability-driven Transformer block, in conjunction with an unbiased representation loss function, for effective representation. We prove that, MISS can eliminate the bias resulting from missingness. Extensive experiments on four public real-world datasets demonstrate that, MISS yields a more than 57 % accuracy gain with competitive efficiency, compared with the state-of-the-art approaches.
Shuwei Liang, Lei Qiang, Xiaoye Miao, Xinkui Zhao, Junlan Cai, Yunjun Gao, Jianwei Yin
ICDE5
2025 A Zero-Training Error Correction System with Large Language Models
abstract
Correcting missing or erroneous data values is an essential task in data cleaning. Traditional pre-configuration error correction (EC) methods rely heavily on predefined rules or constraints, demanding significant domain knowledge and manual effort. While configuration-free EC approaches have been explored, they still demand extensive feature engineering or labeled data for intensive model training. In this paper, we propose a zero-training and interpretable EC system, named ZeroEC, that leverages large language models (LLMs) to generate chain-of-thoughts (CoTs) and correction rules for EC, without the need for model training. ZeroEC consists of two modules, contextual-relevant tuple search (CTS) and training-free explainable correction (TEC). CTS constructs a contextual-relevant tuple retriever using a weighted cosine similarity function to efficiently identify the most relevant tuples for each dirty tuple, reducing redundancy in the LLM prompts and lowering computational costs. TEC employs a clustering-based representative tuple sampling strategy to alleviate “hallucination” risk by exposing LLMs to diverse types of data errors. It further prompts for generating correction CoTs for user-corrected representative tuples, as well as prompts for creating correction rules and explainable ECs, which automatically provide explanations for EC, all without the need for model training. Extensive experiments conducted on various real-world datasets demonstrate that ZeroEC achieves a 66.82% increase in accuracy and a 6.87x speedup compared to state-of-the-art methods. The codes and datasets of this paper are available at https://github.com/YangChen32768/ZeroEC.
Mengying Zhu, Xiaoye Miao, Meng Xi 0002, Xinkui Zhao, Jianwei Yin
ICDE7
2025 Bridging Service Diversity and CPU Heterogeneity Through Program Similarity-Driven Scheduling
Jiayin Luo, Xinkui Zhao, Wei Zhou 0028, Jianwei Yin
ICSOC (2)3
2025 RESTful API Service Discovery via Comprehensive Feature Mining, Deep Neural Networks, and Contrastive Learning
Yueshen Xu, Gairui Bai, Weihao Xiao, Xinkui Zhao, Yuyu Yin, Rui Li 0047, Fanhao Zeng
ICSOC (1)4
2025 Recognition Service for Named Entities via Multilayer Feature Learning for Large Web Knowledge Bases
abstract
In the field of Web knowledge base mining and Web services, the recognition service for named entity faces many challenges such as context complexity, semantic subtlety, and fuzzy entity boundaries, all of which require highly accurate and robust recognition service. The current services usually fail to reach those conditions. To address this issue, this paper proposes a recognition service for named entities for large Web knowledge bases, and the core contrition is the developed dual multilayer feature learning (D-MLFL) service, which combines projected gradient descent (PGD), adversarial learning, and a fused attention mechanism. Our service successfully addresses the challenges of complex context and subtle semantics faced by named entity recognition tasks. Our service integrates deep language models, recurrent neural networks, and conditional random fields, and clearly outperforms existing approaches in many sub-tasks, including in feature extraction, sequence modeling, and label decoding, especially in dealing with complex and diverse entity types and contextual relationships. We performed sufficient experiments, and the results show that our service significantly enhances the robustness against noise and abnormal Web data. The ability to extract entity features is improved, resulting in higher accuracy in identifying entities with fuzzy boundaries and complex semantics.
Chan Li, Rui Li 0047, Yinru Ma, Xinkui Zhao, Lei Hei, Yuyu Yin, Yueshen Xu
ICWS4
2025 Scalable Multi-Stage Influence Function for Large Language Models via Eigenvalue-Corrected Kronecker-Factored Parameterization
abstract
Pre-trained large language models (LLMs) are commonly fine-tuned to adapt to downstream tasks. Since the majority of knowledge is acquired during pre-training, attributing the predictions of fine-tuned LLMs to their pre-training data may provide valuable insights. Influence functions have been proposed as a means to explain model predictions based on training data. However, existing approaches often fail to compute "multi-stage" influence and lack scalability to billion-scale LLMs. In this paper, we propose multi-stage influence functions to attribute the downstream predictions of fine-tuned LLMs to pre-training data under the full-parameter fine-tuning paradigm. To enhance the efficiency and practicality of our multi-stage influence function, we leverage Eigenvalue-corrected Kronecker-Factored (EK-FAC) parameterization for efficient approximation. Empirical results validate the superior scalability of EK-FAC approximation and the effectiveness of our multi-stage influence function. Additionally, case studies on a real-world LLM, dolly-v2-3b, demonstrate its interpretive power, with exemplars illustrating insights provided by multi-stage influence estimates.
Yuntai Bao, Xuhong Zhang 0002, Tianyu Du, Xinkui Zhao, Jiang Zong, Hao Peng 0002, Jianwei Yin
IJCAI4
2025 Horae: A Domain-Agnostic Language for Automated Service Regulation
abstract
Artificial intelligence is rapidly encroaching on the field of service regulation. However, existing AI-based regulation techniques are often tailored to specific application domains and thus are difficult to generalize in an automated manner. This paper presents Horae, a unified specification language for modeling (multimodal) regulation rules across a diverse set of domains. We showcase how Horae facilitates an intelligent service regulation pipeline by further exploiting a fine-tuned large language model named RuleGPT that automates the Horae modeling process, thereby yielding an end-to-end framework for fully automated intelligent service regulation. The feasibility and effectiveness of our framework are demonstrated over a benchmark of various real-world regulation domains. In particular, we show that our open-sourced, fine-tuned RuleGPT with 7B parameters suffices to outperform GPT-3.5 and perform on par with GPT-4o.
Yutao Sun, Mingshuai Chen, Kangjia Zhao, Jintao Chen 0001, Zhongyi Wang 0004, Liqiang Lu, Xinkui Zhao, Shuiguang Deng, Jianwei Yin
IJCAI9
2025 Dual Structure-guided Contrastive Network for Incomplete Multi-view Partial Multi-label Classification
abstract
Incomplete multi-view partial multi-label classification (IMvPMLC), which tackles the combined challenges of incompleteness in both multi-view and multi-label problems, has drawn considerable attention. Existing IMvPMLC methods have made progress but still face several challenges: (i) They mainly focus on the consistency of representations across multiple views but overlook the relationships among instances, leading to suboptimal representations. (ii) They primarily utilize only the available labels for supervised learning, ignoring the missing label distribution and limiting their ability to capture label correlations. In this paper, we propose a novel model named Dual Structure-guided Contrastive Network (DSCN) for IMvPMLC. Specifically, we introduce a similarity-guided instance-level contrastive learning mechanism to achieve multi-view consistent and discriminative representations across instances by leveraging instance structures, while a multi-view attention-based fusion strategy dynamically facilitates the fusion of multi-view representations to derive a robust consensus representation. Then, we design a multi-view shared classifier integrated with a correlation-guided label-level contrastive learning mechanism to enhance predictions by leveraging complementary information across multiple views and capturing label structures, effectively exploiting missing label distribution. Extensive experiments on five benchmark datasets demonstrate that, DSCN yields a more than 13% accuracy, compared with the state-of-the-art approaches. The code and datasets are available at https://anonymous.4open.science/r/DSCN-D471.
Kaixin Xu, Shijun Wu, Xiaoye Miao, Guoqing Chao, Mengying Zhu, Meng Xi 0002, Xinkui Zhao
KDD (2)8
2025 Walking the Schrödinger Bridge: A Direct Trajectory for Text-to-3D Generation
abstract
Recent advancements in optimization-based text-to-3D generation heavily rely on distilling knowledge from pre-trained text-to-image diffusion models using techniques like Score Distillation Sampling (SDS), which often introduce artifacts such as over-saturation and over-smoothing into the generated 3D assets. In this paper, we address this essential problem by formulating the generation process as learning an optimal, direct transport trajectory between the distribution of the current rendering and the desired target distribution, thereby enabling high-quality generation with smaller Classifier-free Guidance (CFG) values. At first, we theoretically establish SDS as a simplified instance of the Schrödinger Bridge framework. We prove that SDS employs the reverse process of an Schrödinger Bridge, which, under specific conditions (e.g., a Gaussian noise as one end), collapses to SDS's score function of the pre-trained diffusion model. Based upon this, we introduce Trajectory-Centric Distillation (TraCe), a novel text-to-3D generation framework, which reformulates the mathematically trackable framework of Schrödinger Bridge to explicitly construct a diffusion bridge from the current rendering to its text-conditioned, denoised target, and trains a LoRA-adapted model on this trajectory's score dynamics for robust 3D optimization. Comprehensive experiments demonstrate that TraCe consistently achieves superior quality and fidelity to state-of-the-art techniques. Our code will be released to the community.
Ziying Li, Xuequan Lu, Xinkui Zhao, Guanjie Cheng, Shuiguang Deng, Jianwei Yin
NeurIPS3
2025 CoreGuard: Safeguarding Foundational Capabilities of LLMs Against Model Stealing in Edge Deployment
abstract
Proprietary large language models (LLMs) exhibit strong generalization capabilities across diverse tasks and are increasingly deployed on edge devices for efficiency and privacy reasons. However, deploying proprietary LLMs at the edge without adequate protection introduces critical security threats. Attackers can extract model weights and architectures, enabling unauthorized copying and misuse. Even when protective measures prevent full extraction of model weights, attackers may still perform advanced attacks, such as fine-tuning, to further exploit the model. Existing defenses against these threats typically incur significant computational and communication overhead, making them impractical for edge deployment. To safeguard the edge-deployed LLMs, we introduce CoreGuard, a computation- and communication-efficient protection method. CoreGuard employs an efficient protection protocol to reduce computational overhead and minimize communication overhead via a propagation protocol. Extensive experiments show that CoreGuard achieves upper-bound security protection with negligible overhead.
Qinfeng Li, Tianyue Luo, Xuhong Zhang 0002, Yangfan Xie, Yier Jin, Hao Peng 0002, Xinkui Zhao, Xianwei Zhu, Jianwei Yin
NeurIPS9
2025 HeatSnap: A Hot Page-Aware Continuous Snapshots System for Virtual Machines in Web Infrastructure
abstract
Snapshot technology is crucial for data protection and system recovery in virtualized environments, particularly with the growing need for continuous snapshots to maintain the integrity of long-running web-based and distributed applications. However, traditional snapshot methods often suffer from performance bottlenecks, and inefficient storage usage. These challenges are closely tied to the way memory pages are accessed during VM execution, where memory access patterns show significant disparities between frequently accessed "hot" pages and less-used "cold" pages. In this paper, we introduce HeatSnap, a continuous snapshot system designed to address these issues by leveraging the uneven access frequencies of memory pages. HeatSnap distinguishes between intensive hot pages and dirty pages, applying specialized snapshotting and storage strategies to optimize the handling of both hot and cold memory regions. This approach aims to optimize snapshot efficiency, minimize performance impact on the VM, and decrease storage costs. Our implementation of HeatSnap on QEMU/KVM demonstrates significant improvements in VM performance loss, snapshot duration, and storage efficiency compared to existing methods, as evidenced by evaluations on common web and cloud-based workloads.
Kangyue Gao, Chuangyu Ouyang, Xinkui Zhao, Miao Ye, Chen Zhi, Guanjie Cheng, Yueshen Xu, Shuiguang Deng, Jianwei Yin
WWW3
2025 MerKury: Adaptive Resource Allocation to Enhance the Kubernetes Performance for Large-Scale Clusters
abstract
As a dominant paradigm in modern web applications, cloud computing has seen a surge in adoption. The deployment of vast and various workloads encapsulated within containers has become ubiquitous across cloud platforms, imposing substantial demands on the supporting infrastructure. However, Kubernetes (k8s), the de facto standard for container orchestration, struggles with low scheduling throughput and high latency in large-scale clusters. The primary challenges are identified as excessive load from read requests and resource contention between co-located components. In this paper, we present MerKury, a general and lightweight framework designed to enhance the Kubernetes performance for large-scale clusters. MerKury employs a dual strategy: first, it preprocesses specific requests to alleviate excessive load; second, it introduces an adaptive resource allocation algorithm to mitigate resource contention. Evaluations across various cluster scales demonstrate that MerKury notably augments node capacity by up to 4.5×, increases scheduling throughput by up to 7.3×, and reduces request latency by 5.6%-57.7%, outperforming vanilla Kubernetes and baseline resource allocation methods.
Jiayin Luo, Xinkui Zhao, Shengye Pang, Jianwei Yin
WWW2
2025 DS-MAE: Dual-Siamese Masked Autoencoders for Point Cloud Analysis
abstract
Masked autoencoders (MAEs) have emerged as a powerful self-supervised approach for point cloud analysis. Nevertheless, existing methods often separately focus on global structures or multi-scale features, ignoring their complementary potential. In this paper, we propose a novel dual-Siamese masked autoencoder (DS-MAE) framework that explores integrating global and hierarchical feature learning in a unified architecture for point cloud analysis. In particular, we introduce a consistent dual-branch patch embedding strategy to partition the point cloud into patches using shared group centers, ensuring both global and hierarchical branches process point patches centered at the same spatial locations. Each branch employs dual-branch Siamese encoders to process original and augmented point patches, learning representations that capture both local details and global context. In addition, we have designed cross-attention Siamese decoders to reconstruct masked point patches and align features both within and between branches with cross-attention mechanisms. Comprehensive experiments demonstrate our method consistently achieves superior results to prior methods. Code is available at https://github.com/shaoandy1211/DS-MAE.git.
Di Shao, Yaping Jing, Xinkui Zhao, Shasha Mao, Lei Lyu 0001, Xiao Liu 0004, Xuequan Lu
Comput. Vis. Media3
2025 BPI: A Novel Efficient and Reliable Search Structure for Hybrid Storage Blockchain
abstract
Hybrid storage solutions have emerged as potent strategies to alleviate the data storage bottlenecks prevalent in blockchain systems. These solutions harness off-chain Storage Services Providers (SP) in conjunction with Authenticated Data Structures (ADS) to ensure data integrity and accuracy. Despite these advancements, the reliance on centralized SPs raises concerns about query correctness, as the integrity of query results depends on the SPs' trustworthiness. Although ADS can verify the integrity of individual data points, they fall short of preventing SPs from omitting valid results. In this paper, we delineate the fundamental distinctions between data retrieval in blockchains and traditional database systems. Drawing upon these insights, we introduce the BPI framework, which employs a suite of validation models that ascertain the inclusion of all valid content in retrieval outcomes, with low overhead. We further present ''Articulated Search'', a query pattern specifically tailored for blockchain environments, which not only enhances retrieval efficiency but also substantially reduces costs during data user updates. Extensive experimental evaluations demonstrate that the BPI framework achieves outstanding scalability and performance in keyword searches within blockchain environments, surpassing EthMB+ and state-of-the-art search databases commonly used in mainstream hybrid storage blockchains (HSB). Notably, the Articulated Search pattern improves query performance by over three orders of magnitude, highlighting its potential as a transformative approach to blockchain query optimization.
Xinkui Zhao, Rengrong Xiong, Guanjie Cheng, Xinhao Jin, Shawn Shi, Xiubo Liang, Gongsheng Yuan, Xiaoye Miao, Jianwei Yin, Shuiguang Deng
Proc. ACM Manag. Data1
2025 Pseudo-labeling with keyword refining for few-supervised video captioning
Ping Li 0006, Xinkui Zhao, Xianghua Xu, Mingli Song
Pattern Recognit.3
2025 Adaptive Scheduling of High-Availability Drone Swarms for Congestion Alleviation in Connected Automated Vehicles
abstract
The Intelligent Transportation System (ITS) serves as a pivotal element within urban networks, offering decision support to users and connected automated vehicles through comprehensive information gathering, sensing, device control, and data processing. Presently, ITS predominantly relies on sensors embedded in fixed infrastructure, notably Roadside Units (RSUs). However, RSUs are confined by coverage limitations and may encounter challenges in prompt emergency responses. On-demand resources, such as drones, present a viable option to supplement these deficiencies effectively. This article introduces an approach where Software-Defined Networking and Mobile Edge Computing technologies are integrated to formulate a high-availability drone swarm control and communication infrastructure framework comprising the cloud layer, edge layer, and device layer. Drones confront limitations in flight duration attributed to battery limitations, posing a challenge in sustaining continuous monitoring of road conditions over extended periods. Effective drone scheduling stands as a promising solution to overcome these constraints. To tackle this issue, we initially utilized Graph WaveNet, a specialized graph neural network structure tailored for spatial-temporal graph modeling, for training a congestion prediction model using real-world dataset inputs. Building upon this, we further propose an algorithm for drone scheduling based on congestion prediction. Our simulation experiments using real-world data demonstrate that, compared to the baseline method, the proposed scheduling algorithm not only yielded superior scheduling gains but also mitigated drone idle rates.
Shengye Pang, Zhen Qin 0004, Xinkui Zhao, Jintao Chen 0001, Fan Wang 0020, Jianwei Yin
ACM Trans. Auton. Adapt. Syst.4
2025 UKFaaS: Lightweight, High-Performance and Secure FaaS Communication With Unikernel
abstract
Unikernel is a promising runtime for serverless computing with its lightweight and isolated architecture. It offers a secure and efficient environment for applications. However, famous serverless frameworks like Knative have introduced heavyweight component sidecars to assist function instance deployment in a non-intrusive manner. But the sidecar not only hinders the throughput of unikernel function services but also consumes excessive memory resources. Moreover, the intricate network communication pathways among various services pose significant challenges for deploying unikernels in production serverless environments. Although shared-memory based communication on the same server can solve the communication bottleneck of unikernel-based function instances. The situation where malicious programs on the server make the shared memory untrustworthy limits the deployment of such technologies.We propose UKFaaS, a lightweight and high-performance serverless framework. UKFaaS leverages the advantages of customized operating systems through unikernel and it non-intrusively integrates sidecar functionality into the unikernel, avoiding the overhead of sidecar request forwarding. Additionally, UKFaaS innovatively implements data communication between unikernels in the same server to eliminate VM-Exit bottlenecks in RPC (remote process call) based on VMFUNC without relying on memory sharing. The preliminary experimental results indicate that UKFaaS can realize 1.8×-3.5× request throughput per second (RPS) compared with the advanced serverless system FaasFlow, UaaF and Nightcore in the Google online boutique microservice benchmark.
Zhenqian Chen, Yuchun Zhan, Xinkui Zhao, Muyu Yang, Siwei Tan, Lufei Zhang, Liqiang Lu, Jianwei Yin, Zuoning Chen
IEEE Trans. Computers4
2025 Cooperative Driving at Multiple Unsignalized Intersections in Fully Autonomous Driving Scenarios
Weihang Pan, Binbin Lin 0001, Zhengxu Yu, Xinkui Zhao, Xiaofei He 0001, Jieping Ye
IEEE Trans. Intell. Transp. Syst.5
2025 Large-Scale Storage Location Assignment via Hierarchical Reinforcement Learning: A Rank and Assign Approach
abstract
The storage location assignment problem for export containers (EC-SLAP) is crucial to the efficiency of cargo turnover in ports. Existing methods fall short in real-world applications due to the challenges of the unpredictability of container arrival sequences and the large-scale problem. We propose a new buffer-wise framework based on hierarchical reinforcement learning for EC-SLAP, aimed at optimizing container turnover efficiency. The framework comprises two processes: 1)Ranking Agentranks containers in the buffer, reducing the uncertainty of random arrival sequences compared to the immediate assignment. 2)Assigning Agentsassign storage locations in a two-step process by block and slot, diminishing the dimensionality of the large-scale discrete action space. We iteratively optimize agents by asynchronously obtaining rewards from the environment. In addition, to address the challenge of sparse rewards in long-sequence decision-making, we have developed a novel immediate reward function to enhance learning efficiency and accelerate convergence. We propose a new large-scale dataset, NZP-SLAD, collected from real-world historical data from the terminal operating system of Ningbo-Zhoushan Port and develop a realistic container terminal simulator. We conducted numerous offline simulations and tests with this dataset. The experimental results demonstrate that our proposed method achieves rapid convergence and significantly surpasses expert methods used in real-world production.
Weihang Pan, Minghao Chen 0001, Binbin Lin 0001, Xinkui Zhao, Xiaofei He 0001, Jieping Ye
IEEE Trans. Knowl. Data Eng.5
2025 TrustPay: A Dual-Layer Blockchain-Based Framework for Trusted Service Transaction
abstract
Web service-oriented transactions have become an integral part of the Internet economy, with the mainstream transaction patterns relying primarily on cloud service markets. However, traditional service transaction methods have deficiencies in terms of both trust and scalability. Distrust between service provider (SP) and consumer (SC), particularly around online payments and data security, impedes the further growth of service transactions. Although blockchain-based transaction mechanisms have made notable progress in addressing trust issues, they still face performance bottlenecks. To tackle these challenges, this paper introduces TrustPay, a service transaction framework that leverages a dual-layer blockchain structure consisting of a parent chain and multiple subchains. The framework partitions the subchain network based on business, with each subchain dedicated to storing service invocation records generated within a specific business unit. The smart contract deployed on the parent chain will settle the invocation records in all subchain networks as transaction records and facilitate automatic transfers among blockchain accounts. This design leverages blockchain's inherent reliability while improving its scalability in large-scale scenarios. Additionally, a novel consensus protocol, REFEREE, is introduced and applied to the subchain network, ensuring efficient recording of invocation data and trusted verification among participants, further enhancing both trust and performance. Comparative experiments and analysis show that TrustPay's dual-layer blockchain structure and REFEREE protocol are not only reliable but also outperform baseline methods in terms of efficiency.
Shengye Pang, Xinkui Zhao, Shuyi Yu, Jintao Chen 0001, Shuiguang Deng, Jianwei Yin
IEEE Trans. Serv. Comput.2
2025 DeSC: Learning Deep Semantic Descriptor for NeRF Registration
abstract
NeRF registration has gained increasing attention recently. While existing research demonstrates considerable potential for this task, most methods primarily focus on either global geometric or rendering photometric information during feature learning, overlooking the rich cross-modal information inherent in the NeRF embedding feature space. In this paper, we propose DeSC, a novel NeRF registration approach that leverages the rich cross-modal features from NeRF to learn robust semantic descriptors. In particular, we propose a Deep Semantic Aggregation module, which employs a weighted graph convolution network to capture high-frequency texture details in NeRF patches. This approach reveals the underlying semantics shared across different NeRFs of the same scene, thereby yielding more robust global feature descriptors that lead to better alignment accuracy and robustness. In addition, we design a density-aware photometric consistency loss that facilitates the learning of robust features. Extensive experimental results on Objaverse datasets demonstrate that our approach produces superior registration performance to state-of-the-art techniques.
Sheldon Fung, Wei Pan 0010, Kui Su, Hui Cui 0002, Xinkui Zhao, Xuequan Lu
IEEE Trans. Vis. Comput. Graph.5
2024 QuFEM: Fast and Accurate Quantum Readout Calibration Using the Finite Element Method
abstract
Quantum readout noise turns out to be the most significant source of error, which greatly affects the measurement fidelity. Matrix-based calibration has been demonstrated to be effective in various quantum platforms. However, existing methodologies are fundamentally limited in either scalability or accuracy. Inspired by the classical finite element method (FEM), a formal method to model the complex interaction between elements, we present our calibration framework named QuFEM. First, we apply a divide-and-conquer strategy that formulates the calibration as a series of tensor products with noise matrices. This matrices are iteratively characterized together with the calibrated probability distribution, aiming to capture the inherent locality of qubit interactions. Then, to accelerate the end-to-end calibration, we propose a sparse tensor-product engine to exploit the sparsity in the intermediate values. Our experiments show that QuFEM achieves 2.5×103× speedup in the 136-qubit calibration compared to the state-of-the-art matrix-based calibration technique [50], and provides 1.2× and 1.4× fidelity improvement on the 18-qubit and 36-qubit real-world quantum devices.
Siwei Tan, Liqiang Lu, Congliang Lang, Yongheng Shang, Xinkui Zhao, Mingshuai Chen, Yun Liang 0001, Jianwei Yin
ASPLOS (2)7
2024 Sharry: An Efficient and Sharing Far Memory System
abstract
Far Memory System(FMS) allows applications to access memory on remote machines(called memory nodes). However, existing FMSs can't deal with large loads and have low efficiency in utilizing far memory, which leads to the inability to share memory nodes among multiple processes, limiting the scalability of FMS. In this paper, we propose Sharry, an efficient Sharing FMS. Sharry manages memory objects from multiple processes within a unified address space, avoiding the overhead of space switching. Sharry also optimizes the utilization of far memory with fine-grained memory management. Additionally, Sharry offloads memory allocation to dedicated CPU core in order to handle larger loads in the sharing scenario. Compared to state-of-the-art FMS, Sharry improves memory utilisation by 45%, causing only 9% performance degradation when multiple processes sharing single memory node.
Yuhang Huang 0005, Shuiguang Deng, Jianwei Yin, Xinkui Zhao
DAC5
2024 EE blockchain: End-to-end service regulation and efficient retrieval and categorization on the underlying level
abstract
In large-scale digital service sharing scenarios, given the large number of participating users, frequent cross-domain service interactions, and high-frequency service transactions, to ensure the trustworthiness of digital services, the architecture of the digital service sharing system usually chooses blockchain as its technological foundation. This not only ensures the security and credible deposit of data, but also achieves the credible traceability of data. However, in the current blockchain-based notarization architecture, there may be potential privacy leakage during data transmission, and users are unable to choose the encryption level of their data for blockchain deposition according to their own needs. Moreover, the underlying databases in current applications using blockchain lack convenient retrieval and categorization functionalities. In this work, we propose a trustworthy blockchain solution based on permission management, which implements hierarchical encryption and enables users to flexibly encrypt data according to their needs. Additionally, through the Double Star storage system, while ensuring the reliability of the system, we have also greatly improved the efficiency of data retrieval and classification. Compared to blockchain platforms like XRP and EOS, our solution achieves superior data retrieval and classification efficiency while implementing layered encryption.
Rengrong Xiong, Guanjie Cheng, DianKai Hu, Yueshen Xu, Xiubo Liang, Xinkui Zhao
ICWS6
2024 LLM-powered Zero-shot Online Log Parsing
abstract
Log parsing is an essential step for log analysis, which transforms raw log messages into structured form by extracting the log templates. Automatic log parsing have been the subject of extensive research. Recently, several studies have explored improving the performance of automatic log parsing via deep-learning-based approaches, especially for the pre-trained language models. However, the use of such large-scale language models for log parsing encounters several challenges, including hallucinations, high-cost and labelling efforts. To address these challenges, this paper introduces YALP, a zero-shot log parsing solution that tackles the aforementioned challenges by utilizing the capabilities of ChatGPT in conjunction with traditional methods, without incorporating user labelling. Our experiments on 16 public log datasets shows that our method outperforms several popular traditional methods in commonly used evaluation metrics. In comparison to directly utilizing GPT for log parsing tasks, our methods demonstrates significant improvements in both efficiency and effectiveness.
Chen Zhi, Liye Cheng, Xinkui Zhao, Yueshen Xu, Shuiguang Deng
ICWS4
2024 PCoTTA: Continual Test-Time Adaptation for Multi-Task Point Cloud Understanding
abstract
In this paper, we present PCoTTA, an innovative, pioneering framework for Continual Test-Time Adaptation (CoTTA) in multi-task point cloud understanding, enhancing the model's transferability towards the continually changing target domain. We introduce a multi-task setting for PCoTTA, which is practical and realistic, handling multiple tasks within one unified model during the continual adaptation. Our PCoTTA involves three key components: automatic prototype mixture (APM), Gaussian Splatted feature shifting (GSFS), and contrastive prototype repulsion (CPR). Firstly, APM is designed to automatically mix the source prototypes with the learnable prototypes with a similarity balancing factor, avoiding catastrophic forgetting. Then, GSFS dynamically shifts the testing sample toward the source domain, mitigating error accumulation in an online manner. In addition, CPR is proposed to pull the nearest learnable prototype close to the testing feature and push it away from other prototypes, making each prototype distinguishable during the adaptation. Experimental comparisons lead to a new benchmark, demonstrating PCoTTA's superiority in boosting the model's transferability towards the continually changing target domain. Our source code is available at: https://github.com/Jinec98/PCoTTA.
Jincen Jiang, Qianyu Zhou 0001, Yuhang Li 0011, Xinkui Zhao, Meili Wang 0001, Lizhuang Ma, Jian Chang 0001, Jian J. Zhang 0001, Xuequan Lu
NeurIPS4
2023 Incentive-Driven Pricing Game for Multi-Edge Service Providers towards Optimal Profits
abstract
The growth of service ecosystems, which include diversified services such as cloud and edge services, has led to a thriving service transaction market. However, service pricing remains a major obstacle to further progress. Without proper pricing guidance, service providers tend to formulate pricing strategies solely based on their own interests, which frequently hinders the maximization of overall market benefits. This problem is even more challenging in edge computing scenarios as different Edge Service Providers (ESPs) are located in distributed regions and influenced by multiple factors, making it difficult to formulate a single pricing model. This paper proposes a multi-participant stochastic game model to formalize the multi-edge service pricing problem. An incentive mechanism based on Pareto improvement is then proposed to drive the game to the Pareto optimal direction with optimal profits. Finally, an improved PSO algorithm is proposed to solve the game model and analyze the equilibrium states under different evolutionary mechanisms. Experimental results indicate that the proposed pricing incentive mechanism can promote a more effective and reasonable pricing allocation, avoiding 21.6% anarchism loss of the overall profits, while showcasing the effectiveness of our algorithm in solving the game.
Shengye Pang, Xinkui Zhao, Jiayin Luo, Bangpeng Zheng, Jianwei Yin
ICWS2
2023 Personalized Repository Recommendation Service for Developers with Multi-modal Features Learning
abstract
Nowadays an increasing number of software developers have joined in open-source software development communities such as GitHub, and develop and share softwares in these communities. Those online communities contain a huge volume of open-source repositories. Developers commonly search from existing repositories and intend to find suitable repositories to their development requirements. However, it is time-and energy-consuming to discover suitable repositories from such a large number of candidates and it may be also hard for developers to choose accurate keywords. So an effective repository recommendation service becomes an indispensable tool for developers. There have been some solutions for repository recommendation, but existing solutions have several defects such as mediocre accuracy and ignorance of useful features. In this paper, we develop a new personalized repository recommendation service with multi-modal features learning. We propose to mine two modes of features and jointly utilize the mined multimodal features. One of the features is the developers’ sequential behavior features and the other is text features of repositories. We design novel features learning mechanisms for the two modes of features. We performed sufficient experiments on a real-world dataset and the experimental results demonstrate that our model generates superior recommendation results and produces an improvement of 15.3% and 14.5% in Precision and Recall compared to well-known existing methods.
Yueshen Xu, Xinkui Zhao, Ying Li 0001, Rui Li 0047
ICWS3
2023 Android malware detection via efficient application programming interface call sequences extraction and machine learning classifiers
abstract
Abstract Malware detection is an important task for the ecosystem of mobile applications (APPs), especially for the Android ecosystem, and is vital to guarantee the user experience of Android APPs. There have been some exiting methods trying to solve the problem of malware detection, but the methods suffer from several defects, such as high time complexity and mediocre accuracy, which seriously decrease the practicability of existing methods. To solve these problems, in this study, we propose a novel Android malware detection framework, where we contribute an efficient Application Programming Interface (API) call sequences extraction algorithm and an investigation of different types of classifiers. In API call sequences extraction, we propose an algorithm for transforming the function call graph from a multigraph into a directed simple graph, which successfully avoids the unnecessary repetitive path searching. We also propose a pruning search, which further reduces the number of paths to be searched. Our algorithm greatly reduces the time complexity. We generate the transition matrix as classification features and investigate three types of machine learning classifiers to complete the malware detection task. The experiments are performed on real‐world Android Packages (APKs), and the results demonstrate that our method significantly reduces the running time and produces high detection accuracy.
Tanjie Wang, Yueshen Xu, Xinkui Zhao, Zhiping Jiang, Rui Li 0047
IET Softw.3
2023 DeepBoot: Dynamic Scheduling System for Training and Inference Deep Learning Tasks in GPU Cluster
abstract
Deep learning tasks (DLT) include training and inference tasks, where training DLTs have requirements on minimizing average job completion time (JCT) and inference tasks need sufficient GPUs to meet real-time performance. Unfortunately, existing work separately deploys multi-tenant training and inference GPU cluster, leading to the high JCT of training DLTs with limited GPUs when the inference cluster is under insufficient GPU utilization due to the periodic inference workload. DeepBoot solves the challenges by utilizing idle GPUs in the inference cluster for the training DLTs. Specifically, 1) DeepBoot designsadaptive task scaling(ATS) algorithm to allocate GPUs in the training and inference clusters for training DLTs and minimize the performance loss when reclaiming inference GPUs. 2) DeepBoot implementsauto-fast elastic(AFE) training based on Pollux to reduce the restart overhead by inference GPU reclaiming. Our implementation on the testbed and large-scale simulation in Microsoft deep learning workload shows that DeepBoot can achieve 32% and 38% average JCT reduction respectively compared with the scheduler without utilizing idle GPUs in the inference cluster.
Zhenqian Chen, Xinkui Zhao, Chen Zhi, Jianwei Yin
IEEE Trans. Parallel Distributed Syst.2
2017 SimMon: a toolkit for simulation of monitoring mechanisms in cloud computing environment
abstract
Summary Monitoring is a precondition for intelligent management in cloud computing environment, such as dynamic resource allocation. Typically, working as an auxiliary tool, a monitoring system is expected to incur the least additional resource usage, thus the strategies to improve the efficiency of monitoring mechanisms become significant. Yet the scale and monitoring requirement of different data centers vary, we cannot determine whether a monitoring mechanism would work well in a new data center before it serves the data center. To evaluate monitoring mechanisms, we proposeSimMon, a toolkit for simulating monitoring mechanisms in cloud computing environments.SimMonis designed to simulate the topologies, actions, and strategies in data collection, dissemination, storage, and management processes.SimMonprovides a controllable and repeatable way to evaluate monitoring mechanisms. In this paper, we describe the requirements analysis, design, implementation, and evaluation ofSimMon. We simulate several different monitoring systems and compare their cost on time and resource to evaluate the efficiency ofSimMon. We reproduce two usage scenarios from former literatures to demonstrate the effectiveness ofSimMonon monitoring mechanisms simulation and evaluation. We build a real‐world working environment to validate the capability ofSimMonon mimicking the characteristics of cloud monitoring systems. Copyright © 2016 John Wiley & Sons, Ltd.
Xinkui Zhao, Jianwei Yin, Chen Zhi, Zuoning Chen
Concurr. Comput. Pract. Exp.1
2017 CloudScout: A Non-Intrusive Approach to Service Dependency Discovery
abstract
Nowadays, numerous enterprises are migrating their applications into cloud computing environments. Typically, the applications are composed of several dependent service components that span many hosts and network devices. In light of this, exploring the dependency between service components can be beneficial for achieving fast network application response time. Moreover, it is significant to consolidate service components according to resource constraints, service dependency, and network structure. However, it is a tedious task to discover the dependency among service components without expert knowledge of the running application. In this paper, we propose CloudScout, a non-intrusive approach that is capable of automatically discovering dependent service components. CloudScout analyzes the correlation among service components based on the time-series information from system monitoring logs. We address two key challenges in CloudScout: service distance calculation and dependent service clustering. We conduct experiments on five applications with 290 service components that span 20 physical hosts across two data centers. The experimental results demonstrate that CloudScout can successfully discover the dependency among service components and facilitate reducing the network latency of network applications and distributed applications.
Jianwei Yin, Xinkui Zhao, Chen Zhi, Zuoning Chen, Zhaohui Wu 0001
IEEE Trans. Parallel Distributed Syst.2
2016 MonValley: An Unified Monitoring and Management Framework for Cloud Services
abstract
Monitoring is the cornerstone for cloud service management, so it is significant for a cloud monitoring tool to support customization for specific monitoring requirements, to indicate the correlation between target services, and to guide adaptive service management. Unfortunately, traditional monitoring tools are always developed independently with service management platforms and are provided as self-contained softwares, which limit their capabilities on addressing these requirements. In this paper, we propose MonValley, an unified monitoring and management framework for cloud services. It consists of four components: 1) a high-level language for practitioners to express monitoring specifications on the services from all three layers, 2) a compiler to translate the expressive program into an executable program, 3) an execution engine to execute the executable program, 4) a runtime to provide supports on basic functionalities, such as data transmission. MonValley provides an innovative approach to facilitate the integrated monitoring and adaptive management for cloud services.
Xinkui Zhao, Jianwei Yin, Chen Zhi, Zuoning Chen
ICWS1
2016 vSpec: workload-adaptive operating system specialization for virtual machines in cloud computing
Xinkui Zhao, Jianwei Yin, Zuoning Chen
Sci. China Inf. Sci.1
2015 Can Cloud Service Get His Family? A Step Towards Service Family Detecting
abstract
In cloud computing environment, an application is always composed of several service components. A collection of service components is called a service family, and we name the cloud service components as service family members. In this paper, we propose a solution named Icebreaker to assemble service components belonging to the same application without sniffing tenants' privacy. Icebreaker characterizes each service component with basic resource consuming information and proposes a new distance calculating algorithm named iEntropy to distinct service components. We adaptively adopt Affinity Propagation (AP) clustering algorithm and maximum Silhouette index to identify the number of service family and assemble the service family members. Experiments are conducted on RUBiS, Hadoop and ApacheBench clusters with 169 VMs. Evaluation results show that Icebreaker can get 96.45% accuracy.
Xinkui Zhao, Jianwei Yin, Chen Zhi, Pengxiang Lin, Zuoning Chen
CLUSTER1
2015 monBench: A Database Performance Benchmark for Cloud Monitoring System
abstract
Monitoring system provides a clear insight into the state and performance of service components in cloud computing platforms. It collects metrics from dispersed sensors and stores them in databases for show and future query. To choose the most suitable database for a monitoring system is complex, since the performance requirement on cloud monitoring system in different data centers differs widely. In this paper, we propose a benchmark named monBench to evaluate the performance of databases in cloud monitoring systems. Monbench extracts structural data from real-world monitoring logs to construct the benchmarking workload. Several performance metrics, such as throughput and query time, are evaluated by monBench.
Xinkui Zhao, Jianwei Yin, Chen Zhi, Pengxiang Lin, Shichun Feng, Zuoning Chen
CLUSTER1
2015 SimMon: A Toolkit for Simulating Monitoring Mechanism in Cloud Computing Environments
Xinkui Zhao, Jianwei Yin, Pengxiang Lin, Chen Zhi, Shichun Feng, Zuoning Chen
ICSOC1
2015 BURSE: A Bursty and Self-Similar Workload Generator for Cloud Computing
abstract
As two of the most important characteristics of workloads, burstiness and self-similarity are gaining more and more attention. Workload generation, which is a key technique for performance analysis and simulations, has also attracted an increasing interest in cloud community in recent years. Though a large number of methods for synthetically generating bursty or self-similar workloads have been proposed in the literature, none of them can deal with workload generation with both of the two characteristics. In this paper, a configurable and intelligible synthetic generator (BURSE) is proposed for bursty and self-similar workloads in cloud computing based on a superposition of two-state Markov Modulated Poisson Processes (MMPP2s). The proposed generator can produce workloads with both specified intension of burstiness and self-similarity. Detailed experimental evaluation demonstrates the accuracy, robustness and good applicability of BURSE.
Jianwei Yin, Xingjian Lu, Xinkui Zhao, Hanwei Chen, Xue (Steve) Liu
IEEE Trans. Parallel Distributed Syst.3
2014 System resource utilization analysis and prediction for cloud based applications under bursty workloads
Jianwei Yin, Xingjian Lu, Hanwei Chen, Xinkui Zhao, Naixue Xiong
Inf. Sci.4
2013 Workload Classification Model for Specializing Virtual Machine Operating System
abstract
There is growing demand on strategies to help cloud computing utilize its scale adaptiveness and cost effectiveness advantages. Previous operating systems(OS) are designed to suit all, leading to that virtual machines with different workloads use indiscriminate processing platform. However, there are conflicts between generality and performance, limited resource utilization and low processing efficiency of common OS penalize system performance. Therefore, we design four kinds of OS optimization strategies corresponding to four primary classes of workloads: CPU-Intensive, Memory-Intensive, I/O-Intensive and Network-Intensive. In this paper, we propose a Feedback-Based Workload Classification(FBWC) model which contains metrics collector, data preprocessor, Training Set Refresh Support Vector Machine(TSRSVM) classifier, decision maker and operating system tuner to classify workloads into appropriate class. TSRSVM combines support vectors of origin training set and correctly classified testing set together as new training set to get higher classification accuracy and efficiency. Comprehensive experiments compared with K Nearest Neighbors(KNN) and SVM demonstrate effectiveness of FBWC model and TSRSVM classification algorithm. Performance comparison between common virtual machine and the tuned one shows high degree performance improvement by OS specialization.
Xinkui Zhao, Jianwei Yin, Zuoning Chen
IEEE CLOUD1
2013 A synthetic bursty workload generation method for web 2.0 benchmark
abstract
As one of most important characteristics of Web-based systems'workloads, burstiness is gaining more and more attentions. And synthetically generating bursty workloads is a key technique for performance analysis. In this paper, a configurable and intelligible synthetic bursty workload generation method for Web 2.0 benchmark Olio has been proposed based on 2-state Markovian arrival processes (MAP2). By comparing the actual value of index of dispersion for counts (IDC) estimated from system logs with the target value deduced from MAP2 model, we show that our method is more accurate than related work.
Jianwei Yin, Hanwei Chen, Xingjian Lu, Xinkui Zhao
CLUSTER4
2013 Distance-aware virtual cluster performance optimization: A hadoop case study
abstract
Cloud computing and big data are becoming two important developing trends in information technology area. However, data-intensive computing has some challenges to work well on virtual machines in cloud computing for virtualized resource competition and complex network communication. Network becomes one of the most notorious bottlenecks, which highlights strategies to lower communication and transmission cost in virtual cluster. In this paper, we present a novel cluster performance optimization strategy named vClusterOpt. vClusterOpt finds out centralized subgraphs of node graph and choose node with the shortest logical distance as kernel node of the subgraph to reduce inter-machine communication and transmission cost under virtual cluster. To calculate logical distance accurately, we define two kinds of logical distance: Logical Communication Distance(LCD) and Logical Transmission Distance(LTD). VM with the shortest LCD with others is used as the communication kernel node who has the most information communication stress, while VM with the shortest LTD is treated as transmission kernel node who has the most data transmission stress. We choose benchmarks running on Hadoop as the represent of data-intensive computing service to demonstrate effectiveness of our approach. Experiments show that an average of 20% performance improvement can get by our distance-aware virtual cluster optimization strategy.
Xinkui Zhao, Jianwei Yin, Zuoning Chen, Xingjian Lu
CLUSTER1
2013 An Approach for Bursty and Self-similar Workload Generation
Xingjian Lu, Jianwei Yin, Hanwei Chen, Xinkui Zhao
WISE (2)4