Ruhui Ma

dblp:01/5518 · DBLP profile ↗
← Back
110ranked-venue papers
9as first author
67since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 32 · 2 first-author · 12 since 2021Artificial intelligence and machine learning · 21 · 17 since 2021Computer networks · 21 · 5 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 14 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 2 first-author · 10 since 2021Security and privacy · 5 · 3 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Software engineering, systems software and programming languages · 4
YearPublicationVenuePosition
2026 DecentFL: A Decentralized Federated Learning System With Multiple Aggregators and Evaluators
Khushal Haider Syed, Ruben Poghosyan, Hongjian Shi, Ruhui Ma, Ejafa Bassam, Ziang Zhou, Zhongle Qu, Haibing Guan
IEEE Internet Things J.4
2026 Ultrahigh-Speed Terminal Access and Collaborative Authentication Scheme in Satellite Networks
Yukun Zhu, Ruhui Ma, Runsheng Fu, Jin Cao 0001, Hui Li 0006, Xiaosong Zhang 0001
IEEE Internet Things J.2
2026 HeShare: Energy-Aware and Efficient Multi-Task GPU Sharing in Heterogeneous GPU-Based Computing Systems
abstract
With the rapid growth of artificial intelligence and large-scale model computing, the demand for GPUs in datacenters continues to increase, especially for large-scale training and inference tasks. Heterogeneous multi-GPU systems, which integrate GPUs with varying types and computational capabilities, have become critical computing resources. This leads to two main challenges. First, due to the differences in GPU performance and power consumption, task scheduling involves a complex multi-objective optimization to balance energy efficiency and performance. More importantly, the lack of coordinated mechanisms for multi-task sharing and energy-efficient resource management across heterogeneous GPUs can result in GPU overload or underutilization, leading to wasted resources and potential system risks.To address these challenges, we propose HESHARE, an energy-aware and efficient heterogeneous GPU framework for datacenters. First, we design an energy-aware task scheduling strategy that optimizes task allocation across different GPUs to achieve a balance between energy consumption and performance. Second, we introduce a GPU sharing optimization mechanism that adaptively configures MPS and DVFS settings for each GPU, enhancing resource utilization, reducing overall energy consumption, and ensuring task performance. Compared to the state-of-the-art framework, we reduce average energy costs by 26% and improve job completion time by 31%, achieving a balance between energy efficiency and performance.
Zhuolong Jiang, Zinuo Cai, Baoheng Zhang, Yiming Qiang, Ruhui Ma, Haibing Guan, Rajkumar Buyya
IEEE Trans. Computers7
2026 MPCPFL: Prototype-Based Multimodal Personalized Federated Learning With Missing Modalities in Medical Domain
abstract
With the global aging population on the rise, the amount of disease data related to the elderly is also increasing. Effectively protecting personal data when analyzing and processing disease data has become crucial. Federated learning (FL), as a proven method, is widely used. In the multimodal FL for disease diagnosis, the intricate challenge of modality missing and heterogeneous data distribution across healthcare clients severely degrades model performance and generalization. Existing frameworks struggle to address these issues in scenarios with complex modality missing and data heterogeneity, particularly in privacy-sensitive medical applications. In this article, we introduce a novel framework,modal prototype and contrast for personalized federated learning, which enhances prototype-based multimodal federated learning with personalized local aggregation and enables the adaptive fusion of global prototypes and local client models during training. By leveraging the modal prototype to compensate for missing modalities and modal contrastive training to align client representations with shared prototypes, the framework mitigates global model degradation caused by modality absence. In addition, we adapt the SLAKE medical multimodal dataset into a VQAv2-style classification task, simulating realistic disease diagnosis scenarios with varying modality missing rates (10$\sim$90%). Experimental results demonstrate that our approach achieves state-of-the-art performance. Compared to the baselines, our method improved inference accuracy by up to 13.7% when 90% of the modality was missing during training.
Qiwei Yu, Hongjian Shi, Zhuolong Jiang, Yuhan Qiu, Ruhui Ma, Shiliang Liao, Haibing Guan
IEEE Trans. Comput. Soc. Syst.5
2026 FedRL: A Reinforcement Learning Federated Recommender System for Efficient Communication Using Reinforcement Selector and Hypernet Generator
abstract
The field of recommender systems aims to predict users’ latent interests by analyzing their preferences and behaviors. However, privacy concerns about user data collection lead to challenges such as incomplete initial information and data sparsity. Federated learning has emerged to address these privacy issues in recommender systems. However, federated recommender systems face heterogeneity among edge devices regarding data features and sample sizes. Moreover, differences in computational and storage capabilities introduce communication overhead and processing delays during parameter aggregation at the third-party server. This article introduces a framework named FedRL , a reinforcement learning federated recommender system for efficient communication using Reinforcement Selector and Hypernet Generator, to address the proposed issues. The Reinforcement Selector dynamically selects participating edge devices and helps to maximize their use of local data resources. Meanwhile, Hypernet Generator optimizes communication bandwidth consumption during the federated learning parameter transmission, enabling rapid deployment and updates of new model architectures or hyperparameters. Furthermore, the framework incorporates item attributes as content embeddings in edge devices’ recommender models, enriching them with global information. Real-world dataset experiments demonstrate that the proposed solution balances recommender quality and communication efficiency. The code for this work is publicly available on GitHub: https://github.com/diyicheng/FedRL .
Yicheng Di, Hongjian Shi, Ruhui Ma, Honghao Gao, Yuan Liu 0021
Trans. Recomm. Syst.3
2026 FaaShare: Enabling Generic Low-Latency and SLO-Aware GPU Sharing in Serverless Computing
Zhuolong Jiang, Zinuo Cai, Yumou Liu, Ruhui Ma, Rajkumar Buyya
IEEE Trans. Parallel Distributed Syst.4
2025 TST: A Schema-Based Top-Down and Dynamic-Aware Agent of Text-to-Table Tasks
abstract
As a bridge between natural texts and information systems like structured storage, statistical analysis, retrieving, and recommendation, the text-to-table task has received widespread attention recently. Existing researches have gone through a paradigm shift from traditional bottom-up IE (Information Extraction) to top-down LLMs-based question answering with RAG (Retrieval-Augmented Generation). Furthermore, these methods mainly adopt end-to-end models or use multi-stage pipelines to extract text content based on static table structures. However, they neglect to deal with precise inner-document evidence extraction and dynamic information such as multiple entities and events, which can not be defined in static table head format and are very common in natural texts.To address this issue, we propose a two-stage dynamic content extraction agent framework called TST (Text-Schema-Table), which uses type recognition methods to extract context evidences with the conduction of domain schema sequentially. Based on the evidence, firstly we quantify the total instances of each dynamic object and then extract them with ordered numerical prompts. Through extensive comparisons with existing methods across different datasets, our extraction framework exhibits state-of-the-art (SOTA) performance. Our codes are available at https://github.com/jiangpw41/TST.
Peiwen Jiang, Haitong Jiang, Ruhui Ma, Yvonne Jie Chen, Jinhua Cheng
ACL (1)3
2025 AARC: Automated Affinity-aware Resource Configuration for Serverless Workflows
abstract
Serverless computing is increasingly adopted for its ability to manage complex, event-driven workloads without the need for infrastructure provisioning. However, traditional resource allocation in serverless platforms couples CPU and memory, which may not be optimal for all functions. Existing decoupling approaches, while offering some flexibility, are not designed to handle the vast configuration space and complexity of serverless workflows. In this paper, we propose AARC, an innovative, automated framework that decouples CPU and memory resources to provide more flexible and efficient provisioning for serverless workloads. AARC is composed of two key components: Graph-Centric Scheduler, which identifies critical paths in workflows, and Priority Configurator, which applies priority scheduling techniques to optimize resource allocation. Our experimental evaluation demonstrates that AARC achieves substantial improvements over state-of-the-art methods, with total search time reductions of $85.8 \%$ and $89.6 \%$, and cost savings of $49.6 \%$ and $61.7 \%$, respectively, while maintaining SLO compliance.
Lingxiao Jin, Zinuo Cai, Ruhui Ma
DAC5
2025 Fine-Grained Global Modeling Learning for Personalized Federated Sequential Recommender
abstract
Personalized sequential recommender has become a key task in the consumer electronics domain. Existing methods for personalized sequential recommenders primarily focus on modeling user behavior and have achieved satisfactory recommender results. However, the inherent quadratic computational complexity in most existing methods often causes typically results in inefficiencies, impeding real-time suggestions. Additionally, these methods cannot be fine-tuned to the personalized needs of users across different scenarios. To address these challenges, we propose the Fine-Grained Global Modeling Learning for Personalized Federated Sequential Recommender (FedSR). Specifically, we design the Associative Mamba Block to model user profiles from a global perspective and enhance prediction efficiency. Furthermore, we introduce the Variable Response Mechanism to fine-tune parameters according to the personalized needs of users. Additionally, we design the Dynamic Magnitude Loss to retain more local personalized information during training. Extensive experiments on three real-world datasets confirms that the proposed FedSR surpasses current approaches in terms of both performance and efficiency, up to 9.48% improvement.
Yicheng Di, Hongjian Shi, Ruhui Ma, Yuan Liu 0021
ICASSP3
2025 FedDA: Resource-adaptive federated learning with dual-alignment aggregation optimization for heterogeneous edge devices
Shaohua Cao, Huixin Wu, Xiwen Wu, Ruhui Ma, Danxin Wang, Zhu Han 0001, Weishan Zhang
Future Gener. Comput. Syst.4
2025 Context-based emotion recognition: A survey
Rizwan Abbas, Bingnan Ni, Ruhui Ma, Yehao Lu, Xi Li 0001
Neurocomputing3
2025 NSAA: A Network Slice Access Authentication and Service Authorization Scheme for Integrated Satellite-Terrestrial Network
abstract
Introducing slicing into integrated satellite-terrestrial networks enables the flexible deployment of network resources and being adaptable for more new applications. However, the heterogeneity of integrated satellite-terrestrial network poses challenges to network resource access control. To ensure users can securely and efficiently access service across multiple management domains, we propose a network slice access authentication and service authorization scheme based on a sharding permissioned blockchain. Slice tenants and wireless network operators with management control act as consortium blockchain nodes, which are divided into shards, and the blockchain is maintained in parallel by multiple shards. First, an efficient public ledger is constructed to establish decentralized trust and manage user identity and service authorization information. Second, utilizing the trapdoor collision resistance of the chameleon hash, users can fully self-select their secret key and generate authentication credentials to register on the blockchain without key escrow problem. When users move into a new network domain, the mutual authentication between users and the visited network can be quickly completed. The session key is negotiated based on the Diffie-Hellman ephemeral protocol with perfect forward secrecy. Then, the editable transaction blocks, storing network slice authorization information and slice templates, are linked using chameleon hashes. This allows the access permissions of slice resources to be dynamically adjusted and easily queried by the service-providing wireless network operators. Performance evaluation and security simulations demonstrate the correctness of the scheme, showing that it can achieve secure access to integrated satellite-terrestrial network slice services with low computational and communication overhead.
Yurong Luo, Jin Cao 0001, Ruhui Ma, Ben Niu 0001, Yinghui Zhang 0002, Hui Li 0006
IEEE Internet Things J.4
2025 Federated Learning With Client Clustering Selection and Quality-Aware Model Aggregation (2024)
abstract
The Industrial Internet of Things (IIoT) is revolutionizing industries through device interconnectivity, enabling real-time data collection and transmission for enhanced monitoring, control, and automation. This has led to improvements in predictive maintenance, production optimization, and supply chain management. The integration of machine learning (ML) and deep learning (DL) within IIoT has been accelerated by increased data availability, computational advancements, and algorithmic innovations, with applications ranging from image processing to autonomous vehicles. Foundation models, such as ChatGPT, are becoming prevalent in IIoT for their capabilities in natural language processing and computer vision. However, the growth of IIoT and foundation models presents challenges, including data volume, real-time processing requirements, computational costs, and security vulnerabilities. Federated learning (FL) addresses these issues by allowing distributed model training without raw data transfer, enhancing privacy and security. FL is particularly beneficial for IIoT’s decentralized architecture and real-time decision-making needs. Despite the advantages, FL faces challenges such as data heterogeneity and communication overhead. To overcome these, we propose a federated learning framework, Fed-CCSQMA, which includes client selection and model aggregation modules to mitigate data heterogeneity’s impact. The client selection module uses Principal Component Analysis (PCA) and clustering to select clients with diverse yet representative data, while the model aggregation module assigns weights based on model accuracy to ensure faster global model convergence. We test our proposed framework on FMNIST and CIFAR-10 datasets, with accuracy surpassing baseline methods and faster convergence, demonstrating an improvement in overall generalizability and learning efficiency.
Y. Peng, Chengjie Wang 0001, Hongjian Shi, Ruhui Ma, Haibing Guan, Hanbo Yang
IEEE Internet Things J.4
2025 PFLlib: A Beginner-Friendly and Comprehensive Personalized Federated Learning Library and Benchmark
abstract
Amid the ongoing advancements in Federated Learning (FL), a machine learning paradigm that allows collaborative learning with data privacy protection, personalized FL (pFL) has gained significant prominence as a research direction within the FL domain. Whereas traditional FL (tFL) focuses on jointly learning a global model, pFL aims to balance each client's global and personalized goals in FL settings. To foster the pFL research community, we started and built PFLlib, a comprehensive pFL library with an integrated benchmark platform. In PFLlib, we implemented 37 state-of-the-art FL algorithms (8 tFL algorithms and 29 pFL algorithms) and provided various evaluation environments with three statistically heterogeneous scenarios and 24 datasets. At present, PFLlib has gained more than 1600 stars and 300 forks on GitHub.
Jianqing Zhang, Yang Liu 0165, Yang Hua 0001, Hao Wang 0022, Tao Song 0003, Zhengui Xue, Ruhui Ma, Jian Cao 0001
J. Mach. Learn. Res.7
2025 HWDSQP: A Historical Weighted and Dynamic Scheduling Quantum Protocol to Enhance Communication Reliability
abstract
Quantum computing holds the promise of solving problems difficult for classical computers. However, we are still in the era of Noisy Intermediate-Scale Quantum (NISQ) computers, necessary to establish effective distributed quantum communication protocols to distribute complex quantum computing tasks across different quantum computers for execution. Significant progress has been made in quantum communication technology, particularly in quantum path creation and resource scheduling. The establishment of quantum paths relies on quantum entanglement and quantum relay technologies, achieving long-distance, high-fidelity quantum state transmission through entanglement swapping between multiple relays. However, resources in quantum communication networks are limited and expensive, making efficient resource scheduling strategies crucial for improving overall network efficiency. To address these issues, we design a network protocol that includes the Historical Weighted Fidelity Routing (HWFR) algorithm and the Dynamic Multi-Priority Quantum Scheduling (DMPQS) algorithm to enhance communication reliability across quantum computers. Both algorithms aim to enhance the reliability of quantum links, optimize resource utilization, and adapt to dynamic changes in the links. The former algorithm dynamically selects the optimal path by considering factors such as link length, noise level, entanglement success rate, and quantum relay resource constraints, ensuring high-fidelity and reliable quantum communication. The latter dynamically adjusts request priorities based on the urgency of quantum service requests and fidelity requirements, optimizing resource utilization. Experimental results show that the proposed protocol performs excellently in terms of an average response time of requests and link utilization, effectively improving the utilization efficiency of network resources and the overall performance of the system.
Rongbo Ma, Zejian Wang, Zinuo Cai, Haochen Xu, Baoheng Zhang, Ruhui Ma, Rajkumar Buyya
IEEE J. Sel. Areas Commun.7
2025 Ephemera: Accelerating I/O-Intensive Serverless Workloads with a Harvested In-memory File System
abstract
Serverless computing has gained popularity for its ability to shift the burden of server management from developers to cloud providers, which allows providers to exercise greater control over resource management, optimizing configurations to enhance efficiency and performance. The diversity of serverless computing tasks, from short-lived, event-driven tasks to more complex workloads, highlights the growing importance of efficient file I/O performance for I/O-intensive workloads, yet effectively handling ephemeral storage for I/O-intensive tasks remains a challenge. Traditional file system approaches often introduce substantial latency and fail to fully leverage available memory resources within the execution environment, limiting performance and efficiency. Our work stems from the observation of the under-utilization of memory resources in serverless computing platforms and the potential efficiency improvement of I/O operations using an in-memory file system. Based on this observation, we propose Ephemera , a system designed to enhance ephemeral storage efficiency and memory utilization. Ephemera satisfies three design goals: transparent memory I/O integration , heterogeneous tasks resource synergy , and harmonized cluster workload orchestration . Ephemera integrates three components: the Runtime Daemon, responsible for managing a container’s in-memory file system; the Tenant Manager, facilitating memory configuration sharing across containers; and the Cluster Controller, optimizing workload balancing. Our experiments demonstrate that Ephemera significantly improves performance for I/O-intensive tasks compared to traditional file systems. Specifically, Ephemera decreases I/O processing time by 50% on average and reduces latency by up to 95.73% in certain scenarios with negligible overhead.
Lingxiao Jin, Zinuo Cai, Haoxin Wang 0005, Zongpu Zhang, Ruhui Ma, Haibing Guan, Yuan Liu 0021, Rajkumar Buyya
ACM Trans. Archit. Code Optim.5
2025 MemoriaNova: Optimizing Memory-Aware Model Inference for Edge Computing
abstract
In recent years, deploying deep learning models on edge devices has become pervasive, driven by the increasing demand for intelligent edge computing solutions across various industries. From industrial automation to intelligent surveillance and healthcare, edge devices are being leveraged for real-time analytics and decision-making. Existing methods face two challenges when deploying machine learning models on edge devices. The first challenge is handling the execution order of operators with a simple strategy, which can lead to a potential waste of memory resources when dealing with directed acyclic graph structure models. The second challenge is that they usually process operators of a model one by one to optimize the inference latency, which may lead to the optimization problem getting trapped in local optima. We present MemoriaNova, comprising BTSearch and GenEFlow, to solve these two problems. BTSearch is a graph state backtracking algorithm with efficient pruning and hashing strategies designed to minimize memory overhead during inference and enlarge latency optimization search space. GenEFlow, based on genetic algorithms (GA), integrates latency modeling, and memory constraints to optimize distributed inference latency. This innovative approach considers a comprehensive search space for model partitioning, ensuring robust and adaptable solutions. We implement BTSearch and GenEFlow and test them on 11 deep-learning models with different structures and scales. The results show that BTSearch can reach 12% memory optimization compared with the widely used random execution strategy. At the same time, GenEFlow reduces inference latency by 33.9% in distributed systems with four-edge devices.
Renjun Zhang, Tianming Zhang, Zinuo Cai, Dongmei Li 0008, Ruhui Ma, Rajkumar Buyya
ACM Trans. Archit. Code Optim.5
2025 FasDL: An Efficient Serverless-Based Training Architecture With Communication Optimization and Resource Configuration
abstract
Deploying distributed training workloads of deep learning models atop serverless architecture alleviates the burden of managing servers from deep learning practitioners. However, when supporting deep model training, the current serverless architecture faces the challenges of inefficient communication patterns and rigid resource configuration that incur subpar and unpredictable training performance. In this paper, we proposeFasDL, an efficient serverless-based deep learning training architecture to solve these two challenges.FasDLadopts a novel training frameworkK-REDUCEto release the communication overhead and accelerate the training. Additionally, FasDL builds a lightweight mathematical model forK-REDUCEtraining, offering predictable performance and supporting subsequent resource configuration. It achieves the optimal resource configuration by formulating an optimization problem related to system-level and application-level parameters and solving it with a pruning-based heuristic search algorithm. Extensive experiments on AWS Lambda verify a prediction accuracy over 94% and demonstrate performance and cost advantages over the state-of-art architecture LambdaML by up to 16.8% and 28.3% respectively.
Xinglei Chen, Zinuo Cai, Hanwen Zhang 0019, Ruhui Ma, Rajkumar Buyya
IEEE Trans. Computers4
2025 Efficient Vehicle Secure Scheduling and Access Authentication Scheme for 5G-Integrated Emergency Rescue Scenario
abstract
With urban population density on the rise, emergency incidents are increasing in both frequency and complexity, placing growing pressure on existing rescue systems. Meanwhile, issues such as slow response times, inadequate coordination mechanisms, and inefficient information exchange further exacerbate the challenges faced by these systems. The integration of Vehicle-to-Everything (V2X) communication and 5G technology offers unprecedented capabilities, such as ultra-low latency and high data throughput, which are critical for real-time coordination and decision-making in emergency rescue scenarios. To establish secure and efficient vehicle communication in 5G and V2X-enabled emergency rescue scenarios, we propose an efficient vehicle secure scheduling and access authentication scheme based on certificateless cryptography and multireceiver signcryption. In this scheme, the command and control center can securely dispatch rescue fleets based on disaster conditions. By enabling mutual authentication and key agreement between rescue vehicles and roadside units, the scheme ensures the reliable and swift exchange of rescue information and instructions. In addition, to address unexpected situations such as traffic congestion, we design a route-switching mechanism. Furthermore, in order to mitigate potential malicious behavior, a vehicle legitimacy revocation mechanism is implemented to ensure the normal operation of the system. The security of the scheme is verified through formal analysis and informal analysis. Performance analysis demonstrates that the scheme offers significant advantages over existing ones in terms of signaling overhead, communication overhead, computational overhead, and energy efficiency.
Jin Cao 0001, Yiqing Xiong, Ruhui Ma, Yinghui Zhang 0002, Ben Niu 0001, Peijie Yin, Hui Li 0006
IEEE Trans. Intell. Transp. Syst.5
2025 Ultra-High-Speed Terminal Secure Access and Intragroup Authentication Scheme in Satellite Networks
abstract
Ultra-High-Speed Terminals (UHSTs) can transport multiple Load Equipment (LEs) to precise locations, such as in a space station resupply mission scenario. In these scenarios, UHSTs need to access ground networks via satellite networks. However, since the connections between UHSTs and ground networks are established through insecure air interface channels, they are susceptible to attacks such as eavesdropping, impersonation, and other. Furthermore, owing to the high-speed mobility of UHSTs, they may not be able to connect successfully to the ground network through a single access point, which is possible for regular terminals. Additionally, UHST may also need to communicate with multiple LEs, which is also connected via insecure air interface channels. Therefore, this paper proposes a secure access and intra-group authentication scheme for UHSTs in satellite network scenarios. In the proposed scheme, based on pre-shared keys and trajectory prediction mechanisms, the UHST can successfully access the ground network through multiple access points and complete key establishment with the access points along its trajectory in advance. Using Shamir’s (t, n) Secret Sharing mechanism, UHST and multiple LEs can share a group key, ensuring secure intra-group data communication. Additionally, when one LE detaches from the UHST, the UHST can authorize the LE to access the ground network. Security and efficiency analysis shows that the proposed scheme achieves comprehensive security features with low overhead.
Yukun Zhu, Ruhui Ma, Jin Cao 0001, Hui Li 0006, Xiaosong Zhang 0001
IEEE Trans. Netw. Serv. Manag.2
2025 Federated Recommender System Based on Diffusion Augmentation and Guided Denoising
abstract
Sequential recommender systems often struggle with accurate personalized recommendations due to data sparsity issues. Existing works use variational autoencoders and generative adversarial network methods to enrich sparse data. However, they often overlook diversity in the latent data distribution, hindering the model’s generative capacity. This characteristic of generative methods can introduce additional noise in many cases. Moreover, retaining personalized user preferences through the generation process remains a challenge. This work introduces DGFedRS, a Federated Recommender System Based on Diffusion Augmentation and Guided Denoising, designed to capture the diversity in the latent data distribution while preserving user-specific information and suppressing noise. In particular, we pre-train the diffusion model using the recommender dataset and use a diffusion augmentation strategy to generate interaction sequences, expanding the sparse user-item interactions in the discrete space. To preserve user-specific preferences in the generated interactions, we employ a guided denoising strategy to guide the generation process during reverse diffusion. Subsequently, we design a noise control strategy to reduce the damage to personalized information during the diffusion process. Additionally, a stepwise scheduling strategy is devised to input generated data into the sequential recommender model based on their challenge levels. The success of the DGFedRS approach is demonstrated by thorough experiments conduct on three real-world datasets.
Yicheng Di, Hongjian Shi, Ruhui Ma, Yuan Liu 0021
ACM Trans. Inf. Syst.4
2025 Efficient Federated Recommender System with Adaptive Model Pruning and Momentum-based Batch Adjustment
abstract
With the development of 5G communication and smart devices, the prosperity of online content has boosted Recommender System (RS) research. Due to the data scarcity problem, researchers employ knowledge transfer techniques to improve the accuracy of RS. Data sharing or data augmentation are promising methods for such a problem, but the data suffers from privacy leakage during sharing. Thus, federated learning has been adopted to collaboratively train recommender models while preserving data privacy. Federated Recommender System (FRS) combines federated learning and RS to provide distributed recommendation services to the users. However, the existing FRS suffers from massive communication and system heterogeneity, where the necessity of model transmission and the diversity of the clients bring significant communication overhead to the system. In this article, we propose Efficient Federated Recommender System with Adaptive Model Pruning and Momentum-based Batch Adjustment ( \(eFRSA^{2}\) ) to reduce the communication overhead of FRS. \(eFRSA^{2}\) contains two modules. Adaptive Model Pruning utilizes magnitude pruning to reduce the communication volume and adaptively modifies the compression ratios of different clients to maintain the model accuracy. Momentum-based Batch Adjustment adjusts the local training batch number by a similar method of gradient descent with momentum to align the local computation time of the clients and reduce the communication overhead. The experimental results demonstrate that \(eFRSA^{2}\) can reduce up to 90% communication volume and mitigate the system heterogeneity by over 75%, demonstrating the priority of \(eFRSA^{2}\) in training efficiency. Source code can be found at https://github.com/shhjwu5/eFRSA2 .
Hongjian Shi, Yicheng Di, Xinyu Ruan, Mingrui Liao, Qian Zhang 0001, Ruhui Ma, Haibing Guan
Trans. Recomm. Syst.6
2025 SMore: Enhancing GPU Utilization in Deep Learning Clusters by Serverless-Based Co-Location Scheduling
Junhan Liu, Zinuo Cai, Yumou Liu, Hao Li 0142, Zongpu Zhang, Ruhui Ma, Rajkumar Buyya
IEEE Trans. Parallel Distributed Syst.6
2024 SKYMASK: Attack-Agnostic Robust Federated Learning with Fine-Grained Learnable Masks
Peishen Yan, Hao Wang 0022, Tao Song 0003, Yang Hua 0001, Ruhui Ma, Ningxin Hu, Mohammad R. Haghighat, Haibing Guan
ECCV (19)5
2024 TKGT: Redefinition and A New Way of Text-to-Table Tasks Based on Real World Demands and Knowledge Graphs Augmented LLMs
abstract
The task of text-to-table receives widespread attention, yet its importance and difficulty are underestimated.Existing works use simple datasets similar to table-to-text tasks and employ methods that ignore domain structures.As a bridge between raw text and statistical analysis, the text-to-table task often deals with complex semi-structured texts that refer to specific domain topics in the real world with entities and events, especially from those of social sciences.In this paper, we analyze the limitations of benchmark datasets and methods used in the text-to-table literature and redefine the textto-table task to improve its compatibility with long text-processing tasks.Based on this redefinition, we propose a new dataset called CPL (Chinese Private Lending), which consists of judgments from China and is derived from a real-world legal academic project.We further propose TKGT (Text-KG-Table), a two stages domain-aware pipeline, which firstly generates domain knowledge graphs (KGs) classes semiautomatically from raw text with the mixed information extraction (Mixed-IE) method, then adopts the hybrid retrieval augmented generation (Hybird-RAG) method to transform it to tables for downstream needs under the guidance of KGs classes.Experiment results show that TKGT achieves state-of-the-art (SOTA) performance on both traditional datasets and the CPL.Our data and main code are available at https://github.com/jiangpw41/TKGT.
Peiwen Jiang, Xinbo Lin, Ruhui Ma, Yvonne Jie Chen, Jinhua Cheng
EMNLP4
2024 Hermes: Memory-Efficient Pipeline Inference for Large Models on Edge Devices
abstract
The application of Transformer-based large models has achieved numerous success in recent years. However, the exponential growth in the parameters of large models introduces formidable memory challenge for edge deployment. Prior works to address this challenge mainly focus on optimizing the model structure and adopting memory swapping methods. However, the former reduces the inference accuracy, and the latter raises the inference latency. This paper introduces PIPELoAD, a novel memory-efficient pipeline execution mechanism. It reduces memory usage by incorporating dynamic memory management and minimizes inference latency by employing parallel model loading. Based on PIPELoAD mechanism, we present Hermes, a framework optimized for large model inference on edge devices. We evaluate Hermes on Transformer-based models of different sizes. Our experiments illustrate that Hermes achieves up to 4.24 x increase in inference speed and 86.7% lower memory consumption than the state-of-the-art pipeline mechanism for BERT and ViT models, 2.58 x increase in inference speed and 90.3% lower memory consumption for GPT-style models.
Xueyuan Han, Zinuo Cai, Yichu Zhang, Chongxin Fan, Junhan Liu, Ruhui Ma, Rajkumar Buyya
ICCD6
2024 SMSS: Stateful Model Serving in Metaverse With Serverless Computing and GPU Sharing
abstract
With the rapid development of information technology, the concept of the Metaverse has swept the world and set off a new wave of the industrial revolution. The construction of living and manufacturing scenes based on the Metaverse requires the joint participation of scientists and engineers from various fields where “human” is at the core. In the Metaverse, predicting human behavior and response based on the deep learning model is meaningful because the prediction results can provide more satisfactory services for participants. Therefore, how to deploy a multi-stage machine learning reasoning model has become the bottleneck to improving the development level of Metaverse. Thanks to its scalability and pay-as-you-go billing model, the emerging serverless computing can effectively cope with the workload of machine learning inference. However, the statelessness of serverless computing and the lack of good GPU resource-sharing support make it difficult to deploy the machine learning model directly on the serverless computing platform to play its advantages. Therefore, we propose SMSS, a stateful model inference service, which is deployed on a serverless computing platform that supports GPU sharing. Since the serverless computing platform does not support stateful workflow execution, SMSS adopts log-based workflow runtime support. We also design a mechanism of two-layer GPU sharing to fully explore the potential of inter-model and intra-model GPU sharing. We evaluate the effectiveness of SMSS with real workloads. Our experimental results show that log-based stateful workflow operation support can ensure the stateful execution of tasks with low overhead but facilitate error location and recovery. Two-layer GPU Sharing can reduce the cold start time of inference tasks to two orders of magnitude at most.
Zinuo Cai, Ruhui Ma, Haibing Guan
IEEE J. Sel. Areas Commun.3
2024 MCAS-GP: Deep Learning-Empowered Middle Cerebral Artery Segmentation and Gate Proposition
abstract
With the fast development of AI technologies, deep learning is widely applied for biomedical data analytics and digital healthcare. However, there remain gaps between AI-aided diagnosis and real-world healthcare demands. For example, hemodynamic parameters of the middle cerebral artery (MCA) have significant clinical value for diagnosing adverse perinatal results. Nevertheless, the current measurement procedure is tedious for sonographers. To reduce the workload of sonographers, we propose MCAS-GP, a deep learning-empowered framework that tackles the Middle Cerebral Artery Segmentation and Gate Proposition. MCAS-GP can automatically segment the region of the MCA and detect the corresponding position of the gate in the procedure of fetal MCA Doppler assessment. In MCAS-GP, a novel learnable atrous spatial pyramid pooling (LASPP) module is designed to adaptively learn multi-scale features. We also propose a novel evaluation metric, Affiliation Index, for measuring the effectiveness of the position of the output gate. To evaluate our proposed MCAS-GP, we build a large-scale MCA dataset, collaborating with the International Peace Maternity and Child Health Hospital of China welfare institute (IPMCH). Extensive experiments on the MCA dataset and two other public surgical datasets demonstrate that MCAS-GP can achieve considerable performance improvement in both accuracy and inference time.
Rui Zhang 0087, Shuo Wang 0008, Ruhui Ma, Yang Hua 0001, Tao Song 0003, Yunyun Cao, Haibing Guan
IEEE Trans. Comput. Biol. Bioinform.3
2024 SLoB: Suboptimal Load Balancing Scheduling in Local Heterogeneous GPU Clusters for Large Language Model Inference
abstract
Large language models (LLMs) are becoming powerful engines for social productivity in the manufacturing lifecycle. Existing application-level LLMs inference services focus on large datacenter and small edge intelligence (EI) scenarios, adopting iteration-level batch schedulers to solve resource utilization and inference speed problems. However, these services are incompatible with the scene of medium-sized local heterogeneous graphics processing unit (GPU) clusters with specific patterns, whose scale is between the two aforementioned scenarios. This type of scene proposes tradeoff problems for inference resource and speed, as well as user satisfaction problems for the semisparse frequency of queries with streaming responses. We propose suboptimal load balancing (SLoB), a distributed LLMs inference service scheduler in medium-sized local heterogeneous GPU clusters. SLoB leverages a multilevel adapter to accommodate LLMs usage patterns of scenes and balance resource utilization with inference efficiency. For semisparse problems, it adopts a mixed-priority pipeline scheduler with the least-padding principle to improve users’ satisfaction, a metric considering the weights of different tokens in streaming responses. Based on the system prototype, our experiments under simulated workloads demonstrate that SLoB gains a maximum improvement of 29.4$\times$under the satisfaction metric compared with the traditional run-to-completion scheduling solution while improving by up to 3.0$\times$compared with the state-of-the-art (SOTA) solution Orca.
Peiwen Jiang, Haoxin Wang 0005, Zinuo Cai, Lintao Gao, Weishan Zhang, Ruhui Ma, Xiaokang Zhou
IEEE Trans. Comput. Soc. Syst.6
2024 SPSC: Stream Processing Framework Atop Serverless Computing for Industrial Big Data
abstract
With the advance of smart manufacturing and information technologies, the volume of data to process is increasing accordingly. Current solutions for big data processing resort to distributed stream processing systems, such as Apache Flink and Spark. However, such frameworks face challenges of resource underutilization and high latency in big data application scenarios. In this article, we propose SPSC, a serverless-based stream computing framework where events are discretized into the atomic stream and stateless Lambda functions are taken as context-irrelevant operators, achieving task parallelism and inherent data parallelism in processing. Also, we implement a prototype of the framework on Amazon Web service (AWS) using AWS Lambda, AWS simple queue service, and AWS DynamoDB. The evaluation shows that compared with Alibaba's real-time computing Flink version, SPSC outperforms by 10.12% when the overhead is close.
Zinuo Cai, Xinglei Chen, Ruhui Ma, Haibing Guan, Rajkumar Buyya
IEEE Trans. Cybern.4
2024 Siren$^+$+: Robust Federated Learning With Proactive Alarming and Differential Privacy
abstract
Federated learning (FL), an emerging machine learning paradigm that trains a global model across distributed clients without violating data privacy, has recently attracted significant attention. However, FL?s distributed nature and iterative training extensively increase the attacking surface for Byzantine and inference attacks. Existing FL defense methods can hardly protect FL from both Byzantine and inference attacks due to their fundamental conflicts. The noise injected to defend against inference attacks interferes with model weights and training data, obscuring model analysis that Byzantine-robust methods utilize to detect attacks. Besides, the practicability of existing Byzantine-robust methods is limited since they heavily rely on model analysis. In this paper, we present SIREN+, a new robust FL system that defends against a wide spectrum of Byzantine attacks and inference attacks by jointly utilizing a proactive alarming mechanism and local differential privacy (LDP). The proactive alarming mechanism orchestrates clients and the FL server to collaboratively detect attacks using distributed alarms, which is free from the noise interference injected by LDP. Compared with the state-of-the-art defense methods, SIREN+can protect FL from Byzantine and inference attacks from a higher proportion of malicious clients in the system while keeping the global model performing normally. Extensive experiments with diverse settings and attacks on real-world datasets show that SIREN+outperforms existing defense methods when attacked by Byzantine and inference attacks.
Hanxi Guo, Hao Wang 0022, Tao Song 0003, Yang Hua 0001, Ruhui Ma, Xiulang Jin, Zhengui Xue, Haibing Guan
IEEE Trans. Dependable Secur. Comput.5
2024 RIDIC: Real-Time Intelligent Transportation System With Dispersed Computing
abstract
Modern transportation big data features high Volume, Velocity, and Variety, making it more and more challenging to develop an intelligent transportation system for data analysis. Current transportation systems resort to cloud computing to deploy their applications but face two bottlenecks, lack of real-time processing and under-utilization of intelligent roadside devices. We observe that dispersed computing—an emerging paradigm of cloud computing—well fits the requirements of modern transportation systems. It provides real-time response by alleviating data transmission between data sources and cloud servers and fully utilizes smart devices by exploring their computing capacity. Therefore, we design RIDIC, an intelligent transportation system with dispersed computing to provide a real-time response when processing transportation big data. RIDIC abstracts all the heterogeneous smart roadside devices as actors, and its workflow consists of three stages, Actor Registration, Resource Application and Task Execution. We conduct experiments on two real-life traffic scenarios—road vehicle detection and traffic signal recognition—and the results show that RIDIC can utilize edge devices to process transportation big data faster while reducing the demand for device computing resources.
Zinuo Cai, Quanmin Xie, Ruhui Ma, Haibing Guan
IEEE Trans. Intell. Transp. Syst.5
2024 A UAV-Assisted UE Access Authentication Scheme for 5G/6G Network
abstract
Unmanned Aircraft Vehicles (UAVs) equipped with base stations can assist ground User Equipments (UEs) in accessing the 5G/6G network. Due to the UAV’s high autonomy, easy configuration, and strong dynamic deployment capabilities, UAV-assisted ground UEs to access the 5G/6G network can effectively expand the communication network coverage. However, some vulnerabilities exist, such as eavesdropping attack, impersonation attack, etc. In addition, the 3rd Generation Partnership Project (3GPP) committee has proposed that the UAV can employ the primary authentication mechanism (i.e., 5G-AKA) to connect to the network. Nevertheless, the primary authentication mechanism 5G-AKA has some security problems. In this paper, we first improve the existing 5G-AKA, which resists quantum attack and traceability attack and consumes moderate signaling overhead and short running time. Then, based on the improved 5G-AKA protocol, we propose a UAV-assisted UE access authentication scheme for the 5G/6G network. In the proposed scheme, the UAV can perform the service access authentication process to access the 5G/6G core network and then execute the UAV-assisted UE access authentication process to assist UE in obtaining network services. Additionally, the ground UE can perform a fast and secure handover process with the target UAV to ensure continuous network services. The automation verification tool Tamarin is employed to verify the security of the proposed scheme. Additionally, we implement the improved 5G-AKA protocol and the existing 5G-AKA protocol on Field Programmable Gate Array (FPGA) to test their running time. The security and performance evaluation results show that the proposed scheme provides robust security with moderate efficiency.
Ruhui Ma, Jin Cao 0001, Shiyang He, Yinghui Zhang 0002, Ben Niu 0001, Hui Li 0006
IEEE Trans. Netw. Serv. Manag.1
2024 RFL-LSU: A Robust Federated Learning Approach with Localized Stepwise Updates
abstract
Distributed intelligence enables the widespread deployment of AI technology, greatly promoting the development of AI. Federated learning is a widely used distributed intelligence technology that allows iterative optimization of global model while protecting user data privacy. Currently, federated learning faces some security threats, as its open architecture provides attackers with opportunities to disrupt the learning process by submitting malicious updates or inserting backdoors. In this article, we propose a robust federated learning method to defend against potential malicious attacks. Specifically, we enhance the algorithm’s performance and stability by implementing localized stepwise updates on the client side and element-wise anomaly detection on the server side. We conducted experiments in a more realistic non-i.i.d. scenario and compared the results with other typical federated learning methods. Our results demonstrate that our approach exhibits strong robustness under non-i.i.d. data distributions, outperforming other methods in terms of test accuracy and resilience to attacks.
Shuming Fan, Hongjian Shi, Chenpei Wang, Ruhui Ma
ACM Trans. Internet Techn.4
2024 Robust Searching-Based Gradient Collaborative Management in Intelligent Transportation System
abstract
With the rapid development of big data and the Internet of Things (IoT), traffic data from an Intelligent Transportation System (ITS) is becoming more and more accessible. To understand and simulate the traffic patterns from the traffic data, Multimedia Cognitive Computing (MCC) is an efficient and practical approach. Distributed Machine Learning (DML) has been the trend to provide sufficient computing resources and efficiency for MCC tasks to handle massive data and complex models. DML can speed up computation with those computing resources but introduces communication overhead. Gradient collaborative management or gradient aggregation in DML for MCC tasks is a critical task. An efficient managing algorithm of the communication schedules for gradient aggregation in ITS can improve the performance of MCC tasks. However, existing communication schedules typically rely on specific physical connection matrices, which have low robustness when a malfunction occurs. In this article, we propose Robust Searching-based Gradient Collaborative Management (RSGCM) in Intelligent Transportation System, a practical ring-based gradient managing algorithm for communication schedules across devices to deal with ITS malfunction. RSGCM provides solutions of communication schedules to various kinds of connection matrices with an acceptable amount of training time. Our experimental results have shown that RSGCM can deal with more varieties of connection matrices than existing state-of-the-art communication schedules. RSGCM also increases the robustness of ITS since it can restore the system’s functionality in an acceptable time when device or connection breakdown happens.
Hongjian Shi, Hao Wang 0022, Ruhui Ma, Yang Hua 0001, Tao Song 0003, Honghao Gao, Haibing Guan
ACM Trans. Multim. Comput. Commun. Appl.3
2024 Sustainable Serverless Computing With Cold-Start Optimization and Automatic Workflow Resource Scheduling
abstract
In recent years, serverless computing has garnered significant attention owing to its high scalability, pay-as-you-go billing model, and efficient resource management provided by cloud service providers. Optimal resource scheduling of serverless computing has become imperative to reduce energy consumption and enable sustainable computing. However, existing serverless platforms encounter two significant challenges: the cold-start problem of containers and the absence of an effective resource allocation strategy for serverless workflows. Existing pre-warm strategies are associated with high computational overhead, while current resource scheduling techniques inadequately account for the intricate structure of serverless workflows. To address these challenges, we present SSC, a pre-warming and automatic resource allocation framework designed explicitly for serverless workflows. We introduce an innovative gradient-based algorithm for pre-warming containers, significantly reducing cold start hit rates. Moreover, leveraging a critical path and priority queue-based algorithm, SSC enables efficient allocation of resources for serverless workflows. In our experimental evaluation, SSC reduces the cold start hit rate by nearly$50\%$and achieves substantial cost savings of approximately$30\%$.
Shanxing Pan, Zinuo Cai, Dongmei Li 0008, Ruhui Ma, Haibing Guan
IEEE Trans. Sustain. Comput.5
2024 Edge Caching Placement Strategy based on Evolutionary Game for Conversational Information Seeking in Edge Cloud Computing
abstract
In Internet applications, network conversation is the primary communication between the user and server. The server needs to efficiently and quickly return the corresponding service according to the conversation sent by the user to improve the users’ Quality of Service. Thus, Conversation Information Seeking (CIS) research has become a hot topic today. In Cloud Computing (CC), a central service mode, the conversation is transmitted between the user and the remote cloud over a long distance. With the explosive growth of Internet applications, network congestion, long-distance communication, and single point of failure have brought new challenges to the centralized service mode. People put forward Edge Cloud Computing (ECC) to meet the new challenges of the centralized service mode of CC. As a distributed service mode, ECC is an extension of CC. By migrating services from the remote cloud to the network edge closer to users, ECC can solve the above challenges in CC well. In ECC, people solve the problem of CIS through edge caching. The current research focuses on designing the edge cache strategy to achieve more predictable caching. In this article, we propose an edge cache placement method Evolutionary Game based Caching Placement Strategy (EG-CPS). This method consists of three modules: the user preference prediction module, the content popularity calculation module, and the cache placement decision module. To maximize the predictability of the cache strategy, we are committed to optimizing the cache hit rate and service latency. The simulation experiment compares the proposed strategy with several other cache strategies. The experimental results illustrate that EG-CPS can reduce up to 2.4% of the original average content request latency, increase the average direct cache hit rate by 1.7%, and increase the average edge cache hit rate by 3.3%.
Hongjian Shi, Ruhui Ma, Rui Zhang 0087, Haibing Guan
ACM Trans. Web3
2023 FedALA: Adaptive Local Aggregation for Personalized Federated Learning
abstract
A key challenge in federated learning (FL) is the statistical heterogeneity that impairs the generalization of the global model on each client. To address this, we propose a method Federated learning with Adaptive Local Aggregation (FedALA) by capturing the desired information in the global model for client models in personalized FL. The key component of FedALA is an Adaptive Local Aggregation (ALA) module, which can adaptively aggregate the downloaded global model and local model towards the local objective on each client to initialize the local model before training in each iteration. To evaluate the effectiveness of FedALA, we conduct extensive experiments with five benchmark datasets in computer vision and natural language processing domains. FedALA outperforms eleven state-of-the-art baselines by up to 3.27% in test accuracy. Furthermore, we also apply ALA module to other federated learning methods and achieve up to 24.19% improvement in test accuracy. Code is available at https://github.com/TsingZ0/FedALA.
Jianqing Zhang, Yang Hua 0001, Hao Wang 0022, Tao Song 0003, Zhengui Xue, Ruhui Ma, Haibing Guan
AAAI6
2023 Information Bound and Its Applications in Bayesian Neural Networks
abstract
Bayesian neural networks have drawn extensive interest because of their distinctive probabilistic representation framework. However, despite its recent success, little work focuses on the information-theoretic understanding of Bayesian neural networks. In this paper, we propose Information Bound as a metric of the amount of information in Bayesian neural networks. Different from mutual information on deterministic neural networks where modification of network structure or specific input data is usually necessary, Information Bound can be easily estimated on current Bayesian neural networks without any modification of network structures or training processes. By observing the trend of Information Bound during training, we demonstrate the existence of the “critical period” in Bayesian neural networks. Besides, we show that the Information Bound can be used to judge the confidence of the model prediction and to detect out-of-distribution datasets. Based on these observations of model interpretation, we propose Information Bound regularization and Information Bound variance regularization methods. The Information Bound regularization encourages models to learn the minimum necessary information and improves the model generality and robustness. The Information Bound variance regularization encourages models to learn more about complex samples with low Information Bound. Extensive experiments on KMNIST, Fashion-MNIST, CIFAR-10, and CIFAR-100 verify the effectiveness of the proposed regularization methods.
Jiaru Zhang, Yang Hua 0001, Tao Song 0003, Hao Wang 0022, Zhengui Xue, Ruhui Ma, Haibing Guan
ECAI6
2023 Online Residual-Based Key Frame Sampling with Self-Coach Mechanism and Adaptive Multi-Level Feature Fusion
abstract
Key frame sampling is a common component in video tasks. Putting more effort into key frames, rather than processing all frames equally, can significantly reduce computational costs and improve processing efficiency. This paper presents ORSampler, an adaptive Online Residual-based key frame Sampler. ORSampler relies on feature residuals to sample key frames and decouples from subsequent video tasks. To facilitate ORSampler, a self-coached mechanism is designed to speed up learning, and an adaptive multi-level feature fusion is proposed to fit the diversity of subsequent video tasks. OR-Sampler has a fast inference speed and can work online. Extensive experiments on two typical video tasks verify the effectiveness and generality of our proposed ORSampler.
Rui Zhang 0087, Yang Hua 0001, Tao Song 0003, Zhengui Xue, Ruhui Ma, Haibing Guan
ICASSP5
2023 GPFL: Simultaneously Learning Global and Personalized Feature Information for Personalized Federated Learning
abstract
Federated Learning (FL) is popular for its privacy-preserving and collaborative learning capabilities. Recently, personalized FL (pFL) has received attention for its ability to address statistical heterogeneity and achieve personalization in FL. However, from the perspective of feature extraction, most existing pFL methods only focus on extracting global or personalized feature information during local training, which fails to meet the collaborative learning and personalization goals of pFL. To address this, we propose a new pFL method, named GPFL, to simultaneously learn global and personalized feature information on each client. We conduct extensive experiments on six datasets in three statistically heterogeneous settings and show the superiority of GPFL over ten state-of-the-art methods regarding effectiveness, scalability, fairness, stability, and privacy. Besides, GPFL mitigates overfitting and outperforms the baselines by up to 8.99% in accuracy.
Jianqing Zhang, Yang Hua 0001, Hao Wang 0022, Tao Song 0003, Zhengui Xue, Ruhui Ma, Jian Cao 0001, Haibing Guan
ICCV6
2023 Adversarial Example Does Good: Preventing Painting Imitation from Diffusion Models via Adversarial Examples
abstract
Recently, Diffusion Models (DMs) boost a wave in AI for Art yet raise new copyright concerns, where infringers benefit from using unauthorized paintings to train DMs and generate novel paintings in a similar style. To address these emerging copyright violations, in this paper, we are the first to explore and propose to utilize adversarial examples for DMs to protect human-created artworks. Specifically, we first build a theoretical framework to define and evaluate the adversarial examples for DMs. Then, based on this framework, we design a novel algorithm to generate these adversarial examples, named AdvDM, which exploits a Monte-Carlo estimation of adversarial examples for DMs by optimizing upon different latent variables sampled from the reverse process of DMs. Extensive experiments show that the generated adversarial examples can effectively hinder DMs from extracting their features. Therefore, our method can be a powerful tool for human artists to protect their copyright against infringers equipped with DM-based AI-for-Art applications. The code of our method is available on GitHub: https://github.com/mist-project/mist.git.
Chumeng Liang, Yang Hua 0001, Jiaru Zhang, Yiming Xue, Tao Song 0003, Zhengui Xue, Ruhui Ma, Haibing Guan
ICML8
2023 FedCP: Separating Feature Information for Personalized Federated Learning via Conditional Policy
abstract
Recently, personalized federated learning (pFL) has attracted increasing attention in privacy protection, collaborative learning, and tackling statistical heterogeneity among clients, e.g., hospitals, mobile smartphones, etc. Most existing pFL methods focus on exploiting the global information and personalized information in the client-level model parameters while neglecting that data is the source of these two kinds of information. To address this, we propose the Federated Conditional Policy (FedCP) method, which generates a conditional policy for each sample to separate the global information and personalized information in its features and then processes them by a global head and a personalized head, respectively. FedCP is more fine-grained to consider personalization in a sample-specific manner than existing pFL methods. Extensive experiments in computer vision and natural language processing domains show that FedCP outperforms eleven state-of-the-art methods by up to 6.69%. Furthermore, FedCP maintains its superiority when some clients accidentally drop out, which frequently happens in mobile settings. Our code is public at https://github.com/TsingZ0/FedCP.
Jianqing Zhang, Yang Hua 0001, Hao Wang 0022, Tao Song 0003, Zhengui Xue, Ruhui Ma, Haibing Guan
KDD6
2023 Self-supervised Multi-object Tracking with Cycle-Consistency
Yuanhang Yin, Yang Hua 0001, Tao Song 0003, Ruhui Ma, Haibing Guan
MMM (2)4
2023 Eliminating Domain Bias for Federated Learning in Representation Space
abstract
Recently, federated learning (FL) is popular for its privacy-preserving and collaborative learning abilities. However, under statistically heterogeneous scenarios, we observe that biased data domains on clients cause a representation bias phenomenon and further degenerate generic representations during local training, i.e., the representation degeneration phenomenon. To address these issues, we propose a general framework Domain Bias Eliminator (DBE) for FL. Our theoretical analysis reveals that DBE can promote bi-directional knowledge transfer between server and client, as it reduces the domain discrepancy between server and client in representation space. Besides, extensive experiments on four datasets show that DBE can greatly improve existing FL methods in both generalization and personalization abilities. The DBE-equipped FL method can outperform ten state-of-the-art personalized FL methods by a large margin. Our code is public at https://github.com/TsingZ0/DBE.
Jianqing Zhang, Yang Hua 0001, Jian Cao 0001, Hao Wang 0022, Tao Song 0003, Zhengui Xue, Ruhui Ma, Haibing Guan
NeurIPS7
2023 FHAP: Fast Handover Authentication Protocol for High-Speed Mobile Terminals in 5G Satellite-Terrestrial-Integrated Networks
abstract
The integration of satellite and terrestrial networks presents new opportunities and challenges for high speed rail (HSR) communications. Since the HSR runs vary fast, user terminals on the HSR have to perform handover authentication when the HSR passes through different terrestrial base stations or satellite coverage areas. To improve the security and robustness of HSR communication services, a fast handover authentication protocol (FHAP) for high-speed mobile terminals in the 5G satellite–terrestrial-integrated networks (STNs) is proposed. In the FHAP, user terminals in the same carriage form a temporary group managed by a relay node. The prehandover authentication mechanism is employed, in which the terrestrial 5G core network configures the preauthentication information to multiple access nodes based on the Chinese Remainder Theorem according to the location information of the HSR. Thus, group members and one of the access nodes can achieve fast handover authentication with the preconfigured information. Considering that HSR has a high running speed, when handover authentication fails, user terminals can perform handover authentication with other access nodes, and the probability of handover authentication failure caused by a single access node can be effectively reduced. We use the protocol verification tool Scyther and the Burrows–Abadi–Needham (BAN) logic to prove the security of the FHAP and compare it with other similar protocols in terms of signaling, bandwidth, and computational overhead. The analysis results show that the FHAP satisfies better security properties and has excellent performance.
Jin Cao 0001, Ruhui Ma, Lifu Cheng, Lilan Chen, Ben Niu 0001, Hui Li 0006
IEEE Internet Things J.3
2023 Automatic Pipeline Parallelism: A Parallel Inference Framework for Deep Learning Applications in 6G Mobile Communication Systems
abstract
With the rapid development of wireless communication, achieving the neXt generation Ultra-Reliable and Low-Latency Communications (xURLLC) in 6G mobile communication systems has become a critical problem. Among many applications in xURLLC, deep learning model inference requires improvement over its efficiency. Due to the heterogeneous hardware environment in 6G, parallel schedules from distributed machine learning and edge computing has been borrowed to tackle the efficiency problem. However, traditional parallel schedules suffer from high latency, low throughput, and low device utility. In this paper, we propose Automatic Pipeline Parallelism ($AP^{2}$), a parallel inference framework for deep learning applications in 6G mobile communication systems, to improve the model inference efficiency while maintaining reliability.$AP^{2}$contains three sub-modules. A task-device affinity predictor predicts a task’s expected execution time on a given device. The parallel inference arrangement optimizer finds the most suitable device for each task. The parallel inference scheduler converts the arrangement to a schedule that can be directly executed in the system. The experimental results show that$AP^{2}$can achieve better latency, throughput, reliability, and device utility than other parallel schedules. Also, the priority of the sub-module designs has been approved through the experiments.
Hongjian Shi, Weichu Zheng, Zifei Liu, Ruhui Ma, Haibing Guan
IEEE J. Sel. Areas Commun.4
2023 OSTTD: Offloading of Splittable Tasks With Topological Dependence in Multi-Tier Computing Networks
abstract
With the arrival of the Internet of Things (IoT) era, multi-tier computing has attracted significant attention. The multi-tier computing can organize many computing devices and provide sufficient computing resources to support various IoT applications. However, due to the complex architecture and the dynamic system status of the multi-tier computing network, task offloading for multi-tier computing is still challenging. This paper proposes a novel task offloading method named OSTTD, to deal with the Offloading of Splittable Tasks with Topological Dependence in multi-tier computing networks. OSTTD formulates the task offloading as a sequential decision-making problem and learns the task offloading policy by Deep Reinforcement Learning (DRL). Compared with existing task offloading approaches, OSTTD is the first method in which the topological dependence among sub-tasks of the splittable task is fully considered. OSTTD makes offloading decisions intelligently based on the dynamic system status and can be applied to various multi-tier network topology structures. To verify the effectiveness of OSTTD, we extend and build a latency-aware multi-tier computing simulation platform. Extensive simulations show that OSTTD can significantly reduce the task processing time, thus, improving the overall task processing efficiency in multi-tier computing networks.
Rui Zhang 0087, Xuesen Chu, Ruhui Ma, Honghao Gao, Haibing Guan
IEEE J. Sel. Areas Commun.3
2023 WH2D2N2: Distributed AI-enabled OK-ASN Service for Web of Things
abstract
Model data-driven ontology and knowledge presentation for evolving semantic Asian social networks (OK-ASN) is a critical strategy for web of things (WoT) services. Meanwhile, Deep Neural Network (DNN)-based OK-ASN service in WoT is growing rapidly. However, most DNN-based services cannot utilize the potential of WoT fully, as heterogeneity exists in WoT. Therefore, this article proposes a novel framework called Web-based Heterogeneous Hierarchical Distributed Deep Neural Network ( WH 2 D 2 N 2 ) to deploy the DNNs for OK-ASN services on WoT, overcoming the heterogeneity. The architecture of the system and the designed Edge-Cloud-Joint execute scheme utilize heterogeneous devices to make DNN inference ubiquitous and output two types of results to meet various requirements. To bring robustness to OK-ASN services, a global scheduling is designed to arrange the workflow dynamically. The results of our experiments prove the efficiency of the execute scheme and the global scheduling in the system.
Ruhui Ma, Yang Hua 0001, Hao Wang 0022, Ningxin Hu, Tao Song 0003, Honghao Gao, Haibing Guan
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2023 Perceptual Data Augmentation for Biomedical Coronary Vessel Segmentation
abstract
Sufficient annotated data is critical to the success of deep learning methods. Annotating for vessel segmentation in X-ray coronary angiograms is extremely difficult because of the small and complex structures to be processed. Although unsupervised domain adaptation methods can be utilized to alleviate the annotation burden by using data in other domains, e.g., eye fundus images, these methods cannot perform well due to the characteristic of medical images. Data augmentation can help improve the similarity of source domain and target domain in unsupervised domain adaptation tasks. Existing data augmentation methods play a limited role in improving domain adaptation performance, especially for special medical image segmentation tasks. In this paper, we propose an effective perceptual data augmentation method to improve the similarity between eye fundus images and coronary angiograms by synthesizing virtual samples. Auto Foreground Augment method is designed to search for geometric transformations that improve the similarity between foreground vessels of eye fundus images and coronary angiograms. The Haar Wavelet-Based Perceptual Similarity Index is utilized to guide the synthesis of virtual samples in foreground and background mixup. Extensive experiments show that our data augmentation method can synthesize high-quality virtual samples and thus improve the domain adaptation performance. To our best knowledge, this is the first work to apply perceptual data augmentation to vessel segmentation in coronary angiograms.
Shuo Wang 0008, Yang Hua 0001, Ruhui Ma, Tao Song 0003, Zhengui Xue, Haibing Guan
IEEE ACM Trans. Comput. Biol. Bioinform.4
2023 GUARDIAN: A Hardware-Assisted Distributed Framework to Enhance Deep Learning Security
abstract
The ubiquity of artificial intelligence (AI) has led to its extensive research and application in various fields, such as computer vision, natural language processing, and medical image analysis. However, responsible AI faces severe security challenges, including the leakage of pretrained models and valuable training data. The existing solutions adopt new algorithm designs (such as federated learning) or cryptography (such as homomorphic encryption) to prevent possible security vulnerabilities. We observe that hardware-assisted trusted execution environments (TEEs) can further improve machine learning responsibility. Intel Software Guard Extension (SGX) is a popular, trusted execution hardware that enables users’ programs to run in an untrusted execution environment, such as a malicious operating system, but ensures the confidentiality and integrity of data. Therefore, we have designed GUARDIAN, a hardware-assisted secure machine learning training framework that protects data security during the training process. We have analyzed the typical characteristics of machine learning applications and characterized GUARDIAN through extensive experiments. Our findings demonstrate that introducing security guarantees causes performance degradation, which provides a feasible optimization direction in the near future.
Zinuo Cai, Bojun Ren, Ruhui Ma, Haibing Guan, Mengke Tian, Yong Wang 0085
IEEE Trans. Comput. Soc. Syst.3
2023 A Software-Based Remote Attestation Scheme for Internet of Things Devices
abstract
With the rapid development of intelligent applications, many Internet of Things (IoT) devices are deployed in various application scenarios, playing an extremely important role. Remote attestation is an important method to ensure the software integrity of these devices and protect them from several attacks. Due to the lack of security hardware and no support of hardware extensions for Class-1 IoT devices, it is particularly important to design a suitable remote attestation scheme for these devices. In this paper, we first propose the delayed observation mechanism to alleviate the problem that the software-based remote attestation scheme is not suitable for wireless networks. At the same time, we propose a ”filling memory at attestation-time” mechanism, which solves the problem that attackers hide malicious code through return-oriented programming. Finally, we introduce a reputation mechanism to assist our attestation, and adopt the principle of ”making higher-performance verification nodes take on more work” to greatly reduce the time-consuming attestation. We analyze the security of the scheme and implement it on a UNO-R3 development board to prove its practicability and effectiveness. Compared with traditional software-based attestation schemes, our scheme can reduce the attestation time and resist proxy attacks.
Jin Cao 0001, Ruhui Ma, Zhenyang Guo, Yinghui Zhang 0002, Hui Li 0006
IEEE Trans. Dependable Secur. Comput.3
2023 A Novel Access and Handover Authentication Scheme in UAV-Aided Satellite-Terrestrial Integration Networks Enabling 5G
abstract
As an outlook of the terrestrial network, the flexible Unmanned Aerial Vehicle (UAV)-aided satellite-terrestrial integration network would have numerous prospective applications such as service enhancement, maritime communication, military, and emergency communication, which is getting significant attention. However, the public and open-access network must result in various imperative risks including impersonation, sensitive data, and privacy disclosure. Due to several unique characteristics including long transmission distance, unstable communication environment, resource-limited and highly dynamic characteristics of UAVs, diversified terminals that lack the ability of accessing the satellite and may be resource-limited, there are new challenges for the access authentication process in the UAV-Aided Satellite-Terrestrial Integration Networks with 5G. In this paper, we present a novel Physically Unclonable Function (PUF)-based access authentication scheme consisting of two access authentication protocols for the UAV and the ground terminals like 5G User Equipments (UEs), respectively. Two access authentication protocols for the drone and the terminals can both achieve mutual authentication, key agreement, and privacy protection with distinct advantages of no need to store secret key and supporting physical attack resistance. Finally, we propose an efficient handover authentication protocol executed by the ground terminals when the UAV is required to switch. We employ different security analysis tools to analyze the security of all proposed protocols, as well as present informal security analysis on various security properties. The performance comparison and evaluations show our protocols have better advantages.
Xiongpeng Ren, Jin Cao 0001, Ruhui Ma, Yurong Luo, Yinghui Zhang 0002, Hui Li 0006
IEEE Trans. Netw. Serv. Manag.3
2022 Improving Bayesian Neural Networks by Adversarial Sampling
abstract
Bayesian neural networks (BNNs) have drawn extensive interest due to the unique probabilistic representation framework. However, Bayesian neural networks have limited publicized deployments because of the relatively poor model performance in real-world applications. In this paper, we argue that the randomness of sampling in Bayesian neural networks causes errors in the updating of model parameters during training and some sampled models with poor performance in testing. To solve this, we propose to train Bayesian neural networks with Adversarial Distribution as a theoretical solution. To avoid the difficulty of calculating Adversarial Distribution analytically, we further present the Adversarial Sampling method as an approximation in practice. We conduct extensive experiments with multiple network structures on different datasets, e.g., CIFAR-10 and CIFAR-100. Experimental results validate the correctness of the theoretical analysis and the effectiveness of the Adversarial Sampling on improving model performance. Additionally, models trained with Adversarial Sampling still keep their ability to model uncertainties and perform better when predictions are retained according to the uncertainties, which further verifies the generality of the Adversarial Sampling approach.
Jiaru Zhang, Yang Hua 0001, Tao Song 0003, Hao Wang 0022, Zhengui Xue, Ruhui Ma, Haibing Guan
AAAI6
2022 An Efficient Deep Learning Approach To IoT Intrusion Detection
abstract
Abstract With the rapid development of the Internet of Things (IoT), network security challenges are becoming more and more complex, and the scale of intrusion attacks against the network is gradually increasing. Therefore, researchers have proposed Intrusion Detection Systems and constantly designed more effective systems to defend against attacks. One issue to consider is using limited computing power to process complex network data efficiently. In this paper, we take the AWID dataset as an example, propose an efficient data processing method to mitigate the interference caused by redundant data and design a lightweight deep learning-based model to analyze and predict the data category. Finally, we achieve an overall accuracy of 99.77% and an accuracy of 97.95% for attacks on the AWID dataset, with a detection rate of 99.98% for the injection attack. Our model has low computational overhead and a fast response time after training, ensuring the feasibility of applying to edge nodes with weak computational power in the IoT.
Ruhui Ma, Haibing Guan, Mengke Tian, Yong Wang 0085
Comput. J.3
2022 RPRIA: Reputation and PUF-Based Remote Identity Attestation Protocol for Massive IoT Devices
abstract
The smart city is a vision for urban development that combines new-generation information technologies to improve the quality of life. The Internet of Things (IoT) technology is a key technology of smart city, which provides the support for the automatic collection and transmission of massive information in smart city. However, there are some challenges for the IoT devices in a smart city in terms of security and performance, such as various security vulnerabilities and a large amount of computational overhead. With the explosive growth of the number of IoT devices and the data they generate, it is very necessary to ensure that a large number of IoT devices can establish secure and reliable communication with the central server. In this article, we design a massive IoT device remote identity attestation protocol based on the reputation mechanism and physically unclonable function (PUF). Our protocol can efficiently and securely accomplish the mutual authentication and key agreement between massive IoT devices and central server, and between IoT devices. The Burrows–Abadi–Needham (BAN) logic and a formal verification tool called Scyther are employed to prove the security of our protocol. We also evaluate the performance by comparing our protocol with other related protocols in terms of computational overhead, communication overhead, etc. The security and performance results show that our protocol has ideal security and performance.
Jin Cao 0001, Ruhui Ma, Yueyu Zhang, Hui Li 0006
IEEE Internet Things J.3
2022 A robust authentication scheme for remote diagnosis and maintenance in 5G V2N
Ruhui Ma, Jin Cao 0001, Dengguo Feng, Hui Li 0006, Xiaowei Li 0001
J. Netw. Comput. Appl.1
2022 Towards Ubiquitous Intelligent Computing: Heterogeneous Distributed Deep Neural Networks
abstract
For the pursuit of ubiquitous computing, distributed computing systems containing the cloud, edge devices, and Internet-of-Things devices are highly demanded. However, existing distributed frameworks do not tailor for the fast development of Deep Neural Network (DNN), which is the key technique behind many intelligent applications nowadays. Based on prior exploration on distributed deep neural networks (DDNN), we propose Heterogeneous Distributed Deep Neural Network (HDDNN) over the distributed hierarchy, targeting at ubiquitous intelligent computing. While being able to support basic functionalities of DNNs, our framework is optimized for various types of heterogeneity, including heterogeneous computing nodes, heterogeneous neural networks, and heterogeneous system tasks. Besides, our framework features parallel computing, privacy protection and robustness, with other consideration for the combination of heterogeneous distributed system and DNN. Extensive experiments demonstrate that our framework is capable of utilizing hierarchical distributed system better for DNN and tailoring DNN for real-world distributed system properly, which is with low response time, high performance, and better user experience.
Zongpu Zhang, Tao Song 0003, Yang Hua 0001, Xufeng He, Zhengui Xue, Ruhui Ma, Haibing Guan
IEEE Trans. Big Data7
2022 ES2: Building an Efficient and Responsive Event Path for I/O Virtualization
abstract
Hypervisor intervention in the virtual I/O event path is a main performance bottleneck for I/O virtualization because of the incurred costly VM exits. The shortcomings of prior software solutions against virtual interrupt delivery, a major source of VM exits, promoted the emergence of the hardware-based Posted-Interrupt (PI) technology. PI can provide non-exit interrupt delivery without compromising any virtualization benefit. However, it only acts on the half of the event path, i.e., the interrupt path, while guests I/O requests may also trigger a large amount of VM exits. Additionally, PI may still suffer a severe latency from the vCPU scheduling while delivering interrupts. Aiming at an optimal event path, we propose ES2 to simultaneously improve bidirectional I/O event delivery between guests and their devices. On the basis of PI, ES2 introduces hybrid I/O handling scheme for efficient I/O request delivery and intelligent interrupt redirection for enhanced I/O responsiveness. It does not require any modification to guest OS. We demonstrate that ES2 greatly reduces I/O-related VM exits with the exit handling time (EHT) below 2.5 percent for TCP streams and 0.1 percent for UDP streams, increases guest throughput by 1.9x for Memcached and 1.6x for Nginx, and keeps guest latency at a low level.
Xiaokang Hu, Jian Li 0021, Ruhui Ma, Haibing Guan
IEEE Trans. Cloud Comput.3
2022 A Group-Based Multicast Service Authentication and Data Transmission Scheme for 5G-V2X
abstract
5G Vehicular to everything (5G-V2X) has grown more vigorously than ever before, and services massive vehicles can obtain provided by content providers via the network in all fields are widely acknowledged, such as listening to news channels and traffic channels. In the 5G-V2X network, content providers can provide service messages to a group of vehicles belonging to a specific area in point-to-multipoint transmission mode. However, various challenges are in front of the way when vehicles obtain multicast services due to the unique features of 5G-V2X, such as massive vehicles and strong mobility. In this paper, we present a typical vehicle multicast service model in 5G-V2X and propose a group-based multicast service authentication and data transmission scheme based on this model. In the scheme, massive vehicles within the same RAN coverage are constructed into a group and connected to the content provider to gain access to a multicast service using the distributed keys securely by the 5G home network. Subsequently, the multicast service key for protecting the multicast service data is distributed to each vehicle so that the multicast service data can be securely transmitted to vehicles in point-to-multipoint mode. The security analysis results using the formal verification tool and informal security analysis show that the proposed scheme supports the multicast services authentication and authorization, multicast service data protection, key distribution protection, anonymity, unlinkability, and protocol attack resistance. The performance analysis results regarding signaling, computational and communication overheads show that the proposed scheme outperforms other related schemes.
Ruhui Ma, Jin Cao 0001, Yinghui Zhang 0002, Lihui Xiong, Hui Li 0006
IEEE Trans. Intell. Transp. Syst.1
2021 Siren: Byzantine-robust Federated Learning via Proactive Alarming
abstract
With the popularity of machine learning on many applications, data privacy has become a severe issue when machine learning is applied in the real world. Federated learning (FL), an emerging paradigm in machine learning, aims to train a centralized model while distributing training data among a large number of clients in order to avoid data privacy leaking, which has attracted great attention recently. However, the distributed training scheme in FL is susceptible to different kinds of attacks. Existing defense systems mainly utilize model weight analysis to identify malicious clients with many limitations. For example, some defense systems must know the exact number of malicious clients beforehand, which can be easily bypassed by well-designed attack methods and become impractical for real-world scenarios.
Hanxi Guo, Hao Wang 0022, Tao Song 0003, Yang Hua 0001, Zhangcheng Lv, Xiulang Jin, Zhengui Xue, Ruhui Ma, Haibing Guan
SoCC8
2021 Robust Bayesian Neural Networks by Spectral Expectation Bound Regularization
abstract
Bayesian neural networks have been widely used in many applications because of the distinctive probabilistic representation framework. Even though Bayesian neural networks have been found more robust to adversarial attacks compared with vanilla neural networks, their ability to deal with adversarial noises in practice is still limited. In this paper, we propose Spectral Expectation Bound Regularization (SEBR) to enhance the robustness of Bayesian neural networks. Our theoretical analysis reveals that training with SEBR improves the robustness to adversarial noises. We also prove that training with SEBR can reduce the epistemic uncertainty of the model and hence it can make the model more confident with the predictions, which verifies the robustness of the model from another point of view. Experiments on multiple Bayesian neural network structures and different adversarial attacks validate the correctness of the theoretical findings and the effectiveness of the proposed approach.
Jiaru Zhang, Yang Hua 0001, Zhengui Xue, Tao Song 0003, Ruhui Ma, Haibing Guan
CVPR6
2021 Fine-Grained Pose Temporal Memory Module for Video Pose Estimation and Tracking
abstract
The task of video pose estimation and tracking has been largely improved with the development of image pose estimation recently. However, there are still many challenging cases, such as body part occlusion, fast body motion, camera zooming, and complex background. Most existing methods generally use the temporal information to get more precise human bounding boxes or just use it in the tracking stage, but they fail to improve the accuracy of pose estimation tasks. To better solve these problems and utilize the temporal information efficiently and effectively, we present a novel structure, called pose temporal memory module, which is flexible to be transferred into top-down pose estimation frameworks. The temporal information stored in the pose temporal memory is aggregated into the current frame feature in our proposed module. We also transfer compositional de-attention (CoDA) to solve the unique keypoint occlusion problem in this task and propose a novel keypoint feature replacement to recover the extreme error detection under fine-grained keypoint-level guidance. To verify the generality and effectiveness of our proposed method, we integrate our module into two widely used pose estimation frameworks and obtain notable improvement on the PoseTrack dataset with only a few extra computing resources.
Yang Hua 0001, Tao Song 0003, Zhengui Xue, Ruhui Ma, Neil Robertson 0002, Haibing Guan
ICASSP5
2021 Self-Supervised Vessel Segmentation via Adversarial Learning
abstract
Vessel segmentation is critically essential for diagnosing a series of diseases, e.g., coronary artery disease and retinal disease. However, annotating vessel segmentation maps of medical images is notoriously challenging due to the tiny and complex vessel structures, leading to insufficient available annotated datasets for existing supervised methods and domain adaptation methods. The subtle structures and con-fusing background of medical images further suppress the efficacy of unsupervised methods. In this paper, we propose a self-supervised vessel segmentation method via adversarial learning. Our method learns vessel representations by training an attention-guided generator and a segmentation generator to simultaneously synthesize fake vessels and segment vessels out of coronary angiograms. To support the research, we also build the first X-ray angiography coronary vessel segmentation dataset, named XCAD. We evaluate our method extensively on multiple vessel segmentation datasets, including the XCAD dataset, the DRIVE dataset, and the STARE dataset. The experimental results show our method suppresses unsupervised methods significantly and achieves competitive performance compared with supervised methods and traditional methods.
Yang Hua 0001, Hanming Deng, Tao Song 0003, Hao Wang 0022, Zhengui Xue, Ruhui Ma, Haibing Guan
ICCV8
2021 Fast and Accurate Scene Parsing via Bi-Direction Alignment Networks
abstract
In this paper, we propose an effective method for fast and accurate scene parsing called Bidirectional Alignment Network (BiAlignNet). Previously, one representative work BiSeNet [1] uses two different paths (Context Path and Spatial Path) to achieve balanced learning of semantics and details, respectively. However, the relationship between the two paths is not well explored. We argue that both paths can benefit each other in a complementary way. Motivated by this, we propose a novel network by aligning two-path information into each other through a learned flow field. To avoid the noise and semantic gaps, we introduce a Gated Flow Alignment Module to align both features in a bidirectional way. Moreover, to make the Spatial Path learn more detailed information, we present an edge-guided hard pixel mining loss to supervise the aligned learning process. Our network achieves 80.1parcent and 78.5parcent mIoU in validation and test set of Cityscapes while running at 30 FPS with full resolution inputs. Code and models will be available at https://github.com/jojacola/BiAlignNet.
Yanran Wu, Xiangtai Li, Yunhai Tong, Yang Hua 0001, Tao Song 0003, Ruhui Ma, Haibing Guan
ICIP7
2021 Themis: A Fair Evaluation Platform for Computer Vision Competitions
abstract
It has become increasingly thorny for computer vision competitions to preserve fairness when participants intentionally fine-tune their models against the test datasets to improve their performance. To mitigate such unfairness, competition organizers restrict the training and evaluation process of participants' models. However, such restrictions introduce massive computation overheads for organizers and potential intellectual property leakage for participants. Thus, we propose Themis, a framework that trains a noise generator jointly with organizers and participants to prevent intentional fine-tuning by protecting test datasets from surreptitious manual labeling. Specifically, with the carefully designed noise generator, Themis adds noise to perturb test sets without twisting the performance ranking of participants' models. We evaluate the validity of Themis with a wide spectrum of real-world models and datasets. Our experimental results show that Themis effectively enforces competition fairness by precluding manual labeling of test sets and preserving the performance ranking of participants' models.
Zinuo Cai, Jianyong Yuan, Yang Hua 0001, Tao Song 0003, Hao Wang 0022, Zhengui Xue, Ningxin Hu, Jonathan Ding, Ruhui Ma, Mohammad R. Haghighat, Haibing Guan
IJCAI9
2021 A Composable Multifactor Identity Authentication and Authorization Scheme for 5G Services
abstract
The fifth-generation (5G) mobile communication technology has already deployed commercially and become a global research focus. The new features of 5G include unlimited information exchange, a large variety of connections with independent energy, and diversified high transmission rate services. Collective synergy of services is expected to change the way of life and future generations and introduce new converged services to the ICT industry. Different application services have to meet differentiated security demands. From the perspective of security, in order to support the multiservice of 5G services, it is necessary to consider the new security mechanism driven by the service. Based on 5G massive data stream, the 5G system can provide customized real-world services for potential users and reduce the user experience gap in different scenarios. However, 3GPP Extensible Authentication Protocol (EAP), which is the present entity authentication mechanism for the 5G service layer, is only an individual authentication architecture and unable to fulfill the flexible security objectives of differentiated services. In this paper, we present a new hierarchical identity management framework as well as an adaptable and composable three-factor authentication and session key agreement protocol for different applications in 5G multiservice systems. Finally, we propose an authorization process by combining with the proposed three-factor authentication mechanism and Service-Based Architecture (SBA) proposed by the 3GPP committee. The proposed mechanism can concurrently provide diverse identity authentication schemes corresponding to four different security levels by easily splitting or assembling three-factor authentication protocol blocks. The proposed scheme can be simultaneously applied to a variety of applications to improve the efficiency and quality of service and reduce the complexity of the whole 5G multiservice system, instead of designing or adopting several different authentication protocols. The performance evaluation results indicate that the proposed scheme can guarantee the multiple security of the system with ideal efficiency.
Yurong Luo, Hui Li 0006, Ruhui Ma, Zhenyang Guo
Secur. Commun. Networks3
2020 Reinforcing Neural Network Stability with Attractor Dynamics
abstract
Recent approaches interpret deep neural works (DNNs) as dynamical systems, drawing the connection between stability in forward propagation and generalization of DNNs. In this paper, we take a step further to be the first to reinforce this stability of DNNs without changing their original structure and verify the impact of the reinforced stability on the network representation from various aspects. More specifically, we reinforce stability by modeling attractor dynamics of a DNN and propose relu-max attractor network (RMAN), a light-weight module readily to be deployed on state-of-the-art ResNet-like networks. RMAN is only needed during training so as to modify a ResNet's attractor dynamics by minimizing an energy function together with the loss of the original learning task. Through intensive experiments, we show that RMAN-modified attractor dynamics bring a more structured representation space to ResNet and its variants, and more importantly improve the generalization ability of ResNet-like networks in supervised tasks due to reinforced stability.
Hanming Deng, Yang Hua 0001, Tao Song 0003, Zhengui Xue, Ruhui Ma, Neil Robertson 0002, Haibing Guan
AAAI5
2020 Dual Adversarial Network for Deep Active Learning
Shuo Wang 0008, Yuexiang Li, Kai Ma 0002, Ruhui Ma, Haibing Guan, Yefeng Zheng 0001
ECCV (24)4
2020 VTT: Long-term Visual Tracking with Transformers
abstract
Long-term visual tracking is a challenging problem. State-of-the-art long-term trackers, e.g., GlobalTrack, utilize region proposal networks (RPNs) to generate target proposals. However, the performance of the trackers is affected by occlusions and large scale or ratio variations. To address these issues, in this paper, we are the first to propose a novel architecture with transformers for long-term visual tracking. Specifically, the proposed Visual Tracking Transformer (VTT) utilizes a transformer encoder-decoder architecture for aggregating global information to deal with occlusion and large scale or ratio variation. Furthermore, it also shows better discriminative power against instance-level distractors without the need for extra labeling and hard-sample mining. We conduct extensive experiments on three large-scale long-term tracking datasets and have achieved state-of-the-art performance.
Tianling Bian, Yang Hua 0001, Tao Song 0003, Zhengui Xue, Ruhui Ma, Neil Robertson 0002, Haibing Guan
ICPR5
2020 A Secure Authentication Scheme for Remote Diagnosis and Maintenance in Internet of Vehicles
abstract
Due to the low latency and high speed of 5G networks, the Internet of Vehicles (IoV) under the 5G network has been rapidly developed and has broad application prospects. The Third Generation Partnership Project (3GPP) committee has taken remote diagnosis as one of the development cores of IoV. However, how to ensure the security of remote diagnosis and maintenance services is also a key point to ensure vehicle safety, which is directly related to the safety of vehicle passengers. In this paper, we propose a secure and efficient authentication scheme based on extended chebyshev chaotic maps for remote diagnosis and maintenance in IoVs. In the proposed scheme, to provide strong security, anyone, such as the vehicle owner or the employee of the Vehicle Service Centre (VSC), must enter the valid biometrics and password in order to enjoy or provide remote diagnosis and maintenance services, and the vehicle and the VSC should authenticate each other to ensure that they are legitimate. The security analysis and performance evaluation results show that the proposed scheme can provide robust security with ideal efficiency.
Ruhui Ma, Jin Cao 0001, Dengguo Feng, Hui Li 0006, Ben Niu 0001, Fenghua Li 0001, Lihua Yin
WCNC1
2020 Online traffic-aware linked VM placement in cloud data centers
David S. L. Wei, Ruhui Ma, Jian Li 0021, Haibing Guan
Sci. China Inf. Sci.3
2020 Balancing Power And Performance In HPC Clouds
abstract
Abstract With energy consumption in high-performance computing clouds growing rapidly, energy saving has become an important topic. Virtualization provides opportunities to save energy by enabling one physical machine (PM) to host multiple virtual machines (VMs). Dynamic voltage and frequency scaling (DVFS) is another technology to reduce energy consumption. However, in heterogeneous cloud environments where DVFS may be applied at the chip level or the core level, it is a great challenge to combine these two technologies efficiently. On per-core DVFS servers, cloud managers should carefully determine VM placements to minimize performance interference. On full-chip DVFS servers, cloud managers further face the choice of whether to combine VMs with different characteristics to reduce performance interference or to combine VMs with similar characteristics to take better advantage of DVFS. This paper presents a novel mechanism combining a VM placement algorithm and a frequency scaling method. We formulate this VM placement problem as an integer programming (IP) to find appropriate placement configurations, and we utilize support vector machines to select suitable frequencies. We conduct detailed experiments and simulations, showing that our scheme effectively reduces energy consumption with modest impact on performance. Particularly, the total energy delay product is reduced by up to 60%.
Lixia Chen, Jian Li 0021, Ruhui Ma, Haibing Guan, Hans-Arno Jacobsen
Comput. J.3
2020 LSAA: A Lightweight and Secure Access Authentication Scheme for Both UE and mMTC Devices in 5G Networks
abstract
As a development of the next generation of mobile communication networks and systems (5G), the Third-Generation Partnership Project (3GPP) committee has standardized a new 5G authentication and key-agreement (5G-AKA) protocol to ensure the access security of a mobile equipment. However, there are still some security vulnerabilities in the 5G-AKA protocol, and there is no authentication protocol proposed for massive device concurrent connection by the 3GPP working groups. In this article, we propose a novel lightweight and secure access authentication scheme named lightweight secure access authentication (LSAA) that contains two lightweight extended Chebyshev chaotic maps-based access authentication protocols for two types of 3GPP standard mobile devices: 1) common user equipment (UE) and 2) massive machine-type communication (mMTC) devices. Our proposed protocols can achieve several security functionalities, including mutual authentication, session-key establishment, identity privacy protection, and perfect forward/backward secrecy (PFS/PBS). In addition, the proposed protocols are lightweight in nature compared with the 5G-AKA. In order to comprehensively and accurately evaluate LSAA, we carry out formal security analysis by employing two formal verification tools Proverif and Scyther, and informal security analysis on the proposed protocols. We further evaluate the performance of the proposed protocols with regard to authentication signaling cost, authentication communication cost, authentication computational cost, and authentication storage cost. The security evaluation and performance analysis results show that our proposed protocols can provide advanced security and high efficiency.
Jin Cao 0001, Zheng Yan 0002, Ruhui Ma, Yinghui Zhang 0002, Hui Li 0006
IEEE Internet Things J.3
2020 LAA: Lattice-Based Access Authentication Scheme for IoT in Space Information Networks
abstract
Space information network (SIN), which has the characteristics of large capacity, high reliability, and wide coverage, can be effectively applied to the Internet of Things (IoT) business. However, SIN is vulnerable to various attacks due to its highly exposed links, and the power and processing capacity of satellites and IoT devices (IoTDs) are usually limited. Meanwhile, massive IoTDs connecting to SIN in an instant incur a severe signaling congestion and there is no authentication protocol proposed for massive IoTDs in SIN. In this article, we first propose a novel semi-aggregated signature mechanism and session key agreement mechanism. Then, based on the above two mechanisms, we propose a novel access authentication scheme named lattice-based access authentication (LAA) which contains two types of lattice-based authentication protocols: 1) LAA for massive IoTDs and 2) LAA for a single IoTD. The security analysis by employing the formal provable security analysis, the automation verification tool named Scyther, the Burrows-Abadi-Needham-logic, and the informal security analysis demonstrate that our proposed protocols successfully achieve these security properties, including mutual authentication, conditional anonymity, unlinkability, data confidentiality, data integrity, unforgeability, undeniability, key establishment, perfect forward/backward secrecy (PFS/PBS), and resistance against protocol attacks and quantum attacks. We further evaluate the performance of our proposed protocols with regard to signaling overhead, transmission overhead, computational overhead, and authentication delay, which shows that our proposed protocols can provide high efficiency.
Ruhui Ma, Jin Cao 0001, Dengguo Feng, Hui Li 0006
IEEE Internet Things J.1
2019 vDARM: Dynamic Adaptive Resource Management for Virtualized Multiprocessor Systems
abstract
Modern data center servers have been enhancing their computing capacity by increasing processor counts. Meanwhile, these servers are highly virtualized to achieve efficient resource utilization and energy savings. However, due to the shifting of server architecture to non-uniform memory access (NUMA), current hypervisor-level or OS-level resource management methods continue to be challenged in their ability to meet the performance requirement of various user applications. In this work, we first build a performance slowdown model to accurate identify the current system overheads. Based on the model, we finally design a dynamic adaptive virtual resource management method (vDARM) to eliminate the runtime NUMA overheads by re-configuring virtual-to-physical resource mappings. Experiment results show that, compared with state-of-art approaches, vDARM can bring up an average performance improvement of 42.3% on an 8-node NUMA machines. Meanwhile, vDARM only incurs extra CPU utilization no more than 4%.
Jianmin Qian, Jian Li 0021, Ruhui Ma, Haibing Guan
DATE3
2019 Object Guided External Memory Network for Video Object Detection
abstract
Video object detection is more challenging than image object detection because of the deteriorated frame quality. To enhance the feature representation, state-of-the-art methods propagate temporal information into the deteriorated frame by aligning and aggregating entire feature maps from multiple nearby frames. However, restricted by feature map's low storage-efficiency and vulnerable content-address allocation, long-term temporal information is not fully stressed by these methods. In this work, we propose the first object guided external memory network for online video object detection. Storage-efficiency is handled by object guided hard-attention to selectively store valuable features, and long-term information is protected when stored in an addressable external data matrix. A set of read/write operations are designed to accurately propagate/allocate and delete multi-level memory feature under object guidance. We evaluate our method on the ImageNet VID dataset and achieve state-of-the-art performance as well as good speed-accuracy tradeoff. Furthermore, by visualizing the external memory, we show the detailed object-level reasoning process across frames.
Hanming Deng, Yang Hua 0001, Tao Song 0003, Zongpu Zhang, Zhengui Xue, Ruhui Ma, Neil Robertson 0002, Haibing Guan
ICCV6
2019 EnclaveCache: A Secure and Scalable Key-value Cache in Multi-tenant Clouds using Intel SGX
abstract
With in-memory key-value caches such as Redis and Memcached being a key component for many systems to improve throughput and reduce latency, cloud caches have been widely adopted for small companies to deploy their own cache systems. However, data security is still a major concern, which affects the adoption of cloud caches. Tenant's data stored in a multi-tenant cloud environment faces threats from both co-located other tenants, as well as the untrusted cloud provider.
Li-Xia Chen, Jian Li 0021, Ruhui Ma, Haibing Guan, Hans-Arno Jacobsen
Middleware3
2019 Unsupervised Video Summarization with Attentive Conditional Generative Adversarial Networks
abstract
With the rapid growth of video data, video summarization technique plays a key role in reducing people's efforts to explore the content of videos by generating concise but informative summaries. Though supervised video summarization approaches have been well studied and achieved state-of-the-art performance, unsupervised methods are still highly demanded due to the intrinsic difficulty of obtaining high-quality annotations. In this paper, we propose a novel yet simple unsupervised video summarization method with attentive conditional Generative Adversarial Networks (GANs). Firstly, we build our framework upon Generative Adversarial Networks in an unsupervised manner. Specifically, the generator produces high-level weighted frame features and predicts frame-level importance scores, while the discriminator tries to distinguish between weighted frame features and raw frame features. Furthermore, we utilize a conditional feature selector to guide GAN model to focus on more important temporal regions of the whole video frames. Secondly, we are the first to introduce the frame-level multi-head self-attention for video summarization, which learns long-range temporal dependencies along the whole video sequence and overcomes the local constraints of recurrent units, e.g., LSTMs. Extensive evaluations on two datasets, SumMe and TVSum, show that our proposed framework surpasses state-of-the-art unsupervised methods by a large margin, and even outperforms most of the supervised methods. Additionally, we also conduct the ablation study to unveil the influence of each component and parameter settings in our framework.
Xufeng He, Yang Hua 0001, Tao Song 0003, Zongpu Zhang, Zhengui Xue, Ruhui Ma, Neil Robertson 0002, Haibing Guan
ACM Multimedia6
2019 PPSHA: Privacy preserving secure handover authentication scheme for all application scenarios in LTE-A networks
Ruhui Ma, Jin Cao 0001, Dengguo Feng, Hui Li 0006, Yinghui Zhang 0002, Xixiang Lv
Ad Hoc Networks1
2019 Implementation of an effective dynamic concolic execution framework for analyzing binary programs
Erzhou Zhu, Kanqi Ni, Ruhui Ma
Comput. Secur.4
2019 vSimilar: A high-adaptive VM scheduler based on the CPU pool mechanism
Ruhui Ma, Jian Li 0021, Dajin Wang, Haibing Guan
J. Syst. Archit.3
2019 LG-RAM: Load-aware global resource affinity management for virtualized multicore systems
Jianmin Qian, Jian Li 0021, Ruhui Ma, Haibing Guan
J. Syst. Archit.3
2019 When I/O Interrupt Becomes System Bottleneck: Efficiency and Scalability Enhancement for SR-IOV Network Virtualization
abstract
High performance networking interface cards (NIC) have become essential networking devices in commercial cloud computing environments. Therefore, efficient and scalable I/O virtualization is one of the primary challenges on virtualized cloud computing platforms. Single Root I/O Virtualization (SR-IOV) is a network interface technology that eliminates the overhead of redundant data copies and the virtual network switches through direct I/O in order to achieve nearly natural I/O performance. However, the SR-IOV still suffers from serious problems due to the high overhead for processing excessive network interrupts as well as the unpredictable and bursty traffic load in high-speed networking connections. In this paper, the defects of SR-IOV with 10 Gigabit Ethernet networking are studied first and two major challenges are identified: excessive interrupt rate and single threaded virtual network driver. Second, two interrupt rate control optimization schemes, called coarse-grained interrupt rate (CGR) control and adaptive interrupt rate (AIR) control are proposed. The proposed control schemes can significantly reduce the overhead and enhance the SR-IOV performance compared with the traditional driver with fixed interrupt throttle rate (FIR). In addition, multi-threaded VF driver (MTVD) is proposed that allows the SR-IOV VFs to leverage multi-core resources in order to achieve high scalability. Finally, these optimizations are implemented and detailed performance evaluations are conducted. The results show that CGR and AIR can improve the throughput by 2.26× and 2.97× while saving the CPU resources by 1.23 core and 1.44 core, respectively. The MTVD can achieve 2.03× performance with additional 1.46 cores consumption for VM using the SR-IOV driver.
Jian Li 0021, Ruhui Ma, Zhengwei Qi, Haibing Guan
IEEE Trans. Cloud Comput.4
2018 Deep Learning based Fetal Middle Cerebral Artery Segmentation in Large-scale Ultrasound Images
Shuo Wang 0008, Yang Hua 0001, Yunyun Cao, Tao Song 0003, Zhengui Xue, Xiaoping Gong, Guanjie Wang, Ruhui Ma, Haibing Guan
BIBM8
2018 Topology-aware virtual resource management for heterogeneous multicore systems
abstract
Virtualization technology consolidates multiple independent workloads on a single physical server with virtual resource management, which can result in a significant utilization improvement and energy saving. However, the management of virtual resource is becoming more and more challenging, due to the lack of accurate performance prediction model for the diverse applications' irregular resource access behaviors as well as the complicated Non-Uniform Memory Access (NUMA) server architecture. These challenges drastically affect the overall consolidation performance. This paper proposes vTRMS, a runtime Topology-aware virtual Resource Management Scheme for heterogeneous NUMA multicore systems. vTRMS can improve application performance based on the comprehensive online monitor of the application resource access behaviors as well as an accurate and platform-independent detected NUMA topology metric. Experiment results show that, compared with state-of-art approach, vTRMS can bring up an average throughput improvement of 28.3% and 36.2% on Intel and AMD NUMA machines respectively, when consolidating 32-VMs. At the same time, vTRMS only incurs a runtime overhead no more than 5%.
Jianmin Qian, Jian Li 0021, Ruhui Ma
DATE3
2018 Optimizing Virtual Resource Management for Consolidated NUMA Systems
abstract
Virtualization consolidates multiple virtual machines (VMs) on a single physical server to achieve high resource utilization and energy conservation. However, as current data center server architectures shifting to Non-Uniform Memory Access (NUMA), the complex interplay between data access affinity and shared resource access overloaded still challenge the consolidation efficiency. Existing NUMA-aware virtual resource management policies lack holistic awareness of the data access affinities as well as unable to manage the system loads, which may result in sub-optimal overall system performance. In this paper, we propose a load-aware global resource access management framework (LG-RAM) that aims to optimize VM consolidation performance on NUMA systems. Our real system based evaluations indicate that, compared with state-of-the-art approaches, LG-RAM exhibits an average throughput improvement of 36.5% on our testbed. Besides, LG-RAM only incurs an extra CPU usage of no more than 9% on average when consolidating 32 VMs.
Jianmin Qian, Jian Li 0021, Ruhui Ma, Haibing Guan
ICCD3
2018 Tracking-assisted Weakly Supervised Online Visual Object Segmentation in Unconstrained Videos
abstract
This paper tackles the task of online video object segmentation with weak supervision, i.e., labeling the target object and background with pixel-level accuracy in unconstrained videos, given only one bounding box information in the first frame. We present a novel tracking-assisted visual object segmentation framework to achieve this. On the one hand, initialized with a given bounding box in the first frame, the auxiliary object tracking module guides the segmentation module frame by frame by providing motion and region information, which is usually missing in semi-supervised methods. Moreover, compared with the unsupervised approach, our approach with such minimum supervision can focus on the target object without bringing unrelated objects into the final results. On the other hand, the video object segmentation module also improves the robustness of the visual object tracking module by pixel-level localization and objectness information. Thus, segmentation and tracking in our framework can mutually help each other in an online manner. To verify the generality and effectiveness of the proposed framework, we evaluate our weakly supervised method on two cross-domain datasets, i.e., the DAVIS and VOT2016 datasets, with the same configuration and parameter setting. Experimental results show the top performance of our method, which is even better than the leading semi-supervised methods. Furthermore, we conduct the extensive ablation study on our approach to investigate the influence of each component and main parameters.
Zongpu Zhang, Yang Hua 0001, Tao Song 0003, Zhengui Xue, Ruhui Ma, Neil Robertson 0002, Haibing Guan
ACM Multimedia5
2018 FastDesk: A remote desktop virtualization system for multi-tenant
Tao Song 0003, Jiewei Wu, Ruhui Ma, Alei Liang, Tao Gu 0001, Zhengwei Qi
Future Gener. Comput. Syst.4
2018 D2FL: Design and Implementation of Distributed Dynamic Fault Localization
abstract
Compromised or misconfigured routers have been a major concern in large-scale networks. Such routers sabotage packet delivery, and thus hurt network performance. Data-plane fault localization (FL) promises to solve this problem. Regrettably, the path-based FL fails to support dynamic routing, and the neighbor-based FL requires a centralized trusted administrative controller (AC) or global clock synchronization in each router and introduces storage overhead for caching packets. To address these problems, we introduce a dynamic distributed and low-cost model, D2FL. Using random two-hop neighborhood authentication, D2FL supports volatile path without the AC or global clock synchronization. Besides, D2FL requires only constant tens of KB for caching which is independent of the packet transmission rate. This is much less than the cache size of DynaFL or DFL which consumes several MB. The simulations show that D2FL achieves low false positive and false negative rate with no more than 3 percent bandwidth overhead. We also implement an open source prototype and evaluate its effect. The result shows that the performance burden in user space is less than 10 percent with the dynamic sampling algorithm.
Fanfu Zhou, Zhengwei Qi, Jianguo Yao 0002, Ruhui Ma, Bin Wang 0062, Athanasios V. Vasilakos, Haibing Guan
IEEE Trans. Dependable Secur. Comput.4
2017 ES2: Aiming at an Optimal Virtual I/O Event Path
abstract
Improving the performance of I/O virtualization is a key issue for cloud and datacenter infrastructures, especially with the rapid increase of network interconnection speeds. Previous efforts have made the performance overhead associated with the virtual I/O data path largely negligible. The remaining bottlenecks mainly lie in the event path: hypervisor interventions trigger costly virtual machine (VM) exits and lead to dramatical performance degradation. Aiming at an optimal virtual I/O event path, we propose ES2, a comprehensive scheme that simultaneously improves bidirectional I/O event delivery between guest VMs and their devices. ES2 can provide efficient I/O request delivery, non-exit interrupt delivery and enhanced I/O responsiveness. Moreover, it does not require any modification to guest operating system (OS) or compromise any virtualization benefit. We demonstrate that ES2 greatly reduces VM exit rate with the time in guest (TIG) for I/O processing above 96% for TCP streams and 99% for UDP streams, increases guest throughput by 1.8× for Memcached and 2× for Apache, and keeps guest latency at a very low level.
Xiaokang Hu, Jian Li 0021, Ruhui Ma, Haibing Guan
ICPP4
2017 Ashman: A Bandwidth Fragmentation-Based Dynamic Flow Scheduling for Data Center Networks
abstract
Current data center (DC) topologies provide redundant paths to satisfy significant communication requirements. Researchers have proposed various approaches including static and dynamic ones to utilize path diversity more effectively. The static methods adopt hash-based way to distribute flows onto multiple paths randomly, while the dynamic ones relying on centralized controller place flows on candidate paths. Unfortunately, the existing flow-scheduling policies have ignored and even resulted in network bandwidth fragmentation which adversely slows down transmission rate of new flows or affects chances of accepting new flow requests, especially when the granularity of flows is larger. This phenomenon is considered as a bottleneck for higher utilization of bandwidth resource in DCs. In this paper, we are the first to identify and define network bandwidth fragmentation within DC. Accordingly, we present Ashman, a flow-based dynamic scheduling approach to reduce bandwidth fragmentation by proactively considering potential large flows. This policy dynamically schedules the existing large flows to resolve congestions and maximize bandwidth utilization and network throughput. We describe our design and implementation in OpenFlow framework with unmodified hosts. Our evaluation on the software-defined networking simulator called Mininet demonstrates that Ashman effectively reduces bandwidth fragmentation, and thereby achieves higher network bandwidth utilization and the overall throughput of DC.
Tao Song 0003, Ruhui Ma, Alei Liang, Zhengwei Qi, Haibing Guan
Comput. J.4
2017 TEES: An Efficient Search Scheme over Encrypted Data on Mobile Cloud
abstract
Cloud storage provides a convenient, massive, and scalable storage at low cost, but data privacy is a major concern that prevents users from storing files on the cloud trustingly. One way of enhancing privacy from data owner point of view is to encrypt the files before outsourcing them onto the cloud and decrypt the files after downloading them. However, data encryption is a heavy overhead for the mobile devices, and data retrieval process incurs a complicated communication between the data user and cloud. Normally with limited bandwidth capacity and limited battery life, these issues introduce heavy overhead to computing and communication as well as a higher power consumption for mobile device users, which makes the encrypted search over mobile cloud very challenging. In this paper, we propose traffic and energy saving encrypted search (TEES), a bandwidth and energy efficient encrypted search architecture over mobile cloud. The proposed architecture offloads the computation from mobile devices to the cloud, and we further optimize the communication between the mobile clients and the cloud. It is demonstrated that the data privacy does not degrade when the performance enhancement methods are applied. Our experiments show that TEES reduces the computation time by 23 to 46 percent and save the energy consumption by 35 to 55 percent per file retrieval, meanwhile the network traffics during the file retrievals are also significantly reduced.
Jian Li 0021, Ruhui Ma, Haibing Guan
IEEE Trans. Cloud Comput.2
2017 Accurate CPU Proportional Share and Predictable I/O Responsiveness for Virtual Machine Monitor: A Case Study in Xen
abstract
In cloud computing, the performance of applications is heavily dependent on resource services provided by the virtualized environment. However, in some virtualized environment, such as Xen, the accuracy of CPU proportional share and the responsiveness of I/O processing are heavily dependent on the proportion of the allocated CPU resource. In this paper, we study how inaccurate share ratio of CPU proportional share and proportion dependent responsiveness of I/O affect the performance of Xen, and discover that they lead to unstable performance and is thus not able to conform service-level agreements (SLA). We conclude that the scheduling scheme and the coarse grained time-slice are the major negative impacts on this issue. Therefore, we propose a novel scheduling scheme, named Predictable Resource Guarantee Scheduler (PRGS), that achieves accurate CPU proportional share and predictable I/O responsiveness. We implement a PRGS prototype on Xen virtualization platform and carry out a thorough evaluation via experimentation. The experimental results show that PRGS achieves accurate CPU proportional share and predictable I/O responsiveness. Also, with only slight overhead, PRGS controls PING packet delay to a specified fixed time threshold (e.g. 30 ms in our experiments).
Jian Li 0021, Ruhui Ma, Haibing Guan, David S. L. Wei
IEEE Trans. Cloud Comput.2
2017 ForenVisor: A Tool for Acquiring and Preserving Reliable Data in Cloud Live Forensics
abstract
Live forensics is an important technique in cloud security but is facing the challenge of reliability. Most of the live forensic tools in cloud computing run either in the target Operating System (OS), or as an extra hypervisor. The tools in the target OS are not reliable, since they might be deceived by the compromised OS. Furthermore, traditional general purpose hypervisors are vulnerable due to their huge code size. However, some modules of a general purpose hypervisor, such as device drivers, are indeed unnecessary for forensics. In this paper, we propose a special purpose hypervisor, called ForenVisor, which is dedicated to reliable live forensics. The reliability is improved in three ways: reducing Trusted Computing Base (TCB) size by leveraging a lightweight architecture, collecting evidence directly from the hardware, and protecting the evidence and other sensitive files with Filesafe module. We have implemented a proof-of-concept prototype on the Windows platform, which can acquire the process data, raw memory, and I/O data, such as keystrokes and network traffic. Furthermore, we evaluate ForenVisor in terms of code size, functionality, and performance. The experiment results show that ForenVisor has a relatively small TCB size of about 13 KLOC, and only causes less than 10 percent performance reduction to the target system. In particular, our experiments verify that ForenVisor can guarantee that the protected files remain untampered, even when the guest OS is compromised by viruses, such as `ILOVEYOU' and Worm.WhBoy. Also, our system can be loaded as a hypervisor without needing to pause the target OS. This allows it to not only avoid destructing but also to gather the live evidence of the target OS. We also posted the source code of ForenVisor on Github.
Zhengwei Qi, Chengcheng Xiang, Ruhui Ma, Jian Li 0021, Haibing Guan, David S. L. Wei
IEEE Trans. Cloud Comput.3
2016 A user mode CPU-GPU scheduling framework for hybrid workloads
Bin Wang 0062, Ruhui Ma, Zhengwei Qi, Jianguo Yao 0002, Haibing Guan
Future Gener. Comput. Syst.2
2016 Optimizations for High Performance Network Virtualization
Fanfu Zhou, Ruhui Ma, Jian Li 0021, Li-Xia Chen, Weidong Qiu, Haibing Guan
J. Comput. Sci. Technol.2
2015 vINT: Hardware-Assisted Virtual Interrupt Remapping for SMP VM with Scheduling Awareness
abstract
Symmetric Multi-Processing (SMP) virtual machine (VM), or virtual SMP for short, enables a single virtual machine to span multiple processors, thereby supporting the virtual machine to run resource-intensive applications. In addition to offering higher computing capacity, virtual SMP also offers the opportunity to alleviate the problem of unpredictable I/O responsiveness. To this end, we propose vINT (scheduling status based virtual INterrupt remapping adapTer), a scheme that leverages hardware-assisted interrupt mapping. vINT obtains high efficiency and flexibility by adding a lightweight module in virtual machine monitor (VMM) with no need of changing VMM scheduler and is transparent to guest OS. We implement the prototype in XEN 4.3.0 and conduct evaluations with both micro-benchmarks and macro-benchmarks. The experimental results show that vINT can increase the networking throughput by 5x and can reduce the required execute time of disk I/O by 17.5%, while introducing only a light overhead.
Jian Li 0021, Ruhui Ma, Haibing Guan, David S. L. Wei
CloudCom2
2015 DaSS: Dynamic Time Slice Scheduler for Virtual Machine Monitor
Ruhui Ma, Jian Li 0021, Haibing Guan
ICA3PP (1)1
2015 A survey on data center networking for cloud computing
Bin Wang 0062, Zhengwei Qi, Ruhui Ma, Haibing Guan, Athanasios V. Vasilakos
Comput. Networks3
2014 Pinso: Precise Isolation of Concurrency Bugs via Delta Triaging
abstract
Concurrent programs are known to be difficult to test and maintain. These programs often fail because of concurrency bugs caused by non-deterministic interleavings among shared memory accesses. Even though a concurrency bug can be detected, it is still hard to isolate the root cause of the bug, due to the challenge in understanding the complex thread interleavings or schedules. In this paper, we propose a practical and precise isolation technique for concurrent bugs called Pinso that seeks to exploit the non-deterministic nature of concurrency bugs and accurately find the root causes of program error, to further help developers maintain concurrent programs. Pinso profiles runtime inter-thread interleavings based on a set of summarized memory access patterns, and then, isolates suspicious interleaving patterns in the triaging phase. Using a filtration-oriented scheduler, Pinso effectively eliminates false positives that are irrelevant to the bug. We evaluate Pinso with 11 real-world concurrency bugs, including single- and multi-variable violation, from sever/desktop concurrent applications (MySQL, Apache, and several others). Experiments indicate that our tool accurately isolates the root causes of all the bugs.
Bo Liu 0001, Zhengwei Qi, Bin Wang 0062, Ruhui Ma
ICSME4
2014 Workload-Aware Credit Scheduler for Improving Network I/O Performance in Virtualization Environment
abstract
Single-root I/O virtualization (SR-IOV) has become the de facto standard of network virtualization in cloud infrastructure. Owing to the high interrupt frequency and heavy cost per interrupt in high-speed network virtualization, the performance of network virtualization is closely correlated to the computing resource allocation policy in Virtual Machine Manager (VMM). Therefore, more sophisticated methods are needed to process irregularity and the high frequency of network interrupts in high-speed network virtualization environment. However, the I/O-intensive and CPU-intensive applications in virtual machines are treated in the same manner since application attributes are transparent to the scheduler in hypervisor, and this unawareness of workload makes virtual systems unable to take full advantage of high performance networks. In this paper, we discuss the SR-IOV networking solution and show by experiment that the current credit scheduler in Xen does not utilize high performance networks efficiently. Hence we propose a novel workload-aware scheduling model with two optimizations to eliminate the bottleneck caused by scheduler. In this model, guest domains are divided into I/O-intensive domains and CPU-intensive domains according to their monitored behaviour. I/O-intensive domains can obtain extra credits that CPU-intensive domains are willing to share. In addition, the total number of credits available is adjusted to accelerate the I/O responsiveness. Our experimental evaluations show that the new scheduling models improve bandwidth and reduce response time, by keeping the fairness between I/O-intensive and CPU-intensive domains. This enables virtualization infrastructure to provide cloud computing services more efficiently and predictably.
Haibing Guan, Ruhui Ma, Jian Li 0021
IEEE Trans. Cloud Comput.2
2013 Cache isolation for virtualization of mixed general-purpose and real-time systems
Ruhui Ma, Alei Liang, Haibing Guan, Jian Li 0021
J. Syst. Archit.1
2013 Performance Enhancement for Network I/O Virtualization with Efficient Interrupt Coalescing and Virtual Receive-Side Scaling
abstract
Virtualization is a key technology in cloud computing; it can accommodate numerous guest VMs to provide transparent services, such as live migration, high availability, and rapid checkpointing. Cloud computing using virtualization allows workloads to be deployed and scaled quickly through the rapid provisioning of virtual machines on physical machines. However, I/O virtualization, particularly for networking, suffers from significant performance degradation in the presence of high-speed networking connections. In this paper, we first analyze performance challenges in network I/O virtualization and identify two problems-conventional network I/O virtualization suffers from excessive virtual interrupts to guest VMs, and the back-end driver does not efficiently use the computing resources of underlying multicore processors. To address these challenges, we propose optimization methods for enhancing the networking performance: 1) Efficient interrupt coalescing for network I/O virtualization and 2) virtual receive-side scaling to effectively leverage multicore processors. These methods are implemented and evaluated with extensive performance tests on a Xen virtualization platform. Our experimental results confirm that the proposed optimizations can significantly improve network I/O virtualization performance and effectively solve the performance challenges.
Haibing Guan, Yaozu Dong, Ruhui Ma, Dongxiao Xu, Jian Li 0021
IEEE Trans. Parallel Distributed Syst.3
2012 Adjustable Credit Scheduling for High Performance Network Virtualization
abstract
Virtualization technology is now widely adopted in cloud computing to support heterogeneous and dynamic workload. The scheduler in a virtual machine monitor (VMM) plays an important role in allocating resources. However, the type of applications in virtual machines (VM) is unknown to the scheduler, and I/O-intensive and CPU-intensive applications are treated the same. This makes virtual systems unable to take full advantage of high performance networks such as 10-Gigabit Ethernet. In this paper, we review the SR-IOV networking solution and show by experiment that the current credit scheduler in Xen does not utilize high performance networks efficiently. For this reason, we propose a novel scheduling model with two optimizations to eliminate the bottleneck caused by scheduler. In this model, guest domains are divided into I/O-intensive domains and CPU-intensive domains according to their monitored behaviour. I/O-intensive domains can obtain extra credits that CPU-intensive domains are willing to share. Besides, the total available credits is adjusted agilely to accelerate the I/O responsiveness. Our experimental evaluation with benchmarks shows that the new scheduling model improves bandwidth even when the system's load is very high.
Zhibo Chang, Jian Li 0021, Ruhui Ma, Haibing Guan
CLUSTER3
2012 Adaptive and Scalable Optimizations for High Performance SR-IOV
abstract
High performance networking interfaces, such as 10-Gigabit Ethernet (10GE), are now widely deployed in commercial Cloud computing environments. Virtualization is a standard technique for these environments, one of whose key challenges is to achieve highly efficient and scalable I/O virtualization. Single Root I/O Virtualization (SR-IOV) eliminates the overhead of redundant data copies and the virtual network switch through direct I/O, but needs more work on performance and scalability. In this paper, we first study the defects of SR-IOV with 10GE networking and find two major challenges. Due to multiplexing of traffic from different virtual machines, SR-IOV may generate redundant interrupts unexpectedly and thus result in high CPU overhead. SR-IOV also suffers from single-threaded NAPI which prevents it from fully utilizing multi-core machines. Then we propose two optimizations for enhancing the SR-IOV performance. The first uses adaptive interrupt rate control (AIRC) to reduce CPU overhead caused by excessive interrupts. The second is a multi-threaded network driver (MTND) which allows SR-IOV to make full use of multi-core resources. We implement these optimizations and carry out a detailed performance evaluation. The results show that AIRC can reduce CPU overhead by up to 143% and MTND can improve SR-IOV performance by up to 38%.
Ruhui Ma, Jian Li 0021, Zhibo Chang, Haibing Guan
CLUSTER2
2012 SINOF: A dynamic-static combined framework for dynamic binary translation
Haibing Guan, Erzhou Zhu, Hongxi Wang, Ruhui Ma, Yindong Yang, Bin Wang 0062
J. Syst. Archit.4
2011 MTCrossBit: A dynamic binary translation system based on multithreaded optimization
Haibing Guan, Ruhui Ma, Yindong Yang, Liang Liu 0010, Ying Chen 0004
Sci. China Inf. Sci.2
2011 Partitioning the Conventional DBT System for Multiprocessors
Ruhui Ma, Haibing Guan, Erzhou Zhu, Yindong Yang, Alei Liang
J. Comput. Sci. Technol.1
2009 MTCrossBit: A Dynamic Binary Translation System Using Multithreaded Optimization Framework
Deen Zheng, Ruhui Ma, Alei Liang, Haibing Guan
ICA3PP3