Peini Liu

dblp:98/5525 · DBLP profile ↗
← Back
20ranked-venue papers
8as first author
13since 2021 · last 2026
0000-0003-0058-8732ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 2 since 2021Computer networks · 4 · 2 first-author · 4 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CloudSkin: AI-Based Learning Plane for Autonomic Management in the Cloud-Edge Continuum
Peini Liu, Joan Oliveras Torra, Marc Palacín, Ramon Nou, Josep Lluís Berral, Jordi Guitart
COMPSAC1
2025 Dynamic In-node Group-Aware Scheduling for Multi-Tenant Machine Learning Services on Kubernetes
abstract
Machine Learning (ML) is becoming pervasive and integrated into different kinds of businesses. Hosting multi-tenant ML services on a unified platform requires efficient orchestration. Service orchestration considers two levels: an in-cluster scheduling to decide the node allocation, followed by an in-node scheduling to manage the resource distribution within the node. Previous works introduced multi-container deployments for ML services, demonstrating that partitioning the ML service and enabling CPU/Memory affinity for each container improves performance. These multi-container deployments have been enabled for Kubernetes at in-cluster level to allocate multiple groups of containers across nodes. However, when the containers are launched in the node, the in-node scheduler lacks awareness of the group information which challenges those containers for a fine-grained resource allocation, especially when multiple groups of containers share the same node. This paper presents an in-node group-aware scheduling mechanism for multiple ML services, where each service scheduling contains a group resource selection and container resource assignment. Moreover, we also provide a dynamic resource controller (DRC) to dynamically reallocate the resource allocation for containers using the mechanism, which monitors the group changes and acts with the real containers' system cgroups adaptation. Our results show that for deploying ML services with the same ML model and different ML models, DRC throughput outperforms other deployment scenarios by up to 258% and 319% respectively. In an experiment deploying a dynamic workload of multiple mixed ML services, DRC outperforms baseline NONE-Single by 242 %, 75 %, and 28 % for the average throughput of services with Mobilenet, Resnet50, and VGG 16, respectively, and also DRC results in a makespan 44% faster than NONE-Single and 4% faster than CM-Multi.
Peini Liu, Jordi Guitart
CLOUD1
2025 Integrating Reliability into Intent-Driven Orchestration on Kubernetes-based Edge Nodes
abstract
Placing workloads across the Edge and Cloud can be performed through containerization. However, conditions in Edge nodes are very different from Cloud, as environmental factors such as temperature, humidity, voltage provisioning or dust, heavily affect execution performance and device health. Reliability is an important factor when deciding to deploy such load onto Edge devices, indicating their capability to achieve a desired Quality of Service (QoS). For this, research on Edge computing must focus on how to monitor, estimate and manage devices, in a distributed, autonomous and reliable manner. Our current efforts on performance analysis for devices under "wild conditions" are moving towards integrating reliability into orchestrator systems. Here we present our vision and roadmap for expanding Cloud orchestration towards the Edge with technologies capable of providing knowledge about node environmental conditions and mitigate its impact. Through Out-Of-Band telemetry, we can retrieve temperature and power consumption variables from node components, indicating its health and estimating its reliability given external stress factors. In particular, using intent-based orchestration for containerized platforms, such factors can be used for enforcing reliability as a key-performance indicator. The current work in progress focuses on industrial and commercial scenarios, e.g., Edge computing for urban mobility, with road-side nodes performing AI-based Video-Analytics, exposed to uncontrolled weather conditions. The principal objective is to achieve an Edge network orchestration that takes into account node health in an automatic manner for reducing operational costs such as energy consumption, device repair and replacement, while maintaining QoS in the Edge.
Josep Lluís Berral, David Aguilera-Luzón, Peini Liu, Ramon Nou, Maria A. Serrano, Angelos Antonopoulos 0001, Javier Santaella Sánchez, Mario José Diván
ICNP3
2025 Coupling Orchestration and DNS for Seamless Service Migration in the Edge-Cloud Continuum
abstract
Modern distributed applications increasingly span an edge-cloud continuum, where services may need to dynamically migrate between far-edge, near-edge, and cloud environments to meet latency, resource, or policy constraints. Ensuring seamless service continuity during such migrations remains a significant challenge, particularly due to delays in DNS resolution and inconsistent client routing. This paper presents a modular DNS-driven orchestration framework designed for Kubernetes-based infrastructures. Our approach integrates service orchestration logic with a shared DNS component to enable fast and autonomous service redirection without relying on public DNS providers or complex service meshes. By coupling orchestrator-triggered migrations with DNS record updates, the system ensures immediate service reachability after migration. An optional analytics and decision engine complements the framework by triggering orchestrator actions based on monitoring insights. We evaluate the DNS reconfiguration performance of our solution compared to traditional ExternalDNS-based architectures, demonstrating significantly lower propagation latency and higher determinism during service migrations across the edge–cloud continuum.
Michail Dalgitsis, Eftychia G. Datsika, Marc Palacín, Peini Liu, Javier Santaella Sánchez, Maria A. Serrano, Angelos Antonopoulos 0001
ICNP4
2025 Mobility Usecase: Intelligent Service Migration in Cloud-Edge Continuum
abstract
As industries increasingly embrace digital and intelligent transformation, enterprises face significant challenges in the containerization upgrades for their Artificial Intelligence(AI) applications and dynamic service migration and management in Cloud-Edge continuum(CEC). This paper presents CloudSkin, an innovative platform designed to realize streamlined, seamless and intelligent service migration in Cloud-Edge Continuum by integrating advanced containerization techniques and AI-driven orchestration capabilities. Our approach, leveraging intelligent algorithms for service migration, can seamlessly transit services between cloud and edge environments, ensuring optimised resource allocation and reducing service latency to assure quality of service(QoS). CloudSkin has been enabled in a Mobility Usecase, empowering Cellnex businesses undergoing digital transformation to achieve higher operational efficiency. The experimental results show that compared to traditional reactive service migration, using intelligent proactive service migration can provide better migration detection, improving F1-Score up to 23.5%, and reducing 28.9% the service running time where the service latency violates SLA.
Peini Liu, Joan Oliveras Torra, Marc Palacín, Michail Dalgitsis, Maria A. Serrano, Eftychia G. Datsika, Angelos Antonopoulos 0001, Javier Santaella Sánchez, Jordi Guitart, Josep Lluís Berral, Ramon Nou
ICNP1
2025 Feature Engineering for Agents: An Adaptive Cognitive Architecture for Interpretable ML Monitoring
Gusseppe Bravo Rocca, Peini Liu, Jordi Guitart, Rodrigo M. Carrillo-Larco, Ajay Dholakia, David Ellison
AAMAS2
2024 Data-Connector: An Agent-Based Framework for Autonomous ML-Based Smart Management in Cloud-Edge Continuum
abstract
Machine Learning (ML) is becoming pervasive and integrated into different kinds of intelligent applications, and the collaborative Cloud-Edge continuum has been introduced as an emerging trend to support their adoption into use cases. However, managing these ML applications in the CloudEdge continuum is challenging due to the ML application's dynamic resource usage with different user loads and Cloud and Edge's dynamic resource availability. We envision machine learning methods that can be used for smart management in this dynamic environment, but how to deploy and utilize them for the adaptation scenario in Cloud-Edge continuum is unknown. This paper proposes an agent-based framework to enable autonomous smart management mechanisms that can be broadly enabled in diverse adaptation scenarios. The agent acts as a data-connector11https://github.com/bsc-scanflow/data-connector, connecting different sources of data, utilizing ML models for decision-making and triggering adaptations in Cloud-edge platforms. The case study shows the feasibility of our proposed data-connector for smart migration of an ML workload in the Cloud-edge continuum. The result shows that the smart migration-enabled Cloud-edge scenario has 11.9% ML application prediction time better than the Cloud scenario without migration. Moreover, with minimal customization, the data connector agent can be adapted for more use cases.
Peini Liu, Joan Oliveras Torra, Marc Palacín, Jordi Guitart, Josep Lluís Berral, Ramon Nou
ICNP1
2024 TADIL: Task-Agnostic Domain-Incremental Learning Through Task-ID Inference Using Transformer Nearest-Centroid Embeddings
Gusseppe Bravo Rocca, Peini Liu, Jordi Guitart, Ajay Dholakia, David Ellison
ICPR (29)2
2023 Performance Characterization of Multi-Container Deployment Schemes for Online Learning Inference
abstract
Online machine learning (ML) inference services provide users with an interactive way to request for predictions in realtime. To meet the notable computational requirements of such services, they are increasingly being deployed in the Cloud. In this context, the efficient provisioning and optimization of ML inference services in the Cloud is critical to achieve the required performance and meet the dynamic queries by end-users. Existing provisioning solutions focus on framework parameter tuning and infrastructure resources scaling, without considering deployments based on containerization technologies. The latter promises reproducibility and portability features for ML inferences services. There is limited knowledge about the impact of distinct deployment schemes at the container-level on the performance of online ML inference services, particularly on how to exploit multi-container deployments and its relation with processor and memory affinity. In light of this, in this paper we investigate experimentally the containerization of ML inference services and analyze the performance of multi-container deployments that partition the threads belonging to an online learning application into multiple containers in each node. This paper shares the findings and lessons learned from conducting realistic client patterns on an image classification model across numerous deployment configurations, especially including the impact of container granularity and its potential to exploit processor and memory affinity. Our results indicate that fine-grained multi-container deployments and affinity are useful for improving performance (both throughput and latency). In particular, our experiments on single-node and four-node clusters show up to 69% and 87% performance improvement compared to the single-container deployment, respectively.
Peini Liu, Jordi Guitart, Amirhosein Taherkordi
CLOUD1
2022 Scanflow-K8s: Agent-based Framework for Autonomic Management and Supervision of ML Workflows in Kubernetes Clusters
abstract
Machine Learning (ML) projects are currently heavily based on workflows composed of some reproducible steps and executed as containerized pipelines to build or deploy ML models efficiently because of the flexibility, portability, and fast delivery they provide to the ML life-cycle. However, deployed models need to be watched and constantly managed, supervised, and debugged to guarantee their availability, validity, and robustness in unexpected situations. Therefore, containerized ML workflows would benefit from leveraging flexible and diverse autonomic capabilities. This work presents an architecture for autonomic ML workflows with abilities for multi-layered control, based on an agent-based approach that enables autonomic management and supervision of ML workflows at the application layer and the infrastructure layer (by collaborating with the orchestrator). We redesign the Scanflow ML framework to support such multi-agent approach by using triggers, primitives, and strategies. We also implement a practical platform, so-called Scanflow-K8s, that enables autonomic ML workflows on Kubernetes clusters based on the Scanflow agents. MNIST image classification and MLPerf ImageNet classification benchmarks are used as case studies to show the capabilities of Scanflow-K8s under different scenarios. The experimental results demonstrate the feasibility and effectiveness of our proposed agent approach and the Scanflow-K8s platform for the autonomic management of ML workflows in Kubernetes clusters at multiple layers.
Peini Liu, Gusseppe Bravo Rocca, Jordi Guitart, Ajay Dholakia, David Ellison, Miro Hodak
CCGRID1
2022 Human-in-the-loop online multi-agent approach to increase trustworthiness in ML models through trust scores and data augmentation
abstract
Increasing a ML model accuracy is not enough, we must also increase its trustworthiness. This is an important step for building resilient AI systems for safety-critical applications such as automotive, finance, and healthcare. For that purpose, we propose a multi-agent system that combines both machine and human agents. In this system, a checker agent calculates a trust score of each instance (which penalizes overconfidence in predictions) using an agreement-based method and ranks it; then an improver agent filters the anomalous instances based on a human rule-based procedure (which is considered safe), gets the human labels, applies geometric data augmentation, and retrains with the augmented data using transfer learning. We evaluate the system on corrupted versions of the MNIST and FashionMNIST datasets. We get an improvement in accuracy and trust score with just few additional labels compared to a baseline approach.
Gusseppe Bravo Rocca, Peini Liu, Jordi Guitart, Ajay Dholakia, David Ellison, Miro Hodak
COMPSAC2
2022 Scanflow: A multi-graph framework for Machine Learning workflow management, supervision, and debuggin
abstract
Machine Learning (ML) is more than just training models, the whole workflow must be considered. Once deployed, a ML model needs to be watched and constantly supervised and debugged to guarantee its validity and robustness in unexpected situations. Debugging in ML aims to identify (and address) the model weaknesses in not trivial contexts. Several techniques have been proposed to identify different types of model weaknesses, such as bias in classification, model decay, adversarial attacks, etc., yet there is not a generic framework that allows them to work in a collaborative, modular, portable, iterative way and, more importantly, flexible enough to allow both human- and machine-driven techniques. In this paper, we propose a novel containerized directed graph framework to support and accelerate end-to-end ML workflow management, supervision, and debugging. The framework allows defining and deploying ML workflows in containers, tracking their metadata, checking their behavior in production, and improving the models by using both learned and human-provided knowledge. We demonstrate these capabilities by integrating in the framework two hybrid systems to detect data drift distribution which identify the samples that are far from the latent space of the original distribution, ask for human intervention, and whether retrain the model or wrap it with a filter to remove the noise of corrupted data at inference time. We test these systems on MNIST-C, CIFAR-10-C, and FashionMNIST-C datasets, obtaining promising accuracy results with the help of human involvement.
Gusseppe Bravo Rocca, Peini Liu, Jordi Guitart, Ajay Dholakia, David Ellison, Jeffrey Falkanger, Miro Hodak
Expert Syst. Appl.2
2021 Performance comparison of multi-container deployment schemes for HPC workloads: an empirical study
Peini Liu, Jordi Guitart
J. Supercomput.1
2019 A field-based service management and discovery method in multiple clouds context
Shuai Zhang 0048, Xinjun Mao, Fu Hou, Peini Liu
Frontiers Comput. Sci.4
2018 Towards Reference Architecture for a Multi-layer Controlled Self-adaptive Microservice System
abstract
With the features of high distribution in deployment and independence in running, the microservice systems that operate in heterogeneous infrastructures and open Internet environment are expected to be self-adaptive to adapt to various changes of both operating contexts and application requirements.This requires the adaptability of the microservice systems to be diverse and flexible, and independent of implementation technologies and platforms.This paper presents a reference architecture for self-adaptive microservice systems with the abilities of multi-layer controlled self-adaptations, including infrastructure-controlled layer and application-controlled layer.Such reference architecture presents a blueprint to cope with diverse changes from different levels in microservice systems and supports the interactions between layers.We have implemented a practical platform called SAMSP based on the reference architecture and Kubernetes and evaluated our approach using a sample.The experimental results are promising, and demonstrate the feasibility and effectiveness of our proposed reference architecture.
Peini Liu, Xinjun Mao, Shuai Zhang 0048, Fu Hou
SEKE1
2018 A Self-Adaptation Framework of Microservice Systems (S)
abstract
Microservice has been more and more applied to build software systems in industry field and research in academic field.And software systems are increasingly expected to dynamically self-adapt to accommodate resource variability, changing user needs, and system faults.Compared with traditional software systems, microservice systems have some characteristics, such as the highly self-contained components and the dynamic running instance, which pose challenges to traditional self-adaptation methods.Therefore, it needs to propose corresponding techniques and methods to cope with the characteristics of microservice systems.This paper analyzes the special selfadaptive requirements of microservice systems, proposes a microservice reference model, which describes basic elements and their relationships of microservice systems.Then we present a microservice system self-adaptation framework MSSAF to support the self-adaptation of microservice systems.We illustrate the feasibility and effectiveness of our approach in the context of a microservice system case.
Shuai Zhang 0048, Xinjun Mao, Peini Liu, Fu Hou
SEKE3
2017 Mapping client messages to a unified data model with mixture feature embedding convolutional neural network
abstract
Data mapping among different data standards in health institutes is often a necessity when data exchanges occur among different institutes. However, no matter rule-based approaches or traditional machine learning methods, none of these methods have achieved satisfactory results yet. In this work, we propose a deep learning method, mixture feature embedding convolutional neural network (MfeCNN), to convert the data mapping to a multiple classification problem. Multi-modal features were extracted from different semantic space with a medical NLP package and powerful feature embeddings were generated by MfeCNN. Classes as many as ten were classified simultaneously by a fully-connected soft-max layer based on multi-view embedding. Experimental results show that our proposed MfeCNN achieved best results than traditional state-of-the-art machine learning models and also much better results than the convolutional neural network of only using bag-of-words as inputs.
Dingcheng Li, Peini Liu, Ming Huang 0006, Yu Gu 0001, Daniel Dean, Jingmin Xu, Hui Lei 0001, Yaoping Ruan
BIBM2
2013 A Bayesian Network-Based Knowledge Engineering Framework for IT Service Management
abstract
Service management is becoming more and more important within the area of IT management. How to efficiently manage and organize service in complicated IT service environments with frequent changes is a challenging issue. IT service and the related information from different sources are characterized as diverse, incomplete, heterogeneous, and geographically distributed. It is hard to consume these complicated services without knowledge assistant. To address this problem, a systematic way (with proposed toolsets and process) is proposed to tackle the challenges of acquisition, structuring, and refinement of structured knowledge. An integrated knowledge process is developed to guarantee the whole engineering procedure which utilizes Bayesian networks (BNs) as the knowledge model. This framework can be successfully applied on key tasks in service management, such as problem determination and change impact analysis, and a real example of Cisco VoIP system is introduced to show the usefulness of this method.
Wei Wang 0033, Hao Wang 0208, Bo Yang 0013, Liang Liu 0010, Peini Liu, Guosun Zeng
IEEE Trans. Serv. Comput.5
2008 A Bayesian knowledge engineering framework for service management
abstract
Service management is becoming more and more important within the area of IT service management. How to efficiently manage and organize service in complicated IT environments with frequent changes is a challenging issue. Service and the related information from different sources are characterized as diverse, incomplete, heterogeneous, and geographically distributed. It is hard to consume these complicated data without knowledge assistant. To address this problem, a knowledge engineering framework is proposed to tackle the challenges of acquisition, structuring and refinement of structured knowledge regarding existing different unstructured information resource, and the Bayesian network is utilized as the knowledge model. This framework can be successfully applied on key tasks in service management, such as problem determination and change impact analysis. And a real example of Cisco VoIP system is introduced to show the usefulness of this method.
Wei Wang 0033, Hao Wang 0208, Bo Yang 0013, Liang Liu 0010, Peini Liu, Guosun Zeng
NOMS5
2004 Predicting the intrusion intentions by observing system call sequences
Xiaohong Guan, Sangang Guo, Peini Liu
Comput. Secur.5