EDBT 2026 Demo / reviewers in the wild / expert
Haibo Mi
dblp:65/8331
· DBLP profile ↗
25ranked-venue papers
7as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | When Server Joins INA: The Resource-Aware Repair Acceleration for Erasure-Coded Storage Systems
Geyao Cheng, Junxu Xia, Hao Fan 0006, Fengzeng Liu, Haibo Mi, Deke Guo |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2026 | Puffer: A Serverless Platform Based on Vertical Memory ScalingabstractThis paper quantitatively analyses the potential of vertical scaling MicroVMs in serverless computing. Our analysis shows that under real-world serverless workloads, vertical scaling can significantly improve execution performance and resource utilization. However, we also find that the memory scaling of MicroVMs is the bottleneck that hinders vertical scaling from reaching the performance ceiling. We propose Faascale, a novel mechanism that efficiently scales the memory of MicroVMs for serverless applications. Faascale employs a series of techniques to tackle this bottleneck: 1) it sizes up/down the memory for a MicroVM by blocks that bind with a function instance instead of general pages; and 2) it pre-populates physical memory for function instances to reduce the delays introduced by the lazy-population. Compared with existing memory scaling mechanisms, Faascale improves the memory scaling efficiency by 2 to 3 orders of magnitude. Based on Faascale, we realize a serverless platform, named Puffer. Experiments conducted on eight serverless benchmark functions demonstrate that compared with horizontal scaling strategies, Puffer reduces time for cold-starting MicroVMs by 89.01%, improves memory utilization by 17.66%, and decreases functions execution time by 23.93% on average. Hao Fan 0006, Kun Wang 0059, Haibo Mi, Song Wu 0001, Chen Yu 0003 |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2026 | A Non-Intrusive Multi-Objective Task Scheduling Method for JointCloud EnvironmentabstractThe advent of advanced technologies, such as large models, has precipitated a surging demand for computational resources, thereby driving the evolution of cloud computing services from single-cloud architectures to multi-cloud paradigms. The accompanying question is how to enable resources from different cloud service providers to collaborate efficiently. Although existing works have explored multi-cloud scheduling methods, most of these works are centralized scheduling, where the decision-maker can schedule computing resources across all clouds. However, this is nearly impossible given the current situation in which computing resources in different clouds come from various cloud service providers. In response to this challenge, JointCloud, a multi-cloud cooperation architecture, has been proposed, which aims at enhancing the cooperation among multiple clouds to provide efficient multi-cluster services. Following the idea of JointCloud, proposes a multi-objective evolutionary algorithm (MOEA) based method for task scheduling in multi-cluster cloud computing environments, without intervening in the intra-cluster scheduling scheme. In the proposed method, we construct a mathematical model with the optimization objectives of minimizing overall waiting time and load imbalance between clusters based on the actual operation data in China Computing NET (C$^{2}$NET). In addition, we also develop an MOEA specifically tailored to address this problem. The performance of the proposed MOEA and existing state-of-the-art MOEAs is examined on the proposed problems. Comparison results highlight the promising performance of the proposed MOEA, the specifically tailored algorithm in effectively addressing the multi-cluster task scheduling problem. In addition, we also compared the results of the MOEAs with the results of three classical scheduling methods, the results proved the effectiveness of the MOEA-based method on this problem. Lianghao Li, Haibo Mi, Bo Ding 0001, Huaimin Wang 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2026 | Efficient Cluster-Based Knowledge Distillation for Deep Face RecognitionabstractKnowledge distillation has been widely used to improve the performance of small compact models for face recognition. However, selecting key knowledge and effectively transferring it from teacher to student remains a challenging problem. In this work, we propose an efficient Cluster-based Knowledge Distillation (CKD) dedicated to aligning the student model with the teacher model in terms of both sample relations and class centers. Specifically, CKD first determines the key sample relations based on the similarities between the sample features extracted by the teacher and their cluster centers generated by existing clustering algorithms. Then, CKD effectively transfers the knowledge of the above relations from the teacher to the student by designing a cluster-based relation distillation loss. Finally, CKD further improves the quality of the student's class centers by constructing a center loss between the above representative cluster centers and the student's class centers. We validate the proposed CKD on multiple face benchmarks. For example, CKD improves the baseline student performance from 91.95% to 94.20% on MegaFace and consistently outperforms recent competitive distillation methods on multiple benchmarks. These results demonstrate the effectiveness and superiority of CKD. Xianwei Lv 0001, Haibo Mi, Xin Niu 0001, Wang Chen 0004, Kun Wang 0059, Chen Yu 0003 |
IEEE Trans. Sustain. Comput. | 2 |
| 2025 | MS2SA: A Federated Learning Method with a Dynamic Straggler Processing Strategy
Zhiyu Xie 0005, Haibo Mi |
ICIC (9) | 4 |
| 2025 | SINA: A Server-Assisted In-Network Repair Acceleration for Erasure-Coded Storage SystemsabstractIn the erasure-coded storage systems, multiple related blocks have to be retrieved from other surviving nodes to repair a failed block. This incurs significant communication overhead with the surging scale of distributed storage systems. To mitigate the bandwidth bottleneck, in-network repair (INR) has emerged as a promising transport paradigm, which migrates the aggregation operations from the repair node to the programmable hardware, such as Intel Tofino switches. However, due to the limited on-chip memory size of these switches, the INR can degrade to the most primitive incast-type transmission, leading to massive traffic volume and hindered repair throughput. While we notice that, there are spare CPU cores in the storage servers that can be leveraged as alternative computing resources. With this intuition, we propose SINA, a Server-assisted In-Network repair Acceleration framework in this paper, which leverages the spare servers to assist aggregation operations when the programming switches' memory size is scarce for failure repair. We formulate this problem by adjusting the aggregation modes across the involved racks and solve this NP-hard problem using the Gurobi optimization solver. For all we know, this is the first work exploring spare servers for assisting the memory-scarce INR in erasure-coded storage systems. We have implemented SINA on an FPGA-based prototype system, and the experimental results show that SINA can ensure fault tolerance and accelerate failure repair by$5.0 \times$compared to the conventional methods. Geyao Cheng, Junxu Xia, Haibo Mi, Deke Guo, Kun Wang 0059 |
IWQoS | 3 |
| 2025 | A Review of Multi-Objective Optimization for Cloud Environment Storage OptimizationabstractThe advent of cloud computing offers a novel mode of managing computing resources, enabling users to flexibly utilize required computing and storage resources through the network. Simultaneously, it allows large-scale computing centers and data centers to more effectively utilize their computing resources. With the rapid development of cloud computing technology, the exponential growth of massive data storage needs from numerous users has brought challenges in storage cost, performance, reliability, and security. Conventional single-objective optimization approaches, which concentrate exclusively on a singular performance metric, are increasingly inadequate to address the intricate requirements of cloud storage systems. In contrast, multi-objective optimization methods can simultaneously optimize multiple aspects of concern to users or operators, providing more comprehensive solutions for cloud computing services. In this paper, we first introduce the basic concepts of multi-objective optimization problems. Subsequently, we present a comprehensive review of multi-objective optimization applications in cloud storage systems, including the setting of optimization objectives, algorithm selection, and comparison methods of experimental results. Finally, this paper summarizes and discusses the current implementation of multi-objective evolutionary algorithms in cloud storage optimization, and provides an outlook on future research directions. Lianghao Li, Haibo Mi, Kun Wang 0059, Bo Ding 0001, Huaimin Wang 0001 |
JCC | 2 |
| 2025 | HyperPart: A Hypergraph-Based Abstraction for Deduplicated Storage SystemsabstractCurrently, deduplication techniques are utilized to minimize the space overhead by deleting redundant data blocks across large-scale servers in data centers. However, such a process exacerbates the fragmentation of data blocks, causing more cross-server file retrievals with plummeting retrieval throughput. Some attempts prefer better file retrieval performance by confining all blocks of a file to one single server, resulting in non-trivial space consumption for more replicated blocks across servers. An ideal network storage system, in effect, should take both the deduplication and retrieval performance into account by implementing reasonable assignment of the detected unique blocks. Such a fine-grained assignment requires an accurate and comprehensive abstraction of the files, blocks, and the file-block affiliation relationships. To achieve this, we innovatively design the weighted hypergraph to profile the multivariate data correlations. With this delicate abstraction in place, we propose HyperPart, which elegantly transforms this complex block allocation problem into a hypergraph partition problem. For more general scenarios with dynamic file updates, we further propose a two-phase incremental hypergraph repartition scheme, which mitigates the performance degradation with minimal migration volume. We implement a prototype system of HyperPart, and the experiment results validate that it saves around 50% of the storage space and improves the retrieval throughput by approximately 30% of state-of-the-art methods under the balance constraints. Geyao Cheng, Junxu Xia, Lailong Luo, Haibo Mi, Deke Guo, Richard T. B. Ma |
IEEE Trans. Cloud Comput. | 4 |
| 2024 | M3ixTS: Mixing of Multi-patch and Multi-view For Time Series Forecasting
Lianghao Li, Haibo Mi, Bo Ding 0001 |
ICONIP (3) | 3 |
| 2023 | WMWatcher: Preventing Workload-Related Misconfigurations in Production EnvironmentabstractAmong the misconfigurations with increasing preva-lence and severity in recent years, workload-related misconfigu-rations, i.e. misconfigurations under certain workloads with valid configuration values, account for a significant portion. Since the runtime constraints of configuration parameters are influenced by workloads, piror researches could not handle workload-related misconfigurations at present. To solve the situation mentioned above, we conducted an empirical study on how configuration variables interact with other program variables, and summarized five handling type of the interactions happen in branch statements. Based on the study, we proposed WMWatcher to help system admins to prevent workload-related misconfigurations in production environment. WMWatcher infers the runtime constraints of configuration parameters under certain workload by instrumenting probes in source code and monitoring the corresponding status. The experiments on seven open-source software systems proved that WMWatcher could automatically instrument proper probes while bringing only 2.33% extra runtime overhead at most. And the case study demonstrates the effectiveness of WMWatcher in preventing workload-related misconfigurations in real-world scenarios. Shulin Zhou, Zhijie Jiang, Shanshan Li 0001, Xiaodong Liu 0004, Zhouyang Jia, Yuanliang Zhang, Jun Ma 0015, Haibo Mi |
APSEC | 8 |
| 2023 | FedH2L: A Federated Learning Approach with Model and Statistical HeterogeneityabstractFederated learning (FL) enables distributed participants to collectively learn a strong global model without sacrificing their individual data privacy. Mainstream FL approaches require each participant to share a common network architecture and further assume that data are sampled IID across participants. However, in real-world deployments, participants may require heterogeneous network architectures; and the data distribution is almost non-uniform. To address these issues we introduce FedH2L, which is agnostic to the model architecture and robust to different data distributions across participants. In contrast to approaches sharing parameters or gradients, FedH2L relies on mutual distillation, exchanging only posteriors on a shared seed set between participants in a decentralized manner. This makes it extremely bandwidth efficient, model agnostic, and crucially produces models capable of performing well on the whole data distribution when learning from heterogeneous silos. Yiying Li, Haibo Mi, Huaimin Wang 0001 |
JCC | 3 |
| 2023 | Bi-level Multi-Agent Actor-Critic Methods with ransformersabstractRecently, deep multi-agent reinforcement learning methods have witnessed great progress, including multi-agent actor-critic methods. However, it’s worth noticing there is a performance gap between multi-agent actor-critic methods and state-of-the-art value-based methods. In this paper, we investigate the causes and attribute inferior performance to issues of contribution-mismatch and indiscriminate guidance. To overcome these problems, we introduce a novel bi-level multi-agent actorcritic reinforcement learning approach with transformers, called BMT. Specifically, we propose a simple but efficient bi-level optimization mechanism to learn both global critic and agentspecific critic, thus jointly guiding the policy update. In addition, we adopt the transformer-based model as the policy network to decouple complicated relationships and generate flexible policy. BMT is also general enough to be plugged into any actor-critic multi-agent reinforcement learning approach, such as MAPPO, and equips it with strong expression. On multiple benchmarks including multi-agent particle environments and a challenging set of StarCraft II micromanagement tasks, large-scale empirical experiments demonstrate that BMT-based multi-agent reinforcement learning methods achieve superior performance over both state-of-the-art actor-critic and value-based approaches. Tianjiao Wan, Haibo Mi, Zijian Gao, Yuanzhao Zhai, Bo Ding 0001 |
JCC | 2 |
| 2021 | OADA: An Online Data Augmentation Method for Raw Histopathology Images
Zhiyue Wu, Yijie Wang 0001, Haibo Mi, Hongzuo Xu, Lanlan Feng |
ICONIP (6) | 3 |
| 2020 | A Distributed Event Extraction Framework for Large-Scale Unstructured TextabstractEvent extraction is an important subtask of information extraction. The goal of event extraction is to quickly extract events of a specified type from a large amount of textual information. Many excellent models and algorithms have been proposed since ACE released the event extraction task in 2005. Most of them are based on the dataset published by ACE and have contributed to the accuracy of event extraction to a certain extent. In practical applications, the processing object of the event extraction task is large-scale text data. However, as far as we know, there is currently no effective model for using multiple computers for event extraction. In this paper, we propose a model for event extraction based on inter-cloud computing technology. The experimental results prove that our method reduces the time consumption and also gets better accuracy than advanced models. Zhigang Kan, Haibo Mi, Sen Yang 0003, Linbo Qiao, Dongsheng Li 0001 |
JCC | 2 |
| 2020 | Collaborative deep learning across multiple data centers
Haibo Mi, Kele Xu, Huaimin Wang 0001, Yiming Zhang 0003, Zibin Zheng, Chuan Chen 0001, Xu Lan |
Sci. China Inf. Sci. | 1 |
| 2019 | Denoising Convolutional Autoencoder Based B-mode Ultrasound Tongue Image Feature ExtractionabstractB-mode ultrasound tongue imaging is widely used in the speech production field. However, efficient interpretation is in a great need for the tongue image sequences. Inspired by the recent success of unsupervised deep learning approach, we explore unsupervised convolutional network architecture for the feature extraction in the ultrasound tongue image, which can be helpful for the clinical linguist and phonetics. By quantitative comparison between different unsupervised feature extraction approaches, the denoising convolutional autoencoder (DCAE)-based method outperforms the other feature extraction methods on the reconstruction task and the 2010 silent speech interface challenge. A Word Error Rate of 6.17% is obtained with DCAE, compared to the state-of-the-art value of 6.45% using Discrete cosine transform as the feature extractor. Our codes are available at https://github.com/DeePBluE666/Source-code1. Bo Li 0005, Kele Xu, Haibo Mi, Huaimin Wang 0001 |
ICASSP | 4 |
| 2019 | A Mobile Application for Sound Event DetectionabstractSound event detection is intended to analyze and recognize the sound events in audio streams and it has widespread applications in real life. Recently, deep neural networks such as convolutional recurrent neural networks have shown state-of-the-art performance in this task. However, the previous methods were designed and implemented on devices with rich computing resources, and there are few applications on mobile devices. This paper focuses on the solution on the mobile platform for sound event detection. The architecture of the solution includes offline training and online detection. During offline training process, multi model-based distillation method is used to compress model to enable real-time detection. The online detection process includes acquisition of sensor data, processing of audio signals, and detecting and recording of sound events. Finally, we implement an application on the mobile device that can detect sound events in near real time. Yingwei Fu, Kele Xu, Haibo Mi, Huaimin Wang 0001, Boqing Zhu |
IJCAI | 3 |
| 2019 | A Quantitative Analysis Platform for PD-L1 Immunohistochemistry based on Point-level Supervision ModelabstractRecently, deep learning has witnessed dramatic progress in the medical image analysis field. In the precise treatment of cancer immunotherapy, the quantitative analysis of PD-L1 immunohistochemistry is of great importance. It is quite common that pathologists manually quantify the cell nuclei. This process is very time-consuming and error-prone. In this paper, we describe the development of a platform for PD-L1 pathological image quantitative analysis using deep learning approaches. As point-level annotations can provide a rough estimate of the object locations and classifications, this platform adopts a point-level supervision model to classify, localize, and count the PD-L1 cells nuclei. Presently, this platform has achieved an accurate quantitative analysis of PD-L1 for two types of carcinoma, and it is deployed in one of the first-class hospitals in China. Haibo Mi, Kele Xu, Yu-Lin He, Huaimin Wang 0001, Yanming Song, Xiaolei Sun |
IJCAI | 1 |
| 2019 | Privacy-Protected Blockchain SystemabstractThe blockchain uses a decentralized consensus mechanism to maintain the books in an immutable way, which ensures the blockchain smart contract system highly secure. In existing blockchain systems, all user information is disclosed in the blockchain. However, currently users begin to pay more and more attention to personal privacy, therefore the future blockchain smart contract system needs not only to keep immutability but also to protect user privacy. To achieve this goal, in this paper we propose a privacy-encrypted blockchain system, where all data is encrypted within a controllable period of time. Although the data is visible from a historical perspective, our design can effectively protect user privacy and against deceivers, making the system more secure and healthy. Ping Zhong 0002, Qikai Zhong, Haibo Mi, Shigeng Zhang |
MDM | 3 |
| 2018 | Sample Dropout for Audio Scene Classification Using Multi-scale Dense Connected Convolutional Neural Network
Kele Xu, Haibo Mi, Feifan Liao |
PKAW | 3 |
| 2013 | An online service-oriented performance profiling tool for cloud computing systems
Haibo Mi, Huaimin Wang 0001, Yangfan Zhou 0002, Michael R. Lyu, Hua Cai, Gang Yin |
Frontiers Comput. Sci. | 1 |
| 2013 | Toward Fine-Grained, Unsupervised, Scalable Performance Diagnosis for Production Cloud Computing SystemsabstractPerformance diagnosis is labor intensive in production cloud computing systems. Such systems typically face many real-world challenges, which the existing diagnosis techniques for such distributed systems cannot effectively solve. An efficient, unsupervised diagnosis tool for locating fine-grained performance anomalies is still lacking in production cloud computing systems. This paper proposes CloudDiag to bridge this gap. Combining a statistical technique and a fast matrix recovery algorithm, CloudDiag can efficiently pinpoint fine-grained causes of the performance problems, which does not require any domain-specific knowledge to the target system. CloudDiag has been applied in a practical production cloud computing systems to diagnose performance problems. We demonstrate the effectiveness of CloudDiag in three real-world case studies. Haibo Mi, Huaimin Wang 0001, Yangfan Zhou 0002, Michael R. Lyu, Hua Cai |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2012 | P-Tracer: Path-Based Performance Profiling in Cloud Computing SystemsabstractIn large-scale cloud computing systems, the growing scale and complexity of component interactions pose great challenges for operators to understand the characteristics of system performance. Performance profiling has long been proved to be an effective approach to performance analysis; however, existing approaches do not consider two new requirements that emerge in cloud computing systems. First, the efficiency of the profiling becomes of critical concern; second, visual analytics should be utilized to make profiling results more readable. To address the above two issues, in this paper, we present P-Tracer, an online performance profiling approach specifically tailored for large-scale cloud computing systems. P-Tracer constructs a specific search engine that adopts a proactive way to process performance logs and generates particular indices for fast queries; furthermore, PTracer provides users with a suite of web-based interfaces to query statistical information of all kinds of services, which helps them quickly and intuitively understand system behavior. The approach has been successfully applied in Alibaba Cloud Computing Inc. to conduct online performance profiling both in production clusters and test clusters. Experience with one real-world case demonstrates that P-Tracer can effectively and efficiently help users conduct performance profiling and localize the primary causes of performance anomalies. Haibo Mi, Huaimin Wang 0001, Hua Cai, Yangfan Zhou 0002, Michael R. Lyu, Zhenbang Chen 0001 |
COMPSAC | 1 |
| 2012 | Performance problems diagnosis in cloud computing systems by mining request trace logsabstractIn cloud computing systems, end-to-end request tracing approach is helpful for developers to understand the runtime behavior of user requests. Based on trace logs, we propose an approach to localize the abnormal methods that are the primary causes of performance problems. Our approach involves three steps: (1) cluster the user requests into different categories according to request call sequences and select major categories; (2) extract the principal methods that might be the causes of performance degradation; (3) pick out abnormal methods from those principal methods in each major category. We conduct four cases of performance degradations to validate our approach over a real-world enterprise-class cloud computing platform. The experimental results show that our approach can locate the prime causes of performance problems with low false-positive rate and false-negative rate. Haibo Mi, Huaimin Wang 0001, Gang Yin, Hua Cai, Tingtao Sun |
NOMS | 1 |
| 2012 | Localizing root causes of performance anomalies in cloud computing systems by analyzing request trace logs
Haibo Mi, Huaimin Wang 0001, Yangfan Zhou 0002, Michael R. Lyu, Hua Cai |
Sci. China Inf. Sci. | 1 |