EDBT 2026 Demo / reviewers in the wild / expert
Yang Bai 0007
dblp:39/6825-7
· DBLP profile ↗
11ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0001-6292-3730ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PaLLOC: Pairwise-based low-latency online coordinated resource manager of last-level cache and memory bandwidth on multicore systems
Yang Bai 0007, Renfa Li |
J. Syst. Archit. | 1 |
| 2023 | UMA-MF: A Unified Multi-CPU/GPU Asynchronous Computing Framework for SGD-Based Matrix FactorizationabstractRecent research has shown that collaborative computing of CPUs and GPUs in the same system can effectively accelerate large-scale SGD-based matrix factorization (MF), but it faces the problem of limited scalability due to parameter synchronization in the server. Theoretically, asynchronous methods can overcome this shortcoming. However, through a series of tests, observations, and analyses, we realize that developing an effective asynchronous multi-CPU/GPU MF framework faces several major design challenges: the underutilized CPUs, high communication overhead, and the asynchronous data safety issue. This article presents a unified multi-CPU/GPU asynchronous computing framework for SGD-based matrix factorization, namedUMA-MF.UMA-MFtreats CPUs and GPUs in the system as distributed workers that train matrix datasets in parallel and update feature parameters asynchronously. It provides a cache-friendly CPU external working mode, which can improve the CPU's cache hit rate, thereby promoting the efficient use of CPUs. It offers an algorithm to find the shortest communication ring topology of heterogeneous CPU/GPU workers and builds computing-communication pipelines to minimize the communication overhead. It implements a wait-free structure and load-balanced data distribution to achieve asynchronous data safety.UMA-MFcan effectively accelerate SGD-based MF on multi-CPU/GPU systems in an asynchronous way. On a physical platform with configurations ranging from single processor system to 2CPUs--4CPUs system, for five common datasets Netfix, R1, R2, Goodreads, and de-dense,UMA-MFachieves up to 3.56x speedup compared with HCC-MF, which is the state-of-the-art multi-CPU/GPU synchronous computing framework for SGD-based MF.UMA-MFalso shows good scalability. When the system is scaled to 2CPUs-4GPUs, the training time speedup ofUMA-MFcan reach 70%--97% of the ideal speedup. Yan Liu 0032, Yang Bai 0007, Renfa Li |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2021 | A Novel Multi-CPU/GPU Collaborative Computing Framework for SGD-based Matrix FactorizationabstractThis paper presents a heterogeneous collaborative computing framework for SGD-based Matrix Factorization, named HCC-MF. HCC-MF can train the feature matrix efficiently using multiple CPUs and GPUs. It performs collaborative computing with data parallelism, where a server CPU is in charge of management and synchronization and other heterogeneous worker CPUs and worker GPUs performs calculation with their data assignments. HCC-MF adopts two data partition strategies, “data partition with heterogeneous load balance” and “data partition with hidden synchronization.” We build a time cost model to guide the data distribution among multiple workers and we design several communication optimization techniques with consideration of datasets’ and processors’ characteristics. Experimental results indicate that HCC-MF can utilize more than 88% of the platform’s computing power, yielding a speedup of 2.9 compared with advanced SGD-based MF, CuMF_SGD, on large-scale data sets. Yanlong Yin, Yan Liu 0032, Shuibing He, Yang Bai 0007, Renfa Li |
ICPP | 5 |
| 2021 | ASDYS: Dynamic Scheduling Using Active Strategies for Multifunctional Mixed-Criticality Cyber-Physical SystemsabstractEmerging cyber-physical systems (CPSs), such as in the domains of automotive, robotics, and industrial automation, often run complex functions with different criticality levels on a heterogeneous and distributed architecture. The ever stronger interactions between the cyber components and the physical environment lead to dynamic and irregular release of these functions. This article investigates dynamic scheduling of such mixed-criticality functions, where each function is modeled by a directed acyclic graph with no assumption on its period or minimum interarrival time. Unlike the existing methods that passively address the mixed criticality with a remedy when deadline misses are observed-this results in a high deadline miss ratio (DMR), and it is particularly undesirable for the high-criticality functions-we propose a novel dynamic scheduling approach using active strategies (ASDYS in short), where the mixed criticality is actively treated throughout the scheduling process. Automotive CPSs are used as an example for illustration. Experimental results show that our approach is significantly better than the existing methods in both the DMR of high-criticality functions and the overall system DMR. Yang Bai 0007, Guoqi Xie, Renfa Li, Wanli Chang 0001 |
IEEE Trans. Ind. Informatics | 1 |
| 2020 | A Survey of Intrusion Detection for In-Vehicle NetworksabstractThe development of the complexity and connectivity of modern automobiles has caused a massive rise in the security risks of in-vehicle networks (IVNs). Nevertheless, existing IVN designs (e.g., controller area network) lack cybersecurity consideration. Intrusion detection, an effective method for defending against cyberattacks on IVNs while providing functional safety and real-time communication guarantees, aims to address this issue. Therefore, the necessity of its research has risen. In this paper, an IVN environment is introduced, and the constraints and characteristics of an intrusion detection system (IDS) design for IVNs are presented. A survey of the proposed IDS designs for the IVNs is conducted, and the corresponding drawbacks are highlighted. Various optimization objectives are considered and comprehensively compared. Lastly, the trend, open issues, and emerging research directions are described. Wufei Wu, Renfa Li, Guoqi Xie, Ji-yao An, Yang Bai 0007, Jia Zhou 0003, Keqin Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2020 | Minimizing Redundancy to Satisfy Reliability Requirement for a Parallel Application on Heterogeneous Service-Oriented SystemsabstractReliability is widely identified as an increasingly relevant issue in heterogeneous service-oriented systems because processor failure affects the quality of service to users. Replication-based fault-tolerance is a common approach to satisfy application's reliability requirement. This study solves the problem of minimizing redundancy to satisfy reliability requirement for a directed acyclic graph (DAG)-based parallel application on heterogeneous service-oriented systems. We first propose the enough replication for redundancy minimization (ERRM) algorithm to satisfy application's reliability requirement, and then propose heuristic replication for redundancy minimization (HRRM) to satisfy application's reliability requirement with low time complexity. Experimental results on real and randomly generated parallel applications at different scales, parallelism, and heterogeneity verify that ERRM can generate least redundancy followed by HRRM, and the state-of-the-art MaxRe and RR algorithm. In addition, HRRM implements approximate minimum redundancy with a short computation time. Guoqi Xie, Yuekun Chen, Yang Bai 0007, Zhili Zhou 0001, Renfa Li, Keqin Li 0001 |
IEEE Trans. Serv. Comput. | 4 |
| 2020 | Quantitative Modeling and Analytical Calculation of Anelasticity for a Cyber-Physical SystemabstractThis paper investigates resource provisioning in cyber-physical systems (CPSs) by developing a new definition of anelasticity. A flat semi-dormant multicontroller (FSDMC) model is established on a special type of CPS platform named arbitrated networked control system with dual communication channels. A novel, quantitative, and formal definition of anelasticity for the FSDMC is proposed. A new finite capacity M/M/c queuing system with N-policy and asynchronous multiple working vacations of partial servers is established, and the FSDMC is modeled as a quasi-birth-and-death process to obtain the stationary probability distribution of the system. Based on the queueing model, we quantify various performance indices of the system to build a nonlinear cost-performance ratio (CPR) function. An optimization model is presented to minimize the CPR. A particle swarm optimization (PSO) algorithm is used to find the optimum solution of the optimization model and obtain the optimal configuration values of the system parameters under stability condition. By changing the system parameters, the sensitivity of the system performance indices and the CPR are analyzed, respectively. The unexpected workload varies randomly over time. Thus, an M/M/1/K queue is constructed in a Markovian environment by employing a three-state, irreducible Markov process. In this queue, the conditional average queue length and the probabilities of the three-state process are calculated. Then, the anelasticity value of the system is precisely determined. When the average arrival rate exceeds the average service rate in the queueing system, an optimal CPR unchanged adaptive algorithm based on PSO is designed to dynamically adjust the controller service rate. Extensive numerical results show the usefulness and effectiveness of the proposed techniques and exhibit that the system can maintain elastic invariance in adaptive adjustment parameters. Hongfang Gong, Renfa Li, Ji-yao An, Yang Bai 0007, Keqin Li 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2019 | Optimal power allocation and load balancing for non-dedicated heterogeneous distributed embedded computing systems
Jing Huang 0012, Yan Liu 0032, Renfa Li, Keqin Li 0001, Ji-yao An, Yang Bai 0007, Fan Yang 0044, Guoqi Xie |
J. Parallel Distributed Comput. | 6 |
| 2019 | Human-Interaction-aware Adaptive Functional Safety Processing for Multi-Functional Automotive Cyber-Physical SystemsabstractThe functional safety research for automotive cyber-physical systems (ACPS) has been studied in recent years; however, these studies merely consider the change in the exposure of the functional safety classification and assume that the driver’s controllability in the functional safety classification is always fixed and uncontrollable. In fact, the driver’s controllability is variable during the runtime phase, such that the execution process of safety-critical automotive functions is a human-interaction-aware process between the driver and ACPS. To adapt to the changes in the driver’s controllability, this article studies the human-interaction-aware adaptive functional safety processing for multi-functional ACPS in two main phases. In the design phase, where the driver’s controllability is fixed at the highest level (i.e., C3), we obtain the approximate optimal priority sequence of safety-critical functions without exhausting all sequences by proposing the refined exploration method. In the runtime phase, where the driver’s controllability level is variable (i.e., C0, C1, C2, or C3), we propose the human-interaction-aware task remapping method to autonomously respond to the change of the driver’s controllability. Examples and experiments confirm that the proposed adaptive functional safety processing can reduce overall task redundancy of safety-critical automotive functions while meeting their functional safety requirements, shorten the overall response time of safety-critical automotive functions, and increase the slack time for non-safety-critical automotive functions. Guoqi Xie, Yang Bai 0007, Yanwen Li, Renfa Li, Keqin Li 0001 |
ACM Trans. Cyber Phys. Syst. | 2 |
| 2018 | Message response time analysis for automotive cyber-physicalsystems with uncertain delay: An M/PH/1 queue approach
Hongfang Gong, Renfa Li, Yang Bai 0007, Ji-yao An, Keqin Li 0001 |
Perform. Evaluation | 3 |
| 2017 | Efficient task scheduling for budget constrained parallel applications on heterogeneous cloud computing systems
Guoqi Xie, Renfa Li, Yang Bai 0007, Chunnian Fan, Keqin Li 0001 |
Future Gener. Comput. Syst. | 4 |