Yuebin Bai

dblp:22/2375 · DBLP profile ↗
← Back
53ranked-venue papers
8as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 20 · 2 first-author · 6 since 2021Computer networks · 12 · 1 first-author · 7 since 2021Software engineering, systems software and programming languages · 7 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4Security and privacy · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Understanding and Optimizing Database Pushdown on Disaggregated Storage
abstract
Database pushdown is a widely adopted technique under compute-storage disaggregation. The rising network and I/O speeds, coupled with stagnated compute and memory subsystems of a disaggregated storage architecture in the past decade, render state-of-the-art policy-driven pushdown designs ineffective. This is because the query performance bottleneck has shifted from network and I/O to compute, where computing power at the storage layer becomes scarce.
Yuebin Bai, Ming Liu 0027
ASPLOS (2)3
2025 Understanding and Profiling CXL.mem Using PathFinder
abstract
CXL.mem and the resulting memory pool are promising and gaining great attention. Unlike local memory, CXL DIMMs stay at the I/O subsystem, whose inferior performance can easily impact the processor pipeline and memory subsystem, yielding performance interference, hardware contention, obscure behaviors, and underutilized communication and computing resources. However, our community lacks a tool to understand and profile the CXL.mem protocol execution end-to-end between CPU and remote DIMM.
Zerui Guo, Yuebin Bai, Mahesh Ketkar, Hugh Wilkinson, Ming Liu 0027
SIGCOMM3
2024 SGCL: Semi-supervised Graph Contrastive Learning with confidence propagation algorithm for node classification
Yuebin Bai
Knowl. Based Syst.2
2024 Probability graph complementation contrastive learning
abstract
Graph Neural Network (GNN) has achieved remarkable progress in the field of graph representation learning. The most prominent characteristic, propagating features along the edges, degrades its performance in most heterophilic graphs. Certain researches make attempts to construct KNN graph to improve the graph homophily. However, there is no prior knowledge to choose proper K and they may suffer from the problem of Inconsistent Similarity Distribution (ISD). To accommodate this issue, we propose Probability Graph Complementation Contrastive Learning (PGCCL) which adaptively constructs the complementation graph. We employ Beta Mixture Model (BMM) to distinguish intra-class similarity and inter-class similarity. Based on the posterior probability, we construct Probability Complementation Graphs to form contrastive views. The contrastive learning prompts the model to preserve complementary information for each node from different views. By combining original graph embedding and complementary graph embedding, the final embedding is able to capture rich semantics in the finetuning stage. At last, comprehensive experimental results on 20 datasets including homophilic and heterophilic graphs firmly verify the effectiveness of our algorithm as well as the quality of probability complementation graph compared with other state-of-the-art methods.
Yuebin Bai
Neural Networks2
2023 LogNIC: A High-Level Performance Model for SmartNICs
abstract
SmartNICs have become an indispensable communication fabric and computing substrate in today’s data centers and enterprise clusters, providing in-network computing capabilities for traversed packets and benefiting a range of applications across the system stack. Building an efficient SmartNIC-assisted solution is generally non-trivial and tedious as it requires programmers to understand the SmartNIC architecture, refactor application logic to match the device’s capabilities and limitations, and correlate an application execution with traffic characteristics. A high-level SmartNIC performance model can decouple the underlying SmartNIC hardware device from its offloaded software implementations and execution contexts, thereby drastically simplifying and facilitating the development process. However, prior architectural models can hardly be applied due to their limited capabilities in dissecting the SmartNIC-offloaded program’s complexity, capturing the nondeterministic overlapping between computation and I/O, and perceiving diverse traffic profiles.
Zerui Guo, Yuebin Bai, Daehyeok Kim, Michael M. Swift, Aditya Akella, Ming Liu 0027
MICRO3
2023 PiPAD: Pipelined and Parallel Dynamic GNN Training on GPUs
abstract
Dynamic Graph Neural Networks (DGNNs) have been widely applied in various real-life applications, such as link prediction and pandemic forecast, to capture both static structural information and temporal characteristics from dynamic graphs. Combining both time-dependent and -independent components, DGNNs manifest substantial parallel computation and data reuse potentials, but suffer from severe memory access inefficiency and data transfer overhead under the canonical one-graph-at-a-time training pattern. To tackle these challenges, we propose PiPAD, a Pipelined and PArallel DGNN training framework for the end-to-end performance optimization on GPUs. From both algorithm and runtime level, PiPAD holistically reconstructs the overall training paradigm from the data organization to computation manner. Capable of processing multiple graph snapshots in parallel, PiPAD eliminates unnecessary data transmission and alleviates memory access inefficiency to improve the overall performance. Our evaluation across various datasets shows PiPAD achieves 1.22 × --9.57× speedup over the state-of-the-art DGNN frameworks on three representative models.
Desen Sun, Yuebin Bai
PPoPP3
2023 LEED: A Low-Power, Fast Persistent Key-Value Store on SmartNIC JBOFs
abstract
The recent emergence of low-power high-throughput programmable storage platforms-SmartNIC JBOF (just-a-bunch-of-flash)-motivates us to rethink the cluster architecture and system stack for energy-efficient large-scale data-intensive workloads. Unlike conventional systems that use an array of server JBOFs or embedded storage nodes, the introduction of SmartNIC JBOFs has drastically changed the cluster compute, memory, and I/O configurations. Such an extremely imbalanced architecture makes prior system design philosophies and techniques either ineffective or invalid.
Zerui Guo, Chenxingyu Zhao, Yuebin Bai, Michael M. Swift, Ming Liu 0027
SIGCOMM4
2023 APGNN : Alarm Propagation Graph Neural Network for fault detection and alarm root cause analysis
Yuebin Bai
Comput. Networks2
2023 Multiple UAVs collaborative traffic monitoring with intention-based communication
Yuebin Bai
Comput. Commun.2
2023 LR-BA: Backdoor attack against vertical federated learning using local latent representations
Yuhao Gu, Yuebin Bai
Comput. Secur.2
2023 LDIA: Label distribution inference attack against federated learning in edge computing
abstract
With the popularity of IoT (Internet of Things) applications, edge computing has received lots of attention. To meet data privacy protection requirements of edge nodes and cope with their unbalanced data distribution , federated learning (FL), a distributed learning framework, is widely used in intelligent edge computing applications. However, recent studies have shown that FL still suffers from privacy leakage problems, including membership inference, data reconstruction, etc. However, these studies mainly focus on the feature information of private data. In this paper, we concern the user-level label privacy in FL. We propose LDIA, a label distribution inference attack against FL in edge computing, exploring the possibility that an honest but curious cloud server can infer the proportions of samples per label in the edge user’s private data. LDIA is inspired by the observation that parameter changes in the output layer of a model can reflect the label distribution of training data . We use a neural network to learn individual features of the output layer updates over different label distributions, and then perform inference from local models uploaded by users. Our comprehensive evaluation shows that LDIA is effective on various datasets in different settings, demonstrating the severe privacy leakage in FL-based edge computing.
Yuhao Gu, Yuebin Bai
J. Inf. Secur. Appl.2
2023 CD-MSA: Cooperative and Deadline-Aware Scheduling for Efficient Multi-Tenancy on DNN Accelerators
abstract
With DNN turning into the backbone of AI cloud services and propelling the emergence of INFerence-as-a-Service (INFaaS), DNN-specific accelerators have become the indispensable components of cloud inference systems. Due to the conservative “one-task-at-a-time” working mode and deadline blindness of those accelerators, implementing multi-tenancy that aims to improve the cost-effectiveness and meet SLA requirements is intractable. Recent studies including the temporal and spatial approaches, employ manifold scheduling mechanisms and sophisticated architecture innovations to address the challenge. However, these researches either still neglect the deadline awareness or render inevitable and expensive hardware overheads such as switches and storage. In this paper, we presentCooperative and Deadline-aware Multi-Systolic-Array scheduling(CD-MSA), a low-cost solution for the cloud inference that utilizes the real time mechanism and task-level parallelism to enable efficient multi-tenancy. Based on our preemptive multi-systolic-array accelerator architecture supporting the simultaneous task co-location, we first construct a fine-grained DNN execution model to lay the groundwork for the lightweight preemption. Second, we design a cooperative, deadline- and laxity-aware scheduler in conjunction with an efficient schedulability test method for better QoS guarantee without introducing additional hardware cost. Finally, to further promote the overall throughput, we proposedynamic task fusion, a software approach that fuses different tasks into the logically “multi-threading” tasks at runtime. We compare CD-MSA with several state-of-the-art researches across three multi-DNN workloads. The evaluation results show CD-MSA improves the latency-bounded throughput, SLA satisfaction rate and weighted system throughput by up to 62%, 63% and 27%, respectively.
Yuebin Bai, Desen Sun
IEEE Trans. Parallel Distributed Syst.2
2022 Multi-UAV Joint Observation, Communication, and Policy in MEC
abstract
The use of multi-agent reinforcement learning methods (MARL) in mobile edge computing (MEC) environments enables multiple unmanned aerial vehicles (multi-UAV) to intelligently provide relay or computational offloading services to mission targets. UAV's observation range and communication methods between UAVs have a significant impact on multi-UAV collaboration strategy. For this purpose, we study the multi-UAV observation range dynamic control method and the optimal inter-UAV communication method. Our approach is to design a multi-UAV joint observation, communication, policy, and service collaboration protocol and study the optimization method of the protocol. We propose an expert-guided deep reinforcement learning framework to optimize this protocol. Each UAV's optimal radar observation range and inter-UAV communication method are learned using an information entropy value decomposition method. Through our observation and communication method, multi-UAV are able to obtain the most valuable information. Experiments demonstrate that our method can improve MEC's service coverage by 9.38%-21.88% compared to the classical MARL algorithm. Our method improves the radar observation efficiency and communication efficiency by 3.05%-38.9% and 8.55%-22.03%, respectively. The results show that this method improves multi-UAV energy utilization.
Yuebin Bai
MSN2
2022 CS-MIA: Membership inference attack based on prediction confidence series in federated learning
Yuhao Gu, Yuebin Bai, Shubin Xu
J. Inf. Secur. Appl.2
2021 A Millimeter-wave Multi-channel MAC with Dynamic Spectrum Access Capability for Mobile Self-organizing Heterogeneous Networks
abstract
With the rapid development of wireless network technology and the demand for interconnection of mobile edge computing, the integration of heterogeneous wireless networks is becoming more and more urgent. At the same time, the increasing scarcity of spectrum resources also requires heterogeneous net-works to effectively improve the utilization of spectrum resources without affecting the use of different licensed spectrum resources. This paper proposed a multi-channel millimeter-wave (mmWave) MAC protocol for mobile self-organizing heterogeneous networks. The mmWave Mac uses a bulk access transmission scheme to improve throughput efficiency and reduce latency for heterogeneous networks. In addition, cognitive radio technology is also introduced into the mmWave Mac to enhance the utilization of spectrum resources. Finally, by comparing the simulation results with other protocol, it is verified that the mmWave MAC protocol has good performance in terms of delay, and throughput efficiency.
Peng Feng 0003, Yuebin Bai, Zerui Guo, Bojian Bai, Junmin Zhang
IPCCC2
2021 A rapid coarse-grained blind wideband spectrum sensing method for cognitive radio networks
Peng Feng 0003, Yuebin Bai, Yuhao Gu, Jun Huang 0001
Comput. Commun.2
2021 Guardauto: A Decentralized Runtime Protection System for Autonomous Driving
abstract
Due to the broad attack surface and the lack of runtime protection, potential safety and security threats hinder the real-life adoption of autonomous vehicles. Although efforts have been made to mitigate some specific attacks, there are few works on the protection of the autonomous driving system, i.e., the control software system performing such as perception, decision making, and motion tracking. This article presents a decentralized self-protection framework called Guardauto to protect the autonomous driving system against runtime threats. First, Guardauto proposes an isolation model to decouple the autonomous driving system and isolate its components with a set of partitions. Second, Guardauto provides self-protection mechanisms for each target component, which combines different methods to monitor the target execution and plan adaption actions accordingly. Third, Guardauto provides cooperation among local self-protection mechanisms to identify the root-cause component in the case of cascading failures affecting multiple components. A prototype has been implemented and evaluated on the open-source autonomous driving system Autoware. Results show that Guardauto could effectively mitigate runtime failures and attacks, and protect the control system with acceptable performance overhead.
Yuan Zhou 0005, Bihuan Chen 0001, Rui Wang 0014, Yuebin Bai, Yang Liu 0003
IEEE Trans. Computers5
2021 MIPSGPU: Minimizing Pipeline Stalls for GPUs With Non-Blocking Execution
abstract
Improving the latency hiding ability is important for GPU performance. Although existing works, which mainly target on either improving thread level parallelism or optimizing memory hierarchy, are effective at improving GPUs’ latency hiding ability, warps are still blocked after executing long latency operations, reducing the number of schedulable warps. This article revisits the recently proposed non-blocking execution for GPUs to improve the latency hiding ability of GPUs. With non-blocking execution, instructions from warps blocked by long latency operations can be pre-executed to make full use of GPU resources. However, we find that the state-of-the-art non-blocking GPU architecture gains limited performance improvement. Through in-depth analysis, we observe that the poor performance is largely due to inefficient pre-execution state management, duplicate instruction extraction, frequent early eviction and severe resource congestion. To make non-blocking execution actually useful for GPUs and minimize hardware overheads, we carefully redesign the non-blocking architecture for GPUs based on our analysis and proposeMIPSGPU. Our evaluations show thatMIPSGPU, relative to the state-of-the-art non-blocking GPU architecture, improves performance of memory intensive applications by 19.05 percent, and reduces memory to SM traffics by 14 percent.
Chao Yu 0001, Yuebin Bai, Rui Wang 0014
IEEE Trans. Computers2
2019 CogMOR-MAC: A cognitive multi-channel opportunistic reservation MAC for multi-UAVs ad hoc networks
Peng Feng 0003, Yuebin Bai, Jun Huang 0001, Yuhao Gu
Comput. Commun.2
2019 Improving Thread-level Parallelism in GPUs Through Expanding Register File to Scratchpad Memory
abstract
Modern Graphic Processing Units (GPUs) have become pervasive computing devices in datacenters due to their high performance with massive thread level parallelism (TLP). GPUs are equipped with large register files (RF) to support fast context switch between massive threads and scratchpad memory (SPM) to support inter-thread communication within the cooperative thread array (CTA). However, the TLP of GPUs is usually limited by the inefficient resource management of register file and scratchpad memory. This inefficiency also leads to register file and scratchpad memory underutilization. To overcome the above inefficiency, we propose a new resource management approach EXPARS for GPUs. EXPARS provides a larger register file logically by expanding the register file to scratchpad memory. When the available register file becomes limited, our approach leverages the underutilized scratchpad memory to support additional register allocation. Therefore, more CTAs can be dispatched to SMs, which improves the GPU utilization. Our experiments on representative benchmark suites show that the number of CTAs dispatched to each SM increases by 1.28× on average. In addition, our approach improves the GPU resource utilization significantly, with the register file utilization improved by 11.64% and the scratchpad memory utilization improved by 48.20% on average. With better TLP, our approach achieves 20.01% performance improvement on average with negligible energy overhead.
Chao Yu 0001, Yuebin Bai, Qingxiao Sun, Hailong Yang 0002
ACM Trans. Archit. Code Optim.2
2018 Towards Fault-Tolerant Task Backup and Recovery in the seL4 Microkernel
abstract
With the improvement of the reliability requirements of modern computer systems, fault-tolerant technology has been widely used in all aspects of computer systems as a key mechanism to ensure system reliability. The seL4 microkernel as the world's first complete, formal, machine-checked operating system kernel, has become one of the ideal platforms for safety-critical systems, but the lack of fault-tolerant mechanism reduces its reliability and availability. This paper designs and implements a fault-tolerant backup and recovery mechanism based on checkpoint in the seL4 microkernel and makes the optimization of resource. The experiment demonstrates the performance of the backup and recovery mechanism in terms of time and space.
Guangqiang Luan, Yuebin Bai, Libin Xu, Chao Yu 0001, Junfang Zeng, Qingbin Chen
COMPSAC (1)2
2018 DTN-Knca: A High Throughput Routing Based on Contact Pattern Detection in DTNs
abstract
In current routing algorithms based on encounter history in Delay-Tolerant Networks (DTNs), packets are always forwarded to nodes with highest probability to reach destination node. However, to the best of our knowledge, no analytical node transient contact pattern detection to achieve high performance, is reported in the literature. In this letter, DTN-Knca - a novel routing which detects frequently encountered nodes' transient contact patterns by correlation analysis is proposed. The trace-driven simulations demonstrate the higher throughput of DTN-Knca in comparison to the state-of-the-art DTN typical routing algorithms based on the encounter history knowledge.
Yuebin Bai, Peng Feng 0003, Yuhao Gu, Jun Huang 0001
COMPSAC (1)2
2018 Research on Asynchronous Inter-VM Communication Mechanism Based on Embedded Hypervisor
abstract
Virtualization technology, which has achieved great success in server and desktop environments in the last few years, is currently extending itself towards a new territory: embedded system. OKL4 from Open Kernel Labs is a leading virtualization software for embedded systems. Its microkernel approach can improve plain virtualization technologies, but it also brings incomplete IPC (Inter-Process Communication) issues. This paper proposes a novel asynchronous communication mechanism that can generate multiple event channels and efficiently manage concurrent communication requests between virtual machines. This mechanism optimizes the original IPC mechanism, and this paper also proposes a shared-memorybased bulk data transmission mechanism. The final experiments prove its feasibility and demonstrate the specific effect.
Rui Wang 0014, Libin Xu, Yuebin Bai, Zhongzhao Wang, Guangqiang Luan, Hailong Yang 0002
COMPSAC (1)3
2018 Network Alarm Flood Pattern Mining Algorithm Based on Multi-dimensional Association
abstract
In the process of network operation, a large number of alerts are generated every day, which reflect the occurrence of some abnormal conditions. Traditional methods depend too much on the knowledge of equipment manufacturers and industry experts, so we need some novel ways to overcome this problem in the network management. The application of data mining technology to alarm pattern analysis has become the focus of current research. Researchers developed many kinds of algorithms fitting different application characteristics. This paper proposes the concept of association matrix pattern mining, which means that before mining the data, we use the multi-dimensional information of the data to construct the association matrices between the items. And we develop a conditional pattern mining algorithm based on the association matrix which aims to find out less but more meaning results. Our experiments validate that with the multi-dimensional information stored in association matrix, the algorithm performs better than traditional pattern mining methods in finding out the detailed alarm pattern from network alarm flood.
Yuebin Bai, Peng Feng 0003, Junfang Zeng, Rui Wang 0014
MSWiM2
2018 Nodes contact probability estimation approach based on Bayesian network for DTN
abstract
Delay tolerant network (DTN) known as suffering from frequent disruption, high latency and heterogeneous, re­sulting in low network availability. To improve DTN availability, routing protocols typically need to predict the probability of encountering the nodes. In this paper, we use the Bayesian Network (BN) to construct the knowledge base, which is an unique tool for creating a representation of the dependence relationships among DTN parameters. Then developed a Bayesian network- based approach to estimate the contact probability among nodes of DTN. We conducted an experiment to compare our approach against its counterparts in PROPHET routing protocol and power law distribution-based method. The experiment shows our approach is superior to other methods in both recall ratio and precision in all four datasets, including HAGGLE, NUS, REALITY and SASSY.
Yuebin Bai, Xu Shao, Wentao Yang 0001, Peng Feng 0003, Rui Wang 0014
NOMS1
2018 A network traffic flow prediction with deep learning approach for large-scale metropolitan area network
abstract
Accurate and timely internet traffic information is important for many applications, such as bandwidth allocation, anomaly detection, congestion control and admission control. Over the last few years, internet flow data have been exploding, and we have truly entered the era of big data. Existing traffic flow prediction methods mainly use simple traffic prediction models and are still unsatisfying for many real-world applications. This situation inspires us to rethink the internet traffic flow prediction problem based on deep architecture models with big traffic data. In this paper, we propose a novel deep-learning-based internet traffic flow prediction method, which is called SDAPM. It consider the spatial and temporal correlations inherently and internet flow data character. A stacked denoising autoencoder prediction model (SDA) is used to learn generic internet traffic flow features, and it is trained in a greedy layer-wise fashion. Moreover, experiments demonstrate that the SDAPM for traffic flow prediction has effective performance. Our prediction model is in production as part of the traffic scheduling system at China Unicom, one of the largest Internet companies in China, helping improving the network bandwidth utilization.
Yuebin Bai, Chao Yu 0001, Yuhao Gu, Peng Feng 0003, Rui Wang 0014
NOMS2
2018 SMGuard: A Flexible and Fine-Grained Resource Management Framework for GPUs
abstract
GPUs have been becoming an indispensable computing platform in data centers, and co-locating multiple applications on the same GPU is widely used to improve resource utilization. However, performance interference due to uncontrolled resource contention severely degrades the performance of co-locating applications and fails to deliver satisfactory user experience. In this paper, we present SMGuard, a software approach to flexibly manage the GPU resource usage of multiple applications under co-location. We also propose a capacity based GPU resource model CapSM, which provisions the GPU resources in a fine-grained granularity among co-locating applications. When co-locating latency-sensitive applications with batch applications, SMGuard can prevent batch applications from occupying resources without constraint using quota based mechanism, and guarantee the resource usage of latency-sensitive applications with reservation based mechanism. In addition, SMGuard supports dynamic resource adjustment through evicting the running thread blocks of batch applications to release the occupied resources and remapping the uncompleted thread blocks to the remaining resources, which avoids the relaunch of the preempted kernel. The SMGuard is a pure software solution that does not rely on special GPU hardware or programming model, which is easy to adopt on commodity GPUs in data centers. Our evaluation shows that SMGuard improves the average performance of latency-sensitive applications by 9.8× when co-located with batch applications. In the meanwhile, the GPU utilization can be improved by 35 percent on average.
Chao Yu 0001, Yuebin Bai, Hailong Yang 0002, Yuhao Gu, Zhongzhi Luan, Depei Qian 0001
IEEE Trans. Parallel Distributed Syst.2
2018 LWPTool: A Lightweight Profiler to Guide Data Layout Optimization
abstract
Memory access latency continues to be a dominant bottleneck in a large class of applications on modern architectures. To optimize memory performance, it is important to utilize the locality in the memory hierarchy. Data layout optimization can significantly improve memory locality. However, pinpointing inefficient code and providing insightful guidance for data layout optimization is challenging. Existing tools typically leverage heavyweight memory instrumentations, which hinders the applicability of these tools for real long-running programs. To address this issue, we develop LWPTool, a profiler to pinpoint top candidates that benefit from data layout optimization. LWPTool makes three unique contributions. First, it adopts lightweight address sampling to collect and analyze memory traces. Second, LWPTool employs a set of novel methods to determine memory access patterns to guide data layout optimization. We also formally prove that our method has high accuracy even with sparse memory access samples. Third, LWPTool scales on multithreaded machines. LWPTool works on fully optimized, unmodified binary executables independently from their compiler and language, incurring around 6.2 percent runtime overhead. To evaluate LWPTool, we study ten sequential and parallel benchmarks. With the guidance of LWPTool, we are able to significantly improve all these benchmarks; the speedup is up to 1.39× on average.
Chao Yu 0001, Probir Roy, Yuebin Bai, Hailong Yang 0002, Xu Liu 0001
IEEE Trans. Parallel Distributed Syst.3
2016 HV2M: A novel approach to boost inter-VM network performance for Xen-based HVMs
Yuebin Bai, Yongwang Zhao, Duo Lu, Yuanfeng Peng, Minxuan Zhou
J. Syst. Softw.2
2015 Optimizing Soft Real-Time Scheduling Performance for Virtual Machines with SRT-Xen
abstract
Multimedia applications are an important part of today's Internet. However, currently most virtualization solutions, including Xen, lack adequate support for soft real-time tasks. Soft real-time applications, e.g. media workloads, are impeded by components of virtualization, such as the increase of scheduling latency. This paper focuses on improving scheduling scheme to support soft real-time workloads in virtualization systems. In this paper, we present an enhanced scheduler SRT-Xen. SRT-Xen can promote the soft real-time domain's performance compared with Xen's existing scheduling. It focuses on not only bringing a new realtime-friendly scheduling framework with corresponding strategies but also improving the management of the virtual CPUs' queuing in order to implement a fair scheduling mechanism for both real-time and non-real-time tasks. Finally, we use PESQ (Perceptual Evaluation of Speech Quality) and other benchmarks to evaluate and compare SRT-Xen with some other works. The results show that SRT-Xen supports soft real-time domains well without penalizing non-real-time ones.
Yuebin Bai, Rui Wang 0014
CCGRID2
2015 An Efficient Transmission Method for Bulk Data Based on Network Coding in Delay Tolerant Network
abstract
With nodes in Delay Tolerant Network(DTN) distributing sparsely and moving rapidly, they usually suffer from intermittent connections and communications, thus bringing about limited message forwarding opportunities. All these could lead to inefficient forwarding, low delivery, long latency and limited transmission capacity in performance. In this paper, efficient encoding and decision methods are presented and integrated into the DTN routing strategy. The custody-encoding-forwarding mode is designed by merging the random linear network coding into the DTN routing, together with the replica re-allocation and memory management, and built on that encoding scheme, the intra/inter flow adaptive collaborative network coding is elaborated to implement the fresh custody-decision-encoding-forwarding mode. A decision-making strategy based on Bayesian Network(BN) measures the ``degree'' in a specific generation in networks by taking comprehensive consideration of current and historical network conditions to enhance the network robustness and self-adaptivity. By evaluating the delivery, delay and overhead performance on ONE and MATLAB platforms, the effectiveness of proposed strategies is validated in the end.
Wancheng Chen, Yuebin Bai, Jiaojiao Liang, Wenjia Liu, Rui Wang 0014, Xiaoyun Mo, Ziming Luo
MSWiM2
2014 BIDS: Bridgehead-Employed Image Distribution System for Cloud Data Centers
Zhongzhao Wang, Yuebin Bai, Jihong Ma, Duo Lv, Yuanfeng Peng
NPC2
2013 A Virtual Network Embedding Algorithm Based on Graph Theory
Zhenxi Sun, Yuebin Bai, Songyang Wang, Shubin Xu
NPC2
2013 A high performance inter-domain communication approach for virtual machines
Yuebin Bai, Duo Lv, Yuanfeng Peng
J. Syst. Softw.1
2012 SAME: A students' daily activity mobility model for campus delay-tolerant networks
abstract
Mobility model plays a crucial part in wireless network simulation or mathematical analysis, especially for human mobility because real mobility network constructed by human is hard to trace, measure or apply to different scenarios. For delay-tolerant networks, mobility model also affects the frequency and duration of opportunities for data transfer between nodes, which is a key factor of network performance, so whether we can choose a model according to realistic scenes is important for the accuracy and practicability of simulation. In this paper, we propose a new mobility model of students' daily activities based on the analysis and conclusion of students' habits and customs in campus environments, which is able to produce inter-contact time, contact time distribution and other contact information. We validate the movement model by using the ONE simulator comparing with other movement models and tracing data from real-world measurement experiments. Results show that SAME is consistent with actual student behavior characteristics.
Yuebin Bai, Wentao Yang 0001, Yuanfeng Peng, Chongguang Bi
APCC2
2012 MOVE: A mobile personalized virtual computing environment
Yuebin Bai, Yanwen Ju
Future Gener. Comput. Syst.1
2012 FAST: Fuzzy Decision-Based Resource Admission Control Mechanism for MANETs
Yuebin Bai, Xu Shao, Wentao Yang 0001
Mob. Networks Appl.1
2011 Building the Knowledge Base through Bayesian Network for Cognitive Wireless Networks
abstract
Tactical communication networking faces complexity, heterogeneity, and reliability requirements. The emerging research area of cognitive networks offers a potential for dealing with these problems. A key feature of cognitive networks is the knowledge base, which is produced during the process of learning and responsible for the decision making. We propose a cognitive network model integrated with the knowledge base, which is a primary part of cognitive networks. And then we focus on the construction of the knowledge base and the expression form of the knowledge in the model. In this paper, we use the Bayesian Network (BN) to construct the knowledge base, which is a unique tool for creating a representation of the dependence relationships among network protocol parameters. The data structure of the dependence relationships of the BN is translated into the knowledge which is expressed by the probability. In the simulation experiments, we create the BN through the sampling data to construct the knowledge base using the mathematical tool MATLAB and prove the efficiency of our cognitive network model for optimizing network performance in the OPENT simulation platform.
Niandong Du, Yuebin Bai, Lianhe Luo, Jianli Guo
ICPADS2
2010 SoftAccel: A Software-Only IP Network Application Accelerator
abstract
Internet technologies are making rapid progress, but achieving high performance of network application over the WAN remains a challenge. When branch office's workers access enterprise applications locating in data center, the performances drop dramatically. To solve this problem, many hardware accelerators are produced. In this paper, we introduce a Software-only IP Network Application Accelerator (Soft Accel), for improving network applications performance over the WAN. To overcome the performance-limiting factors associated with WAN, transport protocol and application protocol, Soft Accelcombines application proxy, delta compress and TCPoptimization techniques. We have implemented Soft Accel in the form of framework, so people can design and test their optimization technologies based on it. Experiments in WAN environment show that Soft Accel can improve the performance of IP network applications at least by 2x.
Huiyong Zhang, Yuebin Bai, Zhi Li 0009, Huixing Peng, Likun Zhao, Lianhe Luo
APSCC2
2010 Affinity-Aware Dynamic Pinning Scheduling for Virtual Machines
abstract
Virtualization provides an effective management in server consolidation. The transparence enables different kinds of servers running in the same platform, making full use of hardware resource. However, virtualization introduces two-level schedulers: one from Guest OS, where the tasks are scheduled to virtual CPUs (VCPUs), the other from the virtual machine monitor (VMM), where VCPUs are scheduled to CPUs. As a result, the lower level scheduler is ignorant of the task information so that it cannot allocate appropriate proportion of CPU resource for every Guest OS in some cases. This paper presents an affinity-aware Dynamic Pinning Scheduling scheduler (DP-Scheduling). We aim at two objects: Bridging the semantic gap between Guest OS and VMM, introducing an affinity-aware method and providing the tasks information about CPU affinity to VMM, Bringing up a novel scheduling, DP-Scheduling, so that VCPU can be pinned or unpinned on one CPU's running queue dynamically. For this purpose, we first get the Machine Address (MA) of process descriptor from the angle of VMM. The affinity information is also acquired before the task is enabled to run. To acknowledge the affinity information, DP-Scheduling calls an API provided by us. Depending on the affinity information, we put forward a series of measures to implement pinning dynamically as well as to keep workload balance. All implementation is confined to Xen VMM and Credit scheduler. Our experiments demonstrate that DP-Scheduling outperforms Credit scheduling by testing various indicators for CPU-bound tasks, without interfering the load balance.
Zhi Li 0009, Yuebin Bai, Huiyong Zhang
CloudCom2
2010 Achieving High Throughput by Transparent Network Interface Virtualization on Multi-core Systems
abstract
Though with the rapid development, there remains a challenge on achieving high performance of I/O virtualization. The Para virtualized I/O driver domain model, used in Xen, provides several advantages including fault isolation, live migration, and hardware independence. However, the high CPU overhead of driver domain leads to low throughput for high bandwidth links. Direct I/O can achieve high performance but at the cost of removing the benefits of the driver domain model. This paper presents software techniques and optimizations to achieve high throughput network I/Ovirtualization by driver domain virtualization model on multicore systems. In our experiments on multi-core system with a quad-port 1GbE NIC, we observe the overall throughput of multiple guest VMs can only be 2.2Gb/s, while the link bandwidth is 4Gb/s in total. The low performance results from the disability of driver domain to concurrently serve multiple guest VMs running bandwidth-intensive applications. Consequently, two approaches are proposed. First, a multi task let net back is implemented to serve multiple net fronts on currently. Second, we implement a new event channel dispatch mechanism to balance event associated with networkI/O over VCPUs of driver domain. To reduce the CPU overhead of the driver domain model, we also propose two optimizations: lower down event frequency in netback and implement LRO in net front. By applying all the above techniques, our experiments show that the overall throughput can be improved from the original 2.2Gb/s to 3.7Gb/s and the multi-core CPU resources can be utilized efficiently. We believe that the approaches of our study can be valuable for high throughput I/O virtualization in the coming multi-core era.
Huiyong Zhang, Yuebin Bai, Zhi Li 0009, Niandong Du, Wentao Yang 0001
CloudCom2
2010 A High Performance Inter-VM Network Communication Mechanism
Yuebin Bai, Huiyong Zhang
ICA3PP (1)1
2010 idsocket: API for Inter-domain Communications Base on Xen
Yuebin Bai
ICA3PP (1)2
2010 A Tracing Approach to Process Migration for Virtual Machine Based on Multicore Platform
Yuebin Bai
ICA3PP (1)2
2009 A Functional Classification Based Inter-VM Communication Mechanism with Multi-core Platform
abstract
With the resurgence of virtualization technologies and the development of multi-core technologies, the combination of the two becomes a trend. Therefore, inter-VM communication becomes a key part in how to improve the performance of virtual machines (VMs) basing on multi-core platform. In this paper, we first analyze the characteristics of multi-core tasks and the properties of virtual machine environment, and then classify processor cores into two categories basing on their different functions. According to the classification, we design an inter-VM communication mechanism with multi-core platform. It discards the traditional communication path between VMs which needs to via a trusted VM, sets up communication channels between virtual CPUs in different VMs and uses shared memory space to implement high-throughput communication of inter-VM. Experiment results have proved the efficiency of them.
Yuebin Bai
ICPADS2
2009 Performance Evaluation of Parallel Programming in Virtual Machine Environment
abstract
As multi-core processors become increasingly mainstream, architects have likewise become more interested in how best to make use of the computing capacity of the CPU, for instance, through multiple simultaneous threads or processes of execution with OpenMP or MPI. At the same time, the increasingly mature and prevailing virtualization technique in server consolidation and HPC promotes the emergence of a large number of virtual SMP servers. Therefore, whether the parallel program can run in the virtual machine environment efficiently or not is a topic of concern. In this paper, we investigate the performance of three typical parallel programming paradigms, including OpenMP, MPI, and Hybrid of OpenMP and MPI in the popular, open-source, Xen virtualization system. The results show that the performance of the traditional parallel program in Xen VMs is close to it in native, non-virtualized environment, if there is little communication or synchronization between threads or processes. In most cases, without excessive IO access, we can get an ideal speedup in a SMP VM or virtual cluster, which is close to linearity when the total virtual CPUs (vCPUs) number is not larger than the number of Physical CPUs (pCPUs). And the pure MPI implementation shows the best scalability and stability in virtual machine environment compared with the other two paradigms.
Yuebin Bai
NPC2
2009 On minimum data replication for delay-bounded query in wireless ad hoc networks
abstract
In this paper, we study the problem of minimizing the number of data replicas for delay-bounded queries in wireless ad hoc networks. We focus our attention on step-by-step expanding ring search, which provides an upper bound on query delay to any expanding ring based search strategies. We analyze the probabilistic behavior of query delay, and develop an analytical approach to approximate the minimum number of data replicas for delay bounded data query in wireless ad hoc networks. We validate our analysis through extensive simulations.
Jun Huang 0001, Yuebin Bai, Xu Shao
WCNC2
2008 Mobile e-Lab: A Mobile Personalized Virtual Research Computing Environment
abstract
In today's scientific research, computers and networks are playing an increasingly important role in the laboratory. It is desirable to researchers that any machines outside the laboratory provide a uniform, consistent, desktop computing environment and the ability to access private laboratory computing resources when outside the familiar work place. This paper proposes a system called mobile e-Lab, which aims to present researchers with such a consistent environment, including customized software, personal data, private network resources accessing and other abilities on any computer attached to an IP network, enabling researchers to work anywhere as if they were at their own laboratories, without the constraints of mobility and geographical location. Based on the virtual machine technology, the user's entire computing environment - including operating system, installed applications and personal data - can be encapsulated to be a virtual disk via network. With the OS browser, a general virtual platform, the user can appoint a virtual disk in the network and create an OS instance with which to operate, which is the same as the one in his laboratory. The mobile e-Lab also provides a virtual network facility, which allows users to operate with their private computing resources in their laboratory LAN transparently.
Yanwen Ju, Yuebin Bai, Depei Qian 0001
eScience2
2008 Link Availability Prediction in Ad Hoc Networks
abstract
Since mobility may cause radio links to break frequently, one pivotal issue for routing in Mobile Ad Hoc Networks is how to select a reliable path that can last longer. Several metrics have been proposed in previous literatures, including link persistence, link duration, link availability, link residual time, and their path equivalents. In this paper, we present a novel algorithm for predicting continuous link availability between two mobile ad hoc nodes. By a rough estimation of the distance between two nodes, our approach is able to accurately predict link availability over a short period of time. Simulation results are given to verify our approach. This study could serve as groundwork for further ad hoc network researches including analyzing and optimizing other network protocols.
Yuebin Bai, Depei Qian 0001
ICPADS2
2008 A Novel Approach of Link Availability Estimation for Mobile Ad Hoc Networks
abstract
Mobile Ad Hoc Networks have inherently dynamic topologies. Due to the distributed, multi-hop nature of these networks, random mobility of nodes affects not only the availability of radio links between particular node pairs, but also impedes the reliability of communication paths. Therefore, to optimize the performance of existing network protocols, it is important to be able to give an accurate prediction of the future link status within a random mobility environment. In this paper, a novel approach is introduced to derive analytical expression of link availability for mobile ad hoc networks. By a rough estimation of the initial distance between two nodes, the prediction algorithm presented in this paper is able to accurately estimate link availability for a given short period of time. Simulation results are reported to verify the correctness of our approach.
Jun Huang 0001, Yuebin Bai
VTC Spring2
2007 Research on Planning and Deployment Platform for Wireless Sensor Networks
Yuebin Bai, Qingmian Han, Yujun Chen, Depei Qian 0001
GPC1
2003 New String Matching Technology for Network Security
abstract
String matching is a comprehensive applicable key technology beyond intrusion detection systems (IDS), and many areas can benefit from faster string matching algorithm. Which can be used in IDS, firewall et al network security applications. These applications are usually deployed at choke points of a network where there is heavily traffic. Using lower efficient string matching algorithm may make these applications to become a performance bottleneck in network. So it is very necessary to develop faster and more efficient string matching algorithms in order to overcome the troubles on performance. On a basis of Boyer-Moore-Horspool algorithm, a new string matching algorithm is presented in this paper. The algorithm is described in detail. The new algorithm has been greatly improved. The algorithm is one simplification of Boyer-Moore-Horspool algorithm. Array NEXT in Preprocessing stage is redesigned. A novel generated rules are presented. Using these rules, a simple NEXT is generated. And based on the concept of reference point, all make the algorithm to have better performance and more efficient. These characteristics will be useful in all these applications. Main features of the algorithm are presented, then explained its work processes. The algorithm also passed test and is validated. The test results show that the algorithm has better performance than Boyer-Moore algorithm and Boyer-Moore-Horspool algorithm, and more simple and efficient.
Yuebin Bai, Hidetsune Kobayashi
AINA1
2003 Intrusion Detection System: Technology and Development
abstract
Attacks on network infrastructure presently are main threats against network and information security. With rapidly growing unauthorized activities in networks, intrusion detection (ID) as a component of defense-in-depth is very necessary because traditional firewall techniques cannot provide complete protection against intrusion. ID is an active and important research area of network security. A survey on ID technology is shown in this paper. It is involved with several main aspects of ID technology. Analyses on intrusion detection techniques and data collection techniques are emphasized. Some novel developments in ID Systems, such as both data mining based ID systems and data fusion based ID systems, are also discussed. Current ID technology faces powerful challenges, major challenges and future promising directions are presented.
Yuebin Bai, Hidetsune Kobayashi
AINA1