Mengjie Lv

dblp:233/8560 · DBLP profile ↗
← Back
53ranked-venue papers
20as first author
45since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 19 · 5 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 6 first-author · 11 since 2021Computer networks · 9 · 3 first-author · 9 since 2021Theory of computation · 8 · 5 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Parallel construction of multiple independent spanning trees on 3-ary n -cube networks
abstract
Abstract High-performance computing utilizes powerful processor clusters to parallel process big data and solve complex problems at extremely high speeds, relying significantly on interconnection networks. As networks grow in scale and complexity, failures become unavoidable. Interconnection networks demand consistent operation and efficient routing algorithms to enable smooth data transmission among processors. Fault-tolerant routing is essential for assessing network reliability. The application of independent spanning trees (ISTs) is an effective method to enhance network fault tolerance. Regarded as a significant extension of the hypercube, the $3$-ary $n$-cube network $(Q^{3}_{n})$ boasts many advantageous such as low vertex degree, regularity, and straightforward implementation. In this paper, we introduce parallel algorithms for generating $2n$ ISTs on $Q^{3}_{n}$, where $2n$ represents the maximum achievable number, enhancing the efficiency and obtaining additional sets of ISTs and disjoint paths. Building upon previously constructed ISTs, a fault-tolerant routing system is developed, utilizing them as the routing table. Subsequently, the effectiveness of this mechanism is assessed through simulated data, showing an increment in transmission success rates as dimensionality grows, nearing near-perfection at almost $100\%$. These results also reveal that the algorithm we proposed demonstrates better performance than traditional classical algorithms.
Weibei Fan, Yuzhen Xu, Mengjie Lv, Xueli Sun
Comput. J.3
2026 Fault-tolerant path and disjoint path construction in data center network based on augmented cube
Weibei Fan, Jingman Pei, Mengjie Lv, Xueli Sun
Frontiers Comput. Sci.3
2026 A Fast Intermittent Fault Diagnosis Algorithm for a Class of Data Center Networks
abstract
As a data center network (DCN) constructed using recursive modules, BCube enables efficient communication for decentralized machine learning systems. Its various variants, such as RCube and RRect, outperform BCube in certain performance metrics. To unify related research, BCube and its variants are integrated into a unified framework known as BCube-based DCNs (BDCN). In practical DCN deployments, efficient fault diagnosis is essential for reliability and stability. However, intermittent faults are more challenging to diagnose than permanent ones due to their randomness and uncertainty. Moreover, existing intermittent fault diagnosis algorithms generally rely on searching for the largest component, which leads to high time complexity. To address this issue, this paper systematically analyzes the intermittent fault diagnosability of BDCN under the PMC model, and proposes a fast intermittent fault diagnosis algorithm (FIFDA). The proposed algorithm significantly improves diagnosis efficiency by avoiding the need to search for the largest component. Extensive experimental results verify the applicability of FIFDA in both BDCN and other high-performance DCNs. Comparative analyses with existing algorithms show that FIFDA achieves higher diagnostic speed. Moreover, under comparable diagnosis times, FIFDA demonstrates superior diagnostic performance. In addition, simulation results demonstrate that FIFDA maintains outstanding performance in large-scale DCNs and under varying noise levels and fault probabilities, fully showcasing its efficiency and scalability in practical DCN environments.
Huaqun Wang, Mengjie Lv, Weibei Fan
IEEE Trans. Computers3
2026 A Highly Cost-Effective and Fault-Tolerant Network Topology for Large-Scale Data Centers
abstract
With the rapid advancement of digital technologies such as cloud computing, big data, and artificial intelligence, large-scale data centers have become critical infrastructure supporting these technologies, imposing increasingly high demands on data center networks (DCNs). Traditional server-centric DCNs face challenges in large-scale distributed systems, such as difficulty in balancing bandwidth and latency, high expansion costs, and conflicts between fault tolerance and communication efficiency. To address these issues, this paper proposes ECQDC, a novel server-centric DCN based on exchanged crossed cube. Specifically, we present its logical structure ECD(s, t) and study the connectivity and edge connectivity of ECD(s, t). Furthermore,we develop efficient fault-free routing algorithm and faulttolerant routing algorithm for the ECD(s, t). The experimental results demonstrate that, compared with Dijkstra and BFS, the proposed ECDR and ECDFTR algorithms reduce the average running time by over 50% and cut the average path length by approximately 20% relative to BFS, while keeping path lengths close to Dijkstra’s optimal performance. Moreover, it exhibits excellent performance in scalability, fault tolerance, and communication efficiency, making it an ideal network topology for large-scale data center deployment.
Weibei Fan, Xiangying Peng, Fu Xiao 0001, Mengjie Lv, Xueli Sun, Sun-Yuan Hsieh
IEEE Trans. Computers4
2026 An Efficient and Fault-Tolerant Data Transmission Scheme in Data Center Networks
abstract
The rapid growth of cloud computing, large-scale distributed systems, and AI-driven applications has placed stringent demands on the performance and reliability of data center networks (DCNs). As DCNs scale in size and structural complexity, they become increasingly vulnerable to multiple concurrent node and/or link failures, which can lead to severe service disruptions and significant performance degradation. Existing data transmission approaches typically address node and link failures in isolation, frequently mitigating one type while overlooking the other, and thus fall short in effectively handling complex multi-failure scenarios. This paper presents a novel and efficient data transmission scheme designed to ensure robust communication under multiple node and/or link failures in DCNs. The proposed solution integrates a proactive path redundancy mechanism with a failure-aware routing strategy to enable rapid identification and avoidance of faulty components. We adopt the generalized hypercube network (GHN), a regular and scalable topology, as the underlying network model. Firstly, leveraging the method of Yang and Chang [44], we construct multiple independent spanning trees (ISTs) in GHNs, which provide structural path diversity and fault isolation. Building upon these ISTs, we propose GFP-IST, an optimized routing algorithm with a time complexity ofO(NlogN), whereNdenotes the number of nodes. GFP-IST enables efficient route computation and resilient packet forwarding in the presence of multiple simultaneous failures. Extensive simulation results demonstrate that our approach outperforms several fault-tolerant routing schemes in terms of average path length, path construction time, and fault recovery success rate, especially in large-scale and high-failure-rate network environments.
Mengjie Lv, Fu Xiao 0001, Weibei Fan, Jian Qiao, Sun-Yuan Hsieh
IEEE Trans. Computers1
2026 EBM: Traffic-Based Differentiated Enhanced Buffer Management in Data Center Networks
abstract
With the rapid advancement of big data processing and artificial intelligence (AI), data center networks (DCNs) must deliver more efficient resource management and data transmission mechanisms. Unfortunately, due to the significant differences in bandwidth requirements, transmission patterns, and temporal characteristics across various traffic types in DCNs (such as short flows, long flows, and bursty flows), traditional buffer allocation strategies fail to adapt flexibly to these disparities. In this paper, we propose Enhanced Buffer Management (EBM), a novel buffer-sharing scheme designed for scenarios that require higher performance from DCNs. Unlike prior approaches, EBM employs a multi-level flow identification and adaptive threshold adjustment mechanism to enhance the flexibility and efficiency of buffer management under varying traffic conditions. Specifically, EBM first performs coarse-grained and fine-grained classification of traffic based on packet size, inter-arrival interval, and other flow characteristics. It then applies an improved threshold computation function to allocate buffer space differentially across traffic classes while maintaining allocation smoothness. Our evaluation results demonstrate that EBM significantly improves performance under realistic workloads. For instance, it reduces the 99th percentile Flow Completion Time (FCT) slowdown by 32.7% for short flows in the web-search workload and by 45.1% for incast flows in the hadoop workload, all without sacrificing overall throughput.
Fu Xiao 0001, Huipeng Huang, Weibei Fan, Mengjie Lv, Xueli Sun, Yiping Zuo, Sun-Yuan Hsieh
IEEE Trans. Computers4
2026 Reliability Assessment of Generalized Hypercube Networks Under a Probabilistic Fault Model
Mengjie Lv, Sixiao Di, Fu Xiao 0001, Weibei Fan, Sun-Yuan Hsieh
IEEE Trans. Dependable Secur. Comput.1
2026 Fault-Tolerant Communication Mechanism Based on Disjoint Paths in Interconnection Networks
abstract
Different interconnection structures exert a significant impact on network communication ability, directly influencing system performance. The half hypercube Network has an excellent topology that can provide high network fault tolerance and communication efficiency while maintaining a low node degree. In this paper, we investigate efficient and reliable communication algorithms for half hypercube networks in distributed system. Firstly, we design a disjoint path construction algorithm for a half hypercube, which enables reliable communication of the optimal number of disjoint paths between any two nodes in the network. Secondly, we present a fault-tolerant path embedding algorithm for a half hypercube. When the number of faulty nodes does not exceed ⌈n/2⌉, this algorithm can obtain a fault-tolerant unicast path between any two non-faulty nodes in ann-dimensional half hypercube network. Finally, we evaluate the performance of communication algorithms through simulation experiments and real testbed. Experimental results demonstrate that the efficiency and buffer utilization rate of the proposed algorithms can be improved by at least 21.8% and 15.6%, respectively. Testbed results show that the data delivery rate increased by 21.8%, and the path interference degree decreased by 32.5%.
Weibei Fan, Xuanli Liu, Fu Xiao 0001, Mengjie Lv, Sun-Yuan Hsieh
IEEE Trans. Netw.4
2026 A Scalable and High-Performance Architecture for Data Center Networks
Xuanli Liu, Weibei Fan, Zhenjiang Dong, Fu Xiao 0001, Mengjie Lv, Xueli Sun, Sun-Yuan Hsieh
IEEE Trans. Netw.5
2026 A Highly Scalable and Fault-Tolerant Topology for Data Center Networks
abstract
As the demand for cloud services and data-intensive applications continues to surge, the design of efficient and reliable data center network (DCN) topologies has become increasingly critical. However, traditional DCNs often face challenges of limited scalability, insufficient fault tolerance, and high communication latency. To address these issues, we introduce SFDC, a novel recursive and modular server-centric network topology. SFDC is built on a hierarchical element-layer structure that enables the construction of highly scalable and fault-tolerant networks. The modular design of SFDC supports flexible expansion, allowing for the integration of servers with varying network interface card (NIC) configurations without requiring significant redesigns. Furthermore, SFDC’s design effectively mitigates the growth of network diameter, ensuring low latency even at massive scales. We also propose a routing algorithm, SFRouting, which leverages SFDC’s hierarchical structure to efficiently compute unicast paths, while minimizing routing complexity and enhancing data transmission efficiency. Additionally, we present a multipath routing scheme based on disjoint path construction, which ensures robust communication by providing alternative paths in case of node or link failures, thus enhancing network fault tolerance. Experimental results demonstrate that SFDC outperforms existing DCN topologies such as BCube, DCell, and HS-DCell, exhibiting superior scalability, reduced network diameter, and enhanced fault tolerance while maintaining low latency and stable performance.
Mengjie Lv, Wenjie Wan, Fu Xiao 0001, Weibei Fan, Sun-Yuan Hsieh
IEEE Trans. Netw.1
2025 Disjoint paths construction algorithm in the data center network DPCell
abstract
Abstract With the development of the fourth industrial revolution, the importance of data centers has significantly increased. Data centers are widely used in many fields due to their ability to provide efficient, secure, and reliable data storage and processing services. However, with the increasing amount of data, traditional data center networks (DCNs) are currently facing various challenges, prompting academia and industry to propose new DCN architectures. As a dual-port server-based DCN, DPCell has excellent scalability and bisection width, enabling it to meet the demands of large-scale data storage, processing, and computation in the digital revolution. In order to ensure the secure and reliable data communication in the DPCell, this paper designs a disjoint paths communication scheme based on the actual DCN routing requirements. This scheme constructs the optimal number of disjoint paths in DPCell, with a maximum path length of $2^{k}+3$, where $k$ represents the dimension of the DPCell. Furthermore, experiments have verified that the time complexity of this scheme is sublinear, making it more efficient than the current optimal maximum flow algorithm. To a certain extent, this scheme provides DPCell with the required high bandwidth, fault tolerance, and security for data communication.
Huaqun Wang, Mengjie Lv, Weibei Fan
Comput. J.3
2025 Fault tolerance assessment of the data center network DPCell based on g-good-neighbor conditions
abstract
Abstract Data center networks (DCNs) provide critical data storage and computing services for cloud computing. The continuous increase in demand for cloud computing has led to a surge in data volume, necessitating the continual expansion of DCNs. However, this expansion also heightens the risk of device failures. Therefore, it is particularly important to study the fault tolerance of DCNs, which refers to their ability to ensure reliable communication even in the presence of device failures. Among DCNs constructed using dual-port servers, DPCell achieves higher scalability and bisection width while maintaining a smaller diameter. This paper assesses the fault tolerance of DPCell using two metrics: connectivity and diagnosability. Recognizing the limitations of traditional connectivity and diagnosability, we investigate the connectivity and diagnosability of DPCell under the condition that each fault-free node in the network has at least $g$ fault-free neighbors. The results indicate that, under this condition, the connectivity and diagnosability of DPCell exceed its traditional metrics by more than $g$ times.
Huaqun Wang, Mengjie Lv, Weibei Fan
Comput. J.3
2025 Efficient fault tolerance and diagnosis mechanism for Network-on-Chips
Mengjie Lv, Weibei Fan
J. Netw. Comput. Appl.1
2025 Reliable Communication Scheme Based on Completely Independent Spanning Trees in Data Center Networks
abstract
With technological advancements, real-time applications have permeated various aspects of human life, relying on fast, reliable, and low-latency data transmission for seamless user experiences. The development of data center networks (DCNs) has greatly advanced real-time applications, with network reliability being a key factor in ensuring high-quality network services. As a switch-centric DCN, DPCell has good scalability and the ability to achieve load balancing at different traffic levels. With the increasing demand for high availability, fault tolerance, and efficient data transmission, highly reliable communication for DPCell is essential. Completely independent spanning trees (CISTs) play a significant role in enhancing reliable communication performance in networks. This paper proposes an algorithm for constructing CISTs in DPCell, which has relatively low time and space consumption compared to other CISTs construction algorithms in DCNs, offering an efficiency advantage. Communication simulations validate the effectiveness of using paths provided by CISTs in DPCell for data transmission. Furthermore, experimental results show that a multi-protection routing scheme configured with multiple CISTs significantly enhances fault tolerance in DPCell.
Huaqun Wang, Mengjie Lv, Weibei Fan
IEEE Trans. Computers3
2025 Reliable and Efficient Multi-Path Transmission Based on Disjoint Paths in Data Center Networks
abstract
Multi-path transmission enables load balancing and improves network performance in data center networks (DCNs). It increases the possibility of network congestion and makes traditional network traffic engineering methods inefficient due to the uneven distribution of network traffic in data centers. In this paper, we present a reliable and efficient Disjoint paths based Multi-Path Transmission scheme (DMPT) that selects distributed requests through topology awareness. Firstly, we propose disjoint path construction algorithms through rigorous theoretical proof, aiming at the different transmission requirements of DCNs. Secondly, we offer an optimal solution to the disjoint multi-path selection problem, which is aimed at the trade-off between link load and transmission time. Furthermore,DMPTcan split the flow over multiple transmission paths based on the link status. Finally, extensive experiments are executed forDMPTon a novel EHDC of DCN that is based on exchanged hypercube. The experimental results show thatDMPTcan reduce the average running time by 18.6%, and the average path length is close to the optimal path. Furthermore, it achieves significant improvements in balancing network link traffic and facilitating deployment, which also reflects the advantages of topology aware multiplexing in practice.
Weibei Fan, Fu Xiao 0001, Mengjie Lv, Shui Yu 0001
IEEE Trans. Computers4
2025 A Highly Reliable Multiplexing Scheme in Hypercube-Structured Hierarchical Networks
abstract
The design and optimization of network topologies play a critical role in ensuring the performance and efficiency of high-performance computing (HPC) systems. Traditional topology designs often fall short in satisfying the stringent requirements of HPC environments, particularly with respect to fault tolerance, latency, and bandwidth. To address these limitations, we propose a novel class of hierarchical networks, termed Hypercube-Structured Hierarchical Networks (HHNs). This architecture generalizes and extends existing architectures such as half hypercube networks and complete cubic networks, while also introducing previously unexplored hierarchical designs. HHNs exhibit several advantages, particularly in high-performance computing. Most notably, their high connectivity enables efficient parallel data processing, and their hierarchical structure supports scalability to accommodate growing computational demands. Furthermore, we present a unicast routing strategy and a broadcast algorithm for HHNs. A fault-tolerant algorithm is also designed based on the construction of disjoint paths. Experimental evaluations demonstrate that HHNs consistently outperform mainstream architectures in critical performance metrics, including scalability, latency, and robustness to failures.
Xuanli Liu, Zhenjiang Dong, Weibei Fan, Mengjie Lv, Xueli Sun, Sun-Yuan Hsieh
IEEE Trans. Computers4
2025 Constructing completely independent spanning trees in the generalized hypercube network
Huaqun Wang, Mengjie Lv, Weibei Fan
J. Supercomput.3
2025 Adaptive system-level fault diagnosis of hierarchical cubic networks
Mengjie Lv, Sixiao Di, Weibei Fan
J. Supercomput.1
2025 Reliability of hierarchical cubic networks based on component fault pattern
Mengjie Lv, Xuanli Liu, Weibei Fan
J. Supercomput.1
2025 Dynamic Topology and Resource Allocation for Distributed Training in Mobile Edge Computing
abstract
In mobile edge computing (MEC), edge servers and mobile terminals use federated learning distributed architecture to build a deep model, so that terminals can cooperate in training without sharing data. Distributed training requires network virtualization to provide high bandwidth and low latency characteristics to support large-scale parallel computing. Traditional virtual network embedding (VNE) relies on a static network topology, which lacks flexibility and incurs high resource costs during model training. To improve the efficiency of embedding distributed training tasks, we propose a novel Node Selection and Dynamic Topology resource allocation scheme for VNE of distributed training, NSDT-VNE, based on reconfigurable network topology. This algorithm divides the underlying network into static and dynamic topologies, enhancing low latency for small flows while providing high bandwidth for large flows as needed. Additionally, we introduce a two-phase coordinated alternating optimization algorithm that optimizes embedding decisions at both computational and topological levels, ensuring optimal node selection. Overall, NSDT-VNE follows demand-aware network design principles, allowing continuous optimization of the underlying topology. Compared to state-of-the-art heuristic and reinforcement learning-based virtual network algorithms, NSDT-VNE achieves superior performance, with request acceptance rates improving by 6.67% to 25.68% and embedding revenue increasing by approximately 7% to 32%.
Weibei Fan, Donglai Wang, Fu Xiao 0001, Yiping Zuo, Mengjie Lv, Sun-Yuan Hsieh
IEEE Trans. Mob. Comput.5
2025 Efficient Parallel Adaptive Diagnosis of a Class of Data Center Networks
abstract
Designing a data center network (DCN) architecture capable of accommodating and managing thousands of servers is crucial for optimizing computing services in fields, such as Big Data, cloud computing, and artificial intelligence. This article introduces a novel class of DCN architecture, termed hypercube-like data center network (HLDN), which extends existing frameworks like high scalability data center network and crossed cube-based scalability data center network, while exploring previously unexamined designs. As the scale of DCNs grows and the number of servers increases, the incidence of network failures also rises, impacting network reliability. To address this challenge, we propose the first adaptive diagnostic schemes specifically designed for DCNs. We investigate the Hamiltonian properties of HLDN and, based on this analysis, present parallel adaptive diagnostic algorithms under the Preparata, Metze, and Chien (PMC) and comparison (COM) models. Simulation experiments validate the effectiveness of these algorithms, showing that the PMC model-based approach significantly outperforms the COM model in terms of diagnostic performance under identical conditions. Furthermore, comparative analyses with existing methods highlight the enhanced diagnostic accuracy of our proposed algorithms. This work not only provides an innovative solution for improving the reliability of DCNs, but also offers valuable theoretical and practical insights for fault detection and management in large-scale network systems.
Mengjie Lv, Weibei Fan
IEEE Trans. Reliab.1
2025 An Incremental Scalable Network Architecture With Fault-Tolerant Communication
abstract
The design of interconnection network topologies significantly impacts the performance and reliability of parallel systems. Enhanced incremental scalability enables networks to expand with reduced hardware overhead. In practice, rather than always adding many nodes at once, a small number of nodes are occasionally added as needed. However, existing topologies struggle to achieve effective incremental scalability. To address this, we propose the incremental scalability exchanged hypercube (ISEH), a novel interconnection network for parallel computing. The significant advantages of ISEH include improved incremental scalability and interconnection flexibility, while maintaining low interconnection complexity. Its diameter remains unchanged as the network size increases linearly and does not exceed the diameter of the exchanged hypercube. First, we present the topological properties of ISEH, including isomorphism, incremental scalability, and diameter. Next, we design an efficient communication method for ISEH to ensure low communication overhead. To support reliable communication, we design algorithms to construct disjoint paths between any two distinct nodes. Furthermore, based on generalized exchangedX-cubes, we propose the incremental scalability generalized exchangedX-cubes, offering better incremental scalability. Finally, we compare the performance of ISEH with other interconnection networks and evaluate the proposed algorithms. The results demonstrate that ISEH achieves a favorable balance among incremental scalability, diameter, and flexibility compared to existing networks.
Weibei Fan, Mengjie Lv, Xueli Sun, Shui Yu 0001
IEEE Trans. Reliab.3
2024 Reliability of Half Hypercube Networks under Cluster Faults
abstract
Malicious attackers frequently aim to partition the network into disjointed segments to facilitate specific attacks. Consequently, enhancing network reliability stands as an effective preventive measure. Connectivity serves as a crucial metric for gauging network reliability, yet classical connectivity inadequately captures a network's fault tolerance in the face of such attacks. To address this, cluster connectivity has been proposed, considering the faults within clusters to improve fault tolerance assessment. In this paper, we establish the cluster connectivity of the half hypercube network HHn. In detail, we show that the K1,1-cluster connectivity of HHnis $\left\lfloor {n/2} \right\rfloor + 1$, where n ≥ 3, and the K1,r- cluster connectivity of HHnis $\left\lceil {\frac{{\left\lceil {n/2} \right\rceil }}{2}} \right\rceil + 1$, where n ≥ 5 and 2 ≤ r ≤ 4, which is almost r times the classical connectivity. This indicates that the network possesses an enhanced capacity to accommodate a greater number of faulty nodes, potentially enabling more effective orchestration of attacks.
Xuanli Liu, Mengjie Lv, Weibei Fan, Xueli Sun, Zhenjiang Dong, Fu Xiao 0001
CSCWD2
2024 Node-disjoint Paths Construction Algorithm in Data Center Network EHDC
abstract
As a centralized location for computer systems, data centers provide high-performance computing hardware, storage devices, and network facilities for collaborative computing. The node-disjoint paths can be used to implement multi-path transmission in data center networks, which provide multiple high-quality transmission paths and improve the performance of the network. Moreover, the disjoint paths can also provide redundant transmission paths, which enhance the fault tolerance of the network. The EHDC network is a novel server-centric and highly scalable data center network based on exchanged hypercube, and its logical structure is ED(s, t). In this paper, we propose the algorithm NDPath to construct the node-disjoint paths between two distinct nodes when the two nodes are in the same EDs in ED(s, t). Moreover, we analyze the maximum length of the disjoint paths. Experimental results show that our proposed algorithm performs better than the classical algorithm Dijkstra in the Average Running Time (ART) and is very close to that in the Average Path Length (APL).
Weibei Fan, Mengjie Lv, Xin He 0010, Fu Xiao 0001
CSCWD3
2024 A protection routing with secure mechanism in the data center network WaveCube
abstract
In the era of information explosion, the scale of data center networks (DCNs) has expanded exponentially, consequently leading to an inevitable increase in server failures. Therefore, how to ensure the efficient and secure operation of the network has emerged as a critically important research topic. WaveCube is a scalable, fault-tolerant, high-performance optical DCN architecture. In this paper, we first propose a local secure model (LS model) of WaveCube. This model segments fault-free nodes within sub-Wavecube by imposing specific constraints, thereby adeptly circumventing potential communication impediments that could arise due to faulty nodes. Secondly, based on this model, we design a protection routing with secure mechanism to ensure stable communication within WaveCube. Finally, we perform a series of experiments, and the results show that when the number of faulty nodes is less than half of the number of total nodes, the hit rate can reach nearly 100%, while the shortest path rate can achieve up to 90%.
Jingman Pei, Mengjie Lv, Weibei Fan, Xueli Sun, Xin He 0010, Fu Xiao 0001
CSCWD2
2024 MBDC: Low Latency and Cost-effective Data Center Network Architecture
abstract
With the rapid development of information technologies such as cloud computing, big data, artificial intelligence, and edge computing, data centers have become essential infrastructure supporting the modern information society. When constructing data center networks, as the network scale increases, both latency and cost also grow. Therefore, it is essential not only to consider network scalability but also to focus on link overhead and communication latency. The hypercube is an excellent base topology for constructing data center networks. The Möbius cube, a version of the hypercube, not only retains the hypercube’s favorable properties, but also outperforms it in terms of link overhead and network diameter. In this paper, we propose a new server-centric data center network architecture, called MBDC, which is based on the Möbius cube. For networks of the same scale, MBDC achieves a smaller diameter than most existing server-centric networks. Additionally, we present an adaptive fault-tolerant routing scheme for MBDC, which is based on an improved local security information model. Extensive evaluations demonstrate that MBDC is an attractive data center network for constructing low-latency and cost-effective data centers.
Jiguo Yu, Anming Dong, Li Zhang 0122, Mengjie Lv
HPCC6
2024 Parallel Construction of Independent Spanning Trees on 3-ary n-cube Networks
Yuzhen Xu, Weibei Fan, Mengjie Lv, Xueli Sun, Fu Xiao 0001
NPC (1)3
2024 BufferConcede: Conceding Buffer for RoCE Traffic in TCP/RoCE Mix-Flows
Lingxuan Meng, Kaiyun Liu, Weibei Fan, Fu Xiao 0001, Mengjie Lv
WASA (1)5
2024 Distributed Dynamic Virtual Network Embedding in Container Networks
Donglai Wang, Weibei Fan, Fu Xiao 0001, Mengjie Lv, Xueli Sun
WASA (2)4
2024 An Efficient Fault-Tolerant Communication Scheme in 3-Ary n-Cube Networks
Yuzhen Xu, Weibei Fan, Mengjie Lv, Xueli Sun, Fu Xiao 0001
WASA (2)3
2024 Fault tolerance of hierarchical cubic networks based on cluster fault pattern
abstract
Abstract Connectivity is a meaningful metric parameter and indicator for estimating network reliability and evaluating network fault tolerance. However, the traditional connectivity and current conditional connectivity do not take into account the association between a certain node and its neighboring nodes. In fact, adjacent nodes are easily influenced by each other so that the failing probability of adjacent nodes around a faulty node is high. Therefore, cluster and super cluster connectivities are proposed to more intuitively measure the fault tolerance of the network. In this paper, we mainly explore the cluster connectivity and super cluster connectivity of the hierarchical cubic network $HCN_{n}$. In detail, we show that $\kappa (HCN_{n}\mid K_{1, 0}(K_{1, 0}^{*}))=n+1$, $\kappa (HCN_{n}\mid K_{1, 1}(K_{1, 1}^{*}))=\kappa ^{\prime}(HCN_{n}\mid K_{1, 1}(K_{1, 1}^{*}))=n+1$, $\kappa (HCN_{n}\mid K_{1, m}(K_{1, m}^{*}))=\lceil n/2\rceil +1$ ($2\leq m\leq 4$), $\kappa ^{\prime}(HCN_{n}\mid K_{1, 0}(K_{1, 0}^{*}))=2n$, and $\kappa ^{\prime}(HCN_{n}\mid K_{1, m}(K_{1, m}^{*}))=n+1$ ($2\leq m\leq 3$) if $n$ is odd and $\kappa ^{\prime}(HCN_{n}\mid K_{1, m}(K_{1, m}^{*}))=n$ ($2\leq m\leq 3$) if $n$ is even, where $n\geq 4$.
Mengjie Lv, Weibei Fan
Comput. J.1
2024 An expandable and cost-effective data center network
Mengjie Lv, Xuanli Liu, Weibei Fan
J. Netw. Comput. Appl.1
2024 Construction algorithms of fault-tolerant paths and disjoint paths in k-ary n-cube networks
Mengjie Lv, Jianxi Fan, Baolei Cheng, Jia Yu 0003, Xiaohua Jia
J. Parallel Distributed Comput.1
2024 Efficient Fault-Tolerant Path Embedding for 3D Torus Network Using Locally Faulty Blocks
abstract
3D tori are significant interconnection architectures in building supercomputers and parallel computing systems. Due to the rapid growth of edge faults and the crucial role of path structures in large-scale distributed systems, fault-tolerant path embedding and correlated issues have drawn widespread researches. However, existing path embedding methods are based on traditional fault models, allowing all faults to be near the same node, so they usually only focus on theoretical proof and generate linear fault-tolerance related to dimension$n$. In order to improve the fault-tolerance of 3D torus, we first propose a novel conditional fault model called the Locally Faulty Block model (LFB model). On the basis of this model, the Hamiltonian paths with large-scale edge defects in torus are investigated. After that, we construct an Hamiltonian path embedding algorithm HP-LFB into torus with$O(N)$under the LFB model, where$N$is the number of nodes in torus. Furthermore, we present an adaptive routing algorithm HoeFA, which is based on the method of distance vector to limit the use of virtual channels (VCs). We also make a comparison with state-of-the-art schemes, indicating that our scheme enhance other comprehensive results. The experiment indicated that HP-LFB can sustain the dynamic degradation of the batting average of establishing Hamiltonian paths, with the added faulty edges exceeding fault-tolerance.
Weibei Fan, Fu Xiao 0001, Mengjie Lv, Shui Yu 0001
IEEE Trans. Computers3
2024 Cluster connectivity and super cluster connectivity of half hypercube networks
Xuanli Liu, Mengjie Lv, Weibei Fan, Xueli Sun
Theor. Comput. Sci.2
2024 Extra connectivity of the data center network - RRect
Ni An, Mengjie Lv, Weibei Fan, Fu Xiao 0001
J. Supercomput.2
2024 Hamiltonian cycle embedding with fault-tolerant edges and adaptive diagnosis in half hypercube
Weibei Fan, Xuanli Liu, Mengjie Lv
J. Supercomput.3
2024 Reliability analysis of complete cubic networks based on extra conditional fault
Mengjie Lv, Xuanli Liu, Weibei Fan
J. Supercomput.1
2024 Fault-Tolerant Communication in HSDC: Ensuring Reliable Data Transmission in Smart Cities
abstract
As the core of cloud computing, the data center network (DCN) provides services and decision support for smart cities by providing powerful data storage and computing capabilities. As a server-centric DCN, the high scalability data center network architecture (HSDC) can cope with the rapid growth of data volume and provide an effective service foundation for smart cities. However, the rapid development of smart cities requires that DCNs can still operate reliably in the presence of faulty while expanding its scale. Therefore, it is very important to design reliable fault-tolerant communication algorithms in DCNs. This article investigates the fault-tolerant communication algorithm in HSDC. Initially, we propose an$O(n^{2})$algorithm to establish$n$disjoint paths between any pair of nodes in the$n$-dimension HSDC. Simulation experiments show that the disjoint paths generated by our algorithm in HSDC have a maximum length only 3 longer than the diameter, which guarantees a small communication delay in the worst case. In addition, we propose an$O(n)$algorithm to establish a fault-tolerant unicast path between any pair of fault-free nodes in the$n$-dimension HSDC. Moreover, simulation experiments indicate that as the scale of HSDC increases, the algorithm performs well in both running efficiency and the length of constructed paths.
Mengjie Lv, Weibei Fan
IEEE Trans. Reliab.2
2023 Fault-tolerant unicast using conditional local safe model in the data center network BCube
Mengjie Lv, Huaqun Wang, Weibei Fan
J. Parallel Distributed Comput.2
2023 Reliability evaluation of half hypercube networks
Mengjie Lv, Weibei Fan
Theor. Comput. Sci.2
2023 Node Essentiality Assessment and Distributed Collaborative Virtual Network Embedding in Datacenters
abstract
Network virtualization (NV) has extensive and significant applications in cloud computing and parallel and distributed systems. Virtual network embedding (VNE) is a key issue in NV, which is an effective means to advance systems’ performance. While existing VNE research lacks resource allocation coordination between mappings of different virtual network requests, resulting in insufficient resource utilization and high overhead. In this article, we propose a novel node essentiality evaluation model for data center networks (DCNs), and design an efficient distributed collaborative virtual network embedding. Firstly, we propose a node essentiality evaluation scheme based on dynamic model, which combines the characteristics of network topology and nodes to make the evaluation results more comprehensive. Secondly, we establish the two-stage node importance evaluation criteria for the deviation mean of the data center dynamic model and the variance based on the deviation mean. Furthermore, we investigate a nodal importance assessment method based on the data center dynamic model for perturbation testing. Finally, we design a distributed coordinated VNE algorithm (CNI-VNE) which calculates the importance index of physical nodes through topology awareness. The proposed algorithm can increase the coordination between different request mappings, thereby reducing the mapping cost of physical node resources and minimizing the cost of VNE. We use the real Fat-tree DCN of 128 servers and 80 switches as testbed, and evaluate them from indicators such as average reliability, average bandwidth consumption, average energy consumption, and average mapping time. Massive simulation results in different scenarios show that our algorithm achieves the best performance on most indicators compared with the existing state-of-the-art proposals, mapping acceptance and average revenue increased by 19.4% and 21.3%, respectively, and DCN reduced bandwidth consumption by about 30%.
Weibei Fan, Fu Xiao 0001, Mengjie Lv, Junchang Wang, Xin He 0010
IEEE Trans. Parallel Distributed Syst.3
2022 The Reliability of k-Ary n-Cube Based on Component Connectivity
abstract
Abstract Connectivity and diagnosability are two crucial subjects for a network’s ability to tolerate and diagnose faulty processors. The $r$-component connectivity $c\kappa _{r}(G)$ of a network $G$ is the minimum number of vertices whose deletion results in a graph with at least $r$ components. The $r$-component diagnosability $ct_{r}(G)$ of a network $G$ is the maximum number of faulty vertices that the system can guarantee to identify under the condition that there exist at least $r$ fault-free components. This paper first establishes that the $(r+1)$-component connectivity of $k$-ary $n$-cube $Q^{k}_{n}$ is $c\kappa _{r+1}(Q^{k}_{n})=-\frac{1}{2}r^{2}+\Big(2n-\frac{1}{2}\Big)r+1$ for $n\geq 2$, $k\geq 4$ and $1\leq r\leq n$. In view of $c\kappa _{r+1}(Q^{k}_{n})$, we prove that the $(r+1)$-component diagnosabilities of $k$-ary $n$-cube $Q^{k}_{n}$ under the PMC model and MM* model are $ct_{r+1}(Q^{k}_{n})=-\frac{1}{2}r^{2}+\Big(2n-\frac{3}{2}\Big)r+2n$ for $n\geq 4$, $k\geq 4$ and $1\leq r\leq n-1$.
Mengjie Lv, Jianxi Fan, Jingya Zhou, Jia Yu 0003, Xiaohua Jia
Comput. J.1
2022 Fault Diagnosis Based on Subsystem Structures of Data Center Network BCube
abstract
Data center networks (DCNs) always strive to ensure high reliability and fault tolerance when Big Data processing and cloud computing are carried out. One effective method is to conduct the fault diagnosis based on subsystem structures (establish the subsystem-based reliability), which can be transformed into solving the problem of the probability that there exists at least one fault-free subnetwork in one network. In this article, we establish the subsystem-based reliability of BCube, which is the first time that fault diagnosis is carried out in the subsystem structures of the DCN. More specifically, we compute the upper and lower bounds on the subsystem-based reliability of BCube and determine the approximation on the subsystem-based reliability of BCube. Furthermore, we conduct some numerical simulations to validate the established analytical formulation. Our results show that the precise value of subsystem-based reliability of BCube can be basically represented by the approximation value of subsystem-based reliability of BCube, which means that it is much easier to evaluate the reliability of the network even when the network scale is sufficient large. Although the analysis is done for a particular network (BCube), the outcome can serve as a useful reference, and can shed light on the effectiveness of the fault diagnosis for other DCNs.
Mengjie Lv, Jianxi Fan, Weibei Fan, Xiaohua Jia
IEEE Trans. Reliab.1
2021 The Conditional Reliability Evaluation of Data Center Network BCDC
abstract
Abstract As the number of servers in a data center network (DCN) increases, the probability of server failures is significantly increased. Traditional connectivity is an important metric to measure the reliability of DCN. However, the traditional connectivity of a DCN based on the condition of arbitrary faulty servers is generally lower. Therefore, it is important to increase the connectivity of a DCN by adding some limited conditions for the faulty server set. As a result, $g$-restricted connectivity and $h$-extra connectivity, which are two crucial subjects for a DCN’s ability to tolerate faulty servers, were proposed in the literature. In this paper, we study the $g$-restricted connectivity and $h$-extra connectivity of a new server-centric DCN, called BCDC, based on crossed cube with excellent performance. We prove that the $g$-restricted connectivity of BCDC is 4 for $n=3$ and $2n+g(n-2)-2$ for $n\geq 4$, where $0\leq g\leq n-3$, and the $h$-extra connectivity of BCDC is 4 for $n=3$ and $2n+h(n-2)-2$ for $n\geq 4$, where $0\leq h\leq n-3$.
Mengjie Lv, Baolei Cheng, Jianxi Fan, Xi Wang 0006, Jingya Zhou, Jia Yu 0003
Comput. J.1
2020 Intermittent Fault Diagnosability of Some General Regular Networks
abstract
Fault tolerance plays an important role in the interconnection networks, where permanent and intermittent faults are two kinds of fault situations. Permanent fault diagnosabilities of regular networks have been proposed widely while the intermittent fault diagnosabilities are also noteworthy. In this paper, we give a sufficient and necessary condition for k-regular k-connected graph Gn to be ti-diagnosable without repair in intermittent fault pattern. Detailly, we show that the intermittent fault diagnosability of Gn under the PMC model is k−⌈g−12⌉−2⁠, where g is the maximum number of common neighbors for any two distinct vertices. As applications, intermittent fault diagnosabilities of many famous networks are explored.
Xueli Sun, Shuming Zhou, Mengjie Lv, Jiafei Liu 0001, Guanqin Lian
Comput. J.3
2020 The reliability analysis of k-ary n-cube networks
Mengjie Lv, Jianxi Fan, Baolei Cheng, Jingya Zhou, Jia Yu 0003
Theor. Comput. Sci.1
2020 The extra connectivity and extra diagnosability of regular interconnection networks
Mengjie Lv, Jianxi Fan, Jingya Zhou, Baolei Cheng, Xiaohua Jia
Theor. Comput. Sci.1
2020 On Reliability of Multiprocessor System Based on Star Graph
abstract
As a critical parameter in evaluating the reliability of a multiprocessor system when processors malfunction, the \boldmath h-extra connectivity (h-EC) of a multiprocessor system modeled by a graph G, denoted by κo(h)(G), is an h-extra vertex-cut with minimum cardinality. Both of the h-extra conditional diagnosability (h-ECD) and the t/h-diagnosability of the multiprocessor system are vital to tolerate and diagnose faulty processors. These two parameters rely on the resolving of hEC. For the multiprocessor system based on star graph Sn, we show that the 5-EC κo(5)(Sn) of Sn(n ≥ 5) is 6n - 18. As a by-product, we present a novel proof of κo(2)(Sn) = 3n - 7 (resp., κo(4)(Sn) = 5n - 14) by relaxing the restriction n ≥ 10 (resp., n ≥ 7) to n ≥ 5 (resp., n ≥ 5). Furthermore, we determine that the h-ECD of Sn(n ≥ 5) under the preparata, metze, and chien (PMC) model is (h + 1)n - 2h - 1 for 1 ≤ h ≤ 3 and (h + 1)n - 3h + 2 for 4 ≤ h ≤ 5. In addition, we show that Snis [(h + 1)n - 4h + 2]/h-diagnosable for 4 ≤ h ≤ 5, which extends the result that Snis [(h + 1)n - 3h - 1]/h-diagnosable for 1 ≤ h ≤ 3 by [Zhou et al. “The t/k-diagnosability of star graph networks,” IEEE Trans. Comput., vol. 64, no. 2, pp. 547-555, Feb. 2015].
Mengjie Lv, Shuming Zhou, Gaolin Chen, Lanxiang Chen, Jiafei Liu 0001, Chin-Chen Chang 0001
IEEE Trans. Reliab.1
2019 Fault diagnosability of DQcube under the PMC model
Mengjie Lv, Shuming Zhou, Jiafei Liu 0001, Xueli Sun, Guanqin Lian
Discret. Appl. Math.1
2019 Reliability of (n, k)-star network based on g-extra conditional fault
Mengjie Lv, Shuming Zhou, Xueli Sun, Guanqin Lian, Jiafei Liu 0001
Theor. Comput. Sci.1
2019 Probabilistic diagnosis of clustered faults for hypercube-based multiprocessor system
Mengjie Lv, Shuming Zhou, Xueli Sun, Guanqin Lian, Jiafei Liu 0001, Dajin Wang
Theor. Comput. Sci.1
2019 Fault tolerance analysis of hierarchical folded cube
Xueli Sun, Qingfeng Dong, Shuming Zhou, Mengjie Lv, Guanqin Lian, Jiafei Liu 0001
Theor. Comput. Sci.4