EDBT 2026 Demo / reviewers in the wild / expert
Yuzhong Sun
dblp:22/5402
· DBLP profile ↗
36ranked-venue papers
7as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 23 · 5 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-authorComputer networks · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 2Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Broad fractional order fuzzy system for noise and outlier resistant data classification
Tong Zhang 0015, Tao Zhang 0103, C. L. Philip Chen, Yuzhong Sun |
Expert Syst. Appl. | 5 |
| 2026 | DeFT: Relaxing data dependencies for efficient communication scheduling in distributed training
Yuzhong Sun, Jie Zhu 0002 |
Future Gener. Comput. Syst. | 2 |
| 2025 | Parallelizing Sharpness-Aware Minimization: A Semi-asynchronous, Small-Batch Approach
Yuzhong Sun |
ICANN (1) | 2 |
| 2025 | Fine-Grained State Sharding and Communication Pipelining for Partially Sharded Data ParallelismabstractPartially sharded data parallel (PSDP) reduces GPU memory footprint in data-parallel training without incurring additional communication overhead. Existing work on PSDP has primarily focused on minimizing memory consumption, such as by offloading model states to CPU, while not enough attention has been paid to poor communication performance. We observe that the imbalanced distribution of model states across GPUs and inter-layer data dependencies lead to severe underutilization of network bandwidth and poor scalability of the current PSDP system. To address these issues, we present Chimera, a system that enhances the communication efficiency of PSDP in both offloading and non-offloading settings. First, Chimera introduces a fine-grained state-sharding strategy to balance both model state and communication load across GPUs, while leveraging tensor fusion to amortize communication launch overhead. Second, we design a multi-stage, fine-grained communication pipelining mechanism that decouples inter-layer dependencies, thereby improving bandwidth utilization. Finally, we develop a practical auto-tuning tensor fusion algorithm to adaptively optimize fusion decisions for end-to-end training performance. We evaluate Chimera on cloud GPU clusters using representative large language models. Experimental results demonstrate that Chimera improves training throughput by$1.1 \times$to$1.8 \times$over state-of-theart PSDP systems. Yuzhong Sun |
ICPADS | 2 |
| 2025 | An effective meta-heuristics for trajectory planning problem in UAV-assisted vessel emission detection system
Jie Zhu 0002, Weizhi Cui, Haiping Huang, Yuzhong Sun |
Peer Peer Netw. Appl. | 5 |
| 2024 | A multi-hierarchy particle swarm optimization-based algorithm for cloud workflow scheduling
Chang Lu 0012, Jie Zhu 0002, Haiping Huang, Yuzhong Sun |
Future Gener. Comput. Syst. | 4 |
| 2024 | An effective trajectory planning heuristics for UAV-assisted vessel monitoring system
Jie Zhu 0002, Kaiyu Guo, Haiping Huang, Reza Malekian, Yuzhong Sun |
Peer Peer Netw. Appl. | 6 |
| 2023 | Near-Linear Scaling Data Parallel Training with Overlapping-Aware Gradient CompressionabstractExisting Data Parallel (DP) trainings for deep neural networks (DNNs) often experience limited scalability in speedup due to substantial communication overheads. While Overlapping technique can mitigate such problem by paralleling communication and computation in DP, its effectiveness is constrained by the high communication-to-computation ratios (CCR) of DP training tasks. Gradient compression (GC) is a promising technique to obtain lower CCR by reducing communication volume directly. However, it is challenging to obtain real performance improvement by applying GC into Overlapping because of (1) severe performance penalties in traditional GCs caused by high compression overhead and (2) decline of Overlapping benefit owing to the possible data dependency in GC schemes. In this paper, we propose COVAP, a novel GC scheme designing a new coarse-grained filter, makes the compression overhead close to zero. COVAP ensures an almost complete overlap of communication and computation by employing adaptive compression ratios and tensor sharding tailored to specific training tasks. COVAP also adopts an improved error feedback mechanism to maintain training accuracy. Experiments are conducted on Alibaba Cloud ECS instances with different DNNs of real-world applications. The results illustrate that COVAP outperforms existent GC schemes in time-to-solution by 1.92x-15.39x and exhibits near-linear scaling. Furthermore, COVAP achieves best scalability under experiments on four different cluster sizes. Yuzhong Sun |
ICPADS | 2 |
| 2021 | Docker Container Networking Based Apache Storm and Flink Benchmark TestabstractMany distributed stream computing engines have emerged to handle big data, and they can be deployed in cloud environments consisting of native networks or container networks. Most of the benchmark research on stream computing engines are carried out under the native network, and the research on the impact on container network on stream computing engines is currently inadequate. However, the use of container network will inevitably lead to performance degradation, which is the disadvantage of all virtual networks. In this work, we build Apache Storm and Apache Flink, which are Streaming Computation Engines in container network and native network environments and conduct performance measurements through experiments processing textual data to verify how much performance decreases in container network. Experiments show that the throughput in a container network environment is 1%-5% lower and CPU utilization is 11%-18% lower than in a local network environment. Zhihong Yang, Yuzhong Sun |
APNOMS | 3 |
| 2017 | Evolution of Cloud Operating System: From Technology to Ecosystem
Zuoning Chen, Kang Chen 0001, Jinlei Jiang, Lufei Zhang, Song Wu 0001, Zhengwei Qi, Chunming Hu, Yongwei Wu 0001, Yuzhong Sun, Aobing Sun, Zilu Kang |
J. Comput. Sci. Technol. | 9 |
| 2014 | MiCA: Real-Time Mixed Compression Scheme for Large-Scale Distributed MonitoringabstractReal-time monitoring, providing the real-time status information of servers, is indispensable for the management of distributed systems, e.g. failure detection and resource scheduling. The scalability of fine-grained monitoring faces more and more severe challenges with scaling up distributed systems. The real-time compression which suppresses remote information update to reduce continuous monitoring cost is a promising approach to address the scalability problem. In this paper, we present the Linear Compression Algorithm (LCA) which is the application of the linear filter to real-time monitoring. To our best knowledge, existing work and LCA only explores the correlations of values of each single metric at various times. We present a novel lightweight REal-time Compression Algorithm (ReCA) which employs discovery methods of the correlation among metrics to suppress remote information update in distributed monitoring. The compression algorithms mentioned above have limited compression power because they only explore either the correlations of values of each single metric at various times or that among metrics. Therefore, we propose the Mixed Compression Algorithm (MiCA) which explores both of the correlations to achieve higher compression ratio. We implement our algorithms and an existing compression algorithm denoted by CCA in a distributed monitoring system Ganglia and conduct extensive experiments. The experimental results show that LCA and ReCA have comparable compression ratios with CCA, that MiCA achieves up to 38.2%, 27% and 44.5% higher compression ratios than CCA, LCA and ReCA with negligible overhead, respectively, and that LCA, and ReCA can both increase the scalability of Ganglia about 1.5 times and MiCA can increase about 2.33 times under a mixed-load circumstance. Bo Wang 0018, Yuzhong Sun, Jun Liu 0002 |
ICPP | 3 |
| 2013 | Insight and reduction of MapReduce stragglers in heterogeneous environmentabstractSpeculative and clone execution are existing techniques to overcome the problems of task stragglers and performance degradation in heterogeneous clusters for big data processing. In this paper, we propose an alternative approach to solving the problems based on analysis results of profiling and the relations of the system parameters. Our approach adjusts the amount of task slots of nodes dynamically to match the processing power of the nodes, according to current task progress rate and resource utilization. It contrasts with the existing techniques by attempting to prevent task stragglers from occurring in the first place through maintaining a balance between resource supply and demand. We have implemented this method in the Hadoop MapReduce platform, and the TPC-H benchmark results show that it achieves 20–30% performance improvement and 35–88% less stragglers than existing techniques. Yuzhong Sun, Yin Song, Minhao Xu |
CLUSTER | 3 |
| 2013 | SecMon: A Secure Introspection Framework for Hardware VirtualizationabstractWith the fusion of cloud computing and virtualization technology, system security under virtualization becomes a key point in recent research. As a foundational technology to construct a secure system, virtual machine introspection receives more attention than ever. Almost all of the existing virtual machine monitors take the privileged virtual machine (Domain-0) as the monitoring machine, which ignore the threats brought by Domain-0 because of its huge code base of user-level tools. Besides, para-virtualized machines cannot provide the basic support for popular security applications of Windows operating system. This paper proposes a secure monitoring framework based on hardware virtualization. We use Windows operating system to build a monitoring virtual machine in hardware virtual machine domain, and set up monitoring mechanism in it. In addition, the security of the Windows monitoring machine itself is ensured all through its lifetime-bootstrap and runtime. The experiments show our secure monitoring system performs well in the secure monitoring process. The performance overhead it brings is considered to be acceptable. Yunwei Gao, Xinhui Tian, Baiming Feng, Yuzhong Sun |
PDP | 7 |
| 2013 | A Two-Tiered On-Demand Resource Allocation Mechanism for VM-Based Data CentersabstractIn a shared virtual computing environment, dynamic load changes as well as different quality requirements of applications in their lifetime give rise to dynamic and various capacity demands, which results in lower resource utilization and application quality using the existing static resource allocation. Furthermore, the total required capacities of all the hosted applications in current enterprise data centers, for example, Google, may surpass the capacities of the platform. In this paper, we argue that the existing techniques by turning on or off servers with the help of virtual machine (VM) migration is not enough. Instead, finding an optimized dynamic resource allocation method to solve the problem of on-demand resource provision for VMs is the key to improve the efficiency of data centers. However, the existing dynamic resource allocation methods only focus on either the local optimization within a server or central global optimization, limiting the efficiency of data centers. We propose a two-tiered on-demand resource allocation mechanism consisting of the local and global resource allocation with feedback to provide on-demand capacities to the concurrent applications. We model the on-demand resource allocation using optimization theory. Based on the proposed dynamic resource allocation mechanism and model, we propose a set of on-demand resource allocation algorithms. Our algorithms preferentially ensure performance of critical applications named by the data center manager when resource competition arises according to the time-varying capacity demands and the quality of applications. Using Rainbow, a Xen-based prototype we implemented, we evaluate the VM-based shared platform as well as the two-tiered on-demand resource allocation mechanism and algorithms. The experimental results show that Rainbow without dynamic resource allocation (Rainbow-NDA) provides 26 to 324 percent improvements in the application performance, as well as 26 percent higher average CPU utilization than traditional service computing framework, in which applications use exclusive servers. The two-tiered on-demand resource allocation further improves performance by 9 to 16 percent for those critical applications, 75 percent of the maximum performance improvement, introducing up to 5 percent performance degradations to others, with 1 to 5 percent improvements in the resource utilization in comparison with Rainbow-NDA. Yuzhong Sun, Weisong Shi |
IEEE Trans. Serv. Comput. | 2 |
| 2012 | LVMCI: Efficient and Effective VM Live Migration Selection Scheme in Virtualized Data CentersabstractVirtualization can provide significant benefits in virtualized data centers by enabling efficient and effective live migration to ensure service level agreement(SLA). Most of existing studies make decision on which bad virtual machines (VMs) should be migrated to which appropriate physical machines (PMs) in terms of resource utilizations. However, migration actions may degrade migrated application performance due to extra CPU and bandwidth consumptions. Furthermore, negative performance interferences amongst applications scheduled to the same PM may arise given the poor performance isolations of VMs on a PM. We design and implement a VM migration selection system with less migration costs and application performance interferences, called LVMCI (Live Virtual machine Migration with less Costs and application Interference). We propose a migration cost evaluation model to analyze quantitatively the aspects (i.e. throughput and response latency) of application performance degradation. Dirty rate and frequent dirty rate are two key factors that affect iteration time and downtime. We implement a tool that measures these parameters before VMs are migrated. We distinguish the performance degradation of migrated applications caused by memory iteration phase and stop-and-copy phase, which helps to select VM migrated. Besides that, we propose a performance interference model which helps to select the destination PM. The experimental results show that our system can estimate memory iteration time and downtime with high accuracy, and ensures a high level of SLAs by minimizing performance degradation during migration process and performance interference among co-located VMs at the destination PM. Wei Zhang 0052, Mingfa Zhu, Yiduo Mei, Yunwei Gao, Yuzhong Sun |
ICPADS | 9 |
| 2012 | Autonomic Resource Allocation in Virtualized Data CentersabstractVirtualization has been widely adopted in data centers for improving efficiency and flexibility. Multiple applications are co-hosted in virtualized data centers. In order to meet the Service Level Agreements (SLA), how to allocate resources for multiple applications is an important and challenging task, especially when dealing with fluctuating workloads and complex server applications. Virtual Machine Monitor provides fine-grained resource allocation and live migration. In this paper, we develop RTCOIN-Qclouds, a response time-aware, cost-aware and interference-aware control framework that tunes resource allocation, which ensures a high level of meeting the SLAs. Every physical machine's resources are assigned to multiple virtual machines which run on it based on application's response time, rather than traditional methods based on resource utilization. Virtual machine migration allows data centers to rebalance workloads across physical machines. However, migration actions may lead to performance impact during the migration process. Current virtualization techniques do not provide effective performance isolation between virtual machines (VMs). Specially, hidden contention for physical resources impacts performance differently in different virtual machines. As to the problem of selecting which virtual machines to be migrated, we consider migration cost. As to the problem of selecting which physical machine to be placed, we consider performance interference. Furthermore, we experimentally validate the effectiveness of response time-aware resource allocation in our framework using microbenchmarks. Wei Zhang 0052, Mingfa Zhu, Qimeng Wu, Yuzhong Sun |
ISPA | 7 |
| 2012 | Performance Degradation-Aware Virtual Machine Live Migration in Virtualized ServersabstractLive migration of virtual machines(VMs) is widely used for system management in virtualized servers. When the loads increase and SLAs of some applications are violated, dynamic migration of virtual machines across physical machines (PMs) has the potential to ensure a high level of meeting the SLAs. Because of consuming extra CPU and bandwidth, application performance may be degraded during the migration process. However, different applications have different performance degradation. We design and implement a VM migration selection method that decides which VMs should be migrated. It can not only eliminate resouce competition on the PM, but also have less performance degradation during the migration process. We propose a performance degration-aware model to analyze applications' performance degradation which is directly sensitive to users. We analyze migration source code and find that memory size, dirty rate and frequent dirty rate are key factors that affect iteration time and downtime. We implement a tool that measures dirty rate and frequent dirty rate before VMs are migrated. we make a distinction between memory iteration phrase and stop-and-copy phrase owing to different performance degradation. The experimental results show that our method is effective. Wei Zhang 0052, Mingfa Zhu, Yiduo Mei, Yuzhong Sun |
PDCAT | 7 |
| 2011 | Green challenges to system software in data centers
Yuzhong Sun, Yiqiang Zhao, Yajun Yang, Haifeng Fang, Hongyong Zang, Yaqiong Li, Yunwei Gao |
Frontiers Comput. Sci. China | 1 |
| 2010 | TRIOB: A Trusted Virtual Computing Environment Based on Remote I/O Binding Mechanism
Haifeng Fang, Yiqiang Zhao, Yuzhong Sun, Zhiyong Liu 0002 |
CANS | 4 |
| 2010 | VMGuard: An Integrity Monitoring System for Management Virtual MachinesabstractA cloud computing provider can dynamically allocate virtual machines (VM) based on the needs of the customers, while maintaining the privileged access to the Management Virtual Machine that directly manages the hardware and supports the guest VMs. The customers must trust the cloud providers to protect the confidentiality and integrity of their applications and data. However, as the VMs from different customers are running on the same host, an attack to the management virtual machine will easily lead to the compromise of the guest VMs. Therefore, it is critical for a cloud computing system to ensure the trustworthiness of management VMs. To this end, we propose VMGuard, an integrity monitoring and detecting system for management virtual machines in a distributed environment. VMGuard utilizes a special VM, Guard Domain, which runs on each physical node to monitor the co-resident management VMs. The integrity measurements collected by the Guard Domains are sent to the VMGuard server for safe store and independent analysis. The experimental evaluation of a Xen-based prototype shows that VMGuard can quickly detect the root kit attacks while the performance overhead is low. Haifeng Fang, Yiqiang Zhao, Hongyong Zang, H. Howie Huang, Yuzhong Sun, Zhiyong Liu 0002 |
ICPADS | 6 |
| 2010 | TRainbow: a new trusted virtual machine based platform
Yuzhong Sun, Haifeng Fang, Hongyong Zang, Yaqiong Li, Yajun Yang, Ran Ao, Yongbing Huang |
Frontiers Comput. Sci. China | 1 |
| 2009 | Multi-Tiered On-Demand Resource Scheduling for VM-Based Data CenterabstractThe trend of using virtualization for server consolidation is more and more popular in enterprise data center. However, on-demand resource allocation among the concurrent hosted services in such a virtualized environment is still a challenge. In order to optimize resource allocation among services in data center, this paper proposes a multi-tiered resource scheduling scheme which automatically provides on-demand capacities to the hosted services via resources flowing among VMs. We model the resource flowing using optimization theory. Based on this model, we present a global re-source flowing algorithm in the multi-tiered resource scheduling scheme. This algorithm preferentially ensures performance of some critical services by degrading of others to some extent when resource competition arises. Using our RAINBOW prototype, we evaluate the multi-tiered resource scheduling scheme with the performance improvements for the most critical services up to 9%~16%, which are 75% of the maximum improvement margin, while performance degradation of others is up to 2%, and leads to 1%~5% improvements in resource utilization than RAINBOW without resource flowing. Compared with the existent scheme, our work leads to 9% less improvements for critical services, while introduces 39% less degradation to low priority services. Yaqiong Li, Binquan Feng, Yuzhong Sun |
CCGRID | 5 |
| 2009 | Utility analysis for Internet-oriented server consolidation in VM-based data centersabstractServer consolidation based on virtualization technology will simplify system administration, reduce the cost of power and physical infrastructure, and improve utilization in today's Internet-service-oriented enterprise data centers. How much power and how many servers for the underlying physical infrastructure are saved via server consolidation in VM-based data centers is of great interest to administrators and designers of those data centers. Various workload consolidations differ in saving power and physical servers for the infrastructure. The impacts caused by virtualization to those concurrent services are fluctuating considerably which may have a great effect on server consolidation. This paper proposes a utility analytic model for Internet-oriented server consolidation in VM-based data centers, modelling the interaction between server arrival requests with several QoS requirements, and capability flowing amongst concurrent services, based on the queuing theory. According to features of those services' workloads, this model can provide the upper bound of consolidated physical servers needed to guarantee QoS with the same loss probability of requests as in dedicated servers. At the same time, it can also evaluate the server consolidation in terms of power and utility of physical servers. Finally, we verify the model via a case study comprised of one e-book database service and one e-commerce Web service, simulated respectively by TPC-W and SPECweb2005 benchmarks. Our experiments show that the model is simple but accurate enough. The VM-based server consolidation saves up to 50% physical infrastructure, up to 53% power, and improves 1.7 times in CPU resource utilization, without any degradation of concurrent services' performance, running on Rainbow - our virtual computing platform. Yuzhong Sun, Weisong Shi |
CLUSTER | 3 |
| 2009 | Optimizing Inter-Domain CommunicationabstractVirtual machine technology has played an important role in data center. Distributed services deployed in multiple virtual machines, may reside on one physical machine. This situation requires an efficient inter-domain communication channel with transparency and security principles ensured. Although current inter-domain mechanism has gained a much better performance compared to traditional inter-domain path offered by hypervisor, shared data channel size limitation and additional copy are still two restrains against a higher performance efficiency. In this paper, we give an analysis on these limitations, overcome these shortcomings, and achieve a higher efficient inter-domain communication channel. By applying virtual address protection mechanism during channel bootstrap, we enlarge the maximum size of shared data channel. By pointing network packet structure to the buffer in the shared data channel, we avoid an extra data copy on the receiver VM side. In our evaluation using a number of standard benchmarks, we have reduced the latency by nearly 40%, increased the throughput by approximately 45% and cut down more than 3500 CPU cycles per packet. Hongyong Zang, Yuzhong Sun, Kuiyan Gu |
ICPADS | 2 |
| 2008 | A Service-Oriented Priority-Based Resource Scheduling Scheme for Virtualized Utility Computing
Yaqiong Li, Binquan Feng, Hongyong Zang, Yuzhong Sun |
HiPC | 7 |
| 2006 | How to avoid herd: a novel stochastic algorithm in grid schedulingabstractGrid technologies promise to bring the grid users high performance. Consequently, scheduling is being becoming a crucial problem. Herd behavior is a common phenomenon, which causes the severe performance decrease in grid environment with respect to bad scheduling behaviors. In this paper, on the basis of the theoretical results of the homogeneous balls and bins model, we proposed a novel stochastic algorithm to avoid herd behavior. Our experiments address that the multi-choice strategy, combined with the advantages of DHT, can decrease herd behavior in large-scale sharing environment, at the same time, providing better schedule performance while burdening much less scheduling overhead than greedy algorithms. In the case of 1000 resources, the simulations show that, for the heavy load(i.e. system utilization rate 0.5), the multi-choice algorithm reduces the number of incurred herds by a factor of 36, the average job waiting time by a factor of 8, and the average job turn-around time by 12% compared to the greedy algorithms Yuzhong Sun |
HPDC | 3 |
| 2005 | Modeling and PerformanceAnalysis of the VEGA Grid SystemabstractIn this paper, we propose four general queueing models based on input and server distributions, to analyze a special grid system, VEGA grid system version 1.1 (VEGA1.1). The mean queue lengths and mean waiting times of these models are deduced. The two classic applications, the computing-oriented application (blast computing) and online transaction processing application (air booking service) are employed to analyze and evaluate the grid system and the models. Meanwhile, we validate VEGA's four prominent characteristics, versatile services, enabling intelligence, global uniformity and autonomous control. Finally, we analyze the experiment results, predict the behavior of VEGA1.1 in equilibrium and point out the bottlenecks of the system Zhiwei Xu 0002, Yuzhong Sun |
e-Science | 3 |
| 2005 | A C/S and P2P Hybrid Resource Discovery Framework in Grid EnvironmentsabstractResource discovery is crucial to efficient deployment of a grid system whose dynamic, heterogeneous characteristics make it difficult. In this paper, Vega Infrastructure for Resource Discovery (VIRD) is developed, then augmented with new features (i.e., some new algorithms) to build a C/S (client/server) and P2P (peer-to-peer) hybrid resource discovery framework. The three layered architecture of the VIRD is developed to make advantage of the physical and logical topologies of the Internet to facilitate resource discovery. With our simulations and theoretical analysis, it is proved that VIRD is of good scalability with respect to the sizes of the underlying backbone. Even when the resource density is low and the max TTL (time-to-live) is small, VIRD still achieves high search success rates in a small amount of hops. Compared with flooding and random walk algorithms via the same search success rates, VIRD outperforms them in both network traffic and response time. Yili Gong, Wei Li 0008, Yuzhong Sun, Zhiwei Xu 0002 |
ICPP | 3 |
| 2005 | Performance Analysis and Prediction on VEGA Grid
Zhiwei Xu 0002, Yuzhong Sun, Zheng Shen, Changshu Liu |
ISPA | 3 |
| 2004 | Grid Replication Coherence ProtocolabstractSummary form only given. Here we present two efficient coherence protocols for data grid applications. The data intensive grid applications often adopt the replica strategy to solve the data access bottleneck. It is important to keep coherence among all replicas of a copy of large amount of data. The first strategy for coherence is called lazy copy coherence protocol and the second aggressive copy coherence. Meanwhile, we consider in which situation we should make a replica for a copy of large amount of data. Yuzhong Sun, Zhiwei Xu 0002 |
IPDPS | 1 |
| 2004 | Design and Implementation of a 3A Accessing Paradigm Supported Grid Application and Programming Environment
Ge He, Donghua Liu, Yuzhong Sun, Zhiwei Xu 0002 |
ISPA | 3 |
| 2004 | Managing Service-Oriented Grids: Experiences from VEGA System Software
Yuzhong Sun, JiPing Cai, Li Zha, Yili Gong |
NPC | 1 |
| 2004 | Mapping Publishing and Mapping Adaptation in the Middleware of Railway Information Grid System
Ganmei You, Huaming Liao, Yuzhong Sun |
NPC | 3 |
| 2001 | Barrier Synchronization on Wormhole-Routed NetworksabstractIn this paper, we propose an efficient barrier synchronization scheme on networks with arbitrary topologies. We first present a distributed method in building a barrier routing tree. The barrier messages can be delivered adaptively according to the hierarchy of the established barrier tree to void congestion and faulty nodes in the network. We then propose a new technique, called bandwidth-preempting technique, for a blocked barrier message to preempt a channel occupied by a data message so that the latency of a barrier message can be controlled without affecting much of the overall system performance. We also propose an analytical performance model and present simulation results for the performance evaluation of the proposed scheme. Performance evaluations show that the proposed scheme outperforms the existing algorithms for barrier synchronization. Yuzhong Sun, Paul Y. S. Cheung, Xiaola Lin |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2000 | Recursive Cube of Rings: A New Topology for Interconnection NetworksabstractIn this paper, we introduce a family of scalable interconnection network topologies, named Recursive Cube of Rings (RCR), which are recursively constructed by adding ring edges to a cube. RCRs possess many desirable topological properties in building scalable parallel machines, such as fixed degree, small diameter, wide bisection width, symmetry, fault tolerance, etc. We first examine the topological properties of RCRs. We then present and analyze a general deadlock-free routing algorithm for RCRs. Using a complete binary tree embedded into an RCR with expansion-cost approximating to one, an efficient broadcast routing algorithm on RCRs is proposed. The upper bound of the number of message passing steps in one broadcast operation on a general RCR is also derived. Yuzhong Sun, Paul Y. S. Cheung, Xiaola Lin |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 1998 | Fault Tolerant All-to-All Broadcast in General Interconnection NetworksabstractWith respect to scalability and arbitrary topologies of the underlying networks in multiprogramming and multithread environments, fault tolerance in acknowledged ATAB and concurrent communications become a challenge to reliable general wormhole routing multicomputers with arbitrary topologies. In this paper, the virtual ring tree (VRT) is proposed to deal with the challenge. A single startup is needed in the two proposed algorithms by a simple virtual node space, which also reduces the complexity of routing at intermediate steps of ATAB algorithms and re-beginning an ATAB, by cacheable virtual channels. The proposed algorithm can automatically handle static faults in networks. Yuzhong Sun, Paul Y. S. Cheung, Xiaola Lin, Keqin Li 0001 |
ICPADS | 1 |