VLDB 2026 Research / reviewers in the wild / expert
Zhiyang Guo
dblp:12/9827
· DBLP profile ↗
30ranked-venue papers
18as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 7 first-authorComputer networks · 11 · 6 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Make-It-Animatable: An Efficient Framework for Authoring Animation-Ready 3D Charactersabstract3D characters are essential to modern creative industries, but making them animatable often demands extensive manual work in tasks like rigging and skinning. Existing automatic rigging tools face several limitations, including the necessity for manual annotations, rigid skeleton topologies, and limited generalization across diverse shapes and poses. An alternative approach is to generate animatable avatars pre-bound to a rigged template mesh. However, this method often lacks flexibility and is typically limited to realistic human shapes. To address these issues, we present Make-It-Animatable, a novel data-driven method to make any 3D humanoid model ready for character animation in less than one second, regardless of its shapes and poses. Our unified framework generates high-quality blend weights, bones, and pose transformations. By incorporating a particle-based shape autoencoder, our approach supports various 3D representations, including meshes and 3D Gaussian splats. Additionally, we employ a coarse-to-fine representation and a structure-aware modeling strategy to ensure both accuracy and robustness, even for characters with non-standard skeleton structures. We conducted extensive experiments to validate our framework’s effectiveness. Compared to existing methods, our approach demonstrates significant improvements in both quality and speed. More demos and code are available at https://jasongzy.github.io/Make-It-Animatable/. Zhiyang Guo, Jinxu Xiang, Wengang Zhou 0001, Houqiang Li |
CVPR | 1 |
| 2025 | EG4D: Explicit Generation of 4D Object without Score DistillationabstractIn recent years, the increasing demand for dynamic 3D assets in design and gaming applications has given rise to powerful generative pipelines capable of synthesizing high-quality 4D objects.
Previous methods generally rely on score distillation sampling (SDS) algorithm to infer the unseen views and motion of 4D objects, thus leading to unsatisfactory results with defects like over-saturation and Janus problem.
Therefore, inspired by recent progress of video diffusion models, we propose to optimize a 4D representation by explicitly generating multi-view videos from one input image.
However, it is far from trivial to handle practical challenges faced by such a pipeline, including dramatic temporal inconsistency, inter-frame geometry and texture diversity, and semantic defects brought by video generation results.
To address these issues, we propose EG4D, a novel multi-stage framework that generates high-quality and consistent 4D assets without score distillation.
Specifically, collaborative techniques and solutions are developed, including an attention injection strategy to synthesize temporal-consistent multi-view videos, a robust and efficient dynamic reconstruction method based on Gaussian Splatting, and a refinement stage with diffusion prior for semantic restoration.
The qualitative comparisons and quantitative results demonstrate that our framework outperforms the baselines in generation quality by a considerable margin. Qi Sun 0005, Zhiyang Guo, Ziyu Wan, Jing Nathan Yan, Shengming Yin, Wengang Zhou 0001, Jing Liao 0001, Houqiang Li |
ICLR | 2 |
| 2025 | AAGS: Appearance-Aware 3D Gaussian Splatting with Unconstrained Photo Collections
Wencong Zhang, Zhiyang Guo, Wengang Zhou 0001, Houqiang Li |
Multim. Syst. | 2 |
| 2025 | HandNeRF++: Modeling Animatable Interacting Hands With Neural Radiance FieldsabstractIn this work, we explore the rendering of photo-realistic free-viewpoint hand pose animation. We present HandNeRF, the first NeRF-based framework to reconstruct accurate appearance and geometry for interacting hands. To overcome the texture contamination and shape artifact problems when dealing with complex interacting scenarios, we further introduce HandNeRF++ to achieve better performance. In our advanced framework, a pose-driven deformation field is designed to establish correspondence from diverse poses to a canonical space, where the pose- and shape-disentangled NeRFs are optimized. To enhance the geometry and texture cues in rarely-observed areas for interacting hands, we establish a connection between the interacting hands by proposing the adaptive hand-sharing technique for cross-hand augmentation. Meanwhile, we further leverage the hand poses to generate fine-grained density priors, serving as valuable guidance for occlusion-aware geometry learning. Furthermore, a neural feature distillation method and a neural refiner are proposed to facilitate color optimization and further polish the renderings. With the collaboration of all the modules and strategies, our HandNeRF++ significantly advances the capabilities of NeRF-based 3D reconstruction in the context of interacting hands. Extensive experiments are conducted to validate the merits of the proposed frameworks. We report a series of state-of-the-art results both qualitatively and quantitatively. Zhiyang Guo, Wengang Zhou 0001, Min Wang 0019, Li Li 0040, Houqiang Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Motion-Aware 3D Gaussian Splatting for Efficient Dynamic Scene Reconstructionabstract3D Gaussian Splatting (3DGS) has become an emerging tool for dynamic scene reconstruction. However, existing methods mainly focus on developing various strategies to extend static 3DGS into a time-variant representation, while overlooking the rich motion information implicitly carried by 2D observations, thus suffering from performance degradation and model redundancy. To address the above problem, we propose a novel motion-aware enhancement framework for dynamic scene reconstruction, which mines useful motion cues from optical flow to improve different paradigms of dynamic 3DGS. Specifically, we first step beyond the vanilla render-based cross-dimensional supervision that suffers from ambiguity and instability, and establish a more robust and effective dense correspondence between 3D Gaussian movements and pixel-level flows. Then a novel flow augmentation method is introduced with additional insights into uncertainty and loss collaboration. Furthermore, for the prevalent deformation-based paradigm that presents a harder optimization problem, a transient-aware deformation auxiliary module is proposed. We conduct extensive experiments on both multi-view and monocular scenes to verify the merits of our work. Compared with the baselines, our method shows significant superiority in both rendering quality and efficiency. The code will be publicly available athttps://github.com/jasongzy/MAGS. Zhiyang Guo, Wengang Zhou 0001, Li Li 0040, Min Wang 0019, Houqiang Li |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Swin-AFF: an improved accuracy 6D pose estimation network for high reflection and texture-less workpieces based on Swin transformer
Zhentao Li, Zhiyang Guo, Panfeng Wang, Wenlei Wu |
Vis. Comput. | 2 |
| 2023 | HandNeRF: Neural Radiance Fields for Animatable Interacting HandsabstractWe propose a novel framework to reconstruct accurate appearance and geometry with neural radiance fields (NeRF) for interacting hands, enabling the rendering of photo-realistic images and videos for gesture animation from arbitrary views. Given multi-view images of a single hand or interacting hands, an off-the-shelf skeleton estimator is first employed to parameterize the hand poses. Then we design a pose-driven deformation field to establish correspondence from those different poses to a shared canonical space, where a pose-disentangled NeRF for one hand is optimized. Such unified modeling efficiently complements the geometry and texture cues in rarely-observed areas for both hands. Meanwhile, we further leverage the pose priors to generate pseudo depth maps as guidance for occlusion aware density learning. Moreover, a neural feature distillation method is proposed to achieve cross-domain alignment for color optimization. We conduct extensive experiments to verify the merits of our proposed HandNeRF and report a series of state-of-the-art results both qualitatively and quantitatively on the large-scale InterHand2.6M dataset. Zhiyang Guo, Wengang Zhou 0001, Min Wang 0019, Li Li 0040, Houqiang Li |
CVPR | 1 |
| 2022 | CMT: Context-Matching-Guided Transformer for 3D Tracking in Point Clouds
Zhiyang Guo, Yunyao Mao, Wengang Zhou 0001, Min Wang 0019, Houqiang Li |
ECCV (22) | 1 |
| 2022 | CT PUF: Configurable Tristate PUF Against Machine Learning Attacks for IoT SecurityabstractPhysical unclonable function (PUF) is a promising lightweight hardware security primitive for resource-limited Internet-of-Things (IoT) devices. Strong PUFs are suitable for lightweight device authentication because it can generate quantities of challenge-response pairs. Unfortunately, while the machine learning (ML) techniques have benefited various areas, such as Internet, industrial automation, robotics and gaming, they pose a severe threat to PUFs by easily modelling their behavior. This article first shows that even a recently reported dual-mode PUF can be cloned by ML (prediction accuracy of up to 95%). To solve this issue, we propose a configurable tristate (CT) PUF which can flexibly perform as an arbiter PUF, a ring oscillator (RO) PUF, or a bistable ring (BR) PUF with a bitwise XOR-based mechanism to obfuscate the relationship between the challenge and the response, hence resisting the ML attacks. An authentication protocol for the use in IoT security is presented. The CT PUF is implemented on Xilinx ZedBoard FPGAs with placement and routing details described. The experimental results show that the modelling accuracy of logistic regression (LR), support vector machine (SVM), covariance matrix adaptation evolutionary strategies (CMA-ES), and artificial neural network (ANN) is close to 60% (50% as the ideal number in theory) while meeting the PUF requirements for uniformity, reliability, and uniqueness. The hardware overhead and power consumption are slight. The entire project has been open sourced. Jiliang Zhang 0002, Chaoqun Shen, Zhiyang Guo, Qiang Wu 0015, Wanli Chang 0001 |
IEEE Internet Things J. | 3 |
| 2016 | BCCC: An Expandable Network for Data CentersabstractDesigning a cost-effective network topology for data centers that can deliver sufficient bandwidth and consistent latency performance to a large number of servers has been an important and challenging problem. Many server-centric data center network topologies have been proposed recently due to their significant advantage in cost efficiency and data center agility, such as BCube, FiConn, and Bidimensional Compound Network (BCN). However, existing server-centric topologies are either not expandable or demanding prohibitive expansion cost. As the scale of data centers increases rapidly, the lack of expandability in existing server-centric data center networks imposes a severe obstacle for data center upgrade. In this paper, we present a novel server-centric data center network topology called BCube connected crossbars (BCCCs), which can provide good network performance using inexpensive commodity off-the-shelf switches and commodity servers with only two network interface card (NIC) ports. A significant advantage of BCCC is its good expandability. When there is a need for expansion, we can easily add new servers and switches into the existing BCCC with little alteration of the existing structure. Meanwhile, BCCC can accommodate a large number of servers while keeping a very small network diameter. A desirable property of BCCC is that its diameter increases only linearly to the network order (i.e., the number of dimensions), which is superior to most of the existing server-centric networks, such as FiConn and BCN, whose diameters increase exponentially with network order. In addition, there are a rich set of parallel paths with similar length between any pair of servers in BCCC, which enables BCCC to not only deliver sufficient bandwidth capacity and predictable latency to end hosts, but also provide graceful performance degradation in case of component failure. We conduct comprehensive comparisons between BCCC with other popular server-centric network topologies, such as FiConn and BCN. We also propose an effective addressing scheme and routing algorithms for BCCC. We show that BCCC has significant advantages over the existing server-centric topologies in many important metrics, such as expandability, server port utilization, and network diameter. Zhenhua Li 0002, Zhiyang Guo, Yuanyuan Yang 0001 |
IEEE/ACM Trans. Netw. | 2 |
| 2015 | Cost efficient and performance guaranteed virtual network embedding in multicast fat-tree DCNsabstractMost of today's data center networks (DCNs) adopt a multi-rooted tree structure called fat-tree, which delivers large bisection bandwidth through rich path multiplicity. In fat-tree DCNs, core switch modules play an important role in providing nonblocking capability, and form a significant part of network cost simultaneously. Reducing core switches while simultaneously guaranteeing performance has been a constant challenge. For example, multicast is an essential communication pattern in cloud services which needs to be supported efficiently. In this paper, we propose virtual network embedding schemes to deal with this problem. In the first scheme, we place the virtual machines (VMs) of a multicast-capable virtual network (MVN) as compact as possible, without any disturbance to existing traffic. In the second scheme, we manage to keep VMs in an even more compact way to reduce cost by allowing a small degree of VM migration. Both schemes are guaranteed to support any multicast communications within MVNs, and simultaneously achieve significant cost saving in terms of core switches, compared to currently best known result. Moreover, we show that our schemes incur only a small overhead in terms of migrations. Finally, we evaluate the performance of proposed schemes and validate the theoretical analysis through extensive simulations. Zhiyang Guo, Yuanyuan Yang 0001 |
INFOCOM | 2 |
| 2015 | Embedding Nonblocking Multicast Virtual Networks in Fat-Tree Data CentersabstractVirtualization of servers and networks is a key technique to resolve the conflict between the increasing demands on computing power and the high cost of hardware in data centers. In order to map virtual networks to physical infrastructure efficiently, designers have to make careful decisions on the allocation of limited resources, which makes network embedding in data centers a very important problem. In this paper, we tackle the network embedding problem in fat-tree data centers. To meet the requirements of instant parallel data transfer between multiple computing units, we propose a model of multicast-capable virtual networks (Mons). We then design three virtual machine (VM) placement schemes with different features for embedding MVNs into fat-tree DCNs, named Most-Vacant-Fit (MVF), Most-Compact-First (MCF) and Mixed-Bidirectional-Fill (MBF). All these VM placement schemes guarantee the no blocking multicast capability of each MVN while simultaneously achieving significant saving on the cost of network hardware. In addition, each VM placement scheme also has its unique features. The MVF scheme has zero interference to existing computing tasks in data centers, the MCF scheme leads to the greatest cost saving, the MBF scheme simultaneously possesses the merits of MVF and MCF, and it provides an adjustable parameter allowing cloud providers to achieve preferred balance between the cost and the overhead. Finally, we compare the performance and overhead of these VM placement schemes, and present simulation results to validate our theoretical results. Zhiyang Guo, Yuanyuan Yang 0001 |
IPDPS | 2 |
| 2015 | On Nonblocking Multicast Fat-Tree Data Center Networks with Server RedundancyabstractFat-tree networks have been widely adopted as network topologies in data center networks (DCNs). However, it is costly for fat-tree DCNs to support nonblocking multicast communication, due to the large number of core switches required. Since multicast is an essential communication pattern in many cloud services and nonblocking multicast communication can ensure the high performance of such services, reducing the cost of nonblocking multicast fat-tree DCNs is very important. On the other hand, server redundancy is ubiquitous in today’s data centers to provide high availability of services. In this paper, we explore server redundancy in data centers to reduce the cost of nonblocking multicast fat-tree data center networks (DCNs). First, we present a multirate network model that accurately describes the communication environment of the fat-tree DCNs. We then show that the sufficient condition on the number of core switches required for nonblocking multicast communication under the multirate model can be significantly reduced when the fat-tree DCNs are 2-redundant, i.e., each server in the data center has exactly one redundant backup. We also study the general redundant fat-tree DCNs where servers may have different numbers of redundant backups depending on the availability requirements of services they provide, and show that a higher redundancy level further reduces the cost of nonblocking multicast fat-tree DCNs. Then, to complete our analysis, we consider a practical faulty data center, where one or more active servers may fail at any time. We give a strategy to re-balance the active servers among edge switches after server failures so that the same nonblocking condition still holds. Finally, we give a multicast routing algorithm with linear time complexity to configure multicast connections in fat-tree DCNs. Zhiyang Guo, Yuanyuan Yang 0001 |
IEEE Trans. Computers | 1 |
| 2015 | Exploring Server Redundancy in Nonblocking Multicast Data Center NetworksabstractClos networks and their variations such as folded-Clos networks (fat-trees) have been widely adopted as network topologies in data center networks. Since multicast is an essential communication pattern in many cloud services, nonblocking multicast communication can ensure the high performance of such services. However, nonblocking multicast Clos networks are costly due to the large number of middle stage switches required. On the other hand, server redundancy is ubiquitous in today's data centers to provide high availability of services. In this paper, we explore such server redundancy in data centers to reduce the cost of nonblocking multicast Clos data center networks (DCNs). To facilitate our analysis, we first consider an ideal fault-free data center with no server failure. We give an algorithm to assign active servers evenly among input stage switches in a multicast Clos DCN where each server has one or more redundant backups depending on the availability requirements of services they provide. We show that the sufficient nonblocking condition on the number of middle stage switches for a multicast Clos DCN can be significantly reduced by exploring server redundancy. Then, to complete our analysis, we consider a practical faulty data center, where one or more active servers may fail at anytime. We give a strategy to re-balance the active servers among input stage switches after server failures so that the same nonblocking condition still holds. Finally, we provide a multicast routing algorithm with linear time complexity to configure multicast connections in Clos DCNs. Zhiyang Guo, Yuanyuan Yang 0001 |
IEEE Trans. Computers | 1 |
| 2015 | Bounded-Reorder Packet Scheduling in Optical Cut-Through SwitchabstractThe recently proposed optical cut-through (OpCut) switch holds a great potential in achieving high energy efficiency, as it allows optical packets to cut through the switch in optical domain whenever possible, which avoids power-hungry O/E/O conversion. In the OpCut switch, to ensure in-order transmission, only optical Head-of-Line (HOL) packet of a switch flow, i.e., the stream of packets sharing the same input and output port, is allowed to cut-through the switch, and optical HOL packets are always prioritized over buffered HOL packets to achieve high cut-through ratio, which is measured by the portion of packets cutting through the switch optically. However, under such priority rule, switch flows with buffered packets are at the risk of starvation, and the OpCut switch fails to achieve 100 percent throughput for all admissible i.i.d. traffics due to the unfairness in packet scheduling. To address this two issues, in this paper we propose a delay threshold rule for packet scheduling, in which buffered packets with delays exceeding a preset delay threshold are prioritized over optical packets. In the meanwhile, the cut-through ratio is very low under heavily congested traffic due to maintaining packet order, whereas the Internet is designed to accommodate a certain degree of packet reorder, which is very common in practice due to path multiplicity. In this paper, we design a bounded-reorder packet scheduling algorithm that significantly increases the cut-through ratio of the OpCut switch while allowing a small degree of out-of-order transmission. Our extensive simulation results show that the energy efficiency of OpCut switch can be significantly improved with only a very small degree of packet reordering, which has little adverse impact on the network application performance. Zhemin Zhang, Zhiyang Guo, Yuanyuan Yang 0001 |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2014 | BCCC: an expandable network for data centersabstractMany server-centric data center network topologies have been proposed recently due to their significant advantage in cost-efficiency and data center agility, such as BCube, FiConn and BCN. However, existing server-centric topologies are either not expandable or demanding prohibitive expansion cost. As the scale is increasing rapidly, the lack of expandability imposes a severe obstacle for data center upgrade. In this paper, we present a novel server-centric data center network topology called BCube Connected Crossbars (BCCC), which can provide good network performance and expandability using commodity off-the-shelf switches and commodity servers with only two NIC ports. BCCC can accommodate a large number of servers while keeping a very small network diameter, as a particular desirable property of BCCC is that its diameter increases only linearly to the network order, which is superior to most of existing server-centric networks, such as FiConn and BCN, whose diameters increase exponentially with network order. Additionally, we propose an effective addressing scheme and routing algorithms for BCCC. We also conduct comprehensive comparisons between BCCC and other popular server-centric networks. We show that BCCC has significant advantages over existing server-centric topologies in many important metrics, such as expandability, port utilization and network diameter. Zhenhua Li 0002, Zhiyang Guo, Yuanyuan Yang 0001 |
ANCS | 2 |
| 2014 | Augmenting data center networks with a fast reconfigurable optical multistage interconnectabstractThe high bandwidth and power efficiency of optical circuit switching (OCS) have motivated the recent development of hybrid packet/circuit switched (Hypac) data center networks (DCNs). However, current Hypac DCNs use a large MEMS optical switch for OCS communication, which offers limited scalability and expandability. In addition, the slow switching speed of MEMS switches imposes considerable network reconfiguration overhead, which results in degraded network performance. In this paper, we first analyze the fundamental challenges in current MEMS-based Hypac DCNs, and show that fast optical network reconfiguration is the key to effectively utilize OCS in data center communication. We then present a Fast Reconfigurable Baseline Optical Network (FARBON), which, by leveraging the ultra-fast optical switching modules and unique topological properties of the baseline multistage network, allows for rapid network configuration. We also design the control plane and a practical low-jitter traffic scheduling algorithm for FARBON. We demonstrate that FARBON has many desirable features, such as good scalability, low traffic jitter, predictable network performance and tolerance to inaccurate traffic information. The performance of FARBON is evaluated via extensive simulations, and the results show that FARBON significantly outperforms MEMS optical switches in terms of average packet delay and jitter, as well as dealing with correlated traffic, thus is a promising candidate for exploiting the potential of OCS in Hypac DCNs. Zhiyang Guo, Yuanyuan Yang 0001 |
GLOBECOM | 1 |
| 2014 | Collaborative Network Configuration in Hybrid Electrical/Optical Data Center NetworksabstractRecently, there has been much effort on introducing optical fiber communication to data center networks (DCNs) because of its significant advantage in bandwidth capacity and power efficiency. However, due to limitations of optical switching technologies, optical networking alone has not yet been able to accommodate the volatile data center traffic. As a result, hybrid packet/circuit (Hypac) switched DCNs, which argument the electrical packet switched (EPS) network with an optical circuit switched (OCS) network, have been proposed to combine the strengths of both types of networks. However, one problem with current Hypac DCNs is that the EPS network is shared in a best effort fashion and is largely oblivious to the accompanying OCS network, which results in severe drawbacks, such as degraded network predictability and deficiency in handling correlated traffic. Since the OCS/EPS networks have unique strengths and weaknesses, and are best suited for different traffic patterns, coordinating and collaborating the configuration of both networks is critical to reach the full potential of Hypac DCNs, which motivates the study in this paper. First, we present a network model that accurately abstracts the essential characteristics of the EPS/OCS networks. Second, considering the recent advances in network control technology, we propose a time-efficient algorithm called Collaborative Bandwidth Allocation (CBA) that configures both networks in a complementary manner. Finally, we conduct comprehensive simulations, which demonstrate that CBA significantly improves the performance of Hypac DCNs in many aspects. Zhiyang Guo, Yuanyuan Yang 0001 |
IPDPS | 1 |
| 2014 | On-Line Multicast Scheduling with Bounded Congestion in Fat-Tree Data Center NetworksabstractMulticast benefits numerous data center applications that require group communication by eliminating sending unnecessary duplicated packets in the network, thus significantly reduces network traffic and improves application throughput. Meanwhile, most data center networks (DCNs) today adopt a multi-rooted tree structure called fat-tree, which utilizes rich path multiplicity to deliver high bisection bandwidth. However, without an efficient flow scheduling algorithm that appropriately routes multicast flows to achieve traffic load balance, heavy congestion may occur throughout the network, which prevents full utilization of such high degree of link parallelism and causes unpredictable network performance. Hence, in this paper we study multicast flow scheduling in fat-tree DCNs, where multicast flow requests arrive one by one without a priori knowledge of future traffic. To address the drastic traffic fluctuation in data centers, we consider a very general traffic model called hose traffic model, where the only assumption is that the total bandwidth demand of traffic that enters (leaves) an ingress (egress) link of each server at any time is bounded by the capacity of its network interface card. We present a low-complexity on-line multicast flow scheduling algorithm for fat-tree DCNs. The algorithm can achieve bounded congestion and efficient bandwidth utilization under any arbitrary sequence of multicast flow requests that satisfy the hose model. We also derive the bound on congestion that the algorithm can achieve in a fat-tree DCN. Finally, we evaluate the algorithm by an event-driven DCN simulator under various types of traffic patterns, and show that the algorithm achieves superior performance in terms of network throughput and evenness of traffic load distribution. Zhiyang Guo, Yuanyuan Yang 0001 |
IEEE J. Sel. Areas Commun. | 1 |
| 2014 | Bufferless Routing in Optical Gaussian Macrochip InterconnectabstractThe ever increasing intra-chip and inter-chip traffic load in computing systems has been pushing traditional electronic interconnects to their limit in communication bandwidth, latency, and energy consumption. In order to achieve the high bandwidth and low latency required by intra-chip and inter-chip communications and mitigate the high interconnect power dissipation, optical interconnects have been considered as a promising candidate for intra-chip and inter-chip interconnections in next generation computing systems. In addition, packet switching is an efficient switching paradigm to fully utilize the communication bandwidth. However, due to lack of random access optical memory, it is challenging to implement all-optical packet switching in optical interconnects. In this paper, we exploit bufferless routing, a special type of packet-switching, to overcome the problem of lack of random access optical buffer. More specifically, we study bufferless routing in a novel optical multichip system, called Gaussian macrochip, where embedded chips are interconnected by an optical Gaussian network. By taking advantage of the underlying Hamiltonian cycles in the Gaussian network, we design a bufferless routing algorithm for the Gaussian macrochip, which routes packets along the shortest path in the absence of deflection, and guarantees that deflected packets reach their destinations within${{N}}$hops. Our extensive simulation results demonstrate that by adopting the proposed routing algorithm, Gaussian macrochip can support much higher inter-chip communication bandwidth, has much shorter average packet delay, and is more power efficient than the previously proposed architectures for optical multichip systems. Zhemin Zhang, Zhiyang Guo, Yuanyuan Yang 0001 |
IEEE Trans. Computers | 2 |
| 2014 | Low-Latency Multicast Scheduling in All-Optical InterconnectsabstractOptical interconnects are considered as a very appealing solution for future high speed interconnections in core networks and parallel computers. In this paper, we study multicast scheduling in all-optical packet interconnects/switches. We first propose a novel optical buffer called multicast-enabled fiber-delay-lines (M-FDLs), which can provide flexible delay for copies of multicast packets using only a small number of FDL segments. We then present a Low Latency Multicast Scheduling (LLMS) Algorithm that considers the schedule of each arriving packet for multiple time slots. We show that LLMS has several desirable features, such as a guaranteed delay upper bound and adaptivity to transmission requirements. To relax the time constraint of LLMS, we further propose a pipeline and parallel architecture for LLMS that distributes the scheduling task to multiple pipelined processing stages, with N processing modules in each stage, where N is the size of the interconnect. Finally, by implementing it with simple combination circuits, we show that each processing module can complete the packet scheduling for a time slot in O(1) time. The performance of LLMS is evaluated extensively against statistical traffic models and real Internet traffic traces, and the results show that the proposed LLMS algorithm can achieve superior performance in terms of average packet delay and packet drop ratio. Zhiyang Guo, Yuanyuan Yang 0001 |
IEEE Trans. Commun. | 1 |
| 2013 | Multicast fat-tree data center networks with bounded link oversubscriptionabstractMany data center networks (DCNs) adopt a multirooted tree structure called fat-tree, which has the potential to deliver large bisection bandwidth through rich path multiplicity. However, unbalanced traffic load distribution may prevent efficient utilization of such high degree of parallelism. Meanwhile, high bandwidth multicast communication is critical to many data center services and applications. Hence, in this paper we consider multicast traffic load balance problem in fat-tree DCNs from a novel angle, aiming to find the most cost-effective way to build a multicast fat-tree DCN with bounded link oversubscription ratio. First, we present a multi-rate network model to accurately describe the communication environment in a fat-tree DCN. Then, we derive the minimum number of core switches required to achieve bounded link oversubscription ratio under arbitrary multicast traffic. Finally, we provide a comprehensive comparison on the cost of different approaches to building such a multicast fat-tree DCN. Zhiyang Guo, Yuanyuan Yang 0001 |
INFOCOM | 1 |
| 2013 | Bounded-reorder packet scheduling in optical cut-through switchabstractEnergy efficiency of optical packet switches (OPS) is the key to ensure the profitability of backbone network providers. However, due to lack of optical random access buffer, most optical packet switches rely on electronic buffer to resolve output contention, which requires power-hungry O/E/O conversion for all packets. The recently proposed optical cut-through (OpCut) switch holds a great potential in achieving high energy efficiency, as it allows optical packets to cut through the switch in optical domain whenever possible. The energy efficiency of OpCut switch hinges on the cut-through ratio, which is the percentage of packets that cut through the switch optically. On the other hand, it is generally desirable to maintain packet order in a switch. To achieve in-order transmission, an optical packet needs to be converted to electronic form and buffered when an earlier packet from the same flow is still in the buffer, which may lead to a low cut-through ratio. In the meanwhile, the Internet is designed to accommodate a certain degree of packet reorder, which is very common in practice due to path multiplicity. In this paper, we introduce a novel reorder metric, reorder degree, to accurately describe the extent of packet reordering, and propose a flow management scheme to bound the reorder degree of transmitted flows. We then design an efficient packet scheduling algorithm that significantly increases the cutthrough ratio of the OpCut switch while allowing a small degree of out-of-order transmission. Our extensive simulation results show that the cut-through ratio can be drastically increased with only a very small reorder degree. Zhemin Zhang, Zhiyang Guo, Yuanyuan Yang 0001 |
INFOCOM | 2 |
| 2013 | Oversubscription Bounded Multicast Scheduling in Fat-Tree Data Center NetworksabstractMulticast benefits numerous data center applications that require group communication by eliminating sending unnecessary duplicated packets in the network, thus significantly reduces network traffic and improves application throughput. Meanwhile, many data center networks (DCNs) adopt a multi-rooted tree structure called fat-tree, which utilizes rich path multiplicity to deliver high bisection bandwidth. However, currently there is no efficient flow scheduling algorithm for the fat-tree that can route multicast flows appropriately to achieve traffic load balance, thus cannot fully take advantage of this high degree of link parallelism. Besides low bandwidth utilization, unbalanced traffic load distribution also leads to unpredictable network performance and degraded data center agility. In this paper, we study multicast traffic load balance problem in fat-tree DCNs. First, we derive a minimum link oversubscription upper bound in multicast fat-tree DCNs based on a network model that accurately describes the DCN communication environment. Then, we present Oversubscription Bounded Multicast Scheduling (OBMS), a low-complexity multicast flow scheduling algorithm that guarantees bounded link oversubscription and efficient network utilization even under the most congested traffic patterns. Finally, we evaluate the performance of OBMS in an event-driven DCN simulator under various types of traffic patterns, and show that OBMS significantly outperforms other load-balance methods in terms of network throughput and evenness of traffic load distribution. Zhiyang Guo, Yuanyuan Yang 0001 |
IPDPS | 1 |
| 2013 | High-Speed Multicast Scheduling for All-Optical Packet SwitchesabstractIn this paper, we study multicast scheduling in all-optical packet switches. We first propose a novel optical buffer called multicast-enabled Fiber-Delay-Lines (M-FDLs), which can provide flexible delay for copies of multicast packets using only a small number of FDL segments. We then present a Delay-Guaranteed Multicast Scheduling (DGMS) algorithm that considers the schedule of each arriving packet for multiple time slots. We show that DGMS has several desirable features, such as guaranteed delay upper bound and adaptivity to transmission requirements. To relax the time constraint of DGMS, we further propose a parallel and pipeline architecture for DGMS that distributes the scheduling task to multiple pipelined processing stages, with N processors in each stage, where N is the switch size. Finally, by using a simple combination logic circuit, we show that each processor can finish the scheduling for one time slot in O(1) time. The performance of DGMS is tested extensively against statistical traffic models and real Internet traffic, and the results show that the proposed DGMS algorithm can achieve ultra-low average packet delay with minimum packet drop ratio. Zhiyang Guo, Yuanyuan Yang 0001 |
NAS | 1 |
| 2013 | High-Speed Multicast Scheduling in Hybrid Optical Packet Switches with Guaranteed LatencyabstractIn this paper, we study multicast scheduling in the OpCut switch, a recently proposed hybrid optical/electronic switching architecture for transmitting high-volume traffic in core networks and parallel computers. First, we present a multicast scheduling algorithm called Guaranteed Latency Multicast Scheduling (GLMS) that considers the schedule of each packet for multiple time slots. We show that GLMS has several desirable features, such as guaranteed latency for all transmitted packets and adaptivity to transmission requirements. To relax the time constraint on computing a schedule, we further propose a parallel and pipeline processing architecture for GLMS that distributes the scheduling task to multiple pipelined processing stages, with N processors in each stage, where N is the switch size. Finally, by implementing it with simple combination logic circuits, we show that each processor can finish the scheduling for one time slot in (O(1)time complexity. We evaluate the performance of GLMS extensively against statistical traffic models and real Internet traffic, and the results show that the proposed GLMS algorithm can achieve very low average packet latency with minimum packet drop ratio. Zhiyang Guo, Yuanyuan Yang 0001 |
IEEE Trans. Computers | 1 |
| 2013 | Efficient All-to-All Broadcast in Gaussian On-Chip NetworksabstractWith the development of multiprocessor system on chips (MPSoCs), it is expected that hundreds of computing cores will be operating on a single chip in the near future. This will require high-performance on-chip networks with very low latency to provide a communication substrate for the increasing number of cores. In this paper, we consider Gaussian on-chip networks that are of significant topological advantages over traditional mesh and torus networks in terms of diameter and average hop distance. Many applications on MPSoCs need global data movement and global control to exchange data and synchronize the execution among cores, which require all-to-all broadcast communication. In this paper, we propose an all-to-all broadcast algorithm suitable for on-chip implementation on the Gaussian network topology. The algorithm utilizes controlled message flooding based on a broadcast pattern, which can be described in a formal, generic way for each node in terms of a few simple operations and can be easily built into router hardware. Furthermore, the generic broadcast pattern also ensures a balanced traffic load in all dimensions in the network so that minimum total latency for all-to-all broadcast can be achieved. The algorithm overlaps message switching time with transmission time in a pipelined fashion to further reduce the total communication latency of all-to-all broadcast. Comparison results demonstrate the topological merits of Gaussian networks and ultralow latency of the proposed all-to-all broadcast algorithm. Zhemin Zhang, Zhiyang Guo, Yuanyuan Yang 0001 |
IEEE Trans. Computers | 2 |
| 2012 | Exploring server redundancy in nonblocking multicast data center networksabstractClos networks and their variations such as folded- Clos networks (fat-trees) have been widely adopted as network topologies in data center networks. Since multicast is an essential communication pattern in many cloud services, nonblocking multicast communication can ensure the high performance of such services. However, nonblocking multicast Clos networks are costly due to the large number of middle stage switches required. On the other hand, server redundancy is ubiquitous in today's data centers to provide high availability of services. In this paper, we explore server redundancy in data centers to reduce the cost of nonblocking multicast Clos data center networks (DCNs). First, we show that the sufficient nonblocking condition on the number of middle stage switches for multicast Clos DCNs can be significantly reduced, when the data center is 2-redundant, i.e., each server in the data center has exactly one redundant backup. We then investigate more general cases that the data center is k-redundant (k >; 2), and show that a higher redundancy level further reduces the cost of nonblocking multicast Clos DCNs. We also extend the result to practical data centers where servers may have different number of redundant backups depending on the availability requirement of services provided. Finally, we provide a multicast routing algorithm with linear time complexity to configure multicast connections in Clos DCNs. Zhiyang Guo, Zhemin Zhang, Yuanyuan Yang 0001 |
INFOCOM | 1 |
| 2012 | On Nonblocking Multirate Multicast Fat-tree Data Center Networks with Server RedundancyabstractFat-tree networks have been widely adopted as network topologies in data center networks (DCNs). However, it is costly for fat-tree DCNs to support nonblocking multicast communication, due to the large number of core switches required. Since multicast is an essential communication pattern in many cloud services and nonblocking multicast communication can ensure the high performance of such services, reducing the cost of nonblocking multicast fat-tree DCNs is very important. On the other hand, server redundancy is ubiquitous in today's data centers to provide high availability of services. In this paper, we explore server redundancy in data centers to reduce the cost of nonblocking multicast fat-tree data center networks (DCNs). First, we present a multirate network model that accurately describes the communication environment of the fat-tree DCNs. Then, we show that the sufficient number of core switches for nonblocking multicast communication under the multirate model can be significantly reduced in arbitrary 2-redundant fat-tree DCNs, i.e., each server has exactly one redundant backup in the data center. We generalize the result to practical fat-tree DCNs where servers may have different number of redundant backups depending on the availability requirements of services they provide, and show that a higher redundancy level further reduces the cost of nonblocking multicast fat-tree DCNs. Finally, we propose a multicast routing algorithm with linear time complexity to configure multicast connections in fat-tree DCNs. Zhiyang Guo, Yuanyuan Yang 0001 |
IPDPS | 1 |
| 2011 | Performance modeling of hybrid optical packet switches with shared bufferabstractAll-optical packet switches (OPS) are considered as a good candidate for future ultra-fast communications as they do not require optical-electronic-optical (O/E/O) conversions. However, currently there is still no practical optical random access memory available, which makes it difficult to reduce packet loss to an acceptable level in OPS. Thus, hybrid optical/electronic switch architectures, such as the switch proposed in which we refer to as the OpCut switch in this paper, are promising alternatives due to their potential to achieve ultra-low packet loss and packet delay. Although there has been extensive work on the performance modeling of different types of electronic and all-optical switches, little work has been done for the performance modeling of hybrid switches. In this paper, we present an efficient analytical model called the aggregation model that comprehensively analyzes various performance metrics of the OpCut switch under different types of traffic. By inductively aggregating more queues in the buffer into a block, the aggregation model can achieve a polynomial complexity to the switch size. We develop the aggregation model for the OpCut switch under both Bernoulli traffic and ON-OFF Markovian traffic. The effectiveness of our model is validated by extensive simulations. The results show that the aggregation model is very accurate in all tested scenarios. Zhiyang Guo, Zhemin Zhang, Yuanyuan Yang 0001 |
INFOCOM | 1 |