Liquan Xiao

dblp:16/1442 · DBLP profile ↗
← Back
40ranked-venue papers
0as first author
17since 2021 · last 2026
0000-0002-3285-2625ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 29 · 14 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 since 2021Computer networks · 3 · 1 since 2021
YearPublicationVenuePosition
2026 Needle: Efficient Host-NIC Memory Mapping Synchronization for Scalable RDMA Virtualization
Zihao Wei, Dezun Dong, Liquan Xiao, Yani Gong
INFOCOM3
2026 End-to-end congestion control in datacenter networks: a survey
Zejia Zhou, Shan Huang 0002, Dezun Dong, Liquan Xiao
Frontiers Comput. Sci.5
2026 PAARD: Proximity-aware all-reduce communication for dragonfly networks
Dezun Dong, Liquan Xiao
J. Parallel Distributed Comput.4
2026 A 61.4 Gb/s/mm Wireline Transceiver Using a 7 bit-Over-8 Lane Symmetric Correlated Coding for High-Density Interconnects
abstract
This paper presents a symmetric correlated coding (SCC) scheme and the corresponding high-density transceiver that deliver 7-bit data over 8 lanes. The SCC method implements the 7-bit data encoding jointly in 8 correlated channels which improves the pin efficiency up to 87.5%. Besides, as an advanced chord signaling, the SCC exhibits the immunity to crosstalk, common-mode noise (CMN) and simultaneous switching noise (SSN). The proposed SCC transceiver innovates a SCC source-series-terminated (SST) driver that maintains an unchanged bandwidth even with a low-voltage power supply, and the CTLE decoder adopts the active inductor to compensate the channel loss. Prototyped in 28-nm CMOS, the proposed wireline transceiver using SCC method supports a maximum data rate of$7\times 10$Gb/s with a bit error rate (BER) <1e−12, achieving a data rate density (DRD) up to 61.4 Gb/s/mm. The SCC transceiver dissipates 99.4 mW with the energy efficiency of 1.42 pJ/bit.
Geng Zhang 0001, Fangxu Lv, Liquan Xiao, Xuqiang Zheng, Heng Huang 0009, Kewei Xin, Liangyong Yuan, Ruixiao Kuai, Bolin Ren, Ruotian Yin, Guohe Zhang
IEEE Trans. Circuits Syst. I Regul. Pap.4
2025 Bridging Metadata Service and CXL: A Metadata-Grained and Directory-Aware Storage Engine for Distributed Storage Systems
abstract
The AI training and inference workloads are particularly metadata-intensive and drive an urgent need for distributed file systems (DFS) with high IOPS metadata service. Meanwhile, the emerging Compute Express Link (CXL) protocol introduces memory semantics to the PCIe-attached storage devices and is compelling for building high-performance metadata storage. However, a fundamental mismatch exists between the metadata access granularity and the internal storage granularity of typical CXL-enabled storage devices. Besides, existing metadata storage engines lack the perception of the DFS directory structures and metadata semantics in distributed storage systems. In this paper, we investigate the way to employ CXL-enabled devices as the storage backend for DFS metadata and propose MDSec, a metadata-grained and directory-aware metadata storage engine that bridges the semantic gaps between the DFS metadata service and CXL-enabled storage devices. MDSec unifies the granularity of all kinds of metadata for better metadata placement across CXL-enabled persistent storage, designs a directory-aware metadata grouping and placement strategy to improve the spatial locality of metadata access, and employs fully parallel metadata handlers to enhance metadata processing parallelism for parallel DFS client accesses. We evaluate MDSec on a cluster with 25 nodes. MDSec improves the throughput of Ext4 and NOVA by 258% and 53%, while reducing their latency by 46% and 11%, respectively. These results indicate that MDSec efficiently integrates CXL storage with DFS metadata service.
Xuchao Xie, Xinghan Qiao, Qiulin Wu, Wenhao Gu, Liquan Xiao
CLUSTER7
2025 Elevating Temporal Prefetching Through Instruction Correlation
Shuiyi He, Zicong Wang, Dezun Dong, Liquan Xiao
MICRO6
2025 A clustering ensemble algorithm for handling deep embeddings using cluster confidence
abstract
Abstract Clustering ensemble, which aims to learn a robust consensus clustering from multiple weak base clusterings, has achieved promising performance on various applications. With the development of big data, the scale and complexity of data is constantly increasing. However, most existing clustering ensemble methods typically employ shallow clustering algorithms to generate base clusterings. When confronted with high-dimensional complex data, these shallow algorithms fail to fully utilize the intricate features present in the latent data space. As a result, the quality and diversity of the generated base clusterings are insufficient, thus affecting the subsequent ensemble performance. To address this issue, we propose a novel clustering ensemble algorithm for handling deep embeddings using cluster confidence (CEDECC) to improve the robustness and performance. Instead of simply combining deep clustering with clustering ensembles, we take into consideration that the performance of existing deep clustering methods heavily relies on the quality of low-dimensional embeddings generated during the pre-training stage. The quality of embeddings is unstable due to the influence of different initialization parameters. In CEDECC, specifically, we first construct a cluster confidence measure to evaluate the quality of low-dimensional embeddings. Typically, high-quality low-dimensional embeddings yield accurate clustering results with the same model parameters. Then, we utilize multiple high-quality embeddings to generate the base partitions. In the ensemble strategy phase, we consider the cluster-wise diversity and propose a novel ensemble cluster estimation to improve the overall consensus performance of the model. Extensive experiments on three benchmark datasets and four real-world biological datasets have demonstrated that the proposed CEDECC consistently outperforms the state-of-the-art clustering ensemble methods.
Lingbin Zeng, Shixin Yao, Xinwang Liu 0002, Liquan Xiao
Comput. J.4
2024 ImSPU: Implicit Sharing of Computation Resources Between Vector and Scalar Processing Units
Hongbing Tan, Guichu Sun, Liquan Xiao, Yuanhu Cheng, Quan Deng 0003, Bingcai Sui, Yongwen Wang, Libo Huang 0002
Euro-Par (2)4
2024 FPGA Implementation of Sequence Detector for High-Speed PAM4 Wireline Transceiver
abstract
To solve the problem of high bit error rate (BER) due to high inter-symbol interference (ISI) in high-speed wireline transceivers, this paper proposes a low-complexity adaptive reduced-state sequence detector (ARSSD). The detector is based on the maximum likelihood sequence detection (MLSD) to reduce the detection bit error rate (BER), adopts the ISI parameter acquisition method based on the zero-forcing algorithm to achieve the adaptive detector parameters, and combines the viterbi algorithm and the set partitioning algorithm to reduce the complexity of operations. The behavioral simulation and the implementation of the hardware circuit are completed in this paper. The experimental results based on the analog front-end and the field programmable gate array (FPGA) show that when the pulse amplitude modulation 4 (PAM4) bit rate is 12 ∼ 56Gbps and the channel loss is -5dB ∼ -17dB@14GHz, the detection BERs of 32x4 parallel ARSSDs are reduced by two orders of magnitude compared to the conventional decision feedback equalization, which is consistent with the results of the behavioral simulation.
Chaolong Xu, Fangxu Lv, Zhengbin Pang, Liquan Xiao, Zhouhao Yang
ACM Great Lakes Symposium on VLSI4
2024 A survey of machine learning for Network-on-Chips
Dezun Dong, Cunlu Li, Liquan Xiao
J. Parallel Distributed Comput.5
2024 A Low-Cost Floating-Point Dot-Product-Dual-Accumulate Architecture for HPC-Enabled AI
abstract
The dot-product$\sum _{i=1}^{N} A_{i}\times B_{i}$is one of the most frequently used operations for a wide variety of high-performance computing (HPC) and artificial intelligence (AI) applications. However, for large-scale algorithms, such as acrshort GEMM and acrshort FFT, independent additions are necessary to accumulate the results of length-limited dot-product in order to form the final result, thus increasing latency and overhead. Hence, we proposed a dot-product-dual-accumulate (DPDAC) architecture capable of performing$\left({\sum _{i=1}^{N=1,2,4} A_{i}\times B_{i} + \sum _{j=1}^{M=1,2} C_{j}}\right)$on a wide range of formats. The proposed architecture supports both single-path and dual-path execution. The single path is designed for performing acrshort DP acrshort FMA or DPDAC of lower formats, while dual-path supports parallel operations for single-precision (SP) addition and 2-term SP or acrshort TF32 dot-product or 4-term acrshort HP or BF16 dot-product. Moreover, numerical precision conversion is also supported by the proposed architecture, allowing for the conversion of numbers to higher or lower formats. The proposed DPDAC has been demonstrated to significantly reduce the overhead in comparison to discrete designs that utilize multiple single-mode acrshort FP units to achieve the same functionalities. Furthermore, when compared to the state-of-the-art multiple-precision designs, the proposed architecture has been shown to support a wide range of formats and a greater variety of operations with lower costs.
Hongbing Tan, Libo Huang 0002, Hui Guo 0004, Qianming Yang, Li Shen 0007, Gang Chen 0023, Liquan Xiao, Nong Xiao 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.8
2023 A Multi-level Parallel Integer/Floating-Point Arithmetic Architecture for Deep Learning Instructions
Hongbing Tan, Libo Huang 0002, Dezun Dong, Yongwen Wang, Liquan Xiao
Euro-Par7
2023 DeTAR: A Decision Tree-Based Adaptive Routing in Networks-on-Chip
Dezun Dong, Cunlu Li, Liquan Xiao
Euro-Par6
2023 Multiple-Mode-Supporting Floating-Point FMA Unit for Deep Learning Processors
abstract
In this article, a new multiple-mode floating-point fused multiply–add (FMA) unit is proposed for deep learning processors. The proposed design supports three functional modes—normal FMA mode, mixed FMA mode, and dual FMA mode—and four types of precision—single-precision (SP), half-precision (HP), BFloat16 (BF16), and TensorFloat-32 (TF32)—based on the practical requirements of deep learning applications. In the normal FMA mode, conventional FMA operations, one SP operation or two parallel HP operations, are performed every clock cycle. In the mixed FMA mode and dual FMA mode, mixed-precision operations, the fused multiply–accumulate and the dot-product, are implemented, respectively. Specifically, the product of lower precision multiplication can be accumulated to a higher precision addend. Compared with the mixed FMA mode, the throughput is doubled in the dual FMA mode due to the full utilization of the multiplier operand bandwidth. In addition to FMA operations, numerical precision conversion (NPCvt) is also supported in this work: higher precision FMA results can be converted into lower precision numbers, corresponding to the datatype transform in the datapath of deep neural network (DNN) training. The FMA design presented herein uses both the segmentation and reusing methods to trade off performance, such as throughput and latency, against area, and power. Compared with the state-of-the-art multiple-precision FMA unit, the proposed design supports more types of floating-point operation and NPCvt, with higher throughput and lower hardware overhead.
Hongbing Tan, Gan Tong, Libo Huang 0002, Liquan Xiao, Nong Xiao 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2022 Efficient Multiple-Precision and Mixed-Precision Floating-Point Fused Multiply-Accumulate Unit for HPC and AI Applications
Hongbing Tan, Run Yan, Ling Yang 0008, Libo Huang 0002, Liquan Xiao, Qianming Yang
ICA3PP5
2021 Evaluation of Topology-Aware All-Reduce Algorithm for Dragonfly Networks
Dezun Dong, Cunlu Li, Ke Wu 0003, Liquan Xiao
NPC5
2021 Communication optimization strategies for distributed deep neural network training: A survey
Shuo Ouyang, Dezun Dong, Yemao Xu, Liquan Xiao
J. Parallel Distributed Comput.4
2020 AquaSee: Predict Load and Cooling System Faults of Supercomputers Using Chilled Water Data
Liquan Xiao, Jinghua Feng
J. Comput. Sci. Technol.2
2019 EC4: ECN and Credit-Reservation Converged Congestion Control
abstract
Bursty traffic and thousands of concurrent flows incur inevitable congestion in data center networks (DCNs) and then affect the overall performance. Various transport protocols are developed to mitigate the network congestion, including reactive and proactive protocols. Reactive schemes to handling congestion after congestion arises are common to current DCNs. However, with the growth of scale and link speed, reactive schemes such as DCTCP encounter the significant problem of slow responding to congestion. On the contrary, proactive protocols are designed to avoid congestion, and they have the advantages of zero data loss, fast convergence and low buffer occupancy (e.g., credit-reservation protocols). But in actual deployment scenario, it is hard to guarantee one protocol to be deployed in every server at one time. When credit-reservation protocol is deployed to DCNs step-by-step, the network is converted to multi-protocol state and faces the following fundamental challenges: (i) unfairness, (ii) high bu er occupancy, and (iii) heavy tail delay. Therefore, we propose EC4, which is for converging ECN-based and credit-reservation protocols with minimal modification. To the best of our knowledge, EC4is the first to address how to harmonize proactive and reactive congestion control. Targeting the common ECN-based protocol-DCTCP, EC4leverages the Forward Explicit Congestion Notification (FECN) to deliver realtime congestion information and redefines feedback control. After evaluation, the results show that EC4e ectively addresses the unfair link allocation. Furthermore, even workloads at 0.6 does not cause buffer overflow, thus largely eliminating the timeouts problem.
Zihao Wei, Dezun Dong, Shan Huang 0002, Liquan Xiao
ICPADS4
2019 Pinpointing and scheduling access conflicts to improve internal resource utilization in solid-state drives
Xuchao Xie, Liquan Xiao, Dengping Wei, Zhenlong Song, Xiongzi Ge
Frontiers Comput. Sci.2
2019 ZoneTier: A Zone-based Storage Tiering and Caching Co-design to Integrate SSDs with SMR Drives
abstract
Integrating solid-state drives (SSDs) and host-aware shingled magnetic recording (HA-SMR) drives can potentially build a cost-effective high-performance storage system. However, existing SSD tiering and caching designs in such a hybrid system are not fully matched with the intrinsic properties of HA-SMR drives due to their lacking consideration of how to handle non-sequential writes (NSWs). We propose ZoneTier, a zone-based storage tiering and caching co-design, to effectively control all the NSWs by leveraging the host-aware property of HA-SMR drives. ZoneTier exploits real-time data layout of SMR zones to optimize zone placement, reshapes NSWs generated from zone demotions to SMR preferred sequential writes, and transforms the inevitable NSWs to cleaning-friendly write traffics for SMR zones. ZoneTier can be easily extended to match host-managed SMR drives using proactive cleaning policy. We implemented a prototype of ZoneTier with user space data management algorithms and real SSD and HA-SMR drives, which are manipulated by the functions provided by libzbc and libaio. Our experiments show that ZoneTier can reduce zone relocation overhead by 29.41% on average, shorten performance recovery time of HA-SMR drives from cleaning by up to 33.37%, and improve performance by up to 32.31% than existing hybrid storage designs.
Xuchao Xie, Liquan Xiao, David Hung-Chang Du
ACM Trans. Storage2
2019 A Cost-Efficient Router Architecture for HPC Inter-Connection Networks: Design and Implementation
abstract
High-radix routers with lower latency and higher bandwidth play an increasingly important role in constructing large-scale interconnection networks such as those used in super-computers and datacenters. The tile-based crossbar approach partitions a single large crossbar into many small tiles and can considerably reduce the complexity of arbitration while providing higher throughput than the conventional switch implementation. However, it is not scalable due to power consumption, placement, and routing problems. Inspired by non-saturated throughput theory, this paper proposes a scalable router microarchitecture, termed Multiport Binding Tile-based Router (MBTR). By aggregating multiple physical ports into a single tile a high-radix router can be flexibly organized into different tile arrays, thus the number of tiles and hardware overhead can be considerably reduced. For a radix-64 router MBTR achieves up to$50 \sim 75\%$reduction in memory consumption as well as wire area compared with a hierarchical switch. We theoretically deduce the sufficient and necessary conditions for the asymmetrical crossbar to achieve un-saturated relative 100 percent throughput. Based on this observation we analyze the MBTR throughput and derive the condition that should be satisfied by the MBTR design parameters to yield 100 percent throughput. We further discuss how to make a trade-off between MBTR parameters based on the constraints of performance, power and area. The simulation results demonstrate MBTR is indistinguishable from the YARC router in terms of throughput and delay, and can even outperform it by reducing potential contention for output ports. We have fabricated a 36-port MBTR chip at 28 nm, providing 100 Gb/s bidirectional bandwidth per port, with a fall-through latency of just 30 ns. Internally it runs at 9.6 Tb/s, thus offering a speedup of$1.34\times$.
Kai Lu 0001, Liquan Xiao, Jinshu Su
IEEE Trans. Parallel Distributed Syst.3
2018 Duchy: Achieving Both SSD Durability and Controllable SMR Cleaning Overhead in Hybrid Storage Systems
abstract
Integrating solid-state drives (SSDs) and shingled magnetic recording (SMR) drives can build cost-effective hybrid storage systems. However, both SSD and SMR drives endure inherent defects that are mutually exclusive. The write endurance of SSD is limited while SMR drives should prevent from writes due to the cleaning-caused performance degradation. In this paper, we propose Duchy, an endurable SSD caching scheme that simultaneously respects SMR constraints. Duchy filters ineffectual write traffic out of SSD without exacerbating the performance degradation of SMR drives. Meanwhile, Duchy leverages SSD to regulate the written zones in SMR drives to achieve controllable cleaning duration. Our experimental results indicate that compared with legacy SSD caching designs, only Duchy can achieve both system performance improvement and SSD write traffic reduction.
Xuchao Xie, Tianye Yang, Dengping Wei, Liquan Xiao
ICPP5
2018 BFRP: Endpoint Congestion Avoidance Through Bilateral Flow Reservation
abstract
In HPC, endpoint congestion is a bottleneck in the network and seriously affects the performance of the system. The endpoint congestion can be effectively mitigated by quickly responding to the network and reducing the injection rate of the source. However, most of the prior works do not consider the impact of flow completion time on system performance, but only focus on the packet latency and perform scheduling at packet granularity. For HPC applications, the flow completion time and throughput are the metrics that determines the speed of application execution. Although prior works reduce package latency, they do not fundamentally reduce the flow latency in flow level. In this paper, we propose the bilateral flow-based reservation protocol (BFRP). BFRP quickly responds to network conditions through light-weight bilateral reservation mechanism and can effectively avoid the formation of endpoint congestion. BFRP also schedules packets based on flows, and smallest flow is preferentially sent to decrease the average flow latency. BFRP ensures the source and destination send or receive flows according to the allocated time slices without any conflict at both ends. We evaluate our BFRP protocol against state-of-the-art reservation-based protocol, speculative reservation protocol(SRP), and the simulation results show that the flow latency can be reduced by 27.68% under hotspot traffic with fixed flow size and 24.59% under uniform traffic with fixed flow size.
Tianye Yang, Dezun Dong, Cunlu Li, Liquan Xiao
IPCCC4
2017 A Scalable and Resilient Microarchitecture Based on Multiport Binding for High-Radix Router Design
abstract
High-radix routers with low latency and high bandwidth play an increasingly important role in the design of large-scale interconnection networks such as those used in super-computers and datacenters. The tile-based crossbar approach partitions a single large crossbar into many small tiles and can considerably reduce the complexity of arbitration while providing throughput higher than the conventional switch implementation. However, it is not scalable due to power consumption, placement, and routing problems. In this paper, we propose a truly scalable router microarchitecture called Multiport Binding Tile-based Router (MBTR). By aggregating multiple physical ports into a single tile a high-radix router can be flexibly organized into a different array of tiles, thus the number of tiles and hardware overhead can be considerably reduced. Compared with a hierarchical crossbar, MBTR achieves up to 50%~75% reduction in memory consumption as well as wire area. Simulation results demonstrate MBTR is indistinguishable from the YARC router in terms of throughput and delay, and can even outperform it by reducing potential contention for output ports. We have fabricated an ASIC MBTR chip with 28nm technology. Internally, it runs at 700MHz and 30ns latency without any speedup. We also discuss how the microarchitecture parameters of MBTR can be adjusted based on the power, area, and design complexity constraints of the arbitration logic.
Kefei Wang, Gang Qu 0001, Liquan Xiao, Dezun Dong, Xingyun Qi
IPDPS4
2016 Graphein: A Novel Optical High-Radix Switch Architecture for 3D Integration
Jie Jian, Liquan Xiao, Weixia Xu 0001
ICA3PP3
2016 The Efficient In-band Management for Interconnect Network in Tianhe-2 System
abstract
Interconnect network plays an important role in high performance computing systems. And its manageability directly affects the RAS (i.e., Reliability, Availability, and Serviceability) of the whole system. The Tianhe-2 system located in NSCC-gz (i.e., National Supercomputing Center of China in Guangzhou) uses proprietary interconnect network, which includes 5,856 high-radix network router chips (i.e., NRC) and 18,304 network interface chips (i.e., NIC). For such a very large-scale interconnect network, it is a great challenge to manage (such as configure, monitor, and debug) the numerous network chips and its network ports in an efficient way. By implementing the in-band management with very few hardware resources, the interconnect network in Tianhe-2 system achieves a highly efficient network management. In this paper, we introduce the design and implementation of the in-band management for interconnect network in Tianhe-2 system, especially emphasizing on several key features, including the set of achieved management functionalities, the architecture of network management, the format of management packets, the data flow and processing of management packets, etc. In this paper, we also evaluate the performance of in-band management by mainly comparing with out-band management scheme. The preliminary results demonstrate the efficiency of the in-band management for interconnect network in Tianhe-2 system.
Jijun Cao, Liquan Xiao, Zhengbin Pang, Kefei Wang
PDP2
2015 CER-IOS: Internal Resource Utilization Optimized I/O Scheduling for Solid State Drives
abstract
Modern Solid State Drives (SSDs) integrate more internal resources to get higher performance and capacity. Improving internal resource utilization by exploiting internal parallelism is important to enhance the performance of SSDs. Unfortunately, the internal resource utilization of SSDs is limited at runtime in practice because of the practical access conflicts to internal resources. In this paper, we propose a Conflict Eliminated Requests Based I/O Scheduler (CER-IOS) to better utilize internal parallelism of flash chips by scheduling I/O requests in a more fine-grained way. We introduce Conflict Eliminated Requests (CERs) in which parallelizable memory requests are grouped during the process of address translation in Flash Translation Layer. To schedule conflicting requests, we propose a small CER size prioritized resource distribution scheme, that ensures internal resources can always be distributed to valuable conflicting requests to further improve the efficiency of resource utilization. Our extensive experimental evaluation results show that CER-IOS provides significant improvement of resource utilization at runtime and reduces average I/O latency largely compared to state-of-the-art I/O schedulers implemented in operating systems.
Xuchao Xie, Dengping Wei, Zhenlong Song, Liquan Xiao
ICPADS5
2014 Selective Extension of Routing Algorithms Based on Turn Model
abstract
Turn-model is a classical method for designing partially adaptive routing algorithms without virtual channels, and can also be the basis of fully adaptive routing algorithms. We propose a novel scheme, Selective Extension of Routing Algorithms based on turn model (SERA), which alleviates restrictions on turn and path selections if possible without adding any new buffers or virtual channels. SERA can improve adaptivity of the original routing algorithms, and maintain the deadlock-free property. Thus, SERA is an important extension of the previous turn model theory. To present the effectiveness of SERA in adaptive algorithms, we redesign two existing routing algorithms, Odd-Even and LEAR. Simulation results show that the SERA scheme achieves an average delay reduction of 6% compared to the original routing algorithms.
Liquan Xiao, Sheng Ma, Zhengbin Pang, Kefei Wang
PDP2
2014 MilkyWay-2 supercomputer: system and application
Xiangke Liao, Liquan Xiao, Canqun Yang, Yutong Lu
Frontiers Comput. Sci.2
2013 ECAM: An Efficient Cache Management Strategy for Address Mappings in Flash Translation Layer
Xuchao Xie, Dengping Wei, Zhenlong Song, Liquan Xiao
APPT5
2013 A truly scalable IP lookup algorithm for next generation internet
abstract
Continuing growth in traffic, the size of routing table as well as the link speeds puts great challenge on the performance of Internet address lookup. The packet forwarding engine need perform 150 million IP address lookups per second for a 100Gbps line card while accommodating up to 500,000 prefixes and very likely 1M in the near future. TCAM based lookup solutions do not scale well for the next-generation due to its high power dissipation and large footprint. Algorithmic solutions can achieve high throughputs with hardware parallelism and compact storage needs, but can hardly scale simultaneously in the three dimensions of increased table size, throughput, and prefix length as IP moves to 64 bit IPv6 addresses. We propose a novel IPv6 lookup algorithm called Prefix Range Based Binary Search (PRB-BS) using binary search over hash tables organized by the range of prefix length. By eliminating the path information of the subtrie our scheme reduces the memory demands considerably. Besides, the binary search scheme based on prefix range reduces the worst case time of hash lookups to log2L, with L being the number of subtrie layers. We also introduce asymmetry binary search to further reduce memory demands and memory accesses. Compared with the tree bitmap algorithm widely used by many modern routers the memory requirements of PRB-BS algorithm is about 55% of the tree bitmap algorithm. The experiment results also show that the average search depth of PRB-BS algorithm is stable at 2 with the increased number of prefixes, which reflects better scalability of PRB-BS algorithm.
Liquan Xiao, Kefei Wang, Heying Zhang, Bao-kang Zhao
ISCC2
2013 Exploiting hierarchy parallelism for molecular dynamics on a petascale heterogeneous system
Canqun Yang, Tao Tang 0001, Liquan Xiao
J. Parallel Distributed Comput.4
2011 A Halting Algorithm to Determine the Existence of the Decoder
abstract
Complementary synthesis automatically synthesizes the decoder circuit of an encoder. It determines the existence of the decoder by checking whether the encoder's input can be uniquely determined by its output. However, this algorithm will not halt if the decoder does not exist. To solve this problem, a novel halting algorithm is proposed. For every path of the encoder, this algorithm first checks whether the encoder's input can be uniquely determined by its output. If yes, the decoder exists; otherwise, this algorithm checks if this path contains loops, which can be further unfolded to prove the non-existence of the decoder for all those longer paths. To illustrate its usefulness and efficiency, this algorithm has been run on several complex encoders, including PCI-E and Ethernet. Experimental results indicate that this algorithm always halts properly by distinguishing correct encoders from incorrect ones, and it is more than three times faster than previous ones.
ShengYu Shen, Liquan Xiao, Kefei Wang, Sikun Li
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2010 TH-1: China's first petaflop supercomputer
Xuejun Yang, Xiangke Liao, Weixia Xu 0001, Junqiang Song, Qingfeng Hu, Jinshu Su, Liquan Xiao, Kai Lu 0001, Qiang Dou, Juping Jiang, Canqun Yang
Frontiers Comput. Sci. China7
2010 Synthesizing Complementary Circuits Automatically
abstract
One of the most difficult jobs in designing communication and multimedia chips is to design and verify the complex complementary circuit pair (E,E-1), in which circuit E transforms information into a format suitable for transmission and storage, and its complementary circuitE-1recovers this information. In order to facilitate this job, we proposed a novel two-step approach to synthesize the complementary circuitE-1fromEautomatically. First, a SAT solver was used to check whether the input sequence ofEcan be uniquely determined by its output sequence. Second, the complementary circuitE-1was built by characterizing its Boolean function, with an efficient all-solution SAT solver based on discovering XOR gates and extracting unsatisfiable cores. To illustrate its usefulness and efficiency, we ran our algorithm on several complex encoders from industrial projects, including PCIE and 10 G Ethernet, and successfully built correct complementary circuits for them.
ShengYu Shen, Kefei Wang, Liquan Xiao, Sikun Li
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2007 Look-Ahead Adaptive Routing on k -Ary n -Trees
Quanbao Sun, Liquan Xiao, Minxuan Zhang
APPT2
2007 Hardware-Based Multicast with Global Load Balance on k-ary n-trees
abstract
The multicast operation is used commonly in parallel applications and can be used to support several other collective communication operations. A significant performance improvement can be achieved by supporting multicast operations at the hardware level. In this paper, we propose two parent selecting strategies which use global information to reduce the conflict among different multicast operations on k-ary n-trees. We first define an equivalence relation to divide the switches at each stage into several equivalence classes. Then we prove that the switches, which are at the same stage and are passed through by the same multicast tree, belong to the same equivalence class. Based on the study, two least loaded parent selecting strategies are developed. The proposed strategies are evaluated through simulation experiments. The results indicate that the proposed strategies lower the multicast latency and increase the multicast throughput significantly.
Quanbao Sun, Minxuan Zhang, Liquan Xiao
ICPP3
2005 Preferential Bandwidth Allocation for Short Flows with Active Queue Management
Heying Zhang, Liu Lu, Liquan Xiao, Wenhua Dou
NPC3
2005 A New Self-tuning Active Queue Management Algorithm Based on Adaptive Control
Heying Zhang, Baohong Liu, Liquan Xiao, Wenhua Dou
NPC3