VLDB 2026 Research / reviewers in the wild / expert
Xuejun An
dblp:67/7555
· DBLP profile ↗
40ranked-venue papers
0as first author
26since 2021 · last 2026
0009-0005-0494-6332ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 33 · 24 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BitRed: Taming Non-Uniform Bit-Level Sparsity with a Programmable RISC-V ISA for DNN AccelerationabstractThe non-uniform and dynamic nature of Bit-Level Sparsity (BLS) poses a critical load-imbalance challenge for parallel hardware accelerators. While the Bit-Interleaving paradigm, represented by state-of-the-art accelerators like Bitlet, shows promise, it is fundamentally constrained by a rigid datapath and severe inter-channel load imbalance. This paper introduces BitRed, an accelerator that embodies a new ''programmable adaptive bit-interleaving'' philosophy. Rather than a monolithic design, BitRed's core is an Adaptive-Sparse Processing Unit (ASPU) that deconstructs the acceleration process into a set of orthogonal RISC-V ISA extensions for pre-processing (cal.pre), adaptive distillation with dynamic load balancing (cal.adis), and PDP-optimal reduction (cal.red). By transforming a rigid hardware problem into a flexible scheduling problem, this ISA-based approach provides a fundamentally more adaptable and extensible solution. Empirical studies on a broad set of benchmarks highlight the following results (normalized to a SCNN baseline): (1) up to 9.4× speedup over the Bitlet, and 5.6× over the latest bit-serial SOTA BitWave; (2) up to 7.6× higher inference efficiency than Bitlet on representative models; (3) 5.072mm2 area and scalable power consumption from 550.43mW (float32) to 495.12mW (16b) and 457.90mW (8b) @28nm TSMC; and (4) high versatility across precisions, and up to 18.9×13.5× higher than NVIDIA A100 and Jetson Orin 32GB, demonstrating significant competitiveness against GPUs. Yanhuan Liu, Kunming Zhang, Yuqun Liu, Siao Wen, Lexin Wang, Tianyu Liu 0007, Zhihua Fan, Xiaochun Ye, Dongrui Fan, Xuejun An |
ASPLOS (2) | 12 |
| 2026 | A2RT: Efficient Ray Tracing Accelerator with Approximate-Accurate Computing and QuantizationabstractRay tracing (RT) has revolutionized photorealistic rendering by simulating light transport, but existing methods face a trade-off between computational efficiency and rendering accuracy. To address this, we present A2RT, a software-hardware co-designed RT accelerator employing the end to end optimization of "quantization → computation". On the software side, we introduce a customized data flow mechanism with type-specific quantization for bounding boxes, ray origins, and directions, and we organize BVH nodes into Group- and Sub-Nodes. At the hardware level, a heterogeneous RT engine allocates resources based on node criticality: accurate computing units handle Group-Nodes, while approximate units process Sub-Nodes. A custom INT-FLOAT approximate multiplier further accelerates the approximate units. Experimental results show that A2RT achieves 45.51% energy consumption and 2.29× speedup over RT Core, and consumes 81.79% of energy while delivering 1.57× performance improvement compared to state-of-the-art accelerators. Zhihua Fan, Yudong Mu, Zhen Wang 0045, Xiaochun Ye, Xuejun An |
DATE | 8 |
| 2026 | DCSR: A Fast Data Structure with Leaf-Oriented Locks for Streaming Graph Processing
Jie Zhang 0130, Huawei Cao, Yuan Zhang 0031, Xuejun An |
EDBT | 5 |
| 2026 | HGNNMap: Heterogeneous Graph Neural Network-Based Mapping for Spatial Accelerators
Shengzhong Tang, Zhihua Fan, Tianyu Liu 0007, Xuejun An, Xiaochun Ye |
Euro-Par (1) | 4 |
| 2026 | B-Graphless: Batch-based serverless graph processing for embodied AI backends
Jie Zhang 0130, Huawei Cao, Yuan Zhang 0031, Xuejun An, Xiaochun Ye |
Future Gener. Comput. Syst. | 6 |
| 2026 | A real-time edge SAR imaging acceleration architecture utilizing multi-level dataflow parallelism
Yinshen Wang, Zhengxuan Hu, Zhihua Fan, Xuejun An, Xiaochun Ye |
J. Syst. Archit. | 6 |
| 2026 | A RISC-V Extended Infrastructure for Edge FHE Through Software and Hardware Co-DesignabstractFully Homomorphic Encryption (FHE) is a foundational technique in privacy-preserving computation, enabling secure data processing without decryption. However, existing acceleration approaches for edge-side FHE suffer from several challenges, including limited speedup, complex hardware designs, and underutilization of CPU resources. To address these issues, we propose a lightweight yet effective hardware-software co-design acceleration scheme based on RISC-V architecture. First, we propose a custom RISC-V instruction set extension tailored for fully homomorphic encryption, enabling efficient and fine-grained acceleration of modular arithmetic operations. Second, we design an efficient FHE acceleration architecture, including customized circuit implementations and a pipelined execution unit, to support low-latency and energy-efficient computation. Third, we introduce a parallel software acceleration strategy based on butterfly computation patterns and multi-threading techniques, fully utilizing the RISC-V vector extension and multi-core resources. Experimental results demonstrate that our solution achieves an average speedup of 11.5× and an energy efficiency improvement of 7.05×. Compared to the current state-of-the-art design, our approach delivers an average performance gain of 2.27× with a significantly reduced design complexity. Zhihua Fan, Xuejun An, Xiaochun Ye |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2026 | A Comprehensive Survey on Dynamic Graph Processing: Storage and AnalyticsabstractDynamic graph processing is becoming increasingly critical across a wide range of domains, including social networks, financial transactions, and business intelligence. Its effectiveness relies heavily on optimizations in both storage and analytics, which are essential for improving system performance, throughput, and scalability. While dynamic graph processing has attracted significant research attention and yielded notable progress, a comprehensive analysis that integrates advancements in both dynamic graph storage and analytics remains lacking. To address this gap, this paper presents a thorough review of stateof-the-art techniques that support dynamic graph processing, with a particular focus on storage and analytical methods. Specifically, we first outline the fundamental challenges and core design principles in the field. Then, we systematically classify and summarize existing approaches, encompassing dynamic graph storage and analytics optimizations across both CPU and GPU platforms. Finally, we identify key research gaps and suggest promising directions for future work. This survey presents a comprehensive and up-to-date review of the literature on dynamic graph processing, offering valuable insights for both new and established researchers and contributing to the advancement of the field. The related materials for this paper are available at:https://github.com/yzhang610/DynGraphSurvey. Yuan Zhang 0031, Huawei Cao, Xuejun An, Xiaochun Ye |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2026 | Toward Resource-Efficient Billion-Scale SpGEMM on CPU-GPU Heterogeneous Server
Ming Dun, Shuhan Song, Huawei Cao, Xuejun An, Xiaochun Ye |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2025 | NFMap: Node Fusion Optimization for Efficient CGRA Mapping with Reinforcement Learning
Yudong Mu, Zhihua Fan, Xuejun An, Xiaochun Ye |
APPT | 5 |
| 2025 | GASgraph: A GPU-Accelerated Streaming Graph Processing System Based on SubHPMAs
Yuan Zhang 0031, Huawei Cao, Xuejun An, Xiaochun Ye |
APPT | 4 |
| 2025 | FDHA: Fusion-Driven Heterogeneous Accelerator for Efficient Diffusion Model Inference
Yudong Mu, Zhihua Fan, Xiaoxia Yao, Honglie Wang, Xuejun An, Xiaochun Ye |
Euro-Par (2) | 7 |
| 2025 | CGP-Graphless: Towards Efficient Serverless Graph Processing via CPU-GPU Pipelined Collaboration
Jie Zhang 0130, Huawei Cao, Xuejun An, Xiaochun Ye |
Euro-Par (1) | 5 |
| 2025 | CacheGuardian: A Timing Side-Channel Resilient LLC DesignabstractIn cloud computing environments, the last-level cache (LLC) shared by multiple tenants is frequently exploited through timing side-channel attacks, enabling unauthorized data leakage. To address this issue, various defense mechanisms have been proposed. However, existing works exhibit deficiencies in terms of performance overhead, coverage of attacks, and detection accuracy. In response to these challenges, we propose CacheGuardian, a hardware-based LLC protection design which aims to provide stronger, broader, and more accurate protection against timing side-channel attacks with low performance overhead. It includes: (1) A behavior-based, generic attack detector capable of identifying multiple timing side-channel attacks in real time; (2) A cache-set-level access control mechanism that strictly restricts cache usage exclusively for the identified attackers instead of influencing all security domains.We implement our design in a gem5 simulator to evaluate both its security and performance. Our proof-of-concept attacks and SPEC 2017 benchmarks show that our design is effective against a wide range of timing side-channel attacks, reducing attack success rates by up to 256×, including camouflaged variants. Moreover, it improves the performance of benign workloads by an average of 2.26% with only 2.4% storage overhead. Ziang Zhou, Huifeng Zhu, Wei Yan 0005, Chenglu Jin, Xuejun An, Xiaochun Ye |
ICCAD | 7 |
| 2025 | StreamDCIM: A Tile-based Streaming Digital CIM Accelerator with Mixed-stationary Cross-forwarding Dataflow for Multimodal TransformerabstractMultimodal Transformers are emerging artificial intelligence (AI) models designed to process a mixture of signals from diverse modalities. Digital computing-in-memory (CIM) architectures are considered promising for achieving high efficiency while maintaining high accuracy. However, current digital CIM-based accelerators exhibit inflexibility in microarchitecture, dataflow, and pipeline to effectively accelerate multimodal Transformer. In this paper, we propose StreamDCIM, a tile-based streaming digital CIM accelerator for multimodal Transformers. It overcomes the above challenges with three features: First, we present a tile-based reconfigurable CIM macro microarchitecture with normal and hybrid reconfigurable modes to improve intra-macro CIM utilization. Second, we implement a mixed-stationary cross-forwarding dataflow with tile-based execution decoupling to exploit tile-level computation parallelism. Third, we introduce a ping-pong-like fine-grained compute-rewriting pipeline to overlap high-latency on-chip CIM rewriting. Experimental results show that StreamDCIM outperforms non-streaming and layer-based streaming CIM-based solutions by geomean 2.63×and 1.28× on typical multimodal Transformer models. Shantian Qin, Ziqing Qiang, Zhihua Fan, Xuejun An, Xiaochun Ye, Dongrui Fan |
ISCAS | 5 |
| 2025 | Accelerating tensor multiplication by exploring hybrid product with hardware and software co-design
Zhihua Fan, Zhen Wang 0045, Xiaochun Ye, Dongrui Fan, Xuejun An |
J. Syst. Archit. | 8 |
| 2025 | GenCNN: A Partition-Aware Multi-Objective Mapping Framework for CNN Accelerators Based on Genetic AlgorithmabstractConvolutional Neural Networks (CNNs) require partitioning to efficiently run on CNN accelerators, which offer multiple parallel processing dimensions, such as Processing Element (PE) array topologies and Single Instruction Multiple Data (SIMD) execution. The choice of parallelization strategy directly impacts accelerator performance. However, the vast search space for CNN partitioning and parallelization makes manual optimization costly and complex, especially when addressing both aspects simultaneously. This highlights the need for an automated framework to efficiently map CNNs onto accelerators. Our key insight is that existing approaches suffer from inadequate accelerator performance modeling and a lack of multi-objective optimization strategies that jointly consider task partitioning and convolution parallelization. To address this, we propose GenCNN, a multi-objective genetic algorithm-based mapping framework for CNN accelerators. GenCNN first constructs a fine-grained performance model that captures both off-chip data access and on-chip data processing. It then applies the Non-dominated Sorting Genetic Algorithm II improved by Multi-Objective Bayesian Optimization to derive a Pareto-optimal partitioning and parallelization strategy that balances off-chip latency and PE utilization. Finally, GenCNN optimizes scheduling and routing to minimize data transfers. Experimental results show that GenCNN achieves up to 17.66× speedup in compilation and 6.47× in execution compared with state-of-the-art mapping frameworks. Yudong Mu, Zhihua Fan, Xuejun An, Dongrui Fan, Xiaochun Ye |
ACM Trans. Archit. Code Optim. | 5 |
| 2025 | PANDA: Adaptive Prefetching and Decentralized Scheduling for Dataflow ArchitecturesabstractDataflow architectures are considered promising architecture, offering a commendable balance of performance, efficiency, and flexibility. Abundant prior works have been proposed to improve the performance of dataflow architectures. Nevertheless, these solutions can be further improved due to the lack of efficient data prefetching and flexible task scheduling. In this article, we propose a novel dataflow architecture with adaptive p refetching an d d ecentr a lized scheduling (PANDA). First, we present an application-adaptive data prefetching method and on-chip memory microarchitecture designed to overlap memory access latency. Second, we introduce a decentralized dataflow scheduling approach and processing element (PE) microarchitecture aimed at improving hardware utilization. Experimental results show that in a wide range of real-world applications, PANDA attains up to 2.53× performance improvement and 1.79× energy efficiency improvement over the state-of-the-art dataflow architectures. Shantian Qin, Zhihua Fan, Zhen Wang 0045, Xuejun An, Xiaochun Ye, Dongrui Fan |
ACM Trans. Archit. Code Optim. | 5 |
| 2025 | CGCGraph: Efficient CPU-GPU Co-execution for Concurrent Dynamic Graph ProcessingabstractWith the continuous growth of user scale and application data, the demand for large-scale concurrent graph processing is increasing. Typically, large-scale concurrent graph processing jobs need to process corresponding snapshots of dynamically changing graph data to obtain information at different time points. To enhance the throughput of such applications, current solutions concurrently process multiple graph snapshots on the GPU. However, when dealing with rapidly changing graph data, transferring multiple snapshots of concurrent jobs to the GPU results in high data transfer overhead between CPU and GPU. Additionally, the execution mode of existing work suffers from underutilization of GPU computational resources. In this work, we introduce CGCGraph, which can be integrated into existing GPU graph processing systems like Subway, to enable efficient concurrent graph snapshot processing jobs and enhance overall system resource utilization. The key idea is to offload unshared graph data of multiple concurrent snapshots to the CPU, reducing CPU-GPU transfer overhead. By implementing CPU-GPU co-execution, there is potential for enhanced utilization of GPU computing resources. Specifically, CGCGraph leverages kernel fusion to process shared graph data concurrently on the GPU, while executing all snapshots in parallel on the CPU, with each snapshot assigned a dedicated thread. This approach enables efficient concurrent processing within a novel CPU-GPU co-execution model, incorporating three optimization strategies targeting storage, computation, and synchronization. We integrate CGCGraph with Subway, an existing system designed for out-of-GPU-memory static graph processing. Experimental results show that the integration of CGCGraph with current GPU-based systems obtains performance improvements ranging from 1.7 to 4.5 times. Jie Zhang 0130, Huawei Cao, Yuan Zhang 0031, Xuejun An, Junying Huang, Xiaochun Ye |
ACM Trans. Archit. Code Optim. | 5 |
| 2025 | A RISC-V Extended Infrastructure for CNNs Through Pipelined Computing and Data Dependence OptimizationabstractWith the rapid development of artificial intelligence (AI), convolutional neural networks (CNNs) have been widely applied in fields like computer vision and recommendation systems. This growth has intensified the demand for hardware acceleration of CNNs. Existing accelerators are either designed as co-processors or improve performance through extended instructions. While these methods can significantly improve performance, they often result in limited programming and execution flexibility. In this paper, we design custom RISC-V instructions specifically for CNNs to maximize data reuse and exploit parallelism. Then, to efficiently execute CNNs instructions, we extend a Pipelined Vector Computing Unit (PPVCU). Finally, we incorporate Pattern Detection Logic (PDL) to identify common data dependence patterns in CNNs, enabling the Data Dependence Computing Unit (DDCU) to process instructions within each pattern in parallel. Experimental results show that our approach achieves, on average, 9.54× performance improvement and 6.7× energy efficiency improvement compared to our baseline, 8.34× performance improvement and 3.1× energy efficiency improvement compared to state-of-the-art designs. Teng Luo, Tengfei Xia, Zhihua Fan, Yudong Mu, Xuejun An, Xiaochun Ye, Dongrui Fan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2024 | OBSD: On-The-Fly Block-Wise Sparse Distillation Accelerating SpGEMMs in DNN ApplicationsabstractThe mainstream deep neural networks (DNNs) widely adopt pruning techniques to alleviate network overfitting and computational complexity. Simultaneously, the weight and activation data exhibit increasingly noticeable block-wise sparsity. Current DNN-specific accelerators mainly focus on element-wise while neglecting sparse features of block granularity. Meanwhile, one of the state-of-the-art work, HIRAC, combining matrix tiling with the fast packing algorithm, SorPack, to achieve effective acceleration of Sparse General Matrix Multiplications (SpGEMMs). However, the SorPack algorithm executes on the host CPU, and its runtime lies on the critical path of the overall execution. In this work, a specific SpGEMM accelerator named OBSD is proposed which achieves high performance and energy efficiency. An on-the-fly block-wise distillation approach is proposed for leveraging the block-wise sparsity in both weight and activation data, which is implemented during the data loading process, consuming minimal additional time. Moreover, a data-flow architecture is devised to improve efficiency of data exchange among data processing elements (DPs) with investigating the relationship for different partition sizes, block-wise sparsity, and block-wise distance. The evaluation results demonstrate that: (1) OBSD achieves average of 1.89× speedup as compared to HIRAC for representative matrices in DNN workloads. (2) An end-to-end evaluation on a DNN model shows a 2.41× speedup and energy efficiency improvement of 1.99× over the HIRAC. (3) OBSD possesses a power consumption 12.9W with 56.6mm2area @28nm TSMC. Yanhuan Liu, Kunming Zhang, Zhihua Fan, Lexin Wang, Tianyu Liu 0007, Zhen Wang 0045, Xiaochun Ye, Dongrui Fan, Xuejun An |
HPCC | 12 |
| 2024 | Improving Utilization of Dataflow Unit for Multi-Batch ProcessingabstractDataflow architectures can achieve much better performance and higher efficiency than general-purpose core, approaching the performance of a specialized design while retaining programmability. However, advanced application scenarios place higher demands on the hardware in terms of cross-domain and multi-batch processing. In this article, we propose a unified scale-vector architecture that can work in multiple modes and adapt to diverse algorithms and requirements efficiently. First, a novel reconfigurable interconnection structure is proposed, which can organize execution units into different cluster typologies as a way to accommodate different data-level parallelism. Second, we decouple threads within each DFG node into consecutive pipeline stages and provide architectural support. By time-multiplexing during these stages, dataflow hardware can achieve much higher utilization and performance. In addition, the task-based program model can also exploit multi-level parallelism and deploy applications efficiently. Evaluated in a wide range of benchmarks, including digital signal processing algorithms, CNNs, and scientific computing algorithms, our design attains up to 11.95× energy efficiency (performance-per-watt) improvement over GPU (V100), and 2.01× energy efficiency improvement over state-of-the-art dataflow architectures. Zhihua Fan, Zhen Wang 0045, Xiaochun Ye, Dongrui Fan, Ninghui Sun, Xuejun An |
ACM Trans. Archit. Code Optim. | 8 |
| 2023 | Improving Utilization of Dataflow Architectures Through Software and Hardware Co-Design
Zhihua Fan, Shengzhong Tang, Xuejun An, Xiaochun Ye, Dongrui Fan |
Euro-Par | 4 |
| 2023 | FSGraph: fast and scalable implementation of graph traversal on GPUs
Yuan Zhang 0031, Huawei Cao, Jie Zhang 0130, Junying Huang, Xiaochun Ye, Xuejun An |
CCF Trans. High Perform. Comput. | 7 |
| 2023 | Accelerating Convolutional Neural Networks by Exploiting the Sparsity of Output ActivationabstractDeep Convolutional Neural Networks (CNNs) are the most widely used family of machine learning methods that have had a transformative effect on a wide range of applications. Previous studies have made great breakthroughs in accelerating CNNs, but they only target on the input sparsity of activation and weight, thus do not eliminate the unnecessary computations due to the fact that more zeros in the output results are not directly caused by the zero-valued positions of the input data. In this paper, we take advantage of the output activation sparsity to reduce the execution time and energy consumption of CNNs. First, we propose an effective prediction method that leverages the output activation sparsity. Our method first predicts the output activation polarity of convolutional layers based on the singular value decomposition (SVD) approach. Then, it uses the predicted negative value to skip invalid computations. Second, an effective accelerator is designed to take advantage of sparsity to achieve CNN inference acceleration. Each PE is equipped with a prediction unit and a non-zero value detection unit to remove invalid computation blocks. And an instruction bypass technique is proposed which further exploits the sparsity of the weights. The efficient dataflow graph mapping approach and pipeline execution ensure high computational resource utilization. Experiments show that our approach achieves up to 1.63× speedup and 55.30% energy reduction compared with dense networks with a slight loss of accuracy. Compared with Eyeriss, our accelerator achieves on average 1.31 × performance improvement and 54% energy reduction. Our accelerator also achieves a similar performance to SnaPEA, but with a better energy efficiency. Zhihua Fan, Zhen Wang 0045, Tianyu Liu 0007, Yanhuan Liu, Meng Wu 0006, Xinxin Wu, Xiaochun Ye, Dongrui Fan, Ninghui Sun, Xuejun An |
IEEE Trans. Parallel Distributed Syst. | 12 |
| 2022 | A Routing-Aware Mapping Method for Dataflow Architectures
Zhihua Fan, Tianyu Liu 0007, Xuejun An, Xiaochun Ye, Dongrui Fan |
NPC | 4 |
| 2019 | SwitchAgg: A Further Step Towards In-Network ComputationabstractMany distributed applications adopt a partition/aggregation pattern to achieve high performance and scalability. The aggregation process, which usually takes a large portion of the overall execution time, incurs large amount of network traffic and bottlenecks the system performance. To reduce network traffic, some researches take advantage of network devices to commit innetwork aggregation. However, these approaches use either special topology or middle-boxes, which cannot be easily deployed in current datacenters. The emerging programmable RMT switch brings us new opportunities to implement in-network computation task. However, we argue that the architecture of RMT switch is not suitable for in-network aggregation since it is designed primarily for implementing traditional network functions. In this paper, we first give a detailed analysis of in-network aggregation, and point out the key factor that affects the data reduction ratio. We then propose SwitchAgg, which is an innetwork aggregation system that is compatible with current datacenter infrastructures. We also evaluate the performance improvement we have gained from SwitchAgg. Our results show that, SwitchAgg can process data aggregation tasks at line rate and gives a high data reduction rate, which helps us to cut down network traffic and alleviate pressure on server CPU. In the system performance test, the job-completion-time can be reduced as much as 50%. Fan Yang 0096, Zhan Wang 0003, Xiaoxiao Ma 0004, Guojun Yuan, Xuejun An |
FPGA | 5 |
| 2019 | T2HT : Traffic-Driven Machine Learning Based Hierarchical Topology Generation ModelabstractIn high-performance computing (HPC) and distributed computing area, network performance greatly influenced application efficiency. However, due to the diversity of traffic patterns, the traditional network with fixed topology may achieve good performance under some applications, while performs poorly under other forms. Network reconfiguration technologies which can change the topology dynamically have been developed to obtain a balanced performance for different traffic patterns. Nonetheless, selecting an appropriate network topology from the wide variety of options remains difficult due to the complexity of analyzing traffic alongside topology performance characteristics. Traditional research focused on congestion estimation and specific parameter adjustment without reconfiguring the global topology. In this paper, we propose a generic Traffic to Hierarchical Topology (T2HT) method to analyze traffic patterns and choose an appropriate network configuration for the given traffic T2HT makes use of actual traffic data with a hierarchical model to predict network performance with a given topology and uses a machine learning (ML) algorithm to score the better options in order to determine the best topology. We performed 8000 simulations of dataset-topology combinations to verify the feasibility of the model. Our results show that T2HT achieved marked improvements with its recommendations, making it feasible for use in hierarchical network design. Under the DOE testbed, the throughput of the topology generated by T2HT can reach above 90% of theoretical limit(full connection), and the latency is improved by about 24.6% compared to typical topology 3D Torus with the same physical restrictions. Hongrui Zhu, Guojun Yuan, Guangming Tan, Zhan Wang 0003, Xuejun An |
ICPADS | 6 |
| 2017 | Regional Congestion Control in Datacenter NetworksabstractThe rapid deployment of cloud computing and online services poses great challenges for data center networks, and congestion control is one of the top concerns. Although numbers of proposals in different network layers have been put forward to alleviate the negative impact of congestion, the short-lived flows, which are latency-sensitive and constitute the majority of total traffic in data centers, still suffer severe performance degradation. Since the existing congestion control methods all rely on end hosts to perceive congestion and then adjust their network sending rate, the response time is relatively long when compared with the duration of short-lived flows, which increases latency significantly. In this paper, we propose RCC, a regional congestion control mechanism, which aims to respond to congestion more quickly and eliminate the mismatch mentioned above. Different from host-based mechanisms, RCC is implemented in the switch, which detects the congestion state and schedule the traffic around the congestion point locally, without sending feedback to the distal host. Evaluation has shown that, compared with the host-based mechanism, our method achieves better performance for short-lived flows and maintains stable buffer occupancy of the switch. In addition, mixed long- and short-lived flows which contend for the same bottleneck link can share the bandwidth more fairly. Fan Yang 0096, Zhan Wang 0003, Xiaoli Liu 0002, Zheng Cao 0003, Guojun Yuan, Xuejun An |
ICPADS | 6 |
| 2014 | Building a large-scale direct network with low-radix routersabstractCommunication locality is an important characteristic of parallel applications. A great deal of research shows that utilizing the characteristic will favor most applications. Aiming at communication locality, we present a hierarchical direct network topology to accelerate neighbor communication. Combining mesh topology and complete graph topology, it can be used to optimize local communication and build large-scale network with low radix routers. Analyzing the characteristic of hierarchical topology, we find the presented topology has high cost performance and excellent expandability. We also design two minimum path routing algorithms and compare them with Mesh, Dragonfly and PERCS topologies. The results show the saturated throughput of hierarchical topology is nearly 40% with uniform random trace and 70% with local communication model of 4K nodes. That indicates high scalability for applications with local communication and cost efficiency for uniform random trace. Zheng Cao 0003, Zhiguo Fan, Zhan Wang 0003, Xiaoli Liu 0002, Li Qiang, Xuejun An, Ninghui Sun |
ICPADS | 8 |
| 2014 | HiNetSim: A Parallel Simulator for Large-Scale Hierarchical Direct Networks
Zhiguo Fan, Zheng Cao 0003, Xiaoli Liu 0002, Zhan Wang 0003, Dawei Zang, Xuejun An |
NPC | 8 |
| 2014 | An Intra-Server Interconnect Fabric for Heterogeneous Computing
Zheng Cao 0003, Xiaoli Liu 0002, Qiang Li 0045, Zhan Wang 0003, Xuejun An |
J. Comput. Sci. Technol. | 6 |
| 2013 | Accelerating Allreduce Operation: A Switch-Based SolutionabstractCollective operations, such as all reduce, are widely treated as the critical limiting factors in achieving high performance in massively parallel applications. Conventional host-based implementations, which introduce a large amount of point-to-point communications, are less efficient in large-scale systems. To address this issue, we propose a design of switch chip to accelerate collective operations, especially the allreduce operation. The major advantage of the proposed solution is the high scalability since expensive point-to-point communications are avoided. Two kinds of allreduce operations, namely block-allreduce and burst-allreduce, are implemented for short and long messages, respectively. We evaluated the proposed design with both a cycle-accurate simulator and a FPGA prototype system. The experimental results prove that switch-based allreduce implementation is quite efficient and scalable, especially in large-scale systems. In the prototype, our switch-based implementation significantly outperforms the host-based one, with a 16 times improvement in MPI time on 16 nodes. Furthermore, the simulation shows that, upon scaling from 2 to 4096 nodes, the switch-based allreduce latency only increases slightly by less than 2 us. Nongda Hu, Zheng Cao 0003, Xuejun An, Ninghui Sun |
ICCCN | 4 |
| 2013 | cHPP controller: A High Performance Hyper-node Hardware AcceleratorabstractThe high-density blade server provides an attractive solution for the rapid increasing demand on computing. The degree of parallelism inside a blade enclosure nowadays has reach up to hundreds of cores. In such parallelism, it is necessary to accelerate communications inside a blade enclosure. However, commercial products seldom set foot in the optimization based on hardware. A hyper-node controller is proposed to provide a low overhead and high performance interconnection based on PCIe, which supports global address space, user-level communication, and efficient communication primitives. Furthermore, the efficient sharing of I/O resource is another goal of this design. The prototype of the hyper-node controller is implemented in FPGA. The testing results show the lowest latency is only 1.242us and the highest bandwidth is 3.19GB/s, which is almost 99.7% of the theoretic peak bandwidth. Zheng Cao 0003, Zhan Wang 0003, Xiaoli Liu 0002, Xuejun An, Ninghui Sun |
PDCAT | 6 |
| 2012 | Design of Hardware-Based Communication Performance Measurement ToolabstractWith the popularity and development of heterogeneous computing, proper communication performance measurement tools are needed to explore new communication patterns under heterogeneous computing systems and optimize program's performance. This paper proposes a hardware-based communication performance measurement tool, named as HCPM, which brings little influence on original program, and can collect communication traces generated by heterogeneous processors which implement PCIe or HT as their system bus. HCPM firstly provides basic communication primitives to set up a communication system. Then based on these primitives, it collects communication trace. Real-time collected traces are transmitted to a dedicated computer for further analysis. Evaluation shows that with the use of proper compression in hardware, HCPM can transmit at least five processors' communication traces with a single Gigabit Ethernet link. Zhan Wang 0003, Zheng Cao 0003, Xiaoli Liu 0002, Xuejun An |
CLUSTER | 6 |
| 2011 | Design of HPC Node with Heterogeneous ProcessorsabstractHeterogeneous Computing is becoming an important technology trend in HPC, where more and more heterogeneous processors are used. However, in traditional node architecture, heterogeneous processors are always used as coprocessors. Such usage increases the communication latency between heterogeneous processors and prevents the node from achieving high density. With the purpose of improving communication efficiency between heterogeneous processors, this paper proposed a new node architecture named HeteNode. In HeteNode, general purpose processors and heterogeneous processors are interconnected by a system controller directly and play the same role in both process of communication and process of computation. The prototype of HeteNode which contains nine processors in 1U chassis is built. Evaluation carried out on the prototype shows that 580ns minimum intra-node latency and 1.78us minimum inter-node latency between heterogeneous processors are achieved. Besides, NPB benchmarks show good scalability in HeteNode. Zheng Cao 0003, Hongwei Tang, Qiang Li 0045, Bo Li 0009, Xuejun An, Ninghui Sun |
CLUSTER | 7 |
| 2010 | Adding an Expressway to Accelerate the Neighborhood CommunicationabstractThe blade system is very popular in high performance computing. In a blade system, the blade is a fundamental element in which are symmetric multi-processors (SMP). About ten blades constitute a blade box, several blade boxes constitute a cabinet and some cabinets constitute a blade system at last. The blades in a blade box are neighbors because they have relatively short distance. Programmers always try to place the tightly related processes into the same blade box. However, there's seldom any optimization made by hardware to accelerate the communication in a blade box. Thus, a single chip design called hyper-node controller is presented to provide ultra low latency and high bandwidth which resembles an expressway between neighbors. All the nodes in a blade box can act as a single hyper node by using the hyper-node controller. It is apparent that the additional controller is a useful supplement to efficiently enhance the communication in a blade box and finally enhance the entire blade system. A FPGA prototype of the hyper-node controller has been implemented and it can connect five blades simultaneously. In the preliminary performance evaluation, the latency for an 8-byte payload between two blades is less than 1us, 1.33GB/s which is nearly 94% of the peak effective bandwidth can be obtained by transferring messages with a payload of only 256 bytes. Zheng Cao 0003, Xuejun An, Ninghui Sun |
HPCC | 4 |
| 2010 | HPP Controller: A System Controller Dedicated for Message PassingabstractThe traditional system controller in symmetric multi-processors (SMP) controls the memory, so it is suitable for the shared memory programming model. With the emergence of the processors which integrate memory controllers, the system controller seems less important than before. However, since the system controller resides in the center of a computer system, it acts as an artery which directly connects to the processors and the high-speed IO devices. Thus making full use of its position advantage can no doubt gain performance enhancement. By now, the message passing programming model has dominated the high performance computing (HPC) field, however the system controller makes little contribution to it. Thus, a system controller called HPP controller which is dedicated for the message passing programming model is presented in this paper. The HPP controller is connected to several processors simultaneously, and the communication between these processors uses the message passing programming model. The HPP controller has powerful DMA engines embedded which can provide flexible and sufficient message passing capability. Two key techniques: supporting arbitrary byte alignment and virtualizing the DMA engine are introduced in detail. The preliminary result of the FPGA prototype shows that the HPP controller has ultra low hardware latency and relatively high bandwidth. Besides, the NPB result shows that it can provide high efficiency for the message passing programming model. Zheng Cao 0003, Xuejun An, Ninghui Sun |
PDCAT | 4 |
| 2010 | HPP controller: a system controller for high performance computing
Zheng Cao 0003, Xuejun An, Ninghui Sun |
Frontiers Comput. Sci. China | 4 |
| 2009 | Gemini NI: An Integration of Two Network InterfacesabstractAccording to the development of the TOP500, the performance of the high performance computers (HPCs) is increasing rapidly. The incredible performance increment of the HPCs should be largely attributed to the development of their communication systems, because the HPCs cannot extend to such a large scale without their excellent communication systems. As an important member of the communication system, the network interface (NI) always plays a significant role. Since the network interface locates on the critical path of the communication system, it can easily become a bottleneck if it cannot provide low latency and high bandwidth for communication. The Gemini NI presented in this paper has a good performance potential in both latency and bandwidth. It has a remote load/store (RLS) mechanism which can provide ultra low latency. Furthermore, it has two HyperTransport (HT) interfaces connected to double processors or symmetric multi-processors (SMP), and it has four proprietary switch interfaces connected to the switches. This approach largely in-creases the throughput of the Gemini NI. Inside the Gemini NI, almost all the components of a network interface are duplicated. The resource sharing within the Gemini NI can provide great flexibility for scheduling. A FPGA prototype of the Gemini NI has been implemented, and the preliminary results prove the validity of our design. Xuejun An, Ninghui Sun |
NAS | 3 |