Weixia Xu 0001

dblp:35/3566-1 · DBLP profile ↗
← Back
36ranked-venue papers
1as first author
20since 2021 · last 2025
0009-0007-6327-7372ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 28 · 1 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Optimizing value prediction for ILP processors: A design space exploration approach
Ling Yang 0008, Libo Huang 0002, Run Yan, Sheng Ma, Yongwen Wang, Weixia Xu 0001
Integr.7
2025 A lightweight RDMA connection protocol based on post-hoc confirmation
Ke Wu 0003, Dezun Dong, Weixia Xu 0001
J. Parallel Distributed Comput.3
2025 Hardware/Software Co-design for spike communication optimization: Leveraging neuron-level communication patterns
Yuanfeng Luo, Weixia Xu 0001, Lei Wang 0011
J. Syst. Archit.6
2025 Bayesian Procedures for Modeling Truck Route Choices
abstract
This study examines logit models applied to the truck route choice problem using GPS trucking data from the Dallas metropolitan area. Instead of assuming a constant coefficient for each variable in the conventional multinomial logit model, the proposed mixed c-logit model assumes a certain probability distribution for each coefficient, in an attempt to better reflect the drivers’ preference heterogeneity. A commonality factor is introduced in the model to address roadway travel time correlations due to route overlaps. Three Bayesian models with different hierarchy levels are introduced and are solved using mean-field variational inference with the block coordinate algorithm. In the reduced subnetwork of the examined area, the drivers are assumed to make their route choices in three groups of routes, referred to as choice groups. Attributes that would affect the truck driver’s route choice decisions are different among choice groups. With this setting, the proposed Bayesian models are then tested with the three truck route choice groups respectively. Generally, the study finds that the factors considered in truckers’ route choice vary with context.
Xiubin Wang, Huibin Tan, Hengzhu Liu, Weixia Xu 0001
IEEE Trans. Intell. Transp. Syst.5
2024 Out-of-Order and Recursive RAS: A Return Address Stack Design on High Performance Processor
abstract
In high-performance processor design, maintaining Return-Address Stack (RAS) integrity is crucial for efficient instruction flow. Yet, separating multi-level branch predictors from L1I-caches brings substantial hurdles, especially when speculative execution corrupts the RAS through out-of-order branches. Past remedies struggle with storage demands and recursive call inefficiencies. Hence, we introduce the Out-of-Order and Recursive RAS (OR-RAS), an innovative architecture. It employs a advanced RAS for out-of-order control and a compressed LUT, enhancing efficiency. OR-RAS targets IPC boost and RAS error reduction (RAS-MPKI). Evaluations on powerful processors reveal a 0.5% IPC uplift, with effectively maintaining the MPKI below 0.02%. In essence, OR-RAS constitutes a holistic strategy against speculative execution issues and RAS corruption, auguring well for peak performance and dependability in contemporary microarchitectures.
Yude Fang, Libo Huang 0002, Yongwen Wang, Weixia Xu 0001
ASAP5
2024 Cost-Effective Value Predictor for ILP processors through Design Space Exploration
abstract
Value prediction is a microarchitectural technique that enhances processor performance by speculatively breaking true data dependencies. It has demonstrated improved performance in both single-threaded and multi-threaded workloads, rendering it an appealing microarchitectural approach. While high-performance value predictors can achieve impressive accuracy, they may also incur significant costs in terms of area, power consumption, and complexity. Therefore, there is a demand for lightweight value prediction techniques capable of striking a favorable balance between performance and overhead. However, designing value predictors with superior performance using limited resources presents an urgent challenge. Consequently, this work proposes a design space exploration framework for the state-of-the-art EVES value predictor, aiming to efficiently configure the design parameters of the value predictor within constrained RAM resources. Additionally, the article evaluates the performance of the explored value predictor across a wide range of workloads. The explored value predictors exhibit high efficiency across RAM sizes ranging from 2KB to 16KB while maintaining acceptable computational complexity. Furthermore, the results indicate that the explored value predictor achieves optimal efficiency under the 2KB constraint, with the highest acceleration-to-cost ratio reaching 4.02%/KB.
Ling Yang 0008, Libo Huang 0002, Run Yan, Sheng Ma, Yongwen Wang, Weixia Xu 0001
ACM Great Lakes Symposium on VLSI7
2024 The Self-adaptive and Topology-aware MPI_Bcast leveraging Collective offload on Tianhe Express Interconnect
abstract
Large parallel applications have heavily used MPI (Massage Passing Interface) collectives that support portable and efficient group communication operations. MPI_Bcast is one of the most commonly used MPI collectives that broadcast data to all processes of the communication domain. However, traditional software-based broadcast algorithms fail to fully utilize modern interconnection networks’ advanced features such as offloading collectives to the network hardware for efficient group communications. Besides, the semantic gap between MPI_Bcast and hardware multicast of underlying interconnects presents challenges for offload-based algorithms to accelerate MPI_Bcast for a wide range of message sizes.In this paper, we propose a hardware-software co-design MPI_Bcast by efficiently leveraging the NIC-based collective offload provided by Tianhe-express interconnect, which completely precludes the involvement of CPU to accelerate message broadcast. We detail this broadcast mechanism that can be adaptively tuned to offload MPI_Bcast operations from the CPU to the NIC for various message and system sizes. In addition, we further propose a topology-aware broadcast design in conjunction with this offload method to significantly reduce the broadcast latency by constructing the optimal global inter-node communication tree. We implement and evaluate the proposed Tianhe-Express Offload-based Broadcast (TOB) design on Tianhe-2A and Tianhe-EP supercomputers. Extensive experiments have been conducted to evaluate TOB performance at both microbenchmark and application levels. Our solution offers up to 4.94x significant performance speedup at the microbenchmark level over state-of-the-art MPI libraries. For the application-level evaluation, our technique accelerates scientific applications by a maximum speedup of 1.34x.
Chongshan Liang, Jinbo Xu, Jintao Peng, Weixia Xu 0001, Jie Liu 0002, Zhiquan Lai, Sheng Ma
IPDPS6
2024 COER: A Network Interface Offloading Architecture for RDMA and Congestion Control Protocol Codesign
abstract
RDMA (Remote Direct Memory Access) networks require efficient congestion control to maintain their high throughput and low latency characteristics. However, congestion control protocols deployed at the software layer suffer from slow response times due to the communication overhead between host hardware and software. This limitation has hindered their ability to meet the demands of high-speed networks and applications. Harnessing the capabilities of rapidly advancing Network Interface Cards (NICs) can drive progress in congestion control. Some simple congestion control protocols have been offloaded to RDMA NICs to enable faster detection and processing of congestion. However, offloading congestion control to the RDMA NIC faces a significant challenge in integrating the RDMA transport protocol with advanced congestion control protocols that involve complex mechanisms. We have observed that reservation-based proactive congestion control protocols share strong similarities with RDMA transport protocols, allowing them to integrate seamlessly and combine the functionalities of the transport layer and network layer. In this article, we present COER, an RDMA NIC architecture that leverages the functional components of RDMA to perform reservations and completes the scheduling of congestion control during the scheduling process of the RDMA protocol. COER facilitates the streamlined development of offload strategies for congestion control techniques —specifically, proactive congestion control —on RDMA NICs. We use COER to design offloading schemes for 11 congestion control protocols, which we implement and evaluate using a network emulator with a cycle-accurate RDMA NIC model that can load Message Passing Interface (MPI) programs. The evaluation results demonstrate that the architecture of COER does not compromise the original characteristics of the congestion control protocols. Compared with a layered protocol stack approach, COER enables the performance of RDMA networks to reach new heights.
Ke Wu 0003, Dezun Dong, Weixia Xu 0001
ACM Trans. Archit. Code Optim.3
2024 Hierarchical Mapping of Large-Scale Spiking Convolutional Neural Networks Onto Resource-Constrained Neuromorphic Processor
abstract
Neuromorphic processors have been designed as non-von Neumann systems for energy-efficient spiking neural network (SNN) execution. Spiking convolutional neural networks (SCNNs), combining the advantage of convolutional neural network (CNN) and SNN, have been widely applied to vision tasks. However, as the scale of SCNNs increases, executing large-scale SCNNs on resource-constrained neuromorphic processor faces many challenges, including massive synapse pruning caused by resource competition, execution performance degradation, etc. Addressing these problems, we propose an efficient approach to map large-scale SCNNs onto resource-constrined neuromorphic processor. The approach consists of three steps: splitting, partitioning, and mapping. We explore three acyclic splitting strategies to divide large-scale SCNNs into subnetworks without cyclic dependency. Axon sharing is the guiding principle to partition subnetworks into multiple clusters. To obtain an optimal cluster-to-core mapping scheme, we use Non-dominated Sorting Genetic Algorithm to collaboratively optimize two metrics. We evaluate our approach with eight realistic SCNN applications. The results show that compared with existing state-of-the-art methods, our approach significantly reduces the synapse pruning and accuracy loss, and increases the execution performance.
Xun Xiao, Yao Wang 0002, Junbo Tie, Lei Wang 0011, Weixia Xu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.9
2023 M-LSM: An Improved Multi-Liquid State Machine for Event-Based Vision Recognition
Lei Wang 0011, Shasha Guo 0001, Lianhua Qu, Shuo Tian, Weixia Xu 0001
J. Comput. Sci. Technol.5
2023 SSD-SGD: Communication Sparsification for Distributed Deep Learning Training
abstract
Intensive communication and synchronization cost for gradients and parameters is the well-known bottleneck of distributed deep learning training. Based on the observations that Synchronous SGD (SSGD) obtains good convergence accuracy while asynchronous SGD (ASGD) delivers a faster raw training speed, we propose Several Steps Delay SGD (SSD-SGD) to combine their merits, aiming at tackling the communication bottleneck via communication sparsification. SSD-SGD explores both global synchronous updates in the parameter servers and asynchronous local updates in the workers in each periodic iteration. The periodic and flexible synchronization makes SSD-SGD achieve good convergence accuracy and fast training speed. To the best of our knowledge, we strike the new balance between synchronization quality and communication sparsification, and improve the tradeoff between accuracy and training speed. Specifically, the core components of SSD-SGD include proper warm-up stage, steps delay stage, and the novel algorithm of global gradient for local update (GLU). GLU is critical for local update operations by using global gradient information to effectively compensate for the delayed local weights. Furthermore, we implement SSD-SGD on MXNet framework and comprehensively evaluate its performance with CIFAR-10 and ImageNet datasets. Experimental results show that SSD-SGD can accelerate distributed training speed under different experimental configurations, by up to 110% (or 2.1× of the original speed), while achieving good convergence accuracy.
Yemao Xu, Dezun Dong, Dongsheng Wang 0004, Enda Yu, Weixia Xu 0001, Xiangke Liao
ACM Trans. Archit. Code Optim.6
2023 Back to Homogeneous Computing: A Tightly-Coupled Neuromorphic Processor With Neuromorphic ISA
abstract
In recent years, neuromorphic processors are widely used in many scenarios, showing extreme energy efficiency over traditional architectures. However, almost all existing neuromorphic hardware are following the heterogeneous computing methodology without Instruction Set Architecture (ISA), leading to inflexibility in programming. In this paper, we first propose a RISC-V Neuromorphic Extension (RVNE) to enable fine-grained and flexible homogeneous programming for neuromorphic algorithms while utilizing SNN sparsity from different levels of granularity and computing flows. Based on RVNE, we next implement a neuromorphic micro-architecture that is tightly coupled to the CPU pipeline to accelerate neuromorphic computing. To demonstrate the proposed homogeneous neuromorphic architecture, we implement a prototype processor called NeuroRVcore based on RISC-V ISA and an open-source RISC-V core. The evaluation results show that RVNE achieves a 2.8 × −4.3 × reduction in code density compared with the general-purpose ISAs. Compared with the state-of-the-art neuromorphic processor, the proposed homogeneous computing reduces energy consumption by 3.4%−22.5% while enabling fine-grained and flexible homogeneous programming.
Lei Wang 0011, Yao Wang 0002, Junbo Tie, Feng Wang 0050, LingHui Peng, Xun Xiao, Gan Zhou, Xuhu Yu, Xia Zhao 0004, Yuhua Tang, Weixia Xu 0001
IEEE Trans. Parallel Distributed Syst.17
2022 Unicorn: a multicore neuromorphic processor with flexible fan-in and unconstrained fan-out for neurons
abstract
Neuromorphic processor is popular due to its high energy efficiency for spatio-temporal applications. However, when running the spiking neural network (SNN) topologies with the ever-growing scale, existing neuromorphic architectures face challenges due to their restrictions on neuron fan-in and fan-out. This paper proposes Unicorn, a multicore neuromorphic processor with a spike train sliding multicasting mechanism (STSM) and neuron merging mechanism (NMM) to support unconstrained fan-out and flexible fan-in of neurons. Unicorn supports 36K neurons and 45M synapses and thus supports a variety of neuromorphic applications. The peak performance and energy efficiency of Unicorn reach 36TSOPS and 424GSOPS/W respectively. Experimental results show that Unicorn can achieve 2×-5.5× energy reduction over the state-of-the-art neuromorphic processor when running an SNN with a relatively large fan-out and fan-in.
Lei Wang 0011, Yao Wang 0002, LingHui Peng, Xun Xiao, Weixia Xu 0001
DAC8
2022 Revisiting network congestion avoidance through adaptive packet-chaining reservation
Ke Wu 0003, Dezun Dong, Cunlu Li, Weixia Xu 0001
Comput. Networks4
2022 Hardware-aware liquid state machine generation for 2D/3D Network-on-Chip platforms
Ziyang Kang, Lei Wang 0011, Lianhua Qu, Weixia Xu 0001
J. Syst. Archit.6
2022 LSMCore: A 69k-Synapse/mm2 Single-Core Digital Neuromorphic Processor for Liquid State Machine
abstract
Neuromorphic processors have gained momentum recently due to their high energy efficiency in artificial intelligence applications compared to DNN accelerators. Most neuromorphic processors are executing SNNs (Spiking Neural Networks). Liquid State Machine (LSM), as the spiking version of reservoir computing, shows advantages and great potential in image classification, speech recognition, language translation, etc.. Comparing with other SNN models, LSM has the characteristics of easy to train and low resource utilization, which is suitable for low-power and resource-constrained edge computing scenarios. In this paper, we propose a novel design of a neuromorphic processor, LSMCore, aiming at LSM acceleration. LSMCore supports both training and inference of LSM. It consists of 256 input neurons, 1024 liquid neurons, and 1.31M synapses. Besides, multiple optimization techniques, including weight quantization for reducing storage, zero-skipping for decreasing dynamic sparsity, and mini-batch training are adopted in this processor. The experimental results show that the frequency of LSMCore achieves 400 MHz, the power is 4.9W and the area is 18.49 mm2with a 40nm library. Comparing with the baseline, LSMCore achieves up to$80.7\times $($49.6\times $),$91.3\times $($56.3\times $), and$83.1\times $($56.8\times $) speedup on MNIST, N-MNIST, and Free Spoken Digital Dataset (FSDD) respectively for training (inference), while the accuracy of LSMCore on these three datasets are 96.8%, 97.6%, and 90% respectively.
Lei Wang 0011, Shasha Guo 0001, Lianhua Qu, Ziyang Kang, Weixia Xu 0001
IEEE Trans. Circuits Syst. I Regul. Pap.7
2021 A neural architecture search based framework for liquid state machine design
Shuo Tian, Lianhua Qu, Lei Wang 0011, Weixia Xu 0001
Neurocomputing6
2021 HashHeat: A hashing-based spatiotemporal filter for dynamic vision sensor
Shasha Guo 0001, Ziyang Kang, Lei Wang 0011, Limeng Zhang, Weixia Xu 0001
Integr.7
2021 A multi-objective LSM/NoC architecture co-design framework
Shuo Tian, Ziyang Kang, Lianhua Qu, Lei Wang 0011, Weixia Xu 0001
J. Syst. Archit.7
2021 Quingo: A Programming Framework for Heterogeneous Quantum-Classical Computing with NISQ Features
abstract
The increasing control complexity of Noisy Intermediate-Scale Quantum (NISQ) systems underlines the necessity of integrating quantum hardware with quantum software. While mapping heterogeneous quantum-classical computing (HQCC) algorithms to NISQ hardware for execution, we observed a few dissatisfactions in quantum programming languages (QPLs), including difficult mapping to hardware, limited expressiveness, and counter-intuitive code. In addition, noisy qubits require repeatedly performed quantum experiments, which explicitly operate low-level configurations, such as pulses and timing of operations. This requirement is beyond the scope or capability of most existing QPLs. We summarize three execution models to depict the quantum-classical interaction of existing QPLs. Based on the refined HQCC model, we propose the Quingo framework to integrate and manage quantum-classical software and hardware to provide the programmability over HQCC applications and map them to NISQ hardware. We propose a six-phase quantum program life-cycle model matching the refined HQCC model, which is implemented by a runtime system. We also propose the Quingo programming language, an external domain-specific language highlighting timer-based timing control and opaque operation definition, which can be used to describe quantum experiments. We believe the Quingo framework could contribute to the clarification of key techniques in the design of future HQCC systems.
Xiang Fu 0003, Hanru Jiang, Fucheng Cheng, Yihang Yang, Chunchao Hu, Anqi Huang 0003, Guangyao Huang 0001, Xiaogang Qiang, Mingtang Deng, Ping Xu 0004, Weixia Xu 0001, Wanwei Liu, Yu Zhang 0086, Yuxin Deng 0001, Junjie Wu 0003, Yuan Feng 0001
ACM Trans. Quantum Comput.18
2020 HashHeat: An O(C) Complexity Hashing-based Filter for Dynamic Vision Sensor
abstract
Neuromorphic event-based dynamic vision sensors (DVS) have much faster sampling rates and a higher dynamic range than frame-based imagers. However, they are sensitive to background activity (BA) events which are unwanted. We propose HashHeat, a hashing-based BA filter with O(C) complexity. It is the first spatiotemporal filter that doesn't scale with the DVS output size N and doesn't store the 32-bits timestamps. HashHeat consumes 100x less memory and increases the signal to noise ratio by 15x compared to previous designs.
Shasha Guo 0001, Ziyang Kang, Lei Wang 0011, Weixia Xu 0001
ASP-DAC5
2020 Application-specific network-on-chip design space exploration framework for neuromorphic processor
abstract
Neuromorphic processors can support the design of various Spiking Neural Networks (SNN) to deal with different tasks, such as recognition and tracking. Neuromorphic processors use Network-on-Chip (NoC) to support communication between neurons in SNN. The different SNN has different communication traffic patterns. It will pose the different challenges of the NoC designing. A reasonable NoC architecture can improve the overall performance such as lower latency of the processor. Hence, it is critical to implement the exploration of NoC architecture design for neuromorphic processors.
Ziyang Kang, Lei Wang 0011, Lianhua Qu, Weixia Xu 0001
CF8
2020 SNEAP: A Fast and Efficient Toolchain for Mapping Large-Scale Spiking Neural Network onto NoC-based Neuromorphic Platform
abstract
Spiking neural network (SNN), as the third generation of artificial neural networks, has been widely adopted in vision and audio tasks. Nowadays, many neuromorphic platforms support SNN simulation and adopt Network-on-Chips (NoC) architecture for multi-cores interconnection. However, a large volume and run-time communication on the interconnection has a significant effect on performance of the platform. In this paper, we propose a toolchain called SNEAP (Spiking NEural network mAPping toolchain) for mapping SNNs to neuromorphic platforms with multi-cores, which aims to reduce the energy and latency brought by spike communication on the interconnection.
Shasha Guo 0001, Limeng Zhang, Ziyang Kang, Lei Wang 0011, Weixia Xu 0001
ACM Great Lakes Symposium on VLSI8
2020 CompressedCache: Enabling Storage Compression on Neuromorphic Processor for Liquid State Machine
Lianhua Qu, Ziyang Kang, Lei Wang 0011, Weixia Xu 0001
NPC7
2020 SIES: A Novel Implementation of Spiking Convolutional Neural Network Inference Engine on Field-Programmable Gate Array
Shuquan Wang, Lei Wang 0011, Yu Deng 0001, Shasha Guo 0001, Ziyang Kang, Yu-Feng Guo, Weixia Xu 0001
J. Comput. Sci. Technol.8
2020 ASIE: An Asynchronous SNN Inference Engine for AER Events Processing
abstract
Neuromorphic computing based on spiking neural network (SNN) shows good energy-efficiency. However, it is inefficient for SNN to perform the convolution based on frame. It may contain a lot of redundant information in the frame. The output of Dynamic Vision Sensors (DVS) is a stream event based on Address Event Representation (AER). The asynchronous nature of AER events makes the event-based convolution reflect the characteristics of SNN low energy consumption. This article presents an SNN hardware inference engine based on an asynchronous Processing Element (PE) array with AER events as input. The engine uses a convolution algorithm based on AER events. This design also uses distributed storage in the PE array to store the state of neurons to reduce the cost of memory access. The experimental results show that the design can achieve a recognition accuracy of 98.0% for the MNIST AER dataset. The design can perform the reference process more efficiently in the case where the accuracy of the loss is negligible. During the filling and draining processes of the systolic array, the number of active PE units in our PE array is reduced and, thus, the average power consumption per PE unit is drastically decreased.
Ziyang Kang, Lei Wang 0011, Shasha Guo 0001, Yu Deng 0001, Weixia Xu 0001
ACM J. Emerg. Technol. Comput. Syst.7
2020 OD-SGD: One-Step Delay Stochastic Gradient Descent for Distributed Training
abstract
The training of modern deep learning neural network calls for large amounts of computation, which is often provided by GPUs or other specific accelerators. To scale out to achieve faster training speed, two update algorithms are mainly applied in the distributed training process, i.e., the Synchronous SGD algorithm (SSGD) and Asynchronous SGD algorithm (ASGD). SSGD obtains good convergence point while the training speed is slowed down by the synchronous barrier. ASGD has faster training speed but the convergence point is lower when compared to SSGD. To sufficiently utilize the advantages of SSGD and ASGD, we propose a novel technology named One-step Delay SGD (OD-SGD) to combine their strengths in the training process. Therefore, we can achieve similar convergence point and training speed as SSGD and ASGD separately. To the best of our knowledge, we make the first attempt to combine the features of SSGD and ASGD to improve distributed training performance. Each iteration of OD-SGD contains a global update in the parameter server node and local updates in the worker nodes, the local update is introduced to update and compensate the delayed local weights. We evaluate our proposed algorithm on MNIST, CIFAR-10, and ImageNet datasets. Experimental results show that OD-SGD can obtain similar or even slightly better accuracy than SSGD, while its training speed is much faster, which even exceeds the training speed of ASGD.
Yemao Xu, Dezun Dong, Weixia Xu 0001, Xiangke Liao
ACM Trans. Archit. Code Optim.4
2019 PRTSM: Hardware Data Arrangement Mechanisms for Convolutional Layer Computation on the Systolic Array
Shuquan Wang, Lei Wang 0011, Shuo Tian, Shasha Guo 0001, Ziyang Kang, Shuzheng Zhang, Weixia Xu 0001
NPC8
2019 SketchDLC: A Sketch on Distributed Deep Learning Communication via Trace Capturing
abstract
With the fast development of deep learning (DL), the communication is increasingly a bottleneck for distributed workloads, and a series of optimization works have been done to scale out successfully. Nevertheless, the network behavior has not been investigated much yet. We intend to analyze the network behavior and then carry out some research through network simulation. Under this circumstance, an accurate communication measurement is necessary, as it is an effective way to study the network behavior and the basis for accurate simulation. Therefore, we propose to capture the deep learning communication (DLC) trace to achieve the measurement. To the best of our knowledge, we make the first attempt to capture the communication trace for DL training. In this article, we first provide detailed analyses about the communication mechanism of MXNet, which is a representative framework for distributed DL. Secondly, we define the DLC trace format to describe and record the communication behaviors. Third, we present the implementation of method for trace capturing. Finally, we make some statistics and analyses about the distributed DL training, including communication pattern, overlap ratio between computation and communication, computation overhead, synchronization overhead, update overhead, and so forth. Both the statistics and analyses are based on the trace files captured in a cluster with six machines. On the one hand, our trace files provide a sketch on the DLC, which contributes to understanding the communication details. On the other hand, the captured trace files can be used for figuring out various overheads, as they record the communication behaviors of each node.
Yemao Xu, Dezun Dong, Weixia Xu 0001, Xiangke Liao
ACM Trans. Archit. Code Optim.3
2016 Graphein: A Novel Optical High-Radix Switch Architecture for 3D Integration
Jie Jian, Liquan Xiao, Weixia Xu 0001
ICA3PP4
2013 Scalable NIC Architecture to Support Offloading of Large Scale MPI Barrier
Weixia Xu 0001, Zhengbin Pang, Pingjing Lu
APPT2
2012 Distributed Coverage in Wireless Ad Hoc and Sensor Networks by Topological Graph Approaches
abstract
Coverage problem is a fundamental issue in wireless ad hoc and sensor networks. Previous techniques for coverage scheduling often require accurate location information or range measurements, which cannot be easily obtained in resource-limited ad hoc and sensor networks. Recently, a method based on algebraic topology is proposed to achieve coverage verification using only connectivity information. The topological method sheds some light on the issue of location-free coverage. Unfortunately, the needs of centralized computation and rigorous restriction on sensing and communication ranges greatly limit the applicability in practical large-scale distributed sensor networks. In this work, we make the first attempt toward establishing a graph theoretical framework for connectivity-based coverage with configurable coverage granularity. We propose a novel coverage criterion and scheduling method based on cycle partition. Our method is able to construct a sparse coverage set in a distributed manner, using purely connectivity information. Compared with existing methods, our design has a particular advantage, which permits us to configure or adjust the quality of coverage by adequately exploiting diverse sensing ranges and specific requirements of different applications. We formally prove the correctness and evaluate the effectiveness of our approach through extensive simulations and comparisons with the state-of-the-art approaches.
Dezun Dong, Xiangke Liao, Kebin Liu 0001, Yunhao Liu 0001, Weixia Xu 0001
IEEE Trans. Computers5
2011 A novel shared-buffer router for network-on-chip based on Hierarchical Bit-line Buffer
abstract
Buffer resources are key components of the on-chip router, shared-buffer structures are proposed to improve performance and reduce power consumption. This paper presents a novel on-chip network router with a shared-buffer based on Hierarchical Bit-line Buffer (HiBB). HiBB can be configured flexibly according to traffics and its inherent characteristic of low power is also noticeable. Moreover, we propose two schemes to further optimize the router. First, a congestion-aware output-port allocation scheme is used to assign higher priority to packets heading to light-loaded directions, and the congestion situation of the total network will be addressed. Second, an efficient run-time Virtual Channel (VC) regulation scheme is proposed to configure the shared buffer, so that VCs are allocated according to the loads of network. Experimental results show that the proposed HiBB router with about 6.9% area savings outperforms the generic router under different traffic patterns. The power consumption of the HiBB router can also be reduced up to about 70% of the generic router under light traffics, while it may exceed that of the generic one up to about 3.7–5.7% under heavy traffics for the increased flit transmissions.
Weixia Xu 0001, Hongguang Ren, Qiang Dou, Zhiying Wang 0003, Li Shen 0007, Cong Liu 0009
ICCD2
2011 Optimizing Linpack Benchmark on GPU-Accelerated Petascale Supercomputer
Feng Wang 0050, Canqun Yang, Yunfei Du 0001, Juan Chen 0001, Huizhan Yi, Weixia Xu 0001
J. Comput. Sci. Technol.6
2010 A Novel Chaining Approach for Direct Control Transfer Instructions
abstract
Software-based code cache systems are the key element in the dynamic translation system or optimization system to store the translated or optimized code for reuse. Translated code is organized in terms of code blocks in the code cache which transfers execution to the next code block through a control transfer instruction. As the target address of the control transfer instruction is in the form of its source program counter, the code cache system has to check the address mapping table for the translated program counter of the target address before entering the required code block. This will cause the performance degradation. As the target address of the direct control transfer instruction is fixed during the execution of a program, its source target address can be replaced with the translated target address. A direct control transfer chaining approach which occupies specific software assists is proposed in this paper. Evaluation of DCTC is conducted on a code cache simulator. The experiment results show the dramatic performance improvement brought by DCTC.
Weixia Xu 0001, Wei Chen 0009, Qiang Dou
ICPADS1
2010 TH-1: China's first petaflop supercomputer
Xuejun Yang, Xiangke Liao, Weixia Xu 0001, Junqiang Song, Qingfeng Hu, Jinshu Su, Liquan Xiao, Kai Lu 0001, Qiang Dou, Juping Jiang, Canqun Yang
Frontiers Comput. Sci. China3