EDBT 2026 Demo / reviewers in the wild / expert
Weigong Zhang
dblp:90/7721
· DBLP profile ↗
31ranked-venue papers
0as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 22 · 7 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Deep Reinforcement Learning and Fuzzy Logic-Based Adaptive Shift Control Method of Driving Robot SystemabstractIn order to improve the fuel economy of the test vehicle manipulated by the driving robot, an adaptive shift control method based on deep reinforcement learning (DRL) and fuzzy logic of driving robot system is proposed. The driving robot system (DRS) is installed in the cockpit of a vehicle, which does not require any modifications to the vehicle. It includes a physical robot that mechanically operates vehicle and a controller for an automated driving. Firstly, the dynamic model of driving robot system is established. Then, a deep reinforcement learning model of gearshift strategy is established. Aiming at fuel economy, the throttle opening, vehicle speed and vehicle acceleration are used as the observation parameters of DRL model for gearshift strategy. Then, a nonlinear disturbance observer is established to observe the parameters of the DRL model. Furthermore, an adaptive shift controller consisting of gearshift strategy and fuzzy logic is proposed considering the Gaussian pulse secondary impact load model during the shifting process. The continuous control of DRS is realized by the fuzzy logic controller. Finally, the stability proof of the proposed shift control method is conducted. The proposed shift control method is verified through testing in different driving cycle tests. Test results demonstrate the effectiveness of the proposed shift control method. Gang Chen 0033, Longqi Tang, Liangmo Wang, Weigong Zhang |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | Rethink Reynolds' rules: flock-inspired network for vehicle trajectory prediction
Qifan Xue, Shengyi Li, Xuanpeng Li, Weigong Zhang |
J. Supercomput. | 6 |
| 2024 | Multi-temporal dependency handling in video smoke recognition: A holistic approach spanning spatial, short-term, and long-term perspectives
Qifan Xue, Yichao Cao, Xuanpeng Li, Weigong Zhang |
Expert Syst. Appl. | 5 |
| 2024 | CEDR: Contrastive Embedding Distribution Refinement for 3D point cloud representation
Yichao Cao, Qifan Xue, Shuai Jin, Xuanpeng Li, Weigong Zhang |
Signal Process. Image Commun. | 6 |
| 2024 | Enhancing Neural Network Reliability: Insights From Hardware/Software Collaboration With Neuron Vulnerability QuantizationabstractEnsuring the reliability of deep neural networks (DNNs) is paramount in safety-critical applications. Although introducing supplementary fault-tolerant mechanisms can augment the reliability of DNNs, an efficiency tradeoff may be introduced. This study reveals the inherent fault tolerance of neural networks, where individual neurons exhibit varying degrees of fault tolerance, by thoroughly exploring the structural attributes of DNNs. We thereby develop a hardware/software collaborative method that guarantees the reliability of DNNs while minimizing performance degradation. We introduce the neuron vulnerability factor (NVF) to quantify the susceptibility to soft errors. We propose two efficient methods that leverage the NVF to minimize the negative effects of soft errors on neurons. First, we present a novel computational scheduling scheme. By prioritizing error-prone neurons, the expedited completion of their computations is facilitated to mitigate the risk of neural computing errors that arise from soft errors without sacrificing efficiency. Second, we propose the NVF-guided heterogeneous memory system. We employ variable-strength error-correcting codes and tailor their error-correction mechanisms to the vulnerability profile of specific neurons to ensure a highly targeted approach for error mitigation. Our experimental results demonstrate that the proposed scheme enhances the neural network accuracy by 18% on average, while significantly reducing the fault-tolerance overhead. Jing Wang 0055, Jinbin Zhu, Xin Fu 0001, Di Zang, Keyao Li, Weigong Zhang |
IEEE Trans. Computers | 6 |
| 2024 | CiPN-TP: a channel-independent pretrained network via tokenized patching for trajectory prediction
Qifan Xue, Shengyi Li, Xuanpeng Li, Weigong Zhang |
J. Supercomput. | 6 |
| 2023 | A Bit Level Acceleration of Mixed Precision Neural NetworkabstractWith the growth of the convolutional neural network (CNN) parameters, the hardware resources become limited when deploying CNN models. Single bit-width quantization may lead to degradation of accuracy, while mixed-precision quantization models maintain higher accuracy. However, current mixed-precision quantization accelerators only consider the algorithm level quantization and fixed-bit-width processing elements (PEs) without fully utilizing the resources. Therefore, we propose a neural network acceleration architecture based on bit-level computational units to improve resource utilization and throughput of mixed-precision quantization accelerators. The 2-bit and 3-bit low-bit computational units are designed to implement the high-bit quantization. We also propose spatio-temporal fusion to satisfy the unique bit widths in each layer with mixed precision quantization. In particular, we use the 2-bit and 3-bit computational units to achieve dynamic layer-level quantization. We also discuss various combinations of different bit widths, which can be dynamically implemented according to the requirements of accuracy, execution time, etc. The proposed acceleration architecture is implemented in Verilog, verified using three network models: ResNet18, VGG7 and LeNet-5, and tested against three accelerators: Eyeriss, Stripes and Bit Fusion. The experimental results show that our accelerators provide more accuracy and the grouping operation reduces the area overhead. It provides a 3.08 to 3.54 times acceleration ratio and 2.4 to 2.57 times reduction in energy consumption on the three networks. Dehui Qiu, Jing Wang 0055, Weigong Zhang, Lan Gao 0004 |
ICPADS | 4 |
| 2023 | Accelerating Look-Up Table based Matrix Multiplication on GPUsabstractMultiplying matrices is among the most fundamental and compute-intensive operations in machine learning. Approximated Matrix Multiplication (AMM) based on table look-ups can significantly reduce the pressure on computing units and memory bandwidth, and has great potential in large-scale machine learning applications. In this work, we speed up table look-ups on GPUs to improve the performance of matrix multiplication. To avoid random memory accesses in table look-ups, we propose a novel warp-wide data sharing execution model. With this execution model, we develop a GPU AMM library to speed up MADDNESS (the state-of-the-art AMM), named GPU-MADDNESS. The experimental results show that GPU-MADDNESS improves the performance by 103X on average, and outperforms the tiling implementation by up to 42%. Lan Gao 0004, Weigong Zhang, Jing Wang 0055, Dehui Qiu |
ICPADS | 3 |
| 2023 | Deep Deterministic Policy Gradient and Active Disturbance Rejection Controller based coordinated control for gearshift manipulator of driving robot
Gang Chen 0033, Liangmo Wang, Weigong Zhang |
Eng. Appl. Artif. Intell. | 4 |
| 2022 | Adaptive Contention Management for Fine-Grained Synchronization on Commodity GPUsabstractAs more emerging applications are moving to GPUs, fine-grained synchronization has become imperative. However, their performance can be severely impaired in case of frequent synchronization failures caused by high data contention. Differently from CPUs, GPUs own thousands of hardware threads and adopt single instruction multiple threads paradigm, making it impractical to deploy the CPU contention management mechanisms directly on GPUs. In this article, we design a Software Warp Controlling Framework (SWCF), which employs producer-consumer execution model and leverages GPU hardware barriers to dynamically control the execution of warps at runtime. On the basis of SWCF, we propose a contention management strategy to decrease frequent synchronization failures while avoiding the over-reducing of parallelism. We evaluate SWCF and the proposed strategy on commodity GPUs using a set of applications with fine-grained synchronization. The results show that on V100 GPU our contention management achieves a 4.7X speedup and outperforms the conventional GPU software backoff solution by 42% on average. Lan Gao 0004, Jing Wang 0055, Weigong Zhang |
ACM Trans. Archit. Code Optim. | 3 |
| 2022 | eRDAC: Efficient and Reliable Remote Direct Access and Control for Embedded SystemsabstractEmerging embedded systems, such as autonomous vehicles, demand highly efficient remote data transfer, whereas existing networking hardware and protocols cause high communication latency and CPU consumption. In this article, we propose embedded RDAC (eRDAC), an efficient and reliable remote direct access and control solution for embedded systems. The proposed remote access controller in eRDAC has a two-layer protocol offload engine that employs the command/response protocol on UDP to ensure the data reliability and security, and a multichannel DMA controller with configurable priority to improve the efficiency. Besides, a reusable hardware Ethernet MAC is implemented to support not only remote access commands but also standard Ethernet communication. We implement eRDAC on FPGA and the corresponding software in the Linux system. Experimental results show that eRDAC can reduce the latency of remote I/O reading/writing by 74.3%/74.9% ($3.76\times /3.98\times $performance improvement) and reduce the latency of remote memory reading and writing with 1024B by 54.2% compared to the socket-based communication. Meanwhile, eRDAC can cut off the consumption of the remote processor and achieve 0.250mJ/Mb energy consumption with only 25-mW power. Xianzhang Chen, Duo Liu 0002, Weigong Zhang, Jiapin Wang, Rongwei Zheng, Yujuan Tan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | Enabling Highly Efficient Capsule Networks Processing Through A PIM-Based Architecture DesignabstractIn recent years, the CNNs have achieved great successes in the image processing tasks, e.g., image recognition and object detection. Unfortunately, traditional CNN's classification is found to be easily misled by increasingly complex image features due to the usage of pooling operations, hence unable to preserve accurate position and pose information of the objects. To address this challenge, a novel neural network structure called Capsule Network has been proposed, which introduces equivariance through capsules to significantly enhance the learning ability for image segmentation and object detection. Due to its requirement of performing a high volume of matrix operations, CapsNets have been generally accelerated on modern GPU platforms that provide highly optimized software library for common deep learning tasks. However, based on our performance characterization on modern GPUs, CapsNets exhibit low efficiency due to the special program and execution features of their routing procedure, including massive unshareable intermediate variables and intensive synchronizations, which are very difficult to optimize at software level. To address these challenges, we propose a hybrid computing architecture design named PIM-CapsNet. It preserves GPU's on-chip computing capability for accelerating CNN types of layers in CapsNet, while pipelining with an off-chip in-memory acceleration solution that effectively tackles routing procedure's inefficiency by leveraging the processing-in-memory capability of today's 3D stacked memory. Using routing procedure's inherent parallellization feature, our design enables hierarchical improvements on CapsNet inference efficiency through minimizing data movement and maximizing parallel processing in memory. Evaluation results demonstrate that our proposed design can achieve substantial improvement on both performance and energy savings for CapsNet inference, with almost zero accuracy loss. The results also suggest good performance scalability in optimizing the routing procedure with increasing network size. Xingyao Zhang 0002, Shuaiwen Song, Chenhao Xie 0001, Jing Wang 0055, Weigong Zhang, Xin Fu 0001 |
HPCA | 5 |
| 2020 | Multi-dimensional optimization for approximate near-threshold computingabstractThe demise of Dennard’s scaling has created both power and utilization wall challenges for computer systems. As transistors operating in the near-threshold region are able to obtain flexible trade-offs between power and performance, it is regarded as an alternative solution to the scaling challenge. A reduction in supply voltage will nevertheless generate significant reliability challenges, while maintaining an error-free system that generates high costs in both performance and energy consumption. The main purpose of research on computer architecture has therefore shifted from performance improvement to complex multi-objective optimization. In this paper, we propose a three-dimensional optimization approach which can effectively identify the best system configuration to establish a balance among performance, energy, and reliability. We use a dynamic programming algorithm to determine the proper voltage and approximate level based on three predictors: system performance, energy consumption, and output quality. We propose an output quality predictor which uses a hardware/software co-design fault injection platform to evaluate the impact of the error on output quality under near-threshold computing (NTC). Evaluation results demonstrate that our approach can lead to a 28% improvement in output quality with a 10% drop in overall energy efficiency; this translates to an approximately 20% average improvement in accuracy, power, and performance. Jing Wang 0055, Wei-wei Liang, Yuehua Niu, Lan Gao 0004, Weigong Zhang |
Frontiers Inf. Technol. Electron. Eng. | 5 |
| 2020 | Enabling Energy-Efficient and Reliable Neural Network via Neuron-Level Voltage ScalingabstractWith the platforms of running deep neural networks (DNNs) move from large-scale data centers to handheld devices, power emerge as one of the most significant obstacles. Voltage scaling is a promising technique that enables power saving. Nevertheless, it raises reliability and performance concerns that may undesirably deteriorate NNs accuracy and performance. Consequently, an energy-efficient and reliable scheme is required for NNs to balance the above three aspects with satisfied user experience. To this end, we propose a neuron-level voltage scaling framework called NN-APP to model the impact of supply voltages on NNs from output accuracy (A), power (P), and performance (P) perspectives. We analyze the error propagation characteristics in NNs at both inter- and intra-network layers to precisely model the impact of voltage scaling on the final output accuracy at neuron-level. Furthermore, we combine a voltage clustering method and the multi-objective optimization to identify the optimal voltage islands and apply the same voltage to neurons with similar fault tolerance capability. We perform three case studies to demonstrate the efficacy of the proposed techniques. Jing Wang 0055, Xin Fu 0001, Lan Gao 0004, Weigong Zhang |
IEEE Trans. Computers | 6 |
| 2019 | Reliability Enhancement of Neural Networks via Neuron-Level Vulnerability Quantization
Keyao Li, Jing Wang 0055, Xin Fu 0001, Xiufeng Sui, Weigong Zhang |
ICA3PP (2) | 5 |
| 2019 | Neuron Fault Tolerance Capability Based Computation Reuse in DNNs
Pengnian Qi, Jing Wang 0055, Weigong Zhang |
ICA3PP (2) | 4 |
| 2019 | Enabling Energy-Efficient and Reliable Neural Network via Neuron-Level Voltage ScalingabstractAs the application scope of deep neural networks (DNNs) moves from large-scale data centers to small-scale mobile devices, power wall has become one of the most important obstacles. Voltage scaling is a typical technique enables power saving, but it causes reliability and performance challenges. Therefore, an energy-efficient and reliable scheme for NNs is required to balance above three aspects according to users' requirements for excellent user experience. In this paper, we innovatively propose neuron-level voltage scaling framework called NN-APP to model the impact of supply voltages on NNs from output accuracy (A), power (P), and performance (P) perspectives. We analyze the error propagation in NNs and precisely model the impact of voltage scaling on the final output accuracy at neuron-level. Multi-objective optimization and clustering method are combined to find the optimal voltage islands. Finally, we conduct experiment to demonstrate the efficacy of the proposed technique. Jing Wang 0055, Xin Fu 0001, Xingyao Zhang 0002, Lan Gao 0004, Weigong Zhang, Tao Li 0006 |
ICPADS | 7 |
| 2019 | DR Refresh: Releasing DRAM Potential by Enabling Read Accesses Under RefreshabstractEmerging data analytic workloads such as graph processing, neural network and edge data preprocesing desire efficient memory read operations. Unfortunately, due to the necessity of dynamic refresh, modern DRAM systems have to stall access during refresh cycles. As DRAM device density continues to grow, refresh operations can be a crucial throughput bottleneck. To fully unleash memory access performance, we revisit conventional refresh mechanism and DRAM architecture. We propose DR refresh, a specific refresh mechanism that enable read and refresh operations to be done simultaneously. We devise DR DRAM, a specific memory hardware system that can efficiently deploy DR refresh. Unlike traditional refresh, DR explores device refresh that only refreshes a designated device at a time. Meanwhile, DR increases read efficiency by recovering the inaccessible data that resides on a device under refreshing. We also propose Hybrid Refresh Main Memory (HRMM) which can designate refresh schemes (DR or traditional refresh) in a specific memory space. We expect that our design can benefit many real-life tasks such as SPEC CPU2006, CNN, LLT and PageRank. Yuhai Cao, Chao Li 0009, Jing Wang 0055, Weigong Zhang, Quan Chen 0002, Jingwen Leng, Bin Yao 0002, Minyi Guo |
IEEE Trans. Computers | 4 |
| 2018 | In-Situ AI: Towards Autonomous and Incremental Deep Learning for IoT SystemsabstractRecent years have seen an exploration of data volumes from a myriad of IoT devices, such as various sensors and ubiquitous cameras. The deluge of IoT data creates enormous opportunities for us to explore the physical world, especially with the help of deep learning techniques. Traditionally, the Cloud is the option for deploying deep learning based applications. However, the challenges of Cloud-centric IoT systems are increasing due to significant data movement overhead, escalating energy needs, and privacy issues. Rather than constantly moving a tremendous amount of raw data to the Cloud, it would be beneficial to leverage the emerging powerful IoT devices to perform the inference task. Nevertheless, the statically trained model could not efficiently handle the dynamic data in the real in-situ environments, which leads to low accuracy. Moreover, the big raw IoT data challenges the traditional supervised training method in the Cloud. To tackle the above challenges, we propose In-situ AI, the first Autonomous and Incremental computing framework and architecture for deep learning based IoT applications. We equip deep learning based IoT system with autonomous IoT data diagnosis (minimize data movement), and incremental and unsupervised training method (tackle the big raw IoT data generated in ever-changing in-situ environments). To provide efficient architectural support for this new computing paradigm, we first characterize the two In-situ AI tasks (i.e. inference and diagnosis tasks) on two popular IoT devices (i.e. mobile GPU and FPGA) and explore the design space and tradeoffs. Based on the characterization results, we propose two working modes for the In-situ AI tasks, including Single-running and Co-running modes. Moreover, we craft analytical models for these two modes to guide the best configuration selection. We also develop a novel two-level weight shared In-situ AI architecture to efficiently deploy In-situ tasks to IoT node. Compared with traditional IoT systems, our In-situ AI can reduce data movement by 28-71%, which further yields 1.4X-3.3X speedup on model update and contributes to 30-70% energy saving. Mingcong Song, Kan Zhong, Jiaqi Zhang 0002, Yang Hu 0001, Duo Liu 0002, Weigong Zhang, Jing Wang 0055, Tao Li 0006 |
HPCA | 6 |
| 2018 | DR DRAM: Accelerating Memory-Read-Intensive ApplicationsabstractToday, many data analytic workloads such as graph processing and neural network desire efficient memory read operation. The need for preprocessing various raw data also demands enhanced memory read bandwidth. Unfortunately, due to the necessity of dynamic refresh, modern DRAM system has to stall memory access during each refresh cycle. As DRAM device density continues to grow, the refresh time also needs to extend to cover more memory rows. Consequently, DRAM refresh operation can be a crucial throughput bottleneck for memory read intensive (MRI) data processing tasks. To fully unleash the performance of these applications, we revisit conventional DRAM architecture and refresh mechanism. We propose DR DRAM, an application-specific memory design approach that makes a novel tradeoff between read and write performance. Simply put, DR has two layers of meaning: device refresh and data recovery. It aims at eliminating stall by enabling read and refresh operations to be done simultaneously. Unlike traditional schemes, DR explores device refresh that only refreshes a specific device at a time. Meanwhile, DR increases read efficiency by recovering the inaccessible data that resides on a device under refreshing. Our design can be implemented on existing redundant data storage area on DRAM. In this paper we detail DR's architecture and protocol design. We evaluate it on a cycle accurate simulator. Our results show that DR can nearly eliminate refresh overhead for memory read operation and brings up to 12% extra maximum read bandwidth and 50~60% latency improvement on present DRR4 device. Yuhai Cao, Chao Li 0009, Quan Chen 0002, Jingwen Leng, Minyi Guo, Jing Wang 0055, Weigong Zhang |
ICCD | 7 |
| 2017 | Processing-in-Memory Enabled Graphics Processors for 3D RenderingabstractThe performance of 3D rendering of Graphics Processing Unit that converts 3D vector stream into 2D frame with 3D image effects significantly impacts users gaming experience on modern computer systems. Due to its high texture throughput requirement, main memory bandwidth becomes a critical obstacle for improving the overall rendering performance. 3D-stacked memory systems such as Hybrid Memory Cube provide opportunities to significantly overcome the memory wall by directly connecting logic controllers to DRAM dies. Although recent works have shown promising improvement in performance by utilizing HMC to accelerate special-purpose applications, a critical challenge of how to effectively leverage its high internal bandwidth and computing capability in GPU for 3D rendering remains unresolved. Based on the observation that texel fetches greatly impact off-chip memory traffic, we propose two architectural designs to enable Processing-In-Memory based GPU for efficient 3D rendering. Additionally, we employ camera angles of pixels to control the performance-quality tradeoff of 3D rendering. Extensive evaluation across several real-world games demonstrates that our design can significantly improve the performance of texture filtering and 3D rendering by an average of 3.97X (up to 6.4X) and 43% (up to 65%) respectively, over the baseline GPU. Meanwhile, our design provides considerable memory traffic and energy reduction without sacrificing rendering quality. Chenhao Xie 0001, Shuaiwen Song, Jing Wang 0055, Weigong Zhang, Xin Fu 0001 |
HPCA | 4 |
| 2017 | Expected Completion Time Aware Message Scheduling for UM-BUS Interconnected SystemabstractIn real-time embedded systems, periodic messages need to be transmitted at the expected time because of timing sensitive requirements. In this paper, we take advantage of the characteristics of UM-BUS, a novel serial bus with the capability of multi-lane concurrent transmissions, and investigate the scheduling problem to reduce the deviation to the expected completion time of messages. By configuring different lanes to change the bus utilization, two sets of experiments were implemented to evaluate the effectiveness of the proposed algorithm. The results show that the heuristic algorithm works effectively and can achieve a deviation within 1.52% which is significantly smaller comparing to the existing scheduling algorithms. Jiqin Zhou, Weigong Zhang, Keni Qiu, Ruiying Bai |
ISORC | 2 |
| 2017 | Data re-allocation enabled cache locking for embedded systems
Chun Jason Xue, Keni Qiu, Weigong Zhang, Jing Wang 0055, Yuanchao Xu 0002, Mengying Zhao |
J. Syst. Archit. | 3 |
| 2017 | Creating Affective Autonomous Characters Using Planning in Partially Observable Stochastic DomainsabstractThe ability to reason about and respond to their own emotional states can enhance the believability of Non-Player Characters (NPCs). In this paper, we use a Partially Observable Markov Decision Process (POMDP)-based framework to model emotion over time. A two-level appraisal model, involving quick and reactive vs. slow and deliberate appraisals, is proposed for the creation of affective autonomous characters based on POMDPs, wherein the probability of goal satisfaction is used in an appraisal and reappraisal process for emotion generation. We not only extend Probabilistic Computation Tree Logic (PCTL) for reasoning about the properties of emotional states based on POMDPs but also illustrate how four reactive (primary) emotions and nine deliberate (secondary) emotions can be derived by combining PCTL with the belief-desire theory of emotion. The results of an empirical study suggest that the proposed model can be used to create characters that appear to be more believable and more intelligent. Xiangyang Huang, Shudong Zhang, Weigong Zhang, Jie Liu 0022 |
IEEE Trans. Comput. Intell. AI Games | 4 |
| 2017 | On the Implication of NTC versus Dark Silicon on Emerging Scale-Out Workloads: The Multi-Core Architecture PerspectiveabstractThe end of Dennard's scaling poses computer systems, especially the datacenters, in front of both power and utilization walls. One possible solution to combat the power and utilization walls is dark silicon where transistors are under-utilized in the chip, but this will result in a diminishing performance. Another solution is Near-Threshold Voltage Computing (NTC) which operates transistors in the near-threshold region and provides much more flexible tradeoffs between power and performance. However, prior efforts largely focus on a specific design option based on the legacy desktop applications, therefore, lacking comprehensive analysis of emerging scale-out applications with multiple design options when dark silicon and/or NTC are/is applied. In this paper, we characterize different perspectives including performance, energy efficiency and reliability in the context of NTC/dark silicon cloud processors running emerging scale-out workloads on various architecture designs. We find NTC is generally an effective way to alleviate the power challenge over scale-out applications compared with dark silicon, it can improve performance by 1.6X, energy efficiency by 50 percent and the reliability problem can be relieved by ECC. Meanwhile, we also observe tiled-OoO architecture improves the performance by 20~370 percent and energy efficiency by 40~600 percent over alternative architecture designs, making it a preferable design paradigm for scale-out workloads. We believe that our observations will provide insights for the design of cloud processors under dark silicon and/or NTC. Jing Wang 0055, Xin Fu 0001, Weigong Zhang, Keni Qiu, Tao Li 0006 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2016 | Refresh-aware loop scheduling for high performance low power volatile STT-RAMabstractThe highlighted advantages of low leakage power, high storage density and immunity to electronic magnetic radiation make STT-RAM a promising candidate to build cache, SPM or main memory in embedded systems. However, write operations on STT-RAM have considerably longer latency and higher energy consumption than conventional SRAM. To solve this problem, researchers have proposed to relax STT-RAM's non-volatility and to have it work in a fast and low power mode. Under this volatile mode, refresh operations are needed to guarantee data correctness if their lifespan is larger than the retention time. It is observed that this refresh overhead is significant for data in stencil loops with the characteristic of constant read and write dependencies. This paper proposes a loop scheduling technique which can traverse loops in a new direction such that data lifespan can be greatly shortened. Therefore, overall refresh overhead can be efficiently mitigated so as to improve performance and reduce power consumption. The experimental results indicate that access latency and dynamic energy can be improved by 21.4~96.0% and 22.0~95.5% respectively by the proposed scheduling scheme. Keni Qiu, Junpeng Luo, Zhiyao Gong, Weigong Zhang, Jing Wang 0055, Yuanchao Xu 0002, Tao Li 0006, Chun Jason Xue |
ICCD | 4 |
| 2016 | An adaptive Non-Uniform Loop Tiling for DMA-based bulk data transfers on many-core processorabstractMesh Network-on-Chip (NoC) is a key fabric to interconnect many cores with desirable scalability, reliability and interoperability. We observe that DMA-based bulk data block transfer exhibits non-negligible NoC latency due to heavy congestions. Loop tiling is an effective way to partition data space for SPM+DMA-based data block transfer. Nevertheless, we observe that the unbalanced NoC latency can degrade the effectiveness of loop tiling in a uniform fashion. In this paper, we propose a NoC-aware Non-Uniform Loop Tiling (NULT) scheme to improve DMA performance. A NULT framework is built on the proposed model to adaptively hide DMA latency into computation time and reduce the overall execution time. The framework first groups cores into different families taking into account their distance-to-data in NoC. Then a heuristic method is presented to solve the near optimal tiling factors for each core family. In this way, different core families are assigned non-uniform tiling sizes. We evaluate the NULT scheme on the NIRGAM platform. Compared to the traditional uniform tiling approach, the proposed NULT technique shows more benefit to overlap memory access time and computation time and thus reduce the overall execution time of a loop nest. Keni Qiu, Yuanhui Ni, Weigong Zhang, Jing Wang 0055, Chun Jason Xue, Tao Li 0006 |
ICCD | 3 |
| 2016 | Exploring Variation-Aware Fault-Tolerant Cache under Near-Threshold ComputingabstractNear threshold voltage computing enables transistor voltage scaling to continue with Moore's Law projection and dramatically improves power and energy efficiency. However, reducing the supply voltage to near-threshold level significantly increases the susceptibility of on-chip caches to process variations, leading to the high error rate. Most existing fault-tolerant schemes significantly sacrifice cache capacity and performance. In this paper, we propose a novel fault-tolerant cache architecture at near-threshold computing, which is suitable for high error rate memories. We first propose a variation-aware skewed-associative cache, and then redirect the faulty blocks to the error-free blocks based on it to explore the fault-tolerance cache design. Unlike previous cache reconfiguration schemes for the fault tolerance, our cache design does not need to sacrifice or disable any fault-free blocks to form a completely functional set. We use all error-free blocks and have the least cache capacity waste. More importantly, since the aging impact could also cause cell failures, our skewed cache takes the aggregated process variation and aging impact into the consideration. Last but not least, our skewed cache design avoids the complex remapping from faulty blocks to the error-free blocks and minimizes the hardware overheads. Our evaluation results show that our variation-aware fault-tolerant cache design exhibits strong capability to tolerate the high error rate, and more excitingly, its effectiveness on reducing the cache miss rate and improving the performance is even more obvious as the supply voltage scales down to the near-threshold region. Jing Wang 0055, Yanjun Liu 0005, Weigong Zhang, Kezhong Lu, Keni Qiu, Xin Fu 0001, Tao Li 0006 |
ICPP | 3 |
| 2016 | Reducing Synchronization Cost for Single-Level Store in Mobile Systems
Yuanchao Xu 0002, Hu Wan 0001, Keni Qiu, Tao Li 0006, Weigong Zhang |
J. Comput. Sci. Technol. | 5 |
| 2016 | Write Mode Aware Loop Tiling for High Performance Low Power Volatile PCM in Embedded SystemsabstractArchitecting PCM, especially MLC PCM, as main memory for MCUs is a promising technique to replace conventional DRAM deployment. However, PCM/MLC PCM suffers from long write latency and large write energy. Recent work has proposed a compiler directed dual-write (CDDW) scheme to combat the drawbacks of PCM by adopting fast or slow mode for different write operations. For large-scale loops, we observe that write instances' lifetime is very long and can only be written by the expensive slow mode. This paper proposes a write mode aware loop tiling approach to effectively reduce the lifetime of write instances and maximize the number of efficient fast writes in loops. The experimental results show that the proposed approach improves performance by 50.8 percent and reduces dynamic energy by 32.0 percent across a set of benchmarks compared to the CDDW approach on average. Keni Qiu, Qing'an Li, Jingtong Hu, Weigong Zhang, Chun Jason Xue |
IEEE Trans. Computers | 4 |
| 2014 | Consensus on compact Riemannian manifolds
Sheng Chen 0010, Lindu Zhao, Weigong Zhang, Peng Shi 0001 |
Inf. Sci. | 3 |