EDBT 2026 Demo / reviewers in the wild / expert
Xin He 0011
dblp:69/1798-11
· DBLP profile ↗
23ranked-venue papers
8as first author
6since 2021 · last 2023
0000-0001-9280-1208ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 7 first-author · 6 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | PEDAL: A Power Efficient GCN Accelerator with Multiple DAtafLowsabstractGraphs are ubiquitous in many application domains due to their ability to describe structural relations. Graph Convolutional Networks (GCNs) have emerged in recent years and are rapidly being adopted due to their capability to perform Machine Learning (ML) tasks on graph-structured data. GCN exhibits irregular memory accesses due to the lack of locality when accessing graph-structured data. This makes it hard for general-purpose architectures like CPUs and GPUs to fully utilize their computing resources. In this paper, we propose PEDAL, a power-efficient accelerator for GCN inference supporting multiple dataflows. PEDAL chooses the best-fit dataflow and phase ordering based on input graph characteristics and GCN algorithm, achieving both efficiency and flexibility. To achieve both high power efficiency and performance, PEDAL features a light-weight processing element design. PEDAL achieves 144.5x, 9.4x, and 2.6x speedup compared to CPU, GPU, and HyGCN, respectively, and 8856x, 1606x, 8.4x, and 1.8x better power efficiency compared to CPU, GPU, HyGCN, and EnGN, respectively. Alireza Khadem, Xin He 0011, Nishil Talati, Tanvir Ahmed Khan 0001, Trevor N. Mudge |
DATE | 3 |
| 2022 | Squaring the circle: Executing Sparse Matrix Computations on FlexTPU - A TPU-Like ProcessorabstractSystolic arrays have been successful to accelerate dense linear algebra for deep neural networks (DNNs), but cannot handle sparse computations efficiently. Though early attempts have been made to perform sparse matrix operations on weight-pruned DNNs, handling highly sparse matrices with skewed nonzero distribution commonly seen in real-world graph analytics remains challenging. In this paper, we propose FlexTPU framework to repurpose tensor processing units (TPUs) to execute sparse matrix-vector operations (SpMV). First, we propose a lightweight Z-shape mapping of sparse matrices onto the systolic array to eliminate the processing of zeros as much as possible, regardless of the sparsity and nonzero distribution. On top of the mapping, we devise an SpMV dataflow executed by an array of PEs, which are a slightly modified version of the conventional TPU PE. Second, in contrast to the excess preprocessing mandatory for prior attempts, the Z-shape mapping facilitates on-the-fly matrix condensing from the widely-used compressed sparse matrix (e.g. CSR) representation. This is accomplished by a proposed sparse data loader that includes an on-chip row decoder and parallel nonzero loaders. We evaluate FlexTPU on a broad set of synthetic and real-world sparse matrices. The experimental result shows that FlexTPU achieves 3.55× speedup and 3.27× energy saving over a state-of-the-art design, Sparse-TPU. It performs even better on sparse matrices with power-law distributions. Compared to state-of-the-art library implementations on a CPU and a GPU, FlexTPU also achieves an average speedup of 2.4× and 4.3×, and energy saving of 130.4× and 495.3×, respectively. FlexTPU is also evaluated against a recent re configurable (chip multi-processor) CMP machine, Transmuter. FlexTPU outperforms Transmuter by achieving 5.12× speedup and 2.65× energy saving. Xin He 0011, Kuan-Yu Chen 0001, Siying Feng, Hun-Seok Kim, David T. Blaauw, Ronald G. Dreslinski, Trevor N. Mudge |
PACT | 1 |
| 2022 | MeNDA: a near-memory multi-way merge solution for sparse transposition and dataflowsabstractNear-memory processing has been extensively studied to optimize memory intensive workloads. However, none of the proposed designs address sparse matrix transposition, an important building block in sparse linear algebra applications. Prior work shows that sparse matrix transposition does not scale as well as other sparse primitives such as sparse matrix vector multiplication (SpMV) and hence has become a growing bottleneck in common applications. Sparse matrix transposition is highly memory intensive but low in computational intensity, making it a promising candidate for near-memory processing. In this work, we propose MeNDA, a scalable near-DRAM multi-way merge accelerator that eliminates the off-chip memory interface bottleneck and exposes the high internal memory bandwidth to improve performance and reduce energy consumption for sparse matrix transposition. MeNDA adopts a merge sort based algorithm, exploiting spatial locality, and proposes a near-memory processing unit (PU) featuring a high-performance hardware merge tree. Because of the wide application of merge sort in sparse linear algebra, MeNDA is an extensible solution that can be easily adapted to support other sparse primitives such as SpMV. Techniques including seamless back-to-back merge sort, stall reducing prefetching and request coalescing are further explored to take full advantage of the increased system memory bandwidth. Compared to two state-of-the-art implementations of sparse matrix transposition on a CPU and a sparse library on a GPU, MeNDA is able to achieve a speedup of 19.1X, 12.0X, and 7.7x, respectively. MeNDA also shows an efficiency gain of 3.8x over a recent SpMV accelerator integrated with HBM. Incurring a power consumption of only 78.6 mW, a MeNDA PU can be easily accommodated by commodity DIMMs. Siying Feng, Xin He 0011, Kuan-Yu Chen 0001, Liu Ke 0001, Xuan Zhang 0001, David T. Blaauw, Trevor N. Mudge, Ronald G. Dreslinski |
ISCA | 2 |
| 2022 | Enabling Software-Defined RF Convergence with a Novel Coarse-Scale Heterogeneous ProcessorabstractRF system development is traditionally constrained by a restrictive trade-off between power efficiency and programmatic flexibility. We outline a path towards achieving both, thereby enabling a range of new system concepts that better utilize limited resources. As an example, for many future applications, we consider RF convergence – reusing the same spectrum and waveforms to achieve multiple distributed system functions and goals, simultaneously. To enable this next step in processing, we develop a novel framework that includes both software and the system-on-chip (SoC) design. Daniel W. Bliss, Tutu Ajayi, Ali Akoglu, Ilkin Aliyev, Toygun Basaklar, Leul Belayneh, David T. Blaauw, John S. Brunhaver, Chaitali Chakrabarti, Liangliang Chang, Kuan-Yu Chen 0001, Ming-Hung Chen, Xing Chen 0004, Alex R. Chiriyath, Alhad Daftardar, Ronald G. Dreslinski, Arindam Dutta, Allen-Jasmin Farcas, Yukang Fu, A. Alper Goksoy, Xin He 0011, Md Sahil Hassan, Andrew Herschfelt, Jacob Holtom, Hun-Seok Kim, Anish Krishnakumar, Owen Ma, Joshua Mack, Saurav Mallik, Sumit K. Mandal, Radu Marculescu, Brittany M. McCall, Trevor N. Mudge, Ümit Y. Ogras, Vishrut Pandey, Saquib Ahmad Siddiqui, Yu-Hsiu Sun, Adarsh A. Venkataramani, Xiangdong Wei, Benjamin R. Willis, Hanguang Yu, Yufan Yue |
ISCAS | 21 |
| 2021 | CoSPARSE: A Software and Hardware Reconfigurable SpMV Framework for Graph AnalyticsabstractSparse matrix-vector multiplication (SpMV) is a critical building block for iterative graph analytics algorithms. Typically, such algorithms have a varying active vertex set across iterations. This variability has been used to improve performance by either dynamically switching algorithms between iterations (software) or designing custom accelerators (hardware) for graph analytics algorithms. In this work, we propose a novel framework, CoSPARSE, that employs hardware and software reconfiguration as a synergistic solution to accelerate SpMV-based graph analytics algorithms. Building on previously proposed general-purpose reconfigurable hardware, we implement CoSPARSE as a software layer, abstracting the hardware as a specialized SpMV accelerator. CoSPARSE dynamically selects software and hardware configurations for each iteration and achieves a maximum speedup of 2.0 × compared to the naïve implementation with no reconfiguration. Across a suite of graph algorithms, CoSPARSE outperforms a state-of-the-art shared memory framework, Ligra, on a Xeon CPU with up to 3.51 × better performance and 877 × better energy efficiency. Siying Feng, Jiawen Sun, Subhankar Pal, Xin He 0011, Kuba Kaszyk, Dong-Hyeon Park, John Magnus Morton, Trevor N. Mudge, Murray Cole, Michael F. P. O'Boyle, Chaitali Chakrabarti, Ronald G. Dreslinski |
DAC | 4 |
| 2021 | System-level Early-stage Modeling and Evaluation of IVR-assisted Processor Power Delivery System
An Zou, Huifeng Zhu, Jingwen Leng, Xin He 0011, Vijay Janapa Reddi, Christopher D. Gill, Xuan Zhang 0001 |
ACM Trans. Archit. Code Optim. | 4 |
| 2020 | Transmuter: Bridging the Efficiency Gap using Memory and Dataflow ReconfigurationabstractWith the end of Dennard scaling and Moore's law, it is becoming increasingly difficult to build hardware for emerging applications that meet power and performance targets, while remaining flexible and programmable for end users. This is particularly true for domains that have frequently changing algorithms and applications involving mixed sparse/dense data structures, such as those in machine learning and graph analytics. To overcome this, we present a flexible accelerator called Transmuter, in a novel effort to bridge the gap between General-Purpose Processors (GPPs) and Application-Specific Integrated Circuits (ASICs). Transmuter adapts to changing kernel characteristics, such as data reuse and control divergence, through the ability to reconfigure the on-chip memory type, resource sharing and dataflow at run-time within a short latency. This is facilitated by a fabric of light-weight cores connected to a network of reconfigurable caches and crossbars. Transmuter addresses a rapidly growing set of algorithms exhibiting dynamic data movement patterns, irregularity, and sparsity, while delivering GPU-like efficiencies for traditional dense applications. Finally, in order to support programmability and ease-of-adoption, we prototype a software stack composed of low-level runtime routines, and a high-level language library called TransPy, that cater to expert programmers and end-users, respectively. Subhankar Pal, Siying Feng, Dong-Hyeon Park, Aporva Amarnath, Chi-Sheng Yang, Xin He 0011, Jonathan Beaumont, Kyle May, Yan Xiong 0002, Kuba Kaszyk, John Magnus Morton, Jiawen Sun, Michael F. P. O'Boyle, Murray Cole, Chaitali Chakrabarti, David T. Blaauw, Hun-Seok Kim, Trevor N. Mudge, Ronald G. Dreslinski |
PACT | 7 |
| 2020 | Sparse-TPU: adapting systolic arrays for sparse matricesabstractWhile systolic arrays are widely used for dense-matrix operations, they are seldom used for sparse-matrix operations. In this paper, we show how a systolic array of Multiply-and-Accumulate (MAC) units, similar to Google's Tensor Processing Unit (TPU), can be adapted to efficiently handle sparse matrices. TPU-like accelerators are built upon a 2D array of MAC units and have demonstrated high throughput and efficiency for dense matrix multiplication, which is a key kernel in machine learning algorithms and is the target of the TPU. In this work, we employ a co-designed approach of first developing a packing technique to condense a sparse matrix and then propose a systolic array based system, Sparse-TPU, abbreviated to STPU, to accommodate the matrix computations for the packed denser matrix counterparts. To demonstrate the efficacy of our co-designed approach, we evaluate sparse matrix-vector multiplication on a broad set of synthetic and real-world sparse matrices. Experimental results show that STPU delivers 16.08X higher performance while consuming 4.39X and 19.79X lower energy for integer (int8) and floating point (float32) implementations, respectively, over a TPU baseline. Meanwhile, STPU has 12.93% area overhead and an average of 4.14% increase in dynamic energy over the TPU baseline for the float32 implementation. Xin He 0011, Subhankar Pal, Aporva Amarnath, Siying Feng, Dong-Hyeon Park, Austin Rovinski, Haojie Ye, Kuan-Yu Chen 0001, Ronald G. Dreslinski, Trevor N. Mudge |
ICS | 1 |
| 2020 | Attacking vision-based perception in end-to-end autonomous driving models
Adith Boloor, Karthik Garimella, Xin He 0011, Christopher D. Gill, Yevgeniy Vorobeychik, Xuan Zhang 0001 |
J. Syst. Archit. | 3 |
| 2020 | NeuADC: Neural Network-Inspired Synthesizable Analog-to-Digital ConversionabstractTraditional analog-to-digital converters (ADCs) employ dedicated analog and mixed-signal (AMS) circuits, requiring time-consuming manual design process. They also exhibit limited configurability to support diverse quantization schemes on the same circuitry. In this paper, we propose NeuADC-an automated design approach to synthesizing an analog-to-digital (A/D) interface that can approximate the desirable quantization function using a neural network (NN) with a single hidden layer. We leverage the mixed-signal resistive random-access memory (RRAM) crossbar architecture to design a novel dual-path configuration for the implementation of the basic NN operations at the circuit level. We exploit alternative bits encoding scheme to the conventional binary encoding to improve the training accuracy. Our method incorporates nonidealities at the device and circuit level into the training process to ensure NeuADC's robustness against variations of process, supply voltage, and temperature (PVT). Results obtained from SPICE simulation based on RRAM and standard 130-nm CMOS technology suggest that not only can NeuADC deliver promising performance compared to the state-of-the-art ADCs and other emerging converter designs across comprehensive design metrics, but it can also intrinsically support multiple configurable quantization schemes using the same hardware substrate, paving ways for future adaptable application-driven signal conversion. Our systematic evaluations on the proposed NeuADC framework also quantify the impacts on the ADC quantization quality from hidden neuron sizes, RRAM resistance imprecision, and PVT variations, and reveal the design tradeoff between speed, power, and area in a NeuADC circuit. Weidong Cao 0001, Xin He 0011, Ayan Chakrabarti, Xuan Zhang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2020 | Voltage-Stacked Power Delivery Systems: Reliability, Efficiency, and Power ManagementabstractIn today's manycore processors, the energy loss of more than 20% may result from inherent inefficiencies of conventional power delivery system (PDS) design. By stacking multiple voltage domains in series to lower the step-down conversion ratio of the off-chip voltage regulator module (VRM) and reduce the energy loss along the path of the power delivery network (PDN), voltage stacking (VS) offers a novel alternative power delivery technique to fundamentally improve power delivery efficiency (PDE). However, VS suffers from aggravated supply voltage noise from the current imbalance, which hinders its adoption. In this article, we investigate practical VS implementation in manycore processors to improve PDE and achieve reliable performance, while maintaining compatibility with advanced power management techniques. We first present the system configuration of a voltage-stacked manycore processor. We then systematically characterize supply voltage noise in VS, identify global, and residual differential currents as its dominant contributors, and calculate the possible worst supply voltage noise. We next propose a hybrid voltage regulation solution, based on a charge-recycling off-chip voltage regulator and distributed integrated voltage regulators, to mitigate supply voltage noise effectively. We also study the compatibility of VS with higher-level power management techniques. Finally, the performance of a voltage-stacked GPU system is comprehensively evaluated. The simulation results show that our approach can achieve 93.5% PDE, reducing the power loss by 13.6% compared to conventional single-layer PDS. An Zou, Jingwen Leng, Xin He 0011, Yazhou Zu, Christopher D. Gill, Vijay Janapa Reddi, Xuan Zhang 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2019 | NeuADC: Neural Network-Inspired RRAM-Based Synthesizable Analog-to-Digital Conversion with Reconfigurable Quantization SupportabstractTraditional analog-to-digital converters (ADCs) employ dedicated analog and mixed-signal (AMS) circuits and require time-consuming manual design process. They also exhibit limited reconfigurability and are unable to support diverse quantization schemes using the same circuitry. In this paper, we propose NeuADC — an automated design approach to synthesizing an analog-to-digital (A/D) interface that can approximate the desired quantization function using a neural network (NN) with a single hidden layer. Our design leverages the mixed-signal resistive random-access memory (RRAM) crossbar architecture in a novel dual-path configuration to realize basic NN operations at the circuit level and exploits smooth bit-encoding scheme to improve the training accuracy. Results obtained from SPICE simulations based on 130nm technology suggest that not only can NeuADC deliver promising performance compared to the state-of-art ADC designs across comprehensive design metrics, but also it can intrinsically support multiple reconfigurable quantization schemes using the same hardware substrate, paving the ways for future adaptable application-driven signal conversion. The robustness of NeuADC’s quantization quality under moderate RRAM resistance precision is also evaluated using SPICE simulations. Weidong Cao 0001, Xin He 0011, Ayan Chakrabarti, Xuan Zhang 0001 |
DATE | 2 |
| 2019 | SparseBNN: Joint Algorithm/Hardware Optimization to Exploit Structured Sparsity in Binary Neural NetworkabstractTo reduce power-hungry floating point operations and memory accesses in deep neural networks, quantized neural networks are proposed that replace floating point multiplications with simplified reduced-precision operations. To compensate for the accuracy loss due to the high degree of quantization, wider neural network layers with three or more times as many feature maps are employed. One by-product from these inflated layers is increased redundancy in the network. To further improve computational efficiency and leverage this inherent redundancy, we propose a joint optimization approach that simultaneously explores hardware-oriented training and efficient accelerator implementation of binary neural networks (BNN) in FPGAs. More specifically, our SparseBNN method consists of two parts. First, SparseBNN-SW is a training algorithm developed to enhance the structured sparsity of BNNs by 1) training for zero-valued ternary weights instead of binary that are more amenable to pruning and 2) regulating the sparsity for more efficient hardware deployment. Next, we present SparseBNN-HW, an accelerator architecture designed to directly execute the inference on the sparse-encoded format to save both memory access and computations. Experimental results on various representative datasets demonstrate that SparseBNN improves the power efficiency (GOPS/Watt) and resource efficiency (GOPS/kLUT) over the baseline BNN FPGA implementation by 1.70X and 2.22X. Xin He 0011, Liu Ke 0001, Xuan Zhang 0001 |
FPGA | 1 |
| 2018 | RiskCap: Minimizing Effort of Error Regulation for Approximate ComputingabstractQuality management, which is responsible for controlling approximation quality to meet user requirement, plays a key role in the applicability of approximate computing. An effective and efficient quality management needs to be accurate to detect intolerable errors meanwhile light-weight in nature. However, it is difficult to design such a quality management satisfying both the two demands and existing work usually optimizes for one demand at the expense of the other. In this paper, we aim to achieve higher energy efficiency of quality management by optimizing detection accuracy and overhead simultaneously. We observe that the detection difficulty varies across inputs and there exists much redundant computation in detection process. Based on this observation, a cascaded quality management which can minimize the overhead and doesn't lower detection accuracy is proposed. The proposed solution pays more proper computation effort according to different detection difficulties of inputs so as to avoid unnecessary energy consumption. What's more, by exploring the design space sufficiently and effectively, we can assure the highest energy-efficiency of the proposed topology. The experiment results demonstrate that our approach can achieve much greater energy-efficiency than existing solutions. Shuhao Jiang, Xin He 0011, Guihai Yan, Xuan Zhang 0001, Xiaowei Li 0001 |
ATS | 3 |
| 2018 | Efficient and reliable power delivery in voltage-stacked manycore system with hybrid charge-recycling regulatorsabstractVoltage stacking (VS) fundamentally improves power delivery efficiency (PDE) by series-stacking multiple voltage domains to eliminate explicit step-down voltage conversion and reduce energy loss along the power delivery path. However, it suffers from aggravated supply noise, preventing its adoption in mainstream computing systems. In this paper, we investigate a practical approach to enabling efficient and reliable power delivery in voltage-stacked manycore systems that can ensure worst-case supply noise reliability without excessive costly over-design. We start by developing an analytical model to capture the essential noise behaviors in VS. It allows us to identify dominant noise contributor and derive the worst-case conditions. With this in-depth understanding, we propose a hybrid voltage regulation solution to effectively mitigate noise with worst-case guarantees. When evaluated with real-world benchmarks, our solution can achieve 93.8% power delivery efficiency, an improvement of 13.9% over the conventional baseline. An Zou, Jingwen Leng, Xin He 0011, Yazhou Zu, Vijay Janapa Reddi, Xuan Zhang 0001 |
DAC | 3 |
| 2018 | AxTrain: Hardware-Oriented Neural Network Training for Approximate InferenceabstractThe intrinsic error tolerance of neural network (NN) makes approximate computing a promising technique to improve the energy efficiency of NN inference. Conventional approximate computing focuses on balancing the efficiency-accuracy trade-off for existing pre-trained networks, which can lead to suboptimal solutions. In this paper, we propose AxTrain, a hardware-oriented training framework to facilitate approximate computing for NN inference. Specifically, AxTrain leverages the synergy between two orthogonal methods---one actively searches for a network parameters distribution with high error tolerance, and the other passively learns resilient weights by numerically incorporating the noise distributions of the approximate hardware in the forward pass during the training phase. Experimental results from various datasets with near-threshold computing and approximation multiplication strategies demonstrate AxTrain's ability to obtain resilient neural network parameters and system energy efficiency improvement. Xin He 0011, Liu Ke 0001, Wenyan Lu, Guihai Yan, Xuan Zhang 0001 |
ISLPED | 1 |
| 2018 | NNest: Early-Stage Design Space Exploration Tool for Neural Network Inference AcceleratorsabstractDeep neural network (DNN) has achieved spectacular success in recent years. In response to DNN's enormous computation demand and memory footprint, numerous inference accelerators have been proposed. However, the diverse nature of DNNs, both at the algorithm level and the parallelization level, makes it hard to arrive at an "one-size-fits-all" hardware design. In this paper, we develop NNest, an early-stage design space exploration tool that can speedily and accurately estimate the area/performance/energy of DNN inference accelerators based on high-level network topology and architecture traits, without the need for low-level RTL codes. Equipped with a generalized spatial architecture framework, NNest is able to perform fast high-dimensional design space exploration across a wide spectrum of architectural/micro-architectural parameters. Our proposed novel date movement strategies and multi-layer fitting schemes allow NNest to more effectively exploit parallelism inherent in DNN. Results generated by NNest demonstrate: 1) previously-undiscovered accelerator design points that can outperform state-of-the-art implementation by 39.3% in energy efficiency; 2) Pareto frontier curves that comprehensively and quantitatively reveal the multi-objective tradeoffs in custom DNN accelerators; 3) holistic design exploration of different level of quantization techniques including recently-proposed binary neural network (BNN). Liu Ke 0001, Xin He 0011, Xuan Zhang 0001 |
ISLPED | 2 |
| 2018 | Voltage-Stacked GPUs: A Control Theory Driven Cross-Layer Solution for Practical Voltage Stacking in GPUsabstractMore than 20% of the available energy is lost in "the last centimeter" from the PCB board to the microprocessor chip due to inherent inefficiencies of power delivery subsystems (PDSs) in today's computing systems. By series-stacking multiple voltage domains to eliminate explicit voltage conversion and reduce loss along the power delivery path, voltage stacking (VS) is a novel configuration that can improve power delivery efficiency (PDE). However, VS suffers from aggravated levels of supply noise caused by current imbalance between the stacking layers, preventing its practical adoption in mainstream computing systems. Throughput-centric manycore architectures such as GPUs intrinsically exhibit more balanced workloads, yet suffer from lower PDE, making them ideal platforms to implement voltage stacking. In this paper, we present a cross-layer approach to practical voltage stacking implementation in GPUs. It combines circuit-level voltage regulation using distributed charge-recycling integrated voltage regulators (CR-IVRs) with architecture-level voltage smoothing guided by control theory. Our proposed voltage-stacked GPUs can eliminate 61.5% of total PDS energy loss and achieve 92.3% system-level power delivery efficiency, a 12.3% improvement over the conventional single-layer based PDS. Compared to the circuit-only solution, the cross-layer approach significantly reduces the implementation cost of voltage stacking (88% reduction in area overhead) without compromising supply reliability under worst-case scenarios and across a wide range of real-world benchmarks. In addition, we demonstrate that the cross-layer solution not only complements on-chip CR-IVRs to transparently manage current imbalance and restore stable layer voltages, but also serves as a seamless interface to accommodate higher-level power optimization techniques, traditionally thought to be incompatible with a VS configuration. An Zou, Jingwen Leng, Xin He 0011, Yazhou Zu, Christopher D. Gill, Vijay Janapa Reddi, Xuan Zhang 0001 |
MICRO | 3 |
| 2018 | CPicker: Leveraging Performance-Equivalent Configurations to Improve Data Center Energy Efficiency
Faqiang Sun, Guihai Yan, Xin He 0011, Huawei Li 0001, Yinhe Han 0001 |
J. Comput. Sci. Technol. | 3 |
| 2017 | ApproxEye: Enabling approximate computation reuse for microrobotic computer visionabstractAiming at real-life problems, microrobotic systems have gained more and more attention. However, limited achievable performance of microrobotic system prevents it from carrying out complex tasks. Current research work propose customize designs for different applications and incorporate dedicated accelerator for high energy efficiency. However, not only such techniques require significant manual effort and expertise for specified applications, but also the accelerator itself dictates unnegligible amount of chip resources. So in this paper we propose ApproxEye, a partial approximate computation reuse framework to accelerate microrobotic computer vision. Leveraging computation locality, ApproxEye reuses previous “similar” computations to reduce redundant computations. To squeeze every piece of computation reuse opportunity, ApproxEye proposes to 1) heuristically define optimal reuse granularity and 2) apply adaptive reuse requirements for different computations. Moreover, to reduce latency of computation reuse, ApproxEye tailors a parallel implemented search scheme for approximate computation reuse. Experimental results show ApproxEye could effectively exploit the potential of computation reuse and achieve 57.05% speedup on average. Xin He 0011, Guihai Yan, Faqiang Sun, Yinhe Han 0001, Xiaowei Li 0001 |
ASP-DAC | 1 |
| 2016 | ACR: Enabling computation reuse for approximate computingabstractApproximate computing, which trades off computation quality (e.g, accuracy) and computation efforts, has becoming a promising technique to improve performance for many mission-non-critical and error-tolerant applications. The computations in such applications usually exhibit superior value locality, i.e, computations performed by a function or code region are very likely to reproduce “similar” results. Reusing the similar results can bypass redundant computations, as long as “exact” results are not mandatory. However, conventional computation reuse techniques are less effective in approximate computing paradigm. The input values of two computation instances have to be identical to reuse one for another, hence “exact” in nature.We propose ACR, an approximate computation reuse framework, to enable computation reuse for approximate computing. ACR relaxes the exact matching requirement in inputs to some extent regulated by “similarity” quantification, thereby shifting the exact computation reuse paradigm to its approximate counterpart. We furthermore propose an input significance-aware similarity quantification scheme through statistical approaches. Experimental result shows ACR could effectively exploit the potential of computation reuse for approximate computing and reduce 47.6% computations on average for a set of approximate applications. Xin He 0011, Guihai Yan, Yinhe Han 0001, Xiaowei Li 0001 |
ASP-DAC | 1 |
| 2016 | Wide Operational Range Processor Power Delivery Design for Both Super-Threshold Voltage and Near-Threshold Voltage Computing
Xin He 0011, Guihai Yan, Yinhe Han 0001, Xiaowei Li 0001 |
J. Comput. Sci. Technol. | 1 |
| 2014 | SuperRange: Wide operational range power delivery design for both STV and NTV computingabstractThe load power range of modern processors is greatly enlarged because many advanced power management techniques like dynamic voltage frequency scaling, Turbo boosting, and Near Threshold Voltage technologies are incorporated. However, the power saving may be offset by power loss in power delivery; moreover, as the efficiency of power delivery varies greatly with different load conditions, conventional power delivery designs cannot maintain high efficiency over the entire voltage range. We propose SuperRange, a wide operational range power delivery scheme. SuperRange complements the power delivery capability of on-chip voltage regulator and off-chip voltage regulator. Experimental results show SuperRange has an average 70% power conversion efficiency over wide operational range which outperforms conventional power delivery schemes. And it also exhibits superior resilience to power-constrained systems. Xin He 0011, Guihai Yan, Yinhe Han 0001, Xiaowei Li 0001 |
DATE | 1 |