Minxuan Zhang

dblp:81/3147 · DBLP profile ↗
← Back
23ranked-venue papers
1as first author
0since 2021 · last 2020
0009-0001-1340-6638ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
5 papers
Memory systems · 30% Hardware accelerators and domain-specific architectures · 24% GPUs and heterogeneous computing · 19%
Computer networks
1 paper
Internet architecture and protocols · 50% Routing and switching · 50%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.722019
LAcc: Exploiting Lookup Table-based Fast and Accurate Vector Multiplication in DRAM-based CNN Accelerator · DAC 2019
DrAcc: a DRAM based accelerator for accurate CNN inference · DAC 2018
Memory systems
processing-in-memory
0.522019
LAcc: Exploiting Lookup Table-based Fast and Accurate Vector Multiplication in DRAM-based CNN Accelerator · DAC 2019
DrAcc: a DRAM based accelerator for accurate CNN inference · DAC 2018
GPUs and heterogeneous computing
GPU architecture
0.412020
FRF: Toward Warp-Scheduler Friendly STT-RAM/SRAM Fine-Grained Hybrid GPGPU Register File Design · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Memory systems
non-volatile memory
0.412020
FRF: Toward Warp-Scheduler Friendly STT-RAM/SRAM Fine-Grained Hybrid GPGPU Register File Design · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Processor architecture and microarchitecture › register file
register file design
0.412020
FRF: Toward Warp-Scheduler Friendly STT-RAM/SRAM Fine-Grained Hybrid GPGPU Register File Design · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Routing and switching › data plane › router data plane
high-speed packet processing
0.212015
Towards high-performance packet processing on commodity multi-cores: current issues and future directions · Sci. China Inf. Sci. 2015
Internet architecture and protocols
packet processing
0.212015
Towards high-performance packet processing on commodity multi-cores: current issues and future directions · Sci. China Inf. Sci. 2015
GPUs and heterogeneous computing › GPU scheduling
warp scheduling
0.112020
FRF: Toward Warp-Scheduler Friendly STT-RAM/SRAM Fine-Grained Hybrid GPGPU Register File Design · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2020
Integrated circuit design › emerging device technologies
carbon nanotube field-effect transistor
0.112010
Opitimization of tunneling carbon nanotube-FETs based on stair-case doping strategy · Sci. China Inf. Sci. 2010
Integrated circuit design
digital circuit design
0.112010
Opitimization of tunneling carbon nanotube-FETs based on stair-case doping strategy · Sci. China Inf. Sci. 2010
Parallel and multicore computing › parallel architecture
multicore packet processing
0.112015
Towards high-performance packet processing on commodity multi-cores: current issues and future directions · Sci. China Inf. Sci. 2015
Parallel and multicore computing
parallel programming models
0.112015
Towards high-performance packet processing on commodity multi-cores: current issues and future directions · Sci. China Inf. Sci. 2015

Methods — techniques the papers use, named apart from their topics

on-demand register remapping · 0.4interleaved register mapping · 0.4lookup table · 0.4DRAM-based acceleration · 0.3staircase doping · 0.1TCAD simulation · 0.1
YearPublicationVenuePosition
2020 FRF: Toward Warp-Scheduler Friendly STT-RAM/SRAM Fine-Grained Hybrid GPGPU Register File Design
abstract
Modern graphics processing units (GPUs) exhibit increasing demands for register files (RFs) with larger capacity and bank sizes, which jeopardize the traditional SRAM-based RF designs due to their large die area and long access latency. Recent hybrid RF designs, e.g., SRAM and spin-transfer torque random access memory (STT-RAM)-based RFs, mitigate the issue by exploiting the density and performance advantages in STT-RAM and SRAM, respectively. However, existing hybrid RF designs adopt coarse integration that has limited write bandwidth between SRAM and STT-RAM, which restricts the adoption of different warp schedulers at runtime. In this article, we propose FRF, a warp-scheduler friendly fine-grained hybrid RF design using SRAM/STT-RAM hybrid cell (HC) structures. By integrating one SRAM cell and N STT-RAM cells as one HC, FRF exploits internal write paths to enlarge the access bandwidth between SRAM and STT-RAM and thus greatly optimizes the area and performance. FRF enables the concurrent context-switching such that different warp schedulers may be adopted at runtime. FRF adopts interleaved register mapping (IRM) and on-demand register remapping to further improve the utilization of SRAM in each HC. Our experimental results show that, on average, FRF achieves 50% performance improvement and 40% energy consumption reduction over the coarse-grained hybrid design when adopting loose round-robin (LRR), and achieves 159% efficiency improvement over pure STT-RAM-based RF.
Quan Deng 0003, Youtao Zhang, Shuzheng Zhang, Minxuan Zhang, Jun Yang 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2019 LAcc: Exploiting Lookup Table-based Fast and Accurate Vector Multiplication in DRAM-based CNN Accelerator
abstract
PIM (Processing-in-memory)-based CNN (Convolutional neural network) accelerators leverage the characteristics of basic memory cells to enable simple logic and arithmetic operations so that the bandwidth constraint can be effectively alleviated. However, it remains a major challenge to support multiplication operations efficiently on PIM accelerators, in particular, DRAM-based PIM accelerators. This has prevented PIM-based accelerators from being immediately adopted for accurate CNN inference.
Quan Deng 0003, Youtao Zhang, Minxuan Zhang, Jun Yang 0002
DAC3
2018 DrAcc: a DRAM based accelerator for accurate CNN inference
abstract
Modern Convolutional Neural Networks (CNNs) are computation and memory intensive. Thus it is crucial to develop hardware accelerators to achieve high performance as well as power/energy-efficiency on resource limited embedded systems. DRAM-based CNN accelerators exhibit great potentials but face inference accuracy and area overhead challenges.
Quan Deng 0003, Lei Jiang 0001, Youtao Zhang, Minxuan Zhang, Jun Yang 0002
DAC4
2017 Towards warp-scheduler friendly STT-RAM/SRAM hybrid GPGPU register file design
abstract
Modern Graphics Processing Units (GPUs) widely adopt large SRAM based register file (RF) to enable fast context-switch. A large SRAM RF may consume 20% to 40% GPU power, which has become one of the major design challenges for GPUs. Recent studies mitigate the issue through hybrid RF designs that architect a large STT-RAM (Spin Transfer Torque Magnetic memory) RF and a small SRAM buffer. However, the long STT-RAM write latency throttles the data exchange between STT-RAM and SRAM, which deprecates warp scheduler with frequent context switches, e.g., round robin scheduler. In this paper, we propose HC-RF, a warp-scheduler friendly hybrid RF design using novel SRAM/STT-RAM hybrid cell (HC) structure. HC-RF exploits cell level integration to improve the effective bandwidth between STT-RAM and SRAM. By enabling silent data transfer from SRAM to STT-RAM without blocking RF banks, HC-RF supports concurrent context-switching and decouples its dependency on warp scheduler. Our experimental results show that, on average, HC-RF achieves 50% performance improvement and 44% energy consumption reduction over the coarse-grained hybrid design when adopting LRR(Loose Round Robin) warp scheduler.
Quan Deng 0003, Youtao Zhang, Minxuan Zhang, Jun Yang 0002
ICCAD3
2015 Towards high-performance packet processing on commodity multi-cores: current issues and future directions
Jinli Yan, Zhigang Sun 0002, Tao Li 0008, Minxuan Zhang
Sci. China Inf. Sci.5
2014 Software/Hardware Parallel Long-Period Random Number Generation Framework Based on the WELL Method
abstract
This paper presents a hardware architecture for efficient implementation of the well equidistributed long-period linear (WELL) algorithm. Our design achieves a throughput of one sample-per-cycle and runs as fast as 423 MHz on a Xilinx XC5VFX130T field-programmable gate array (FPGA) device. This performance is 7.1-fold faster than a dedicated software implementation. The proposed architecture is also implemented on targeting different devices for the comparison of other types of pseudorandom number generators. In addition, we design a software/hardware framework that is capable of dividing the WELL stream into an arbitrary number of independent parallel substreams. With support from software, this framework can obtain speedup roughly proportional to the number of parallel cores. The sequences produced by the single design are verified to be consistent with the standard software generator. In addition, the statistical tests of interleaved sequences are also performed to check for correlations between different substreams of the parallel framework. We apply our framework to two applications. Experimental results verify the correctness of our framework as well as the better characteristics of the WELL algorithm compared with the Mersenne Twister method.
Paul Chow, Minxuan Zhang, Shaojun Wei
IEEE Trans. Very Large Scale Integr. Syst.4
2013 Addressing Transient and Permanent Faults in NoC With Efficient Fault-Tolerant Deflection Router
abstract
Continuing decrease in the feature size of integrated circuits leads to increases in susceptibility to transient and permanent faults. This paper proposes a fault-tolerant solution for a bufferless network-on-chip, including an on-line fault-diagnosis mechanism to detect both transient and permanent faults, a hybrid automatic repeat request, and forward error correction link-level error control scheme to handle transient faults and a reinforcement-learning-based fault-tolerant deflection routing (FTDR) algorithm to tolerate permanent faults without deadlock and livelock. A hierarchical-routing-table-based algorithm (FTDR-H) is also presented to reduce the area overhead of the FTDR router. Synthesized results show that, compared with the FTDR router, the FTDR-H router can reduce the area by 27% in an 88 network. Simulation results demonstrate that under synthetic workloads, in the presence of permanent link faults, the throughput of an 8 8 network with FTDR and FTDR-H algorithms are 14% and 23% higher on average than that with the fault-on-neighbor (FoN) aware deflection routing algorithm and the cost-based deflection routing algorithm, respectively. Under real application workloads, the FTDR-H algorithm achieves 20% less hop counts on average than that of the FoN algorithm. For transient faults, the performance of the FTDR router can achieve graceful degradation even at a high fault rate. We also implement the fault-tolerant deflection router which can achieve 400 MHz in TSMC 65-nm technology.
Chaochao Feng, Zhonghai Lu, Axel Jantsch, Minxuan Zhang, Zuocheng Xing
IEEE Trans. Very Large Scale Integr. Syst.4
2012 Software/hardware framework for generating parallel Gaussian random numbers based on the Monty Python method
abstract
We present a hardware architecture for efficient implementation of a Gaussian random number generator (GRNG), using the Monty Python method. To maximize the performance/complexity efficiency, an efficient word-length optimization model is proposed to find out both the optimal integer and fractional word-lengths for signals. Experimental results show that our optimized Fixed-Point design achieves a throughput of almost 1 sample-per-cycle and runs as fast as 375.9 MHz on a Xilinx XC6VLX240T FPGA device. This performance is 23.4-fold faster than a dedicated software version running on a 2.67-GHz Intel core i5 processor. It takes 1976 LUTs, 1785 Flip-Flops, 12 BRAMs and 35 DSPs, which is only about 1% of the device as well as a great reduction compared to its corresponding Floating-Point implementations. Furthermore, we develop a framework that is capable of partitioning the Gaussian distribution stream into an arbitrary number of parallel sub-streams. With support from software, this framework can obtain speedup roughly linearly with the number of parallel cores. The quality of the variables produced by our design are verified via the standard Gaussian statistical test suit, the chi-square (X2) test.
Paul Chow, Minxuan Zhang, Shaojun Wei
FPT4
2012 PSA-NUCA: A Pressure Self-Adapting Dynamic Non-uniform Cache Architecture
abstract
The constantly widening processor-memory speed gap substantially exacerbates the dependence of program performance on the on-chip memory hierarchy design and data management in chip multiprocessors. However, traditional data management mechanisms take neither the characteristic of asymmetric distribution of on-chip memory accesses nor the property of non-uniform access latency into consideration in large distributed cache. It is difficult to make an intelligent trade-off between the hit rate and the hit latency, which has an important impact on the memory efficiency. To tackle this problem, this paper presents a novel pressure self-adapting dynamic non-uniform cache architecture (PSA-NUCA). By integrating the replica and activity aware pseudo-LRU replacement policy (RAA-LRU), the enhanced first-touch mapping policy based on selective victim retention (FT-SVR), and the pressure aware adaptive replication policy (PA-ARP) into a unified intelligent data management framework, PSA-NUCA alleviates the contradiction between the miss rate and hit latency effectively with concern for both the characteristic of asymmetric distribution of memory access and the property of non-uniform access latency. Simulation results using a full system simulator demonstrate that PSA-NUCA outperforms the baseline shared non-uniform cache architecture by an average of 7.78% for the multi-thread benchmark programs we examined, while the hardware overhead is negligible.
Anwen Huang, Wenqiang Shi, Minxuan Zhang
NAS5
2011 Timing-Driven Routing of High Fanout Nets
abstract
It has been observed in the past that the PathFinder routing algorithm runtime could be hampered by high fan out nets, primarily due to the time spent on the initialization of the priority queue. However, a solution has only been reported for routability/wirelength driven routers. In this paper, we report two heuristics that address the same issue for timing-driven routers. We show that on standard MCNC benchmarks, the proposed techniques can achieve 1.53 and 1.56 time speed up against the versatile placement and router (VPR), while achieving the same quality of result.
Jianwen Zhu, Minxuan Zhang
FPL3
2011 Software/Hardware Framework for Generating Parallel Long-Period Random Numbers Using the WELL Method
abstract
The Well Equidistributed Long-period Linear (WELL) algorithm is proven to have better characteristics than the Mersenne Twister (MT), one of the most widely used long-period pseudo-random number generators (PRNGs). In this paper, we propose a hardware architecture for efficient implementation of WELL. Our design achieves a throughput of 1 sample-per-cycle and runs as fast as 449.4 MHz on a Xilinx XC6VLX240T FPGA. This performance is 7.6-fold faster than a dedicated software implementation, and is comparable to a MT hardware generator built on the same device. It takes up 633 LUTs, 537 Flip-Flops and 4 BRAMs, which is only 0.5% of the device. Furthermore, we design a software/hardware framework that is capable of dividing the WELL stream into an arbitrary number of independent parallel sub-streams. With support from software, this framework can obtain speedup roughly proportional to the number of parallel cores. The quality of the random numbers generated by our design is verified by the standard statistical test suites Diehard and TestU01. We also apply our framework to a Monte-Carlo simulation for estimating p. Experimental results verify the correctness of our framework as well as the better characteristics of the WELL algorithm.
Paul Chow, Minxuan Zhang
FPL4
2011 Accelerating the Extraction of Representative Behaviors of Programs with Dynamic Binary Translation
abstract
Program behavior analysis is the foundation of computer architecture research. Therefore it is vital to be able to extract the representative behaviors of programs in an efficient manner. Representative behaviors of programs are usually extracted through the SimPoint methodology. However, generating BBV (Basic Block Vector) profiles for SimPoint is usually quite slow. This paper evaluates the effectiveness of accelerating BBV profile generation with dynamic binary translation technique. First, A general framework for BBV profile generation using dynamic binary translation is presented. Then several optimization techniques and accuracy enhancements are proposed. Based on the framework and the optimizations, a highly efficient BBV profile generator, QPoint, is presented. The performance, overhead and accuracy of QPoint is evaluated using the SPEC2006 benchmark set. Experimental results show that the optimization method proposed can improve the performance by up to 147%, on average 56%. The speed of the optimized QPoint is up to 40x, and on average 10.5x compared with a functional simulation based BBV profile generator. The overhead incurred by BBV profile gathering is less than 4% which is the lowest among existing tools. The accuracy of QPoint is also validated against a functional simulation based tool. Compared with existing tools, the proposed QPoint tool has two main advantages. First, the performance of QPoint is tremendous, with a speed of up to 292 MIPS, on average 109 MIPS, on an ordinary PC. Second, QPoint supports most architectures, including x86/x86 64, ARM, POWER, SPARC, MIPS et al., and can be used to generate cross-platform BBV profiles.
Tianlei Zhao, Guitao Fu, Shubo Qi, Xiaomin Jia, Minxuan Zhang
HPCC6
2010 Phase Characterization and Classification for Micro-architecture Soft Error
abstract
Transient faults have become a key challenge to modern processor design. Processor designers take Architectural Vulnerability Factor (AVF) as an estimation method of micro-architectures soft error rate. Dynamic, phase-based system reliability management, which tunes system hardware and software parameters at runtime for different phases, has become a focus in the field of processor design. Phase characterization technique (PCT) and phase classification algorithm (PCA) determine the accuracy of phase identification, which is the foundation of dynamic, phase-based system management. To our knowledge, this paper is the first to give a comprehensive evaluation and comparison of PCTs and PCAs for micro-architecture soft error. We first compare the efficiency of basic block vectors (BBV) and performance metric counters (PMC) based PCTs in reliability-oriented phase characterization on three micro-architectural structures (i.e. instruction queue, function unit and reorder buffer). Experimental results show that PMC based PCT performs better than BBV based PCT for most programs studied. Also, we compare the accuracy of three clustering algorithms (i.e. hierarchical clustering, k-means clustering and regression tree) in reliability-oriented phase classification. Regression tree method is demonstrated to improve the accuracy of classification by 30% compared with other two PCAs on average. Furthermore, based on the comparisons of PCTs and PCAs, we propose the optimal combination of PCT and PCA for soft error reliability-oriented phase identification - the combination of PMC and regression tree. In addition, we quantify the upper bound of predictability of AVF using BBV/PMC. Overall, an average of 82% AVF can be explained by PMC, while BBV can explain 78% AVF averagely.
Anguo Ma, Yuxing Tang, Minxuan Zhang
EUC4
2010 Towards Online Application Cache Behaviors Identification in CMPs
abstract
On chip multiprocessors (CMPs) platforms, multiple co-scheduled applications can severely degrade performance and quality of service (QoS) when they contend for last-level cache (LLC) resources. Whether an application will impose destructive interference on co-scheduled applications is largely dependent on its own inherent cache access behavior characteristics. In this work, we first present case studies that show how inter-application interferences result in undesirable performance in both shared and private cache based LLC designs. We then propose a new online approach for application cache behavior identification on the basis of detailed simulation and analysis with SPEC CPU2006 benchmarks. We demonstrate that our approach can more concisely identify application cache behaviors. Moreover, the proposed approach can be implemented directly in hardware to dynamically identify the application cache behaviors at runtime. Finally, we show with two case studies that how the proposed approach can be adopted by both shared and private based cache sharing mechanisms, i.e. cache partitioning algorithms (CPAs) and cache spilling techniques, for more concise cache resource management.
Xiaomin Jia, Tianlei Zhao, Shubo Qi, Minxuan Zhang
HPCC5
2010 A high performance router with dynamic buffer allocation for on-chip interconnect networks
abstract
With the number of processor cores increasing in chip multi-processors (CMPs) and global wire delays increasing, networks on chip have been gaining wide acceptance for on-chip inter-core communication. This paper introduces a low latency Dynamic Virtual Output Queues Router (DVOQR), which can reduce the router latency to two cycles by leveraging look-ahead routing computation and virtual output address queues scheme. Simulation results show that network throughput on a 4×4 mesh increases by up to 46.9% and 28.6%, compared to wormhole router and virtual channel router, and that DVOQR outperforms doubled buffer virtual channel router by 1.9% under same input speedup. Network zero-load-latency also decreases by 25.6% and 41% respectively under random traffic. The results with place and route used by Cadence Encounter in TSMC 65nm technology display that the frequency of DVOQR can reach 1.4 GHz, the cell area of the router is only 0.424mm2and the power consumption is 274 mw under the 50% injection rate.
Shubo Qi, Minxuan Zhang, Tianlei Zhao, Shaoqing Li
ICCD2
2010 Opitimization of tunneling carbon nanotube-FETs based on stair-case doping strategy
Hailiang Zhou, Minxuan Zhang
Sci. China Inf. Sci.3
2008 Dimensional Bubble Flow Control and Fully Adaptive Routing in the 2-D Mesh Network on Chip
abstract
In this paper, the novel flow control strategy called dimensional bubble flow control (DBFC) is presented. The flow control strategy of DBFC builds on virtual cut-through switching and credit-based flow control mechanism and analyzes the credit value of port and the routing information of the packets to realize the point-point flow control. In the 2-D mesh network on chip, when the flow control strategy of DBFC is accepted, the adaptive dimensional bubble routing (ADBR) algorithm designed in this paper can get the goals including deadlock-free and minimal distance even if the cyclic dependencies exist. In this paper, the detail proof is provided for these conclusions. Lastly, we adapt the source code of NOXIM that is a popular simulator of on-chip networks and realize the flow control of DBFC and ADBR algorithm in NOXIM. We test the performance of ADBR on NOXIM. The simulation performance shows our scheme is superior to the usual approach such as XY dimension-order routing, with nearly 17.5% improvement in the packets latency and throughput.
Canwen Xiao, Minxuan Zhang, Yong Dou, Zhitong Zhao
EUC (1)2
2007 Look-Ahead Adaptive Routing on k -Ary n -Trees
Quanbao Sun, Liquan Xiao, Minxuan Zhang
APPT3
2007 A Parallel Infrastructure on Dynamic EPIC SMT
Qingying Deng, Minxuan Zhang
ICA3PP2
2007 Hardware-Based Multicast with Global Load Balance on k-ary n-trees
abstract
The multicast operation is used commonly in parallel applications and can be used to support several other collective communication operations. A significant performance improvement can be achieved by supporting multicast operations at the hardware level. In this paper, we propose two parent selecting strategies which use global information to reduce the conflict among different multicast operations on k-ary n-trees. We first define an equivalence relation to divide the switches at each stage into several equivalence classes. Then we prove that the switches, which are at the same stage and are passed through by the same multicast tree, belong to the same equivalence class. Based on the study, two least loaded parent selecting strategies are developed. The proposed strategies are evaluated through simulation experiments. The results indicate that the proposed strategies lower the multicast latency and increase the multicast throughput significantly.
Quanbao Sun, Minxuan Zhang, Liquan Xiao
ICPP2
2007 A Parallel Infrastructure on Dynamic EPIC SMT and Its Speculation Optimization
Qingying Deng, Minxuan Zhang
ISPA2
2006 Controlling Performance of a Time-Criticial Thread in SMT Processors by Instruction Fetch Policy
abstract
In simultaneous multithreading (SMT) processors, the instruction fetch policy affects the speed at which each thread runs and overall throughput. However, current fetch policies almost focus on overall throughput optimization, and provide no control over how fast individual threads run. As a result, the performance of a thread varies with fetch policy and the workload it is executed. This performance unpredictability means that the execution time of a thread is unpredictable. So only depending on the operating system (OS) thread scheduler to guarantee the execution time constraint of a time critical thread is not enough even fails. The hardware must ensure that the performance of the time critical thread is predictable in any timeslice. In this paper, we propose a novel fetch policy to control performance of a time critical thread in SMT processors. We evaluate our policy using many different workloads, and results show that for more than 94% of all cases measured, our policy can achieve the desired performance. For the failing cases, the average variance is within 1.25%. Furthermore, our policy does not sacrifice overall throughput severely. Compared to fetch policies orienting towards throughput maximization such as ICOUNT, the average degradation of overall throughput is less than 3%. Especially, our policy makes efforts to maximize the throughput of all coscheduled threads other than the time critical one, and gives 98.25% of the throughput achieved by ICOUNT on average
Caixia Sun, Hong-Wei Tang, Minxuan Zhang
PDCAT3
2005 Enhancing DCache Warn Fetch Policy for SMT Processors
Minxuan Zhang, Caixia Sun
ISPA1