Zuocheng Xing

dblp:06/3113 · DBLP profile ↗
← Back
16ranked-venue papers
0as first author
1since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9Computer networks · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Processor architecture and microarchitecture · 62% High-performance computing · 38%
Software engineering, system software, and programming languages
2 papers
Compilers and program optimization · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Processor architecture and microarchitecture › dataflow architecture
stream architecture
0.222009
Fei Teng 64 Stream Processing System: Architecture, Compiler, and Programming · IEEE Trans. Parallel Distributed Syst. 2009
A 64-bit stream processor architecture for scientific applications · ISCA 2007
Processor architecture and microarchitecture › data-parallel architecture
stream processor
0.222009
Fei Teng 64 Stream Processing System: Architecture, Compiler, and Programming · IEEE Trans. Parallel Distributed Syst. 2009
A 64-bit stream processor architecture for scientific applications · ISCA 2007
High-performance computing
scientific computing
0.122009
A 64-bit stream processor architecture for scientific applications · ISCA 2007
Fei Teng 64 Stream Processing System: Architecture, Compiler, and Programming · IEEE Trans. Parallel Distributed Syst. 2009
Compilers and program optimization › domain-specific compilation
stream compiler
0.112009
Fei Teng 64 Stream Processing System: Architecture, Compiler, and Programming · IEEE Trans. Parallel Distributed Syst. 2009
High-performance computing
scientific computing systems
0.112007
A 64-bit stream processor architecture for scientific applications · ISCA 2007
High-performance computing › performance optimization
scientific application acceleration
0.012009
Fei Teng 64 Stream Processing System: Architecture, Compiler, and Programming · IEEE Trans. Parallel Distributed Syst. 2009

Methods — techniques the papers use, named apart from their topics

stream programming · 0.2compiler design · 0.2
YearPublicationVenuePosition
2021 Low latency group-sorted QR decomposition algorithm for larger-scale MIMO systems
abstract
Abstract Sorted QR decomposition (SQRD) has been extensively adopted for various multiple‐input‐multiple‐output (MIMO) detectors, in which the sorting process incurs severe latency when it comes to larger‐scale MIMO situations. This paper proposes a group‐SQRD (GSQRD) algorithm to alleviate the latency problem of general SQRD architectures for larger‐scale MIMO systems. Via predictively sorting a group of 4 columns at one stage, the GSQRD could eliminate the processing latency by 41% for decomposing 1616 complex‐valued matrices. Additionally, this percentage even rises up to 68% for decomposing 128128 matrices. To analyse the side effects, the GSQRD is applied in various MIMO detectors in a simulation link, which exhibits a negligible performance degradation for MIMO detection. Moreover, GSQRD is a hardware‐friendly algorithm because the division and square root operations in GSQRD are converted to multiplications for simplifying the hardware implementation. Based on this algorithm, two corresponding hardware architectures, which contains 2 and 4 columns respectively in a sorting group, are also implemented with 65‐nm CMOS technology. These architectures can work at 513 MHz to decompose 1616 complex‐valued matrices. The processing latencies are respectively 0.32 and 0.26 s, superior to the state‐of‐art designs.
Lirui Chen, Yu Wang 0068, Zuocheng Xing, Shikai Qiu, Yongzhong Li
IET Commun.3
2020 A Low-Latency Successive Cancellation Hybrid Decoder for Convolutional Polar Codes
abstract
By adopting successive cancellation list decoding (SCL), polar codes demonstrate competitive error correction performance over LDPC and Turbo codes. However, SCL decoding suffers from high computational complexity and long decoding latency, especially when the list size is very large. Successive cancellation flip (SCF), as another decoding algorithm that can achieve high error correction performance, has a complexity that is close to that of successive cancellation (SC) decoding. With the observation that SCL and SCF decoding are similar at giving more chances to inspect possible codewords simultaneously or sequentially, a novel hybrid decoder is proposed in this paper, which essentially combines the ideas of SCF and SCL decoders. Moreover, in order to compensate for the degradation of performance caused by the reduction of path splitting and further reduce the decoding latency, the convolutional polar codes are adopted with a designed bit-flipping set. Simulation results demonstrate that the proposed decoder achieves the reduction of decoding latency while attaining better performance than conventional CRC-aided SCL decoder.
Yu Wang 0068, Shikai Qiu, Lirui Chen, Yang Zhang 0026, Cang Liu, Zuocheng Xing
ICASSP7
2020 A Paralleled Greedy LLL Algorithm for 16×16 MIMO Detection
abstract
This brief proposes a paralleled greedy Lenstra-Lenstra-Lovsz (PGLLL) algorithm for 16×16 MIMO detection. First, a paralleled constant-throughput scheme is designed for LLL algorithm. Then, greedy algorithm is adopted on this scheme to select the most urgent iterations for each stage. This selecting criterion outperforms others in that numerous iterations can be concurrently selected to reduce latency, and that the two factors of LLL potential and MIMO detection strategy are comprehensively considered by this criterion to improve bit-error-rate (BER) performance. Simulation indicates that the PGLLL can realize a comparable performance to the non-greedy algorithm and LLL algorithm with less iterations. Finally, this brief is the first to propose a hardware architecture with greedy LLL algorithm. This architecture is implemented with 65-nm 1P9M CMOS technology, which can work at a maximum frequency of 625 MHz to process 16×16 complex-valued matrices every 16 clocks. The latency is 362 ns. Comparison indicates that the proposed PGLLL architecture is superior to other existing works in terms of throughput and latency performance.
Lirui Chen, Yu Wang 0068, Zuocheng Xing, Shikai Qiu, Yang Zhang 0026
ISCAS3
2020 Practical AMC model based on SAE with various optimisation methods under different noise environments
abstract
Automatic modulation classification (AMC) has recently attracted widespread attention nowadays due to its desirable features of generalisability and requirement of little prior knowledge through artificial intelligence (AI) technology. The authors propose a stacked auto‐encoder (SAE) based on various optimisation methods structure to intelligently process a feature space that includes spectral‐based features and high‐order cumulants. To unify the dimensionality of the features, they apply different normalisation methods to the feature space before training the SAE model to decide corresponding normalisations under different noise environments. Linear normalisation is superior when signal‐to‐noise ratio (SNR) is low, and standardisation is superior when SNR is between ‐1 and 4 dB. Regularisation works best when SNR is greater than 5 dB. To increase the recognition accuracy of the proposed model, they introduce the unconstrained optimisation theory to adjust the proposed SAE model, including Nelder‐Mead method, Newton optimisation method, conjugate gradient method and quasi‐Newton method. They observe that the quasi‐Newton method offers desirable performance when optimising SAE model. It is the first time to compare these data normalisation methods and discuss unconstrained optimisation theory together to recognise modulation types. The recognition accuracy of this model for eight modulation types can reach 99.8% when SNR ranges from to 10 dB.
Zerun Li, Weisong Liu, Zuocheng Xing, Yongzhong Li
IET Commun.5
2019 Algorithm and Architecture for Path Metric Aided Bit-Flipping Decoding of Polar Codes
abstract
Polar codes attract more and more attention of researchers in recent years, since its capacity achieving property. However, their error-correction performance under successive cancellation (SC) decoding is inferior to other modern channel codes at short or moderate blocklengths. SC-Flip (SCF) decoding algorithm shows higher performance than SC decoding by identifying possibly erroneous decisions made in initial SC decoding and flipping them in the sequential decoding attempts. However, it performs not well when there are more than one erroneous decisions in a codeword. In this paper, we propose a path metric aided bit-flipping decoding algorithm to identify and correct more errors efficiently. In this algorithm, the bit-flipping list is generated based on both log likelihood ratio (LLR) based path metric and bit-flipping metric. The path metric is used to verify the effectiveness of bit-flipping. In order to reduce the decoding latency and computational complexity, its corresponding pipeline architecture is designed. By applying these decoding algorithm and pipeline architecture, an improvement on error-correction performance can be got up to 0.25dB compared with SCF decoding at frame error rate of 10-4, with low average decoding latency.
Yu Wang 0068, Lirui Chen, Yang Zhang 0026, Zuocheng Xing
WCNC5
2018 Locality based warp scheduling in GPGPUs
Yang Zhang 0026, Zuocheng Xing, Cang Liu, Chuan Tang
Future Gener. Comput. Syst.2
2018 CWLP: coordinated warp scheduling and locality-protected cache allocation on GPUs
abstract
As we approach the exascale era in supercomputing, designing a balanced computer system with a powerful computing ability and low power requirements has becoming increasingly important. The graphics processing unit (GPU) is an accelerator used widely in most of recent supercomputers. It adopts a large number of threads to hide a long latency with a high energy efficiency. In contrast to their powerful computing ability, GPUs have only a few megabytes of fast on-chip memory storage per streaming multiprocessor (SM). The GPU cache is inefficient due to a mismatch between the throughput-oriented execution model and cache hierarchy design. At the same time, current GPUs fail to handle burst-mode long-access latency due to GPU’s poor warp scheduling method. Thus, benefits of GPU’s high computing ability are reduced dramatically by the poor cache management and warp scheduling methods, which limit the system performance and energy efficiency. In this paper, we put forward a coordinated warp scheduling and locality-protected (CWLP) cache allocation scheme to make full use of data locality and hide latency. We first present a locality-protected cache allocation method based on the instruction program counter (LPC) to promote cache performance. Specifically, we use a PC-based locality detector to collect the reuse information of each cache line and employ a prioritised cache allocation unit (PCAU) which coordinates the data reuse information with the time-stamp information to evict the lines with the least reuse possibility. Moreover, the locality information is used by the warp scheduler to create an intelligent warp reordering scheme to capture locality and hide latency. Simulation results show that CWLP provides a speedup up to 19.8% and an average improvement of 8.8% over the baseline methods.
Yang Zhang 0026, Zuocheng Xing, Cang Liu, Chuan Tang
Frontiers Inf. Technol. Electron. Eng.2
2017 Approximate iteration detection with iterative refinement in massive MIMO systems
abstract
To improve energy efficiency and spectral efficiency, massive multiple‐input–multiple‐output (MIMO) is proposed and becomes a promising technology in the next generation mobile communication. However, massive MIMO systems equip with scores of or hundreds of antennas which induce large‐scale matrix computations with tremendous complexity, especially for matrix inversion in data detection. Thus, many detection methods have been proposed using approximate matrix inversion algorithms, which satisfy the demand of precision with low complexity. In this study, the authors focus on the approximate detection method based on Newton iteration (NI), and propose upgraded methods named NI method with iterative refinement (NIIR) and diagonal band NIIR (DBNIIR) which combine NI method and DBNI method with iterative refinement (IR). The results show that their proposals provide about 2 dB improvement on bit error rate (BER) for 16‐quadrature amplitude modulation (QAM), and could even break the error floor existing in NI and DBNI methods for 64‐QAM modulation. Furthermore, the BER of their proposals could provide almost the same performance as the exact method. Moreover, in contrast with NI and DBNI methods, NIIR and DBNIIR methods require quite few extra complexity cost and no extra hardware resource which is quite suitable for data detection in massive MIMO.
Chuan Tang, Cang Liu, Luechao Yuan, Zuocheng Xing
IET Commun.4
2017 Hardware Architecture Based on Parallel Tiled QRD Algorithm for Future MIMO Systems
abstract
QR decomposition (QRD) has been a vital component in the transceiver processor of future multiple-input multiple-output (MIMO) systems, in which antenna configuration will be more and more flexible. Therefore, the QRD hardware architecture in the future MIMO systems should be more flexible to meet various antenna configurations. Unfortunately, the existing QRD hardware architectures mainly focus on the matrix of one or several fixed sizes. This paper presents a new triangular systolic array QRD hardware architecture based on parallel tiled QRD algorithm to decompose an 8 × 8 real matrix. The designed hardware architecture is flexible and can be used in various MIMO systems, in which the number of antennas is smaller than 4. This paper also proposes a modified algorithm for the bottleneck operations of parallel tiled QRD algorithm to reduce the hardware overhead. To further reduce the hardware overhead, the Newton-Raphson algorithm is adopted in the proposed algorithm. The implementation results show that the normalized processing latency performance and the normalized processing efficiency performance of the designed QRD hardware architecture both are better than most of the existing QRD hardware architectures. To the best of our knowledge, the hardware architecture presented in this paper achieves the superior normalized QRD rate performance to the existing QRD hardware architectures.
Cang Liu, Chuan Tang, Zuocheng Xing, Luechao Yuan, Yang Zhang 0026
IEEE Trans. Very Large Scale Integr. Syst.3
2017 A Flexible Divide-and-Conquer MPSoC Architecture for MIMO Interference Cancellation
abstract
The fast-evolving standards of the wireless communication systems drive the demand for flexible baseband processing platforms. However, with the proliferation of MIMO technologies, traditional single-core-based solutions are hardly able to fulfill requirements with acceptable power and area cost. The reliance on multi-/many-core system is increasing. Different from the computation-limited single-core-based solutions, multi/many-core systems are often communication-limited. In this paper, aiming at MIMO interference cancellation algorithms, we propose a flexible master-slave-based multiprocessor system-on-chiparchitecture based on a systematically divide-and-conquer approach to optimize the communication problems from the application-, architecture- and programming-levels. First, a comprehensively analysis of several typical applications in terms of parallelism, communication patterns and computation patterns is presented. According to the analysis results, a low-complexity and flexible ad hoc point-to-point interconnected fine-grained programmable-element (f -PE) is proposed to execute the arithmetic calculation. In order to reduce the communication traffic, an f-PE-based slave-node is constructed to exploit the data and instruction localities of applications, and a master node that is used to schedule and serve data for the slave nodes is also integrated. Furthermore, to improve the ease of use of the architecture, a multiple instruction multiple data like programming model is adopted and an optimizing mapping strategy is developed. In order to show its flexibility potential, seven linear and nonlinear IC algorithms with distinct computation natures are implemented on the proposed architecture. Finally, the gate-level synthesis and postlayout results are presented to demonstrate the strength and weaknesses of our design.
Luechao Yuan, Cang Liu, Chuan Tang, Anupam Chattopadhyay, Gerd Ascheid, Zuocheng Xing
IEEE Trans. Very Large Scale Integr. Syst.7
2014 Flexible Virtual Channel Power-Gating for High-Throughput and Low-Power Network-on-Chip
abstract
Power-gating is a representative circuit level technique to mitigate leakage power. While in low-power Network-on-Chip (NoC) design, the former fine-grained power-gating methods will decrease network performance due to serial wake-up latency and head-of-line blocking. Therefore, we propose a flexible Virtual Channel (VC) management scheme for fine-grained power-gating to achieve high throughput and low-power. The proposed power-gating method with the early wake-up is evaluated by using some synthetic workloads. When compared with an optimized early wake-up power-gating technique, it can improve performance effectively in medium and high network loads, and increases the network throughput by 15.7%~44.1% for different synthetic loads, while keeps network power consumption as low as the optimized method. For the PARSEC application traces of token based protocol, it can significantly decrease packet latency by 20.3% on average, however only increases less than 3.6% peak power when compared with the optimized method.
Xiantuo Tang, Zuocheng Xing, Hengzhu Liu
DSD4
2013 Addressing Transient and Permanent Faults in NoC With Efficient Fault-Tolerant Deflection Router
abstract
Continuing decrease in the feature size of integrated circuits leads to increases in susceptibility to transient and permanent faults. This paper proposes a fault-tolerant solution for a bufferless network-on-chip, including an on-line fault-diagnosis mechanism to detect both transient and permanent faults, a hybrid automatic repeat request, and forward error correction link-level error control scheme to handle transient faults and a reinforcement-learning-based fault-tolerant deflection routing (FTDR) algorithm to tolerate permanent faults without deadlock and livelock. A hierarchical-routing-table-based algorithm (FTDR-H) is also presented to reduce the area overhead of the FTDR router. Synthesized results show that, compared with the FTDR router, the FTDR-H router can reduce the area by 27% in an 88 network. Simulation results demonstrate that under synthetic workloads, in the presence of permanent link faults, the throughput of an 8 8 network with FTDR and FTDR-H algorithms are 14% and 23% higher on average than that with the fault-on-neighbor (FoN) aware deflection routing algorithm and the cost-based deflection routing algorithm, respectively. Under real application workloads, the FTDR-H algorithm achieves 20% less hop counts on average than that of the FoN algorithm. For transient faults, the performance of the FTDR router can achieve graceful degradation even at a high fault rate. We also implement the fault-tolerant deflection router which can achieve 400 MHz in TSMC 65-nm technology.
Chaochao Feng, Zhonghai Lu, Axel Jantsch, Minxuan Zhang, Zuocheng Xing
IEEE Trans. Very Large Scale Integr. Syst.5
2011 Accurate and Simplified Prediction of AVF for Delay and Energy Efficient Cache Design
Anguo Ma, Zuocheng Xing
J. Comput. Sci. Technol.3
2009 Performance Optimization Strategies of High Performance Computing on GPU
Anguo Ma, Xiaoqiang Ni, Yuxing Tang, Zuocheng Xing
APPT6
2009 Fei Teng 64 Stream Processing System: Architecture, Compiler, and Programming
abstract
The stream architecture is a novel microprocessor architecture with wide application potential. It is critical to study how to use the stream architecture to accelerate scientific computing programs. However, existing stream processors and stream programming languages are not designed for scientific computing. To address this issue, we design and implement a 64-bit stream processor, Fei Teng 64 (FT64), which has a peak performance of 16 Gflops. FT64 supports two kinds of communications, message passing and stream communications, based on which, an interconnection architecture is designed for a FT64-based high-performance computer. This high-performance computer contains multiple modules, with each module containing eight FT64s. We also design a novel stream programming language, stream Fortran 95 (SF95), together with the compiler SF95 compiler, so as to facilitate the development of scientific applications. We test nine typical scientific application kernels on our FT64 platform to evaluate this design. The results demonstrate the effectiveness and efficiency of FT64 and its compiler for scientific computing.
Xuejun Yang, Xiaobo Yan, Zuocheng Xing, Yu Deng 0001, Jing Du 0002, Ying Zhang 0032
IEEE Trans. Parallel Distributed Syst.3
2007 A 64-bit stream processor architecture for scientific applications
abstract
Stream architecture is a novel microprocessor architecture with wide application potential. But as for whether it can be used efficiently in scientific computing, many issues await further study. This paper first gives the design and implementation of a 64-bit stream processor, FT64 (Fei Teng 64), for scientific computing. The carrying out of 64-bit extension design and scientific computing oriented optimization are described in such aspects as instruction set architecture, stream controller, micro controller, ALU cluster, memory hierarchy and interconnection interface here. Second, two kinds of communications as message passing and stream communications are put forward. An interconnection based on the communications is designed for FT64-based high performance computers. Third, a novel stream programming language, SF95 (Stream FORTRAN95), and its compiler, SF95Compiler (Stream FORTRAN95 Compiler), are developed to facilitate the development of scientific applications. Finally, nine typical scientific application kernels are tested and the results show the efficiency of stream architecture for scientific computing.
Xuejun Yang, Xiaobo Yan, Zuocheng Xing, Yu Deng 0001, Ying Zhang 0032
ISCA3