Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Xianyi Zhang

dblp:63/7449 · DBLP profile ↗
← Back
14ranked-venue papers
2as first author
3since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
High-performance computing · 77% Processor architecture and microarchitecture · 23%
Software engineering, system software, and programming languages
1 paper
Compilers and program optimization · 50% Program synthesis and code generation · 50%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Compilers and program optimization
code generation
0.212013
AUGEM: automatically generate high performance dense linear algebra kernels on x86 CPUs · SC 2013
Program synthesis and code generation › generative programming
template-based code generation
0.212013
AUGEM: automatically generate high performance dense linear algebra kernels on x86 CPUs · SC 2013
High-performance computing › linear algebra library
BLAS
0.212013
AUGEM: automatically generate high performance dense linear algebra kernels on x86 CPUs · SC 2013
High-performance computing › numerical linear algebra
dense linear algebra
0.212013
AUGEM: automatically generate high performance dense linear algebra kernels on x86 CPUs · SC 2013
Processor architecture and microarchitecture
instruction set architecture
0.012013
AUGEM: automatically generate high performance dense linear algebra kernels on x86 CPUs · SC 2013
Processor architecture and microarchitecture › SIMD
SIMD instructions
0.012013
AUGEM: automatically generate high performance dense linear algebra kernels on x86 CPUs · SC 2013

Methods — techniques the papers use, named apart from their topics

template-based optimization · 0.3parameterized code templates · 0.3low-level c optimization · 0.3
YearPublicationVenuePosition
2025 Integrating contrastive learning and adversarial learning on graph denoising encoder for recommendation
Wei Zhou 0028, Xianyi Zhang, Junhao Wen 0001, Xibin Wang
Expert Syst. Appl.2
2025 Multi-task intent recommendation based on dynamic and static intent integration and disentanglement
abstract
Recommender systems that mine users’ intentions to explore their potential interaction preferences have received increasing attention. However, the existing research on intent recommendation has some limitations. On one hand, the extant studies only consider users’ historical interaction information and the sparse interactions cannot reflect users’ potential interaction intention; on the other hand, they do not consider the changes in users’ kinematic and static intentions over time and the importance of users’ intention disentanglement representation, which makes it impossible for the general intention recommendation model to obtain a better representation of the intention. We propose a multitask recommendation model with dynamic and static intent integration and de-entanglement. The model mines users’ dynamic and static intents and then combines them with regularization to model the independence of the intents, encouraging the differences between the intents. Meanwhile, to further alleviate the data sparsity problem, this study additionally constructs user–user and item–item graphs using four different similarity measures, such as cosine similarity and mutual information, applies graph convolutional networks to learn about the three graphs, and then captures the complementarity between different graphs using a graph-level cross-attention mechanism. Extensive comparative experiments and ablation studies on three public datasets demonstrate that DSI-ID consistently outperforms all baseline methods, achieving a 3.5%–7.4% improvement in recommendation performance over the best baseline.
Xianyi Zhang, Junhao Wen 0001
Intell. Data Anal.2
2023 An Energy Control Strategy Based on Adaptive Fuzzy Logic for Onboard Hybrid Energy Storage System
abstract
This paper proposes an energy control strategy based on adaptive fuzzy logic for onboard hybrid energy storage system (HESS) with lithium-ion batteries (LIB) and electric double-layer capacitors (EDLC). Firstly, adaptive fuzzy logic energy control method for the system is proposed. The fuzzy rules are modified and reorganized according to the system deviation and deviation change rate to improve the energy-saving and voltage-stabilizing effect. Secondly, “soft connection out” and dead zone control method are implemented in closed-loop control to improve system stability and reduces voltage oscillation. Finally, the proposed strategy is compared with the classical strategy through simulation, where it shows better performance in the increase of 0.62% in voltage stabilization rate and 0.59% in energy conservation rate.
Tao Peng 0010, Rongchun Wan, Chao Yang 0017, Jinqiu Gao, Xianyi Zhang
IECON6
2016 The BLIS Framework: Experiments in Portability
abstract
BLIS is a new software framework for instantiating high-performance BLAS-like dense linear algebra libraries. We demonstrate how BLIS acts as a productivity multiplier by using it to implement the level-3 BLAS on a variety of current architectures. The systems for which we demonstrate the framework include state-of-the-art general-purpose, low-power, and many-core architectures. We show, with very little effort, how the BLIS framework yields sequential and parallel implementations that are competitive with the performance of ATLAS, OpenBLAS (an effort to maintain and extend the GotoBLAS), and commercial vendor implementations such as AMD’s ACML, IBM’s ESSL, and Intel’s MKL libraries. Although most of this article focuses on single-core implementation, we also provide compelling results that suggest the framework’s leverage extends to the multithreaded domain.
Field G. Van Zee, Tyler M. Smith, Bryan Marker, Tze Meng Low, Robert A. van de Geijn, Francisco D. Igual, Mikhail Smelyanskiy, Xianyi Zhang, Michael Kistler, Vernon Austel, John A. Gunnels, Lee Killough
ACM Trans. Math. Softw.8
2014 Optimizing and Scaling HPCG on Tianhe-2: Early Experience
Xianyi Zhang, Chao Yang 0002, Fangfang Liu 0004, Yiqung Liu 0005, Yutong Lu
ICA3PP (1)1
2014 Accelerating HPCG on Tianhe-2: A hybrid CPU-MIC algorithm
abstract
In this paper, we propose a hybrid algorithm to enable and accelerate the High Performance Conjugate Gradient (HPCG) benchmark on a heterogeneous node with an arbitrary number of accelerators. In the hybrid algorithm, each subdomain is assigned to a node after a three-dimensional domain decomposition. The subdomain is further divided to several regular inner blocks and an outer part with a flexible inner-outer partitioning strategy. Each inner task is assigned to a MIC device and the size is adjustable to adapt the accelerator's computational power. The only outer part is assigned to CPU and the thickness of boundary size is also adjustable to maintain load balance between CPU and MICs. By properly fusing the computational kernels with preceding ones, we present an asynchronous data transfer scheme to better overlap local computation with the PCI-express data transfer. All basic HPCG kernels, especially the time-consuming sparse matrix-vector multiplication (SpMV) and the symmetric Gauss-Seidel relaxation (SymGS), are extensively optimized for both CPU and MIC, on both algorithmic and architectural levels. On a single node of Tianhe-2 which is composed of an Intel Xeon processor and three Intel Xeon Phi coprocessors, we successfully obtain an aggregated performance of 50.2 Gflops, which is around 1.5% of the peak performance.
Yiqung Liu 0005, Xianyi Zhang, Chao Yang 0002, Fangfang Liu 0004, Yutong Lu
ICPADS2
2014 Memory Efficient Two-Pass 3D FFT Algorithm for Intel® Xeon PhiTM Coprocessor
Yiqung Liu 0005, Yan Li 0005, Yunquan Zhang, Xianyi Zhang
J. Comput. Sci. Technol.4
2013 AUGEM: automatically generate high performance dense linear algebra kernels on x86 CPUs
abstract
Basic Liner algebra subprograms (BLAS) is a fundamental library in scientific computing. In this paper, we present a template-based optimization framework, AUGEM, which can automatically generate fully optimized assembly code for several dense linear algebra (DLA) kernels, such as GEMM, GEMV, AXPY and DOT, on varying multi-core CPUs without requiring any manual interference from developers. In particular, based on domain-specific knowledge about algorithms of the DLA kernels, we use a collection of parameterized code templates to formulate a number of commonly occurring instruction sequences within the optimized low-level C code of these DLA kernels. Then, our framework uses a specialized low-level C optimizer to identify instruction sequences that match the pre-defined code templates and thereby translates them into extremely efficient SSE/AVX instructions. The DLA kernels generated by our template-based approach surpass the implementations of Intel MKL and AMD ACML BLAS libraries, on both Intel Sandy Bridge and AMD Piledriver processors.
Xianyi Zhang, Yunquan Zhang, Qing Yi
SC2
2012 Model-driven Level 3 BLAS Performance Optimization on Loongson 3A Processor
abstract
Every mainstream processor vendor provides an optimized BLAS implementation for its CPU, as BLAS is a fundamental math library in scientific computing. The Loongson 3A CPU is a general-purpose 64-bit MIPS64 quad-core processor, developed by the Institute of Computing Technology, Chinese Academy of Sciences. To date, there has not been a sufficiently optimized BLAS on the Loongson 3A CPU. The purpose of this research is to optimize level 3 BLAS performance on the Loongson 3A CPU. We analyzed the Loongson 3A architecture and built a performance model to highlight the key point, L1 data cache misses, which is different from level 3 BLAS optimization on the mainstream x86 CPU. Therefore, we employed a variety of methods to avoid L1 cache misses in single thread optimization, including cache and register blocking, the Loongson 3A 128-bit memory accessing extension instructions, software prefetching, and single precision floating-point SIMD instructions. Furthermore, we improved parallel performance by reducing bank conflicts among multiple threads in the shared L2 cache. We created an open source BLAS project, OpenBLAS, to demonstrate the performance improvement on the Loongson 3A quad-core processor.
Xianyi Zhang, Yunquan Zhang
ICPADS1
2011 CRSD: Application Specific Auto-tuning of SpMV for Diagonal Sparse Matrices
Xiangzheng Sun, Yunquan Zhang, Guoping Long, Xianyi Zhang, Yan Li 0005
Euro-Par (2)5
2011 Optimizing SpMV for Diagonal Sparse Matrices on GPU
abstract
Sparse Matrix-Vector multiplication (SpMV) is an important computational kernel in scientific applications. Its performance highly depends on the nonzero distribution of sparse matrices. In this paper, we propose a new storage format for diagonal sparse matrices, defined as Compressed Row Segment with Diagonal-pattern (CRSD). In CRSD, we design diagonal patterns to represent the diagonal distribution. As the Graphics Processing Units (GPUs) have tremendous computation power and OpenCL makes them more suitable for the scientific computing, we implement the SpMV for CRSD format on the GPUs using OpenCL. Since the OpenCL kernels are complied at runtime, we design the code generator to produce the codelets for all diagonal patterns after storing matrices into CRSD format. Specifically, the generated codelets already contain the index information of nonzeros, which reduces the memory pressure during the SpMV operation. Furthermore, the code generator also utilizes property of memory architecture and thread schedule on the GPUs to improve the performance. In the evaluation, we select four storage formats from prior state-of-the-art implementations (Bell and Garland, 2009) on GPU. Experimental results demonstrate that the speedups reach up to 1.52 and 1.94 in comparison with the optimal implementation of the four formats for the double and single precision respectively. We also evaluate on a two-socket quad-core Intel Xeon system. The speedups reach up to 11.93 and 12.79 in comparison with CSR format under 8 threads for the double and single precision respectively.
Xiangzheng Sun, Yunquan Zhang, Xianyi Zhang
ICPP4
2009 RCC: A New Programming Language for Reconfigurable Computing
abstract
Reconfigurable computing is one of the most important developing directions of future high-performance computing, it combines the flexibility of general processors and the high efficiency of ASIC. Reconfigurable computing technology can make use of system resources sufficiently and exert the efficiency of the applications. Reconfigurable programming environment is the key to popularize reconfigurable computing technology. Software engineers can program FPGA by using reconfigurable programming language, and moreover, the reconfigurable compiler can provide an architecture-independent developing platform for high-level language programmers in order that they can use the reconfigurable computing system flexibly. Three typical reconfigurable computing systems are analyzed in this paper, including SGI RASC RC100, Cray XD1 and SRC-6E. Some typical reconfigurable programming environments are introduced, such as the Mitrion-C platform, the Impulse-C platform, etc. However, neither of them is convenient enough for software engineers to program FPGA. In this paper, a more useful reconfigurable programming language called RCC (Reconfigurable Computing C) is proposed. It is based on the subset of ANSI-C and extends some data types and functions. A RCC compiler is built to translate the RCC applications into VHDL description. Users can program FPGA easily by using the RCC language if they are familiar with ANSI-C. On FPGA simulator, we test 3 applications and get about 19X speedup. Furthermore, it is easier to use than some typical reconfigurable programming languages.
Fengbin Qi, Xianyi Zhang, Xingquan Mao
HPCC2
2009 QuantWiz: A Parallel Software Package for LC-MS-based Label-Free Protein Quantification
abstract
Nowadays proteomics becomes more and more popular in life science. Protein quantification, especially based on mass spectrometry (short for MS) method, is perceived as an essential part of research on proteomics. There have been some algorithms and software for protein quantification based on MS. But they have difficulties on portability, applicability and longtime running. To solve these problems, we developed a new domestic parallel software package called QuantWiz for high performance liquid chromatography (short for LC)-MS-based label-free protein quantification. In this paper, we described the framework design and prototype development of this high performance software package firstly. Also, user interface developed for the visualization of QuantWiz is introduced. Finally, we showed implementation of the parallelization version and performance of some experiments on this software package.
Yunquan Zhang, Xianyi Zhang, Xiangzheng Sun, Zelin Hu, Sujun Li
HPCC3
2009 Early Performance Evaluation of Dawning 5000A and DeepComp 7000
abstract
In this paper, we present our early performance evaluation results with the NPB benchmark and two scientific computing applications program, i.e., a HFFT package developed by our lab and a CFDO application software, on two 100 Teraflops-scale Dawning 5000A and DeepComp 7000. We compared the NPB performance evaluation results of Dawning 5000A and DeepComp 7000, with their corresponding predecessor, Dawning 4000A and DeepComp 6800, which demonstrating their performance improvements across a variety of NPB benchmark problems. From our evaluation results, we can find that the NPB benchmark can keep its scalability up to 16384 cores on Dawning 5000A while on DeepComp 7000 this number becomes 4096. We also find that the HFFT scales well over 2048 cores while the CFDO scales well over 512 cores on Dawning 5000A.
Yunquan Zhang, Jiachang Sun, Xianyi Zhang
ICPADS5