EDBT 2026 Demo / reviewers in the wild / expert
Jianxun Yang
dblp:206/8058
· DBLP profile ↗
6ranked-venue papers
3as first author
5since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MCBP: A Memory-Compute Efficient LLM Inference Accelerator Leveraging Bit-Slice-enabled Sparsity and Repetitiveness
Huizheng Wang, Zichuan Wang, Zhiheng Yue, Yousheng Long, Taiquan Wei, Jianxun Yang, Yang Wang 0089, Chao Li 0009, Shaojun Wei, Yang Hu 0001, Shouyi Yin |
MICRO | 6 |
| 2024 | A Fast Solution for the Eikonal Equation Based on Quadratic Function in Weakly Tilted Transversely Isotropic MediaabstractCalculating the first-arrival traveltimes of quasi-compressional (qP) waves has important applications in geophysics. In practice, geophysical problems often involve extensive calculations of traveltimes, especially when dealing with a large number of source-receiver pairs in complex geological structures. This necessitates a high level of computational efficiency. However, it is challenging to solve the anisotropic eikonal equation due to its complex nonlinearity. Conventional approaches to solving this equation in tilted transversely isotropic (TTI) media involve the use of various iterative algorithms combined with a fast sweeping method (FSM). When tackling large-scale geophysical problems, optimizing efficiency and minimizing the number of iterations is a crucial challenge. To address this problem, we developed a fitting algorithm to solve the anisotropic eikonal equation. We analyze the corresponding slowness surface equations of the qP wave in a local solver and observe a small difference between the slowness function and its quadratic fitting function. Therefore, we propose using the quadratic function to approximate the slowness equation based on three known traveltime points in the local solver. Subsequently, we obtain two roots from the fitting quadratic equation as approximate solutions. After obtaining the traveltime solutions, we check the traveltime causality to preserve the solution satisfying this condition. Three numerical tests are used to further illustrate the validity and computational efficiency of the proposed algorithm in computing traveltimes for 2-D/3-D weakly TTI models, which has been improved by more than three times compared with the conventional method. Yongming Lu, Jianxun Yang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | GQNA: Generic Quantized DNN Accelerator With Weight-Repetition-Aware Activation AggregatingabstractQuantization is a prominent approach to compress model sizes of deep neural networks (DNNs), which clusters high-precision weights into a smaller set of quantization levels and represents high-precision weights by low-precision indexes. To achieve the same accuracy, nonuniform quantized DNNs (NUQ-DNNs) with unequal quantization intervals need lower index precision than uniform quantized DNNs (UQ-DNNs) with equal intervals, achieving smaller model sizes. Hence, deploying NUQ-DNNs on accelerators costs less on- and off-chip memory accesses than UQ-DNNs, which are more valuable for edge devices. However, accelerating NUQ-DNNs is nontrivial, since weight indexes cannot be directly used for computations. Previous NUQ-DNN accelerators adopt standard convolutions by decoding weight indexes into actual-weights multiplied with activations, causing abundant look-up overhead and redundant computations. In this work, we propose a weight-repetition-aware activation aggregating (WPAA) convolution approach to accelerate inference of variable-precision NUQ- and UQ-DNNs. By merging convolutions of multiple kernels, WPAA requires no look-up operation and removes redundant computations. Based on WPAA, we design a generic quantized DNN accelerator (GQNA). Furthermore, we propose a layer-adaptive kernel-reordering merging scheme to off-line adjust merging order of kernels for minimizing energy consumption of GQNA. Implemented under TSMC 28-nm technology, GQNA achieves 31.9 and 32.6 TOPS/W energy efficiency for 1-b UQ- and NUQ-VGG-16, respectively. Jianxun Yang, Fengbin Tu, Yiqi Wang 0005, Leibo Liu, Shaojun Wei, Shouyi Yin |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2022 | Reconstruction of the S-Wave Velocity via Mixture Density Networks With a New Rayleigh Wave Dispersion FunctionabstractHow to determine the velocity of an S-wave from measured seismic data is an important topic in seismology. A modern technique for obtaining the S-wave velocity is to solve an inverse problem so that the simulated dispersion curve (the relation between the frequencies and phase velocities) coincides with the actual experimental results. In this work, by using the seismic impedance tensor, we propose a new function describing Rayleigh wave dispersion in the layered medium model of the Earth, which offers an efficient way to compute the dispersion curve. With this newly established forward model, based on mixture density networks (MDNs), we develop a physics-informed neural network, named MDN based on a physics informed forward model (FW-MDN), to estimate the S-wave velocity from dispersion curves. The FW-MDN method deals with the nonuniqueness issue encountered in the inversion of dispersion curves for the crust and upper mantle models, and attains satisfactory performance on an artificial dataset with various noise structures. Numerical simulations are performed to show that the FW-MDN offers easy calculation, efficient computation, and high precision for model characterization. Jianxun Yang, Ye Zhang 0017 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | FuseKNA: Fused Kernel Convolution based Accelerator for Deep Neural NetworksabstractBit-serial computation has been a prevailing convolution method to accelerate varying-precision DNNs by slicing a multi-bit data into multiple 1-bit data and transforming a multiplication into multiple additions, where additions of zero bits are ineffectual, while additions of non-zero bits are repetitive since multiple kernels are quite possible to possess non-zero bits at the same kernel positions. Previous bit-serial accelerators only remove ineffectual additions by skipping computation of zero bits, however, repetitive additions are unable to be eliminated since they compute convolution of each kernel independently. In this work, we propose fused kernel convolution algorithm to eliminate both ineffectual and repetitive additions in bit-serial computation by exploiting bit repetition and bit sparsity in weights, for both convolutional and fully-connected layers. It unifies convolutions of multiple kernels into convolution of one fused kernel by firstly grouping additions into different patterns and secondly reconstructing convolution results, minimizing addition count. Meantime, the memory accesses of activations and partial sums are decreased due to less convolution count. Then a fused kernel convolution based accelerator, FuseKNA, is designed with compact compute logic, which fully exploits value sparsity of activations and bit sparsity of weights. Benchmarked with a set of mainstream DNNs, FuseKNA improves performance by $4.47 \times$, $2.31 \times$ and $1.81 \times$, energy efficiency by $4.13 \times$, $3.06 \times$ and $2.53 \times$ over state-of-the-art Stripes, Pragmatic and Bit-Tactical. Jianxun Yang, Zhuangzhi Liu, Leibo Liu, Shaojun Wei, Shouyi Yin |
HPCA | 1 |
| 2017 | A two-stage variation-aware task mapping scheme for fault-tolerant multi-core Network-on-ChipsabstractWith technology scaling, process variations influence the performance, power and reliability significantly, especially for multi-core and many-core systems. Faulty-tolerant multi-core architectures are paid widely attention to, which integrate redundant cores to improve the manufacturing yield. In this paper, a two-stage variation-aware task mapping scheme is proposed for multi-core NoCs with redundant cores. Firstly, a static genetic task mapping algorithm is presented to generate multiple task mapping solutions to cover a maximum range of chips. Then, at runtime, one optimal mapping solution is selected, and logical cores are mapped to physical available cores. Both core asymmetry and topology changes are considered in the proposed approach. Experimental results demonstrate that the proposed approach increases the performance yield by 56%, and the communication cost is reduced by 11.3%. Jianxun Yang, Chengbo Xue |
ISCAS | 2 |