Weizhi Xu 0001

dblp:71/8419 · DBLP profile ↗
← Back
33ranked-venue papers
6as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 1 first-author · 13 since 2021Systems, architecture and hardware · 11 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-authorDatabases, data management, data science and information retrieval · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Dynamic Phased Pruning with Similarity-Aware Filter Selection
Yingxian Jiang, Yaoyao Yan, Feiyue Diao, Weizhi Xu 0001, Hui Yu 0010
ISCAS5
2026 TensorFHE+: Fully Homomorphic Encryption Acceleration Based on Linear Algebra
abstract
Fully Homomorphic Encryption (FHE) enables encrypted data processing on untrusted cloud servers, crucial for privacy-sensitive applications. Despite its potential, performance overheads (about 10, 000× slower) limit adoption. ASIC accelerators outperform GPUs/FPGAs by optimizing specific operations but rely on costly 7nm processes and large on-chip memory, hindering cost-effective deployment. Balancing efficiency with manufacturing constraints remains critical. This paper presents TensorFHE+, a GPU-optimized FHE acceleration framework leveraging Tensor Cores to accelerate Number Theoretic Transform (NTT) operations. Key innovations include: 1) Decomposing CKKS kernels into vector/matrix operations for hardware utilization; 2) Vectorized modulo arithmetic; 3) Data layout optimization for memory efficiency. Evaluated on NVIDIA A100, TensorFHE+ outperforms TensorFHE [1] by 1.44× in average (up to 1.69× on ResNet-20) and surpasses prior GPU implementations [2], [3]. The design also demonstrates compatibility with commercial linear algebra accelerators, enabling efficient FHE deployment.
Yintai Sun, Shengyu Fan, Zhenhua Yin, Xinkai Song, Xing Hu 0001, Zidong Du, Qi Guo 0001, Weizhi Xu 0001, Rui Hou 0001, Dan Meng 0002, Song Bian 0001, Mingzhe Zhang 0005
IEEE Trans. Computers8
2025 Attribution-Driven Adaptive Token Pruning for Transformers
abstract
Transformers have been widely adopted in natural language processing, computer vision, and other domains due to their exceptional performance across a variety of tasks. However, the computational cost of Transformers is prohibitively high, particularly when handling long input sequences, significantly increasing both training and inference time. Although various token pruning methods have been proposed to reduce the computational burden of Transformers, most approaches overlook critical differences in sequences in terms of length and complexity, leading to suboptimal compression efficiency. In this paper, we propose AD-TP, an Attribution-Driven Adaptive Token Pruning method designed to retain only the most informative tokens. We analyze the performance of using accumulated attention values to measure token importance and find that attention values do not accurately reflect the actual contribution of each token to text understanding. Additionally, we observe significant variations in the length and complexity of different sequences within the dataset. Based on these insights, we adopt Integrated Gradients to evaluate token importance and introduce a lightweight adaptive token retainer module that dynamically generates pruning configurations for each input sequence. In addition, we incorporate both teacher supervision and self-supervised learning objectives to enhance the training efficiency, accuracy, and robustness of the model. Experiments conducted on GLUE, SQuAD, and 20News demonstrate that AD-TP outperforms state-of-the-art token pruning and model compression methods in both accuracy and computational efficiency. On GLUE, AD-TP reduces FLOPs by an average of 7.8× while improving performance by 0.6%.
Yaoyao Yan, Hui Yu 0010, Weizhi Xu 0001
NeurIPS3
2025 STP: Special token prompt for parameter-efficient tuning of pre-trained language models
Yaoyao Yan, Hui Yu 0010, Fang'ai Liu, Weizhi Xu 0001
Expert Syst. Appl.6
2025 DCHF_T: A multi-dimensional adaptive compression approach for transformer-based models
Yaoyao Yan, Hui Yu 0010, Dianjie Lu, Weizhi Xu 0001, Fang'ai Liu
Neurocomputing7
2025 Label-specific multi-label text classification based on dynamic graph convolutional networks
Yaoyao Yan, Fang'ai Liu, Kenan Liu, Weizhi Xu 0001, Xuqiang Zhuang
Soft Comput.4
2024 Obtaining Optimal Spiking Neural Network in Sequence Learning via CRNN-SNN Conversion
Jiahao Su, Kang You, Weizhi Xu 0001, Zhezhi He
ICANN (10)4
2024 Efficient Selection Based on Integrated Information for Dialogue State Tracking
abstract
Dialogue State Tracking (DST) is a critical component in Task-Oriented Dialogue (TOD) systems, responsible for generating the dialogue state at each turn. Current approaches often struggle with complicated conversational contexts, primarily due to issues of information redundancy and insufficiency, which adversely affects accuracy. To address these challenges, we propose a novel approach termed Selection based on Integrated Information (SII). This method comprises three key components: an Information Integrator, which distills core information from the dialogue; an Information Selector, which identifies the most pertinent core information for each slot; and a State Predictor, which executes predictions based on the selected information. By focusing on selected information, SII demonstrates enhanced performance, achieving joint goal accuracies of 55.44% and 59.89% on the MultiWOZ2.0 and MultiWOZ2.1 datasets, respectively.
Hongyun Du, Jikun Dong, Shengyu Fan, Shengjie Jia, Feiyue Diao, Jiran Zhu, Hui Yu 0010, Weizhi Xu 0001
IJCNN9
2024 Semantic-aware enhancement: Integrating semantic compensation with 3-Dimensional Lookup Tables for low-light image enhancement
Weizhi Xu 0001, Chen Lyu 0001
Eng. Appl. Artif. Intell.2
2023 TensorFHE: Achieving Practical Computation on Encrypted Data Using GPGPU
abstract
In the cloud computing era, privacy protection is becoming pervasive in a broad range of applications (e.g., machine learning, data mining, etc). Fully Homomorphic Encryption (FHE) is considered the perfect solution as it enables privacy-preserved computation on untrusted servers. Unfortunately, the prohibitive performance overhead blocks the wide adoption of FHE (about 10, 000× slower than the normal computation). As heterogeneous architectures have gained remarkable success in several fields, achieving high performance for FHE with specifically designed accelerators seems to be a natural choice. Until now, most FHE accelerators have focused on efficiently implementing one FHE operation at a time based on ASIC and with significantly higher performance than GPU and FPGA. However, recent state-of-the-art FHE accelerators rely on an expensive and large on-chip storage and a high-end manufacturing process (i.e., 7nm), which increase the cost of FHE adoption.In this paper, we propose TensorFHE, an FHE acceleration solution based on GPGPU for real applications on encrypted data. TensorFHE utilizes Tensor Core Units (TCUs) to boost the computation of Number Theoretic Transform (NTT), which is the part of FHE with highest time-cost. Moreover, TensorFHE focuses on performing as many FHE operations as possible in a certain time period rather than reducing the latency of one operation. Based on such an idea, TensorFHE introduces operation-level batching to fully utilize the data parallelism in GPGPU. We experimentally prove that it is possible to achieve comparable performance with GPGPU as with state-of-the-art ASIC accelerators. TensorFHE performs 913 KOPS and 88 KOPS for NTT and HMULT (key FHE kernels) within NVIDIA A100 GPGPU, which is 2.61× faster than state-of-the-art FHE implementation on GPGPU; Moreover, TensorFHE provides comparable performance to the ASIC FHE accelerators, which makes it even 2.9× faster than the F1+ with a specific workload. Such a pure software acceleration based on commercial hardware with high performance can open up usage of state-of-the-art FHE algorithms for a broad set of applications in real systems.
Shengyu Fan, Weizhi Xu 0001, Rui Hou 0001, Dan Meng 0002, Mingzhe Zhang 0005
HPCA3
2023 Application of Data Encryption in Chinese Named Entity Recognition
Jikun Dong, Kaifang Long, Hui Yu 0010, Weizhi Xu 0001
ICANN (8)4
2023 Recurrent Update Representation Based on Multi-head Attention Mechanism for Joint Entity and Relation Extraction
Shengjie Jia, Jikun Dong, Kaifang Long, Jiran Zhu, Hongyun Du, Guijuan Zhang, Hui Yu 0010, Weizhi Xu 0001
ICONIP (13)8
2023 KSRE-CNER: A Knowledge and Semantic Relation Enhancement Framework for Chinese NER
Jikun Dong, Kaifang Long, Jiran Zhu, Hui Yu 0010, Zengzhen Shao, Weizhi Xu 0001
PRICAI (2)7
2023 Team Recruitment of Collaborative Crowdsensing under Joint Constraints of Willingness and Trust
abstract
Collaborative crowdsensing (CCS) requires the recruited team to collaborate closely to complete sensing tasks with high quality of service (QoS). The team recruitment of CCS is mainly influenced by the subjective willingness of participants and the objective trust evaluation of the sensing platform; that is, the higher the subjective mutual willingness to work together and the objective mutual trust among participants, the more efficiency with which the CCS tasks will be achieved. However, the existing research lacks comprehensive consideration of mutual willingness and mutual trust among recruited participants. This results in poor QoS. To address this problem, we propose a novel team recruitment method for CCS that jointly considers the willingness and trust to recruit optimal teams. First, we build a graph convolutional network‐based willingness‐trust network (GCN‐WTN) model for CCS to obtain mutual willingness and trust among participants more accurately. Second, we propose a willingness and trust‐based team recruitment (WT‐TR) method to recruit the optimal teams for CCS. This method introduces the consensus and similarity constraints into the willingness and trust networks to better meet the collaboration needs of CCS. Finally, we implement a recruitment simulation platform for CCS to simulate the team recruitment process and validate the effectiveness of our proposed method. The experimental results show that the teams recruited by the proposed method can significantly improve QoS for CCS.
Nianyun Song, Dianjie Lu, Chunyu Hu 0001, Weizhi Xu 0001, Guijuan Zhang
Int. J. Intell. Syst.4
2023 Tell me your position: Distantly supervised biomedical entity relation extraction using entity position marker
Jiran Zhu, Jikun Dong, Hongyun Du, Yanfang Geng, Shengyu Fan, Hui Yu 0010, Zengzhen Shao, Yaping Yang, Weizhi Xu 0001
Neural Networks10
2023 Accelerating Convolutional Neural Network by Exploiting Sparsity on GPUs
abstract
The convolutional neural network (CNN) is an important deep learning method, which is widely used in many fields. However, it is very time consuming to implement the CNN where convolution usually takes most of the time. There are many zero values in feature maps and filters, which leads to redundant calculations and memory accesses if dense methods are used to compute convolution. Many works recently have made use of sparsity to skip the calculations for zero values to reduce the inference time of the CNN. On the graphics processing unit platform, current works cannot fully exploit the sparsity of the feature map and achieve satisfactory performance. Therefore, we design a new parallel strategy to transform the feature map into a new storage format to avoid the redundant computation of zero values on graphics processing units. Also considering the sparsity in the feature map, we propose a fused storage format to combine the convolution operation with the following pooling operation, to further improve the performance. We carry out experiments with mainstream CNN models and achieve better performance compared with cuDNN and cuSPARSE. For VGG-19, ResNet-50, DenseNet-121, and RegNetX-16GF, 1.97×, 2.23×, 2.74×, and 1.58× speedups respectively are obtained over cuDNN. The speedups over cuSPARSE respectively are 2.10×, 1.83×, 2.35×, and 1.35× when only using the first method.
Weizhi Xu 0001, Yintai Sun, Shengyu Fan, Hui Yu 0010, Xin Fu 0001
ACM Trans. Archit. Code Optim.1
2023 Deep Neural Network with Embedding Fusion for Chinese Named Entity Recognition
abstract
Chinese Named Entity Recognition (NER) is an essential task in natural language processing, and its performance directly impacts the downstream tasks. The main challenges in Chinese NER are the high dependence of named entities on context and the lack of word boundary information. Therefore, how to integrate relevant knowledge into the corresponding entity has become the primary task for Chinese NER. Both the lattice LSTM model and the WC-LSTM model did not make excellent use of contextual information. Additionally, the lattice LSTM model had a complex structure and did not exploit the word information well. To address the preceding problems, we propose a Chinese NER method based on the deep neural network with multiple ways of embedding fusion. First, we use a convolutional neural network to combine the contextual information of the input sequence and apply a self-attention mechanism to integrate lexicon knowledge, compensating for the lack of word boundaries. The word feature, context feature, bigram feature, and bigram context feature are obtained for each character. Second, four different features are used to fuse information at the embedding layer. As a result, four different word embeddings are obtained through cascading. Last, the fused feature information is input to the encoding and decoding layer. Experiments on three datasets show that our model can effectively improve the performance of Chinese NER.
Kaifang Long, Zengzhen Shao, Yanfang Geng, Yintai Sun, Weizhi Xu 0001, Hui Yu 0010
ACM Trans. Asian Low Resour. Lang. Inf. Process.7
2022 Multi-attention deep neural network fusing character and word embedding for clinical and biomedical concept extraction
Shengyu Fan, Hui Yu 0010, Xiaoya Cai, Yanfang Geng, Guangzhen Li, Weizhi Xu 0001, Yaping Yang
Inf. Sci.6
2020 Character-level neural network model based on Nadam optimization and its application in clinical concept extraction
Lantian Li, Weizhi Xu 0001, Hui Yu 0010
Neurocomputing2
2019 Machine Translation Evaluation Metric Based on Dependency Parsing Model
abstract
Most of the syntax-based metrics obtain the similarity by comparing the sub-structures extracted from the trees of hypothesis and reference. These sub-structures cannot represent all the information in the trees because their lengths are limited. To sufficiently use the reference syntax information, a new automatic evaluation metric is proposed based on the dependency parsing model. First, a dependency parsing model is trained using the reference dependency tree for each sentence. Then, the hypothesis is parsed by this dependency parsing model and the corresponding hypothesis dependency tree is generated. The quality of hypothesis can be judged by the quality of the hypothesis dependency tree. Unigram F-score is included in the new metric so that lexicon similarity is obtained. According to experimental results, the proposed metric can perform better than METEOR and BLEU on system level and get comparable results with METEOR on sentence level. To further improve the performance, we also propose a combined metric which gets the best performance on the sentence level and on the system level.
Hui Yu 0010, Weizhi Xu 0001, Shouxun Lin, Qun Liu 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2018 A Clustering Algorithm of High-Dimensional Data Based on Sequential Psim Matrix and Differential Truncation
Gongming Wang, Wenfa Li, Weizhi Xu 0001
ICA3PP (2)3
2016 CWFP: Novel Collective Writeback and Fill Policy for Last-Level DRAM Cache
abstract
Stacked DRAM used as the last-level caches (LLCs) in multicore systems delivers performance enhancement due to its capacity benefit. While the performance of LLC depends heavily upon its block replacement policy, the conventional replacement policy needs redesigning to exploit the best of DRAM cache and avoid its drawbacks. The existing DRAM cache insertion policy blindly forwards victim lines replaced from L3 to the off-chip memory, regardless of the potential for increased hits by placing a fraction of them in the DRAM cache. Nevertheless, a naïve design that steers all dirty victims to the DRAM cache introduces excessive writeback traffic, which aggravates capacity misses and DRAM interference. To leverage insertions in terms of writeback or fill requests, we propose a collective writeback and fill policy that adapts to the distinct access patterns of heterogeneous workloads based on runtime misses and writeback efficiency. The synthesis result shows that the new policy has only a small hardware overhead. According to the experimental results on the gem5 simulator, the new policy increases harmonic instruction per cycle throughput by 18%, 11.6%, and 11.7%, respectively, compared with the Always Writeback (AW)-Always Fill policy, Never Writeback Adaptive DRAM Placement policy, and AW Adaptive DRAM Placement policy on 64-MB alloy cache, while the improvement is 19.3%, 13.7%, and 14.5% on 64-MB MissMap cache.
Shouyi Yin, Weizhi Xu 0001, Leibo Liu, Shaojun Wei
IEEE Trans. Very Large Scale Integr. Syst.2
2015 An Effective TSV Self-Repair Scheme for 3D-Stacked ICs
abstract
Various types of defects are prone to be occurred inside the TSV during the manufacturing and bonding steps, thereby severely impacting the yield of 3D-stacked ICs. Moreover, several types of TSV defects are latent and may easily escape detection during the manufacturing test. However, these latent TSVs are prone to degrade during the field operation and may eventually become faulty and then destroy the entire 3D-stacked IC. To tackle the above problems, in this paper, we present an effective TSV self-repair scheme for 3D-stacked ICs. By designing redundant TSVs and a TSV self-repair architecture, the proposed scheme can effectively repair faulty TSVs detected by manufacturing test for improving the yield of 3D-stacked ICs. Moreover, the latent TSVS failed and then detected during the in-field operation can also be self-repaired, thereby elevating the 3D ICs' quality and reliability. Experimental results are presented to validate the proposed method.
Songwei Pei, Weizhi Xu 0001
ACM Great Lakes Symposium on VLSI6
2015 Memory bandwidth optimization of SpMV on GPGPUs
Chenggang Yan 0001, Hui Yu 0010, Weizhi Xu 0001, Yingping Zhang, Bochuan Chen, Zhu Tian, Jian Yin 0003
Frontiers Comput. Sci.3
2015 Corrigendum to "Fast and scalable lock methods for video coding on many-core architecture" [J. Visual Communication and Image Representation 25(7) (2014) 1758-1762]
Weizhi Xu 0001, Hui Yu 0010, Dianjie Lu, Fenglong Song, Xiaochun Ye, Songwei Pei, Dongrui Fan, Hongtao Xie 0001
J. Vis. Commun. Image Represent.1
2015 Corrigendum to "Fast and scalable lock methods for video coding on many-core architecture" [J. Visual Communication and Image Representation 25 (7) (2014) 1758-1762]
Weizhi Xu 0001, Hui Yu 0010, Dianjie Lu, Fenglong Song, Xiaochun Ye, Songwei Pei, Dongrui Fan, Hongtao Xie 0001
J. Vis. Commun. Image Represent.1
2014 Fast and scalable lock methods for video coding on many-core architecture
Weizhi Xu 0001, Hui Yu 0010, Dianjie Lu, Fenglong Song, Xiaochun Ye, Songwei Pei, Dongrui Fan, Hongtao Xie 0001
J. Vis. Commun. Image Represent.1
2014 Highly parallel GEMV with register blocking method on GPU architecture
Jian Yin 0003, Hui Yu 0010, Weizhi Xu 0001, Zhu Tian, Yingping Zhang, Bochuan Chen
J. Vis. Commun. Image Represent.3
2012 Supporting User-directed Fault Tolerance over Standard MPI
abstract
User-directed means the process of carrying out fault tolerance is dynamic and the fault tolerance mode is chosen by users based on application requirements. In this paper, we introduce a general scheme based on standard MPI to provide the user directed support for application level algorithmic fault tolerance. The user-directed fault tolerance plays the role as a connection between applications and algorithmic fault tolerance. As a case study, our scheme has been incorporated to HPL combined with a non-blocking ABFT technique. We have tested the functional availability of our scheme for fault tolerance in real circumstance. We also evaluated that when there is no failure occurring, our support only brings 2.5 percent overhead. When failure occurs, with our scheme, the scalability of algorithmic fault tolerance maintains well.
Zhimin Wu, Weizhi Xu 0001, Mingyu Chen 0001, Erlin Yao
ICPADS3
2012 Auto-Tuning GEMV on Many-Core GPU
abstract
GPUs provide powerful computing ability especially for data parallel algorithms. However, the complexity of the GPU system makes the optimization of even a simple algorithm difficult. Different parallel algorithms or optimization methods on a GPU often lead to very different performances. The matrix-vector multiplication routine for general dense matrices (GEMV) is a building block for many scientific and engineering computations. We find that the implementations of GEMV in CUBLAS 4.0 or MAGMA are not efficient, especially for small matrix or fat matrix (a matrix with small number of rows and large number of columns). In this paper, we propose two new algorithms to optimize GEMV on Fermi GPU. Instead of using only one thread, we use a warp to compute an element of vector y. We also propose a novel register blocking method to accelerate GEMV on GPU further. The proposed optimization methods for GEMV are comprehensively evaluated on the matrices with different sizes. Experiment results show that the new methods can achieve over 10x speedup for small square matrices and fat matrices compared to CUBLAS 4.0 or MAGMA, and the new register blocking method can also perform better than CUBLAS 4.0 or MAGMA for large square matrices. We also propose a performance-tuning framework on how to choose an optimal algorithm of GEMV for an arbitrary input matrix on GPU.
Weizhi Xu 0001, Zhiyong Liu 0002, Xiaochun Ye, Shuai Jiao, Fenglong Song, Dongrui Fan
ICPADS1
2012 A SAT-based diagnosis pattern generation method for timing faults in scan chains
abstract
Scan is a widely used DFT technique to improve test and diagnosis quality. However, failures on scan chain itself account for up to 30% of chip failures. In this paper, a SAT-based technique is proposed to generate patterns to diagnose four types of timing faults in scan chains. The proposed method can efficiently generate high quality diagnostic patterns while achieving high diagnosis resolution. Further more, the computation overhead of equivalent faults proving is reduced. Experimental results on ISCAS'89 benchmark circuits show that the proposed method can reduce at least 70% diagnostic patterns' volume and 60% CPU time compared with other works.
Lunkai Zhang, Weizhi Xu 0001, Dongrui Fan
ISCAS3
2012 Optimizing Sparse Matrix Vector Multiplication Using Cache Blocking Method on Fermi GPU
abstract
It is an important task to tune performance for sparse matrix vector multiplication (SpMV), but it is also a difficult task because of its irregularity. In this paper, we propose a cache blocking method to improve the performance of SpMV on the emerging GPU architecture. The sparse matrix is partitioned into many sub-blocks, which are stored in CSR format. With the blocking method, the corresponding part of vector x can be reused in the GPU cache, so the time spent on accessing the global memory for vector x is reduced heavily. Experimental results on GeForce GTX 480 show that SpMV kernel with the cache blocking method is 5x faster than the unblocked CSR kernel in the best case.
Weizhi Xu 0001, Hao Zhang 0009, Shuai Jiao, Fenglong Song, Zhiyong Liu 0002
SNPD1
2010 Efficient Address Mapping of Shared Cache for On-Chip Many-Core Architecture
Fenglong Song, Dongrui Fan, Zhiyong Liu 0002, Junchao Zhang 0004, Lei Yu 0012, Weizhi Xu 0001
Euro-Par (1)6