Xiangwei Li

dblp:120/9389 · DBLP profile ↗
← Back
17ranked-venue papers
9as first author
7since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 2 since 2021Systems, architecture and hardware · 6 · 4 first-author · 2 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 2 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Making vulnerability prediction more practical: Prediction, categorization, and localization
Xiang Chen 0005, Xiangwei Li, Yinxing Xue
Inf. Softw. Technol.3
2023 Prediction of Vulnerability Characteristics Based on Vulnerability Description and Prompt Learning
abstract
Identifying which vulnerabilities need to be prioritized is a long-term challenge in IT security, especially as the number of vulnerabilities grows. Faced with a large number of vulnerability reports, there is an urgent need for automated tools or models to assess the potential severity and exploitability of vulnerabilities. This will help security experts screen vulnerabilities that should be focused on. In this study, we aim to predict vulnerability severity and exploitability characteristics using only vulnerability descriptions. Some previous studies are based on traditional deep learning models, and their performance is relatively backward in the current era of pre-trained language models (PLMs). Therefore, we introduce a prompt learning method based on PLMs to predict vulnerability characteristics. The conventional fine-tuning PLMs method is difficult to make full use of the domain knowledge in PLMs and performs poorly with less training data. Unlike the fine-tuning paradigm, prompt learning imitates the pre-training process of PLM by reconstructing the task input and adding prompts, and uses the output of PLM itself as the prediction output. Combined with prompt ensembling and transfer learning, the performance of prompt learning in the above tasks is further improved. Our experiments show that prompt learning can make more effective use of the knowledge in PLMs. Compared with fine-tuning PLMs and other deep learning models, prompt learning based on BERT or RoBERTa achieves better performance in the above tasks. This advantage is more significant in predicting exploitability with few samples, which proves the ability of prompt learning in few-sample scenarios. In addition, prompt learning also shows the transferability between different tasks in the domain.
Xiangwei Li, Xiaoning Ren, Yinxing Xue, Zhenchang Xing, Jiamou Sun
SANER1
2023 Fixed-point FPGA Implementation of the FFT Accumulation Method for Real-time Cyclostationary Analysis
abstract
The spectral correlation density (SCD) is an important tool in cyclostationary signal detection and classification. Even using efficient techniques based on the fast Fourier transform (FFT), real-time implementations are challenging because of the high computational complexity. A key dimension for computational optimization lies in minimizing the wordlength employed. In this article, we analyze the relationship between wordlength and signal-to-quantization noise in fixed-point implementations of the SCD function. A canonical SCD estimation algorithm, the FFT accumulation method (FAM) using fixed-point arithmetic, is studied. We derive closed-form expressions for SQNR and compare them at wordlengths ranging from 14 to 26 bits. The differences between the calculated SQNR and bit-exact simulations are less than 1 dB. Furthermore, an HLS-based FPGA design is implemented on a Xilinx Zynq UltraScale+ XCZU28DR-2FFVG1517E RFSoC. Using less than 25% of the logic fabric on the device, it consumes 7.7 W total on-chip power and has a power efficiency of 12.4 GOPS/W, which is an order of magnitude improvement over an Nvidia Tesla K40 graphics processing unit (GPU) implementation. In terms of throughput, it achieves 50 MS/sec, which is a speedup of 1.6 over a recent optimized FPGA implementation.
Carol Jingyi Li, Xiangwei Li, Binglei Lou, Craig T. Jin, David Boland, Philip H. W. Leong
ACM Trans. Reconfigurable Technol. Syst.2
2023 A Scalable Systolic Accelerator for Estimation of the Spectral Correlation Density Function and Its FPGA Implementation
abstract
The spectral correlation density (SCD) function is the time-averaged correlation of two spectral components used for analyzing periodic signals with time-varying spectral content. Although the analysis is extremely powerful, it has not been widely adopted in real-time applications due to its high computational complexity. In this article, we present an efficient FPGA implementation of the FFT accumulation method (FAM) for estimating the SCD function and its alpha profile. The implementation uses a linear systolic array with a bi-directional datapath consisting of DSP-based processing elements (PEs) with a dedicated instruction schedule, achieving a PE utilization of 88.2%. The 128-PE implementation achieves a clock frequency in excess of 530 MHz and consumes 151K LUTs, 151K FFs, 264 BRAMs, 4 URAMs, and 1,054 DSPs, which is less than 36% of the logic fabric on a Zynq UltraScale+ XCZU28DR-2FFVG1517E RFSoC device. It has a modest 12.5W power consumption and an energy efficiency of 4,832 MOPS/W, which is 20.6× better than the published state-of-the-art GPU implementation. In terms of throughput, it achieves 15,340 windows/s (15,340 windows/s × 2,048 samples/window = 31.4 MS/s), which is a 4.65× improvement compared to the above-mentioned GPU implementation and 807× compared to an existing hybrid FPGA-GPU implementation.
Xiangwei Li, Douglas L. Maskell, Carol Jingyi Li, Philip H. W. Leong, David Boland
ACM Trans. Reconfigurable Technol. Syst.1
2021 Weakly supervised salient object detection via double object proposals guidance
abstract
Abstract The weakly supervised methods for salient object detection are attractive, since they greatly release the burden of annotating time‐consuming pixel‐wise masks. However, the image‐level annotations utilized by current weakly supervised salient object detection models are too weak to provide sufficient supervision for this dense prediction task. To this end, a weakly supervised salient object detection method is proposed via double object proposals guidance, which is generated under the supervision of double bounding boxes annotations. With the double object proposals, the authors' method is capable of capturing both accurate but incomplete salient foreground and background information, which contributes to generating saliency maps with uniformly highlighted saliency regions and effectively suppressed background. In addition, an unsupervised salient object segmentation method is proposed, taking advantage of the non‐parametric statistical active contour model (NSACM), for segmenting salient objects with complete and compact boundaries. Experiments on five benchmark datasets show that the authors' weakly supervised salient object detection approach consistently outperforms other weakly supervised and unsupervised methods by a considerable margin, and even has comparable performance to the fully supervised ones.
Zhiheng Zhou 0001, Yongfan Guo, Junchu Huang, Xiangwei Li
IET Image Process.5
2021 QoS-oriented joint optimization of concurrent scheduling and power control in millimeter wave mesh backhaul network
Zhongyu Ma, Jie Cao 0014, Qun Guo 0001, Xiangwei Li, Hongfeng Ma
J. Netw. Comput. Appl.4
2021 Asymmetric alignment joint consistent regularization for multi-source domain adaptation
Junyuan Shang, Chang Niu, Zhiheng Zhou 0001, Junchu Huang, Zhiwei Yang 0010, Xiangwei Li
Multim. Tools Appl.6
2020 High Throughput Accelerator Interface Framework for a Linear Time-Multiplexed FPGA Overlay
abstract
Coarse-grained FPGA overlays improve design productivity through software-like programmability and fast compilation. However, the effectiveness of overlays as accelerators is dependent on suitable interface and programming integration into a typically processor-based computing system, an aspect which has often been neglected in evaluations of overlays. We explore the integration of a time-multiplexed FPGA overlay over a server-class PCI Express interface. We show how this integration can be optimised to maximise performance, and evaluate the area overhead. We also propose a user-friendly programming model for such an overlay accelerator system.
Xiangwei Li, Kizheppatt Vipin, Douglas L. Maskell, Suhaib A. Fahmy, Abhishek Kumar Jain
ISCAS1
2020 Common-specific feature learning for multi-source domain adaptation
abstract
Multi‐source domain adaptation (MDA) aims to leverage knowledge from multiple source domains to improve the classification performance on target domains. Different degrees of distribution discrepancies between every two domains pose a huge challenge to MDA tasks. Most works focus on extracting features shared by all domains, which is critical but not enough to reduce distribution discrepancies. In this paper, we propose a method named as common‐specific feature learning (CSFL). Constituting a framework of feature learning, CSFL explores a subspace where the combination of common and specific features makes learned representations comprehensive. Based on this framework, we conduct a metric learning method for learning a discriminative feature representation. Considering redundant information caused by source domains is likely to hurt the performance, we impose an effective low‐rank constraint to remove the redundant information. Further, we adopt structure consistent constraint to preserve the local structure in each domain. CSFL has obtained about 1–5% improvement of mean accuracy, compared to the state‐of‐the‐art shallow methods. Further, compared with 90.2% and 89.4% of the best baseline deep method, CSFL achieves mean accuracy of 90.8% and 89.7% on the Office‐31 and ImageCLEF‐DA datasets respectively. The encouraging results validate the effectiveness of our method.
Chang Niu, Junyuan Shang, Zhiheng Zhou 0001, Junchu Huang, Tianlei Wang, Xiangwei Li
IET Image Process.6
2019 Time-Multiplexed FPGA Overlay Architectures: A Survey
abstract
This article presents a comprehensive survey of time-multiplexed (TM) FPGA overlays from the research literature. These overlays are categorized based on their implementation into two groups: processor-based overlays, as their implementation follows that of conventional silicon-based microprocessors, and; CGRA-like overlays, with either an array of interconnected processor-based functional units or medium-grained arithmetic functional units. Time-multiplexing the overlay allows it to change its behavior with a cycle-by-cycle execution of the application kernel, thus allowing better sharing of the limited FPGA hardware resource. However, most TM overlays suffer from large resource overheads, due to either the underlying processor-like architecture (for processor-based overlays) or due to the routing array and instruction storage requirements (for CGRA-like overlays). Reducing the area overhead for CGRA-like overlays, specifically that required for the routing network, and better utilizing the hard macros in the target FPGA are active areas of research.
Xiangwei Li, Douglas L. Maskell
ACM Trans. Design Autom. Electr. Syst.1
2018 A time-multiplexed FPGA overlay with linear interconnect
abstract
Coarse-grained overlays improve FPGA design productivity by providing fast compilation and software like programmability. Soft processor based overlays with well-defined ISAs are attractive to application developers due to their ease of use. However, these overlays have significant FPGA resource overheads. Time multiplexed (TM) CGRA-like overlays represent an interesting alternative as they are able to change their behavior on a cycle by cycle basis while the compute kernel executes. This reduces the FPGA resource needed, but at the cost of a higher initiation interval (II) and hence reduced throughput. The fully flexible routing network of current CGRA-like overlays results in high FPGA resource usage. However, many application kernels are acyclic and can be implemented using a much simpler linear feed-forward routing network. This paper examines a DSP block based TM overlay with linear interconnect where the overlay architecture takes account of the application kernels' characteristics and the underlying FPGA architecture, so as to minimize the II and the FPGA resource usage. We examine a number of architectural extensions to the DSP block based functional unit to improve the II, throughput and latency. The results show an average 70% reduction in II, with corresponding improvements in throughput and latency.
Xiangwei Li, Abhishek Kumar Jain, Douglas L. Maskell, Suhaib A. Fahmy
DATE1
2017 A new compressive sensing video coding framework based on Gaussian mixture model
Xiangwei Li, Xuguang Lan, Meng Yang 0002, Jianru Xue, Nanning Zheng 0001
Signal Process. Image Commun.1
2017 Online Variable Coding Length Product Quantization for Fast Nearest Neighbor Search in Mobile Retrieval
abstract
Quantization methods are crucial for efficient nearest neighbor search in many applications such as image, music, or product search. As mobile devices are becoming increasingly more popular, the quantization methods on mobile devices are more important, because a large portion of the search queries are becoming performed on mobile devices. One important characteristic of the communication on mobile devices is the inherent unreliability of their communication channels. In order to adapt the quality changes of the communication channels, we need to change the coding length of the quantization accordingly. The existing quantization methods use fixed-length codebooks, and it is expensive to retrain another codebook with different coding length. In this paper, we propose a novel variable length product quantization framework that consists of a set of fast universal scalar quantizers. The framework is capable of producing variable length quantization without retraining the codebook. Each data vector is transformed into a new space to reduce the correlation across dimensions. A proper number of bits is allocated to represent the scalar component in each dimension according to the given coding length. For each component, we estimate its probability density function (PDF) and design an efficient universal scalar quantizer based on the PDF and the allocated bits. To reduce distortion, we learn a Gaussian mixture model for the data. The experimental results show that, compared to state-of-the-art product quantization methods, our approach can construct the codebooks online for variable coding lengths and achieve the comparable performance.
Jin Li 0011, Xuguang Lan, Xiangwei Li, Jiang Wang 0001, Nanning Zheng 0001, Ying Wu 0001
IEEE Trans. Multim.3
2016 DeCO: A DSP Block Based FPGA Accelerator Overlay with Low Overhead Interconnect
abstract
Coarse-grained FPGA overlay architectures paired with general purpose processors offer a number of advantages for general purpose hardware acceleration because of software-like programmability, fast compilation, application portability, and improved design productivity. However, the area overheads of these overlays, and in particular architectures with island-style interconnect, negate many of these advantages, preventing their use in practical FPGA-based systems. Crucially, the interconnect flexibility provided by these overlay architectures is normally over-provisioned for accelerators based on feed-forward pipelined datapaths, which in many cases have the general shape of inverted cones. We propose DeCO, a cone shaped cluster of FUs utilizing a simple linear interconnect between them. This reduces the area overheads for implementing compute kernels extracted from compute-intensive applications represented as directed acyclic dataflow graphs, while still allowing high data throughput. We perform design space exploration by modeling programmability overhead as a function of overlay design parameters, and compare to the programmability overhead of island-style overlays. We observe 87% savings in LUT requirements using the proposed approach compared to DSP block based island-style overlays. Our experimental evaluation shows that the proposed overlay exhibits an achievable frequency of 395 MHz, close to the DSP theoretical limit on the Xilinx Zynq. We also present an automated tool flow that provides a rapid and vendor-independent mapping of the high level compute kernel code to the proposed overlay.
Abhishek Kumar Jain, Xiangwei Li, Pranjul Singhai, Douglas L. Maskell, Suhaib A. Fahmy
FCCM2
2016 Efficient compressive sensing video compression method based on Gaussian mixture models
abstract
In this paper, we propose an efficient lossy compression method for the compressive sensing video that utilizes Gaussian mixture models (GMM). The GMM is used to model the compressive sensing video (CSV) frames. Then we design an efficient lossy compression method based on the GMM. Each CSV frame can be efficiently compressed by the proposed method. The proposed method is for better compromise of compression efficiency and computational complexity. And it achieves a significant Bjontegaard-Delta (BD)-PSNR improvement about 8.84~11.81dB in average compared with existing low complexity compression solutions for compressing the CSV sequence.
Xiangwei Li, Xuguang Lan, Meng Yang 0002, Jianru Xue, Nanning Zheng 0001
VCIP1
2015 Optimized truncation model for adaptive compressive sensing acquisition of images
abstract
The sparsity of the input signal is important for compressive sensing (CS) reconstruction in CS system. In this paper, we establish an optimized truncation model to determine the number of the sparsified coefficients to be truncated in CS acquisition according to the sampling rate. The proposed truncation model suits for signals of any dimension. With the truncation model, the sparsity of the signal can be optimized by properly truncating the small elements of the sparsified coefficients. Furthermore we propose an adaptive CS acquisition solution based on the truncation model to reduce the noise folding effect. The proposed solution is verified for CS acquisition of natural images. Simulation results show that the proposed solution achieves significant improvement of the reconstructed image quality by 0.7~1.4 dB on average compared with existing solutions.
Xiangwei Li, Xuguang Lan, Meng Yang 0002, Jianru Xue, Nanning Zheng 0001
VCIP1
2013 Universal and low-complexity quantizer design for compressive sensing image coding
abstract
Compressive sensing imaging (CSI) is a new framework for image coding, which enables acquiring and compressing a scene simultaneously. The CS encoder shifts the bulk of the system complexity to the decoder efficiently. Ideally, implementation of CSI provides lossless compression in image coding. In this paper, we consider the lossy compression of the CS measurements in CSI system. We design a universal quantizer for the CS measurements of any input image. The proposed method firstly establishes a universal probability model for the CS measurements in advance, without knowing any information of the input image. Then a fast quantizer is designed based on this established model. Simulation result demonstrates that the proposed method has nearly optimal rate-distortion (R~D) performance, meanwhile, maintains a very low computational complexity at the CS encoder.
Xiangwei Li, Xuguang Lan, Meng Yang 0002, Jianru Xue, Nanning Zheng 0001
VCIP1