EDBT 2026 Demo / reviewers in the wild / expert
Jinwei Xu
dblp:55/11360
· DBLP profile ↗
40ranked-venue papers
9as first author
30since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 16 · 2 first-author · 12 since 2021Software engineering, systems software and programming languages · 14 · 3 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Computer networks · 2 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AMXStencil: Boosting the Performance of Stencil Computations on AMX-Powered CPUs via Fusion
Zitong An, Kangkang Chen, Huayou Su, Jinwei Xu, Xi Yang 0020 |
APPT | 4 |
| 2026 | Dual cuts for efficiency: Streamlining NAS with train-time pruning of supernet and search space
Di Niu 0001, Hengyue Pan, Jingfei Jiang, Jinwei Xu |
Knowl. Based Syst. | 4 |
| 2026 | One Size Does Not Fit All: Investigating Efficacy of Perplexity in Detecting LLM-Generated CodeabstractLarge Language Model-Generated Code (LLMgCode) has become increasingly common in software development. So far LLMgCode has more quality issues than Human-Authored Code (HaCode). It is common for LLMgCode to mix with HaCode in a code change, while the change is signed by only human developers, without being carefully examined. Many automated methods have been proposed to detect LLMgCode from HaCode, in which the perplexity-based method ( Perplexity for short) is the state-of-the-art method. However, the efficacy evaluation of Perplexity has focused on detection accuracy. Yet it is unclear whether Perplexity is good enough in a wider range of realistic evaluation settings. To this end, we carry out a family of experiments to compare Perplexity against feature- and pre-training-based methods from three perspectives: detection accuracy , detection speed , and generalization capability . The experimental results show that Perplexity has the best generalization capability while having limited detection accuracy and detection speed. Based on that, we discuss the strengths and limitations of Perplexity , e.g., Perplexity is unsuitable for high-level programming languages. Finally, we provide recommendations to improve Perplexity and apply it in practice. As the first large-scale investigation on detecting LLMgCode from HaCode, this article provides a wide range of findings for future improvement. Jinwei Xu, He Zhang 0001, Yanjing Yang, Lanxin Yang, Zeru Cheng, Bohan Liu 0003, Xin Zhou 0016, Alberto Bacchelli, Yin Kia Chiam, Thiam Kian Chiew |
ACM Trans. Softw. Eng. Methodol. | 1 |
| 2026 | Automated Localization of Affected Libraries and Versions from Vulnerability Reports
Jinwei Xu, He Zhang 0001, Xin Zhou 0016, Yanjing Yang, Jinghao Hu 0001, Lanxin Yang, Bohan Liu 0003 |
IEEE Trans. Software Eng. | 1 |
| 2025 | Multi-modal Parallelism Scheduling for Heterogeneous Multicore Computing Systems
Jingfei Jiang, Jinwei Xu |
ICA3PP (2) | 3 |
| 2025 | FQuant: Fast Quantization with Adaptive Resolution via the Clustering Algorithm
Linagwei Li, Jingfei Jiang, Jinwei Xu, Shunan Zhou, Minghua Zhu |
ICIC (21) | 3 |
| 2025 | Securing Self-Managed Third-Party LibrariesabstractModern software development reuses third-party libraries to cut costs but may introduce vulnerabilities. A critical practice is to verify the security of third-party libraries against public vulnerability reports. Many automated methods have been proposed to identify vulnerable libraries from vulnerability reports. Existing methods are designed for the generic identification of vulnerable libraries, considering the security of all software libraries. Generic identification is inherently challenging, resulting in limited accuracy. However, organizations only consider the security of libraries they trust and use, by self-managing a library whitelist. Therefore, we propose LibGuard, a framework to adapt existing methods to help organizations secure the libraries they use. LibGuard supplies a library whitelist for existing methods and filters the results according to a threshold, facilitating the discovery of risks overlooked by organizations while controlling false alarms. LibGuard is implemented in two ways. The first attaches the whitelist after existing methods. The second integrates the whitelist into existing methods. We evaluated LibGuard using 5,107 vulnerability reports and the library whitelist built from 79 Google projects and 29 Huawei projects. The results show that the two implementations of LibGuard increase the average F1 score by 10.25% and 11.77%, respectively. Moreover, LibGuard performs stably during the extension of whitelists. To our knowledge, this paper is the first study dedicated to securing self-managed third-party libraries, offering insights into adapting generic software security management to self-managed contexts. Xin Zhou 0016, Jinwei Xu, He Zhang 0001, Yanjing Yang, Lanxin Yang, Bohan Liu 0003, Hongshan Tang |
ASE | 2 |
| 2025 | Automated detection of affected libraries from vulnerability reports
Jinwei Xu, He Zhang 0001, Xin Zhou 0016, Yanjing Yang, Runfeng Mao, Lanxin Yang, Haifeng Shen |
Autom. Softw. Eng. | 1 |
| 2025 | Prioritizing code review requests to improve review efficiency: a simulation study
Lanxin Yang, Bohan Liu 0003, Junyu Jia, Jinwei Xu, Junming Xue, He Zhang 0001, Alberto Bacchelli |
Empir. Softw. Eng. | 4 |
| 2025 | Correction to: A preliminary investigation on using multi-task learning to predict change performance in code reviews
Lanxin Yang, He Zhang 0001, Jinwei Xu, Xin Zhou 0016, Dong Shao, Shan Gao 0009, Alberto Bacchelli |
Empir. Softw. Eng. | 3 |
| 2025 | DLAP: A Deep Learning Augmented Large Language Model Prompting framework for software vulnerability detection
Yanjing Yang, Xin Zhou 0016, Runfeng Mao, Jinwei Xu, Lanxin Yang, Haifeng Shen, He Zhang 0001 |
J. Syst. Softw. | 4 |
| 2025 | SAPFIS: a parallel fuzzy inference system for air combat situation assessment
Jingfei Jiang, Jinwei Xu, Pengbo Wu |
J. Supercomput. | 3 |
| 2025 | Research and application of smart insole assisted gait recognition technology
Jinwei Xu, Fansen Wei, Jingsong Mu, Pin Chen |
J. Supercomput. | 2 |
| 2025 | SPDFA: A Novel Dataflow Fusion Sparse Deep Neural Network AcceleratorabstractUnstructured sparse pruning significantly reduces the computational and parametric complexities of deep neural network models. Nevertheless, the highly irregular nature of sparse models limits their performance and efficiency on traditional computing platforms, thereby prompting the development of specialized hardware solutions. To improve computational efficiency, we introduce the Sparse Dataflow Fusion Accelerator (SPDFA), a specialized architecture meticulously designed for sparse deep neural networks. Firstly, we present a non-blocking data distribution-computing engine that integrates inner product and column product. This engine boosts computational efficiency by decomposing matrix multiplication and convolution into rectangular matrix-vector multiplications. Secondly, we implement a computation array to further exploit the parallelism, and design an on-chip buffer structure that supports multi-line memory access mode. Lastly, to bolster the adaptability of our accelerator, we propose an innovative macroinstruction set coupled with a micro-kernel scheme. Furthermore, we refine the macroinstruction issue strategy, thereby further enhancing computational efficiency. Our evaluation results demonstrate that SPDFA achieves an average 1.29 \(\times\) –2.38 \(\times\) improvement in computational efficiency compared to the state-of-the-art SpMM accelerators when applied to unstructured sparse deep neural network models. Furthermore, its performance outperforms existing sparse neural network accelerators by a factor of 1.03 \(\times\) –1.83 \(\times\) . Additionally, SPDFA exhibits excellent scalability with a scaling efficiency exceeding 80%. Jinwei Xu, Jingfei Jiang, Xifu Qian, Yong Dou |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2024 | HPFIA: A High-Performance Fuzzy Inference Accelerator for Situation Assessment on Airborne EquipmentabstractThe ever-increasing multi-source fusion information perceived by situation assessment system pose a computational challenge to current airborne equipment. Fuzzy inference method introduced in situation assessment could effectively adapt to the incompleteness and uncertainty of situational information, but still struggling to meet the high-performance requirements under limited hardware resources on airborne equipment. Leveraging hardware accelerators (GPUs, FPGAs, etc.) to accelerate intensive computation like situation factor evaluation has become paramount. Since the lack of relevant acceleration methods in state-of-the-art researches, this paper presents the first-ever FPGA accelerator of the HPFIA to optimize calculation performance of situation assessment. Our designed accelerator delivers up to 237.48× performance improvement and 3230.15× better energy efficiency ratio over the software implementation on three universal computing platforms. Jingfei Jiang, Jinwei Xu |
HPCC | 3 |
| 2024 | Mining Pull Requests to Detect Process Anomalies in Open Source Software DevelopmentabstractTrustworthy Open Source Software (OSS) development processes are the basis that secures the long-term trustworthiness of software projects and products. With the aim to investigate the trustworthiness of the Pull Request (PR) process, the common model of collaborative development in OSS community, we exploit process mining to identify and analyze the normal and anomalous patterns of PR processes, and propose our approach to identifying anomalies from both control-flow and semantic aspects, and then to analyze and synthesize the root causes of the identified anomalies. We analyze 17531 PRs of 18 OSS projects on GitHub, extracting 26 root causes of control-flow anomalies and 19 root causes of semantic anomalies. We find that most PRs can hardly contain both semantic anomalies and control-flow anomalies, and the internal custom rules in projects may be the key causes for the identified anomalous PRs. We further discover and analyze the patterns of normal PR processes. We find that PRs in the non-fork model (42%) are far more likely than the fork model (5%) to bypass the review process, indicating a higher potential risk. Besides, we analyzed nine poisoned projects whose PR practices were indeed worse. Given the complex and diverse PR processes in OSS community, the proposed approach can help identify and understand not only anomalous PRs but also normal PRs, which offers early risk indications of suspicious incidents (such as poisoning) to OSS supply chain. Bohan Liu 0003, He Zhang 0001, Weigang Ma, Hongyu Kuang, Jinwei Xu, Shan Gao 0009 |
ICSE | 6 |
| 2024 | Funnel: An Efficient Sparse Attention Accelerator with Multi-Dataflow FusionabstractThe self-attention mechanism is the core component of Transformer, which provides a powerful ability to understand the sequence context. However, the self-attention mechanism also suffers from a large amount of redundant computation. Model sparsification can effectively reduce computational load, but the irregularity of non-zeros introduced by sparsification significantly decreases hardware efficiency. This paper proposes Funnel, an accelerator that dynamically predicts sparse attention patterns and efficiently processes unstructured sparse data. Firstly, we adopt a fast quantization method based on lookup table to minimize the cost of sparse patterns prediction. Secondly, we propose Funnel Computing Unit (FCU), a hardware architecture that efficiently handles sparse attention through multi-dataflow fusion. Sampled Dense-Dense Matrix Multiplication (SDDMM) and Sparse-Dense Matrix Multiplication (SpMM) are core components of sparse attention mechanism. FCU unifies the computation ways of matrix inner product and row-wise product to support SDDMM and SpMM at the same time, which greatly reduces the storage and movement overhead of intermediate results. Lastly, we devise a lightweight buffer and data tiling strategy tailored to the proposed accelerator, aimed at enhancing data reuse. Experiments demonstrate that our accelerator achieves 0.10-0.25 sparsity with small accuracy loss. When computing the self-attention layer, it attains hardware efficiency ranging from 60% to 85%. Compared to CPU and GPU, it achieves 5.60x and 8.20x speedup. Compared to the state-of-the-art attention accelerators A3, SpAtten, FTRANS, and Sanger, it achieves 7.37x, 4.52x, 9.58x, and 3.08x speedup. Shenghong Ma, Jinwei Xu, Jingfei Jiang, Dongsheng Li 0001 |
ISPA | 2 |
| 2024 | GPP: A Graph-Powered Prioritizer for Code Review RequestsabstractPeer code review has become a must-have in modern software development. However, many code review requests (CRRs) could be a backlog for large-scale and active projects, blocking continuous integration and continuous delivery (CI/CD). Prioritizing CRRs to make the relevant ones to be reviewed first is a critical method for addressing this issue. Early studies have shown that many factors affect the review priority of a CRR, including its properties and relationships with other CRRs. However, the relationships, e.g., modifying the same files and sharing the same authors, are rarely considered when developing CRR prioritizers. In this paper, we propose a Graph-Powered Prioritizer (namely GPP) to make full use of the properties and relationships of CRRs. GPP uses the multi-graph structure to develop an initial representation of a collection of CRRs and uses the graph neural network algorithm to learn the prioritization-adapted representation, and eventually, outputs an ordered list of CRRs based on it. With experimental evaluation, we define relevant CRRs in the context of CI/CD as those that are likely to achieve three objectives, i.e., being merged while undergoing a few iterations in a short duration. We compare GPP against two rule-based and six learning-based prioritizers on 15 open-source software projects with more than 420K CRRs. The experimental results indicate that GPP outperforms the baselines on three basic ranking-aware evaluation metrics, including NDCG (82.94%), MRR (36.52%), and MAP (63.80%); while providing benefits in recommending the most relevant CRRs and balancing multiple objectives. Data&materials: https://figshare.com/s/133f23da558b7b254041 Lanxin Yang, Jinwei Xu, He Zhang 0001, Fanghao Wu, Yue Li 0047, Alberto Bacchelli |
ASE | 2 |
| 2024 | A preliminary investigation on using multi-task learning to predict change performance in code reviews
Lanxin Yang, He Zhang 0001, Jinwei Xu, Xin Zhou 0016, Dong Shao, Shan Gao 0009, Alberto Bacchelli |
Empir. Softw. Eng. | 3 |
| 2024 | A blockchain-based and microservices-architected software composition analysis systemabstractAbstract “Shift To Left” is the cornerstone of the successful implementation of DevSecOps. By testing projects for vulnerabilities in the early stages of development, teams can save overall costs before security issues reach the build phase. As one of the popular practices in “Shift To Left,” the Software Composition Analysis (SCA) system aims to leverage the Software Bill of Materials (SBOM) to enhance software supply chain security. However, the SBOM lacks mature generation and distribution mechanisms, requiring incentive measures to drive industry consensus. Additionally, the data and tools associated with the SBOM lack effective record‐keeping and monitoring, making it challenging to ensure data integrity and tool security. Traditional SCA systems treat SBOM as a regular data format for external service provision, yet fail to solve problems such as lack of shared platforms, inability to guarantee data integrity and tool security, as well as issues with poor interoperation compatibility. This paper introduces blockchain technology into the SCA system, utilizing smart contracts to provide core SBOM tool services and microservices to improve the operational efficiency of smart contract deployment and maintenance. The proposed SCA system effectively provides a shared platform for SBOM with reliable data integrity, guaranteed tool security, and good interoperability. Xin Zhou 0016, Jinwei Xu, Lingli Cao, Shanshan Li 0002 |
J. Softw. Evol. Process. | 2 |
| 2024 | Efficient SpMM Accelerator for Deep Learning: Sparkle and Its Automated GeneratorabstractDeep learning (DL) technology has made breakthroughs in a wide range of intelligent tasks, such as vision, language, recommendation systems, and so on. Sparse matrix multiplication (SpMM) is the key computation kernel of most sparse models. Conventional computing platforms, such as CPUs, GPUs, and AI chips with regular processing units, are unable to effectively support sparse computation due to their fixed structure and instruction sets. This work extends Sparkle, an accelerator architecture, which is developed specifically for processing SpMM in DL. During the balanced data loading process, some modifications are implemented to enhance the flexibility of the Sparkle architecture. Additionally, a Sparkle generator is proposed to accommodate diverse resource constraints and facilitate adaptable deployment. Leveraging Sparkle’s structural parameters and template-based design methods, the generator enables automatic Sparkle circuit generation under varying parameters. An instantiated Sparkle accelerator is implemented on the Xilinx xqvu11p FPGA platform with a specific configuration. Compared to the state-of-the-art SpMM accelerator SIGMA, the Sparkle accelerator instance improves the sparse computing efficiency by about 10 to 20 \(\%\) . Furthermore, the Sparkle instance achieved 7.76 \(\times\) higher performance over the Nvidia Orin NX GPU. More instances of accelerators with different parameters were evaluated, demonstrating that the Sparkle architecture can effectively accelerate SpMM. Shiyao Xu, Jingfei Jiang, Jinwei Xu, Xifu Qian |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2023 | Coded Caching Scheme for Two-Dimensional Caching-Aided Ultra-Dense NetworksabstractIn this paper, we consider a two-dimensional caching-aided ultra-dense network caching system, which consists of a server containing N files, K1K2cache nodes that are arranged neatly on the grid with K1rows and K2columns, and U cacheless users randomly distributed around cache nodes. The server connects to users through an error-free shared-link, and the users can be served by nearby cache nodes and freely retrieve the cache content from it. Our goal is to minimize the transmission load in the worst case while meeting all possible users’ demands. We propose a coded caching scheme based on Maddah-Ali and Niesen scheme (MN scheme), which uses the placement strategy of MN scheme to the cache nodes and uses the delivery strategy of MN scheme multiple rounds for different types of users according to their geometrical locations. We prove that our scheme is order optimal and can greatly improve the transmission performance compared with conventional uncoded caching schemes. Minquan Cheng, Jinwei Xu, Mingming Zhang 0003, Youlong Wu |
ISIT | 2 |
| 2023 | EvaCRC: Evaluating Code Review CommentsabstractIn code reviews, developers examine code changes authored by peers and provide feedback through comments. Despite the importance of these comments, no accepted approach currently exists for assessing their quality. Therefore, this study has two main objectives: (1) to devise a conceptual model for an explainable evaluation of review comment quality, and (2) to develop models for the automated evaluation of comments according to the conceptual model. To do so, we conduct mixed-method studies and propose a new approach: EvaCRC (Evaluating Code Review Comments). To achieve the first goal, we collect and synthesize quality attributes of review comments, by triangulating data from both authoritative documentation on code review standards and academic literature. We then validate these attributes using real-world instances. Finally, we establish mappings between quality attributes and grades by inquiring domain experts, thus defining our final explainable conceptual model. To achieve the second goal, EvaCRC leverages multi-label learning. To evaluate and refine EvaCRC, we conduct an industrial case study with a global ICT enterprise. The results indicate that EvaCRC can effectively evaluate review comments while offering reasons for the grades. Data and materials: https://doi.org/10.5281/zenodo.8297481 Lanxin Yang, Jinwei Xu, He Zhang 0001, Alberto Bacchelli |
ESEC/SIGSOFT FSE | 2 |
| 2023 | Evaluating Learning-to-Rank Models for Prioritizing Code Review Requests using Process SimulationabstractIn large-scale, active software projects, one of the main challenges with code review is prioritizing the many Code Review Requests (CRRs) these projects receive. Prior studies have developed many Learning-to-Rank (LtR) models in support of prioritizing CRRs and adopted rich evaluation metrics to compare their performances. However, the evaluation was performed before observing the complex interactions between CRRs and reviewers, activities and activities in real-world code reviews. Such a pre-review evaluation provides few indications about how effective LtR models contribute to code reviews. This study aims to perform a post-review evaluation on LtR models for prioritizing CRRs. To establish the evaluation environment, we employ Discrete-Event Simulation (DES) paradigm-based Software Process Simulation Modeling (SPSM) to simulate real-world code review processes, together with three customized evaluation metrics. We develop seven LtR models and use the historical review orders of CRRs as baselines for evaluation. The results indicate that employing LtR can effectively help to accelerate the completion of reviewing CRRs and the delivery of qualified code changes. Among the seven LtR models, LambdaMART and AdaRank are particularly beneficial for accelerating completion and delivery, respectively. This study empirically demonstrates the effectiveness of using DES-based SPSM for simulating code review processes, the benefits of using LtR for prioritizing CRRs, and the specific advantages of several LtR models. This study provides new ideas for software organizations that seek to evaluate LtR models and other artificial intelligence-powered software techniques.Data&materials: https://figshare.com/s/a033e99cd2a61e64c8bc. Lanxin Yang, Bohan Liu 0003, Junyu Jia, Junming Xue, Jinwei Xu, Alberto Bacchelli, He Zhang 0001 |
SANER | 5 |
| 2023 | Coded Caching Schemes for Two-Dimensional Caching-Aided Ultra-Dense NetworksabstractCoded caching technique is an efficient approach to reduce the transmission load in networks. In this paper, we consider a new widespread caching system called$(K_{1},K_{2},U,r,M,N)$two-dimensional (2D) caching-aided ultra-dense networks (UDNs) with a server containing$N$files,$K_{1}K_{2}$cache nodes arranged neatly on a grid with$K_{1}$rows and$K_{2}$columns, and$U$cache-less users randomly distributed around cache nodes. Each cache node can cache at most$M\leq N$files and has a certain service region by Euclidean distance. The server connects to users through the error-free shared link and the users in the service region of a cache node can freely retrieve all cached contents of this cache node. We aim to design a coded caching scheme for 2D caching-aided UDN system to reduce the transmission load in the worst case while meeting all possible users’ demands. First, we divide all possible users into four classes according to their geographical locations. Then our first order optimal scheme is proposed based on the Maddah-Ali and Niesen scheme. Furthermore, by compressing the transmitted signals of our first scheme based on Maximum Distance Separable (MDS) code, we obtain an improved order optimal scheme with a smaller transmission load. Minquan Cheng, Jinwei Xu, Mingming Zhang 0003, Youlong Wu |
IEEE Trans. Commun. | 2 |
| 2022 | Sparkle: A High Efficient Sparse Matrix Multiplication Accelerator for Deep LearningabstractDeep learning (DL) technology is applied to a wide range of intelligent tasks across vision, language, recommendation systems, etc. Large DL models with high sparsity become critical for various intelligent applications and require an energy-efficient hardware accelerator. Sparse-dense matrix multiplication (SpMM) is a key computation kernel widely used in most sparse and large DL workloads. However, traditional computing platforms such as CPU, GPU, and Al chips with regular processing units are limited to support sparsity by their fixed structures. In this work, a specific SpMM accelerator named Sparkle is proposed which achieves high performance and high computational efficiency. A block-wise arrangement approach is proposed in Sparkle to process matrix multiplications. A novel compressed sparsity format, the pointer-bitmap, is designed to simplify the decoding process and improve the efficiency of data loading. Grouped PEs and configurable hierarchical reduction network are deployed to leverage sparsity, further enhancing the utilization of the compute resources. Sparkle is implemented using the Xilinx xqvu11p FPGA. A diverse set of matrices in DL workloads are evaluated and Sparkle achieves 2.1× higher energy efficiency over the NVIDIA TITAN X GPU. Our experiments also show that Sparkle roughly promotes 26% compute efficiency better than state-of-the-art sparse accelerators SIGMA. Shiyao Xu, Jingfei Jiang, Jinwei Xu, Chaorun Liu, Yuanhong He |
ICCD | 3 |
| 2022 | Evaluating a New Attention Framework Based on Matrix Blocking for Attention Models on FPGAsabstractThe attention mechanism has recently shown superior performance in natural language processing and computer vision tasks. But its complex dataflow and large-scale matrix calculation with huge computing and memory overhead pose a great challenge for the design of hardware accelerators. And previous solutions that benefited from matrix partitioning are bounded by the softmax function. In this paper, we propose a new attention framework that can dramatically improve the performance of attention model inference for long sequence tasks on FPGAs. We design a novel accelerator architecture that employs two systolic arrays and a ping-pong structure to accelerate attention calculation. Meanwhile, we propose an analytical model to predict resource usage and performance, which guides a fast design space exploration. Experiments using the state-of-the-art BERT demonstrate the design achieves 4.61 and 1.24× improvement in speed and energy efficiency compared to CPU and GPU on the Xilinx XCZU11EG platform. Jingfei Jiang, Jinwei Xu |
ICTAI | 3 |
| 2021 | RFC-HyPGCN: A Runtime Sparse Feature Compress Accelerator for Skeleton-Based GCNs Action Recognition Model with Hybrid PruningabstractSkeleton-based Graph Convolutional Networks (GCNs) models for action recognition have achieved excellent prediction accuracy in the field. However, limited by large model and computation complexity, GCNs for action recognition like 2s-AGCN have insufficient power-efficiency and throughput on GPU. Thus, the demand of model reduction and hardware acceleration for low-power GCNs action recognition application becomes continuously higher.To address challenges above, this paper proposes a runtime sparse feature compress accelerator with hybrid pruning method: RFC-HyPGCN. First, this method skips both graph and spatial convolution workloads by reorganizing the multiplication order. Following spatial convolutions channel-pruning dataflow, a coarse-grained pruning method on temporal filters is designed, together with sampling-like fine-grained pruning on time dimension. Later, we come up with an architecture where all convolutional layers are mapped on chip to pursue high throughput. To further reduce storage resource utilization, online sparse feature compress format is put forward. Features are divided and encoded into several banks according to presented format, then bank storage is split into depth-variable mini-banks. Furthermore, this work applies quantization, input-skipping and intra-PE dynamic data scheduling to accelerate the model. In experiments, proposed pruning method is conducted on 2s-AGCN, acquiring 3.0x-8.4x model compression ratio and 73.20% graph-skipping efficiency with balancing weight pruning. Implemented on Xilinx XCKU-115 FPGA, the proposed architecture has the peak performance of 1142 GOP/s and achieves up to 9.19x and 3.91x speedup over high-end GPU NVIDIA 2080Ti and NVIDIA V100, respectively. Compared with latest accelerator for action recognition GCNs models, our design reaches 22.9x speedup and 28.93% improvement on DSP efficiency. Dong Wen 0004, Jingfei Jiang, Jinwei Xu, Yang Zhao 0003, Yong Dou |
ASAP | 3 |
| 2021 | A high-throughput scalable BNN accelerator with fully pipelined architecture
Jingfei Jiang, Jinwei Xu, Peng Zhang 0035, Dong Wen 0004, Yong Dou |
CCF Trans. High Perform. Comput. | 3 |
| 2021 | An energy-efficient convolutional neural network accelerator for speech classification based on FPGA and quantization
Dong Wen 0004, Jingfei Jiang, Yong Dou, Jinwei Xu |
CCF Trans. High Perform. Comput. | 4 |
| 2020 | A Dynamic Mapping Model for General CNN Accelerator Based on FPGA
Jingfei Jiang, Jinwei Xu |
NPC | 4 |
| 2019 | RF Aerially Charging Scheduling for UAV Fleet : A Q-Learning ApproachabstractIn recent years, unmanned aerial vehicles (UAVs) have attracted extensive interests from both academia and industry due to the potential wide applications with universal applicable nature of the deployment. However, currently the bottleneck for UAVs is the limited carried energy resources (e.g. oil box, battery), especially for electric-driven UAVs. For a system consisting of multiple UAVs using batteries, its stability depends on each UAV. Therefore, the lifetime of each UAV is expected to be extended. In this paper, we propose the concept of RF charging aerially for the UAV fleet. Specifically, in order to ensure the stability of the system, wireless charging is considered for enhancing the lifetime of each UAV. However, it may be unbalanced. Accordingly, the issue of charging scheduling arises. The problem is formulated as a Q-Learning problem in this paper. Agent constantly explores and optimizes its scheduling policy. Finally, it can adapt to different UAV distribution situations. We take the energy levels of UAVs as input, which is easy for implementation. We have compared with two other algorithms (RSA and LESA) and compared with the case of no-charging. The results show that comparing with no-charging, the stability of the system can be improved by up to 78%. Compared with RSA and LESA, system stability is increased by up to 30%-40%. In addition, our method is more flexible and applicable to fleet than other ways (such as return to base station, landing to power line, ground laser, etc) to supplement energy. Jinwei Xu, Kun Zhu 0001, Ran Wang 0004 |
MSN | 1 |
| 2018 | mmCNN: A Novel Method for Large Convolutional Neural Network on Memory-Limited DevicesabstractDeep learning recently has been widely used in many interactive application fields including but not limited to object recognition, speech recognition, natural language processing and so on. At the same time more and more attractive interactive applications (face recognition and augmented reality) are available on wearable and mobile devices. However, traditional deep learning methods such as CNN cost a lot of memory resources. This challenge makes it difficult to apply the powerful deep learning method on mobile memory limited platforms. In this paper we present a novel memory management strategy called mmCNN to solve this problem. This method helps us deploy a trained large size CNN on an any memory size platform including GPU, FPGA and memory-limited mobile devices. In our experiments, we run a feed-forward CNN process in an extremely small memory size (as low as 5MB) on a GPU platform. The result shows that our method saves more than 98% memory compared to a traditional CNN algorithm and further saves more than 90% compared to the sate-of-the-art related work "vDNN". Our work improve the computing scalability of interaction applications and break the memory bottleneck of using deep learning method on a memory-limited devices. Shijie Li 0002, Yong Dou, Jinwei Xu, Qiang Wang 0006, Xin Niu 0002 |
COMPSAC (1) | 3 |
| 2018 | A target image-oriented dictionary learning-based method for fully automated latent fingerprint forensicabstractAbstract Several fully automated latent print forensic techniques have been reported. In this paper, we propose a fully automated latent print segmentation module for the partition of the fingerprint region in a query latent image, which can help find a corresponding match of the suspect reliably. Being different from the existing methods that build the prelearned dictionary from the high‐quality fingerprint image patches, the proposed dictionary learning procedure is conducted on the target images. The advantages of the proposed method are the following: (i) it does not require a large number of high‐quality “ridge‐valley” atoms and (ii) not only the structure similarity but also the pattern scale has been kept consistent between the target image patches and learned dictionary atoms. Because no commercial latent fingerprint matcher is publicly available and the latent matcher reported in the literature is not accessible to the public either, a latent fingerprint matching platform is implemented to evaluate the obtained segmentation results and automated latent fingerprint matching performance. On the basis of this platform, experimental comparisons are conducted to assess closeness to the system performance upper bound when different segmentation modules are deployed. Moreover, matcher‐independent criteria such as genuine minutiae preservation rate and the segmented region of interest's accuracy are used. All the experimental results demonstrate that the proposed segmentation approach outperforms state‐of‐the‐art techniques in terms of finding the correct fingerprint of the suspect subject. Jinwei Xu, Jiankun Hu, Xiuping Jia |
Comput. Intell. | 1 |
| 2017 | Throughput-Optimized FPGA Accelerator for Deep Convolutional Neural NetworksabstractDeep convolutional neural networks (CNNs) have gained great success in various computer vision applications. State-of-the-art CNN models for large-scale applications are computation intensive and memory expensive and, hence, are mainly processed on high-performance processors like server CPUs and GPUs. However, there is an increasing demand of high-accuracy or real-time object detection tasks in large-scale clusters or embedded systems, which requires energy-efficient accelerators because of the green computation requirement or the limited battery restriction. Due to the advantages of energy efficiency and reconfigurability, Field-Programmable Gate Arrays (FPGAs) have been widely explored as CNN accelerators. In this article, we present an in-depth analysis of computation complexity and the memory footprint of each CNN layer type. Then a scalable parallel framework is proposed that exploits four levels of parallelism in hardware acceleration. We further put forward a systematic design space exploration methodology to search for the optimal solution that maximizes accelerator throughput under the FPGA constraints such as on-chip memory, computational resources, external memory bandwidth, and clock frequency. Finally, we demonstrate the methodology by optimizing three representative CNNs (LeNet, AlexNet, and VGG-S) on a Xilinx VC709 board. The average performance of the three accelerators is 424.7, 445.6, and 473.4GOP/s under 100MHz working frequency, which outperforms the CPU and previous work significantly. Yong Dou, Jingfei Jiang, Jinwei Xu, Shijie Li 0002, Yongmei Zhou, Yingnan Xu |
ACM Trans. Reconfigurable Technol. Syst. | 4 |
| 2016 | Automatic code generation of convolutional neural networks in FPGA implementationabstractConvolutional neural networks (CNNs) have gained great success in various computer vision applications. However, state-of-the-art CNN models are computation-intensive and hence are mainly processed on high performance processors like server CPUs and GPUs. Owing to the advantages of high performance, energy efficiency and reconfigurability, Field-Programmable Gate Arrays (FPGAs) have been widely explored as CNN accelerators. In this paper, we propose parallel structures to exploit the inherent parallelism and efficient computation units to perform operations in convolutional and fully-connected layers. Further, an automatic generator is proposed to generate Verilog HDL source code automatically according to high-level hardware description language. Execution time, DSP consumption and performance are analytically modeled based on some critical design variables. We demonstrate the automatic methodology by implementing two representative CNNs (LeNet and AlexNet) and evaluate the execution time models by comparing estimated and measured values. Our results show that the proposed automatic methodology yields hardware design with good performance and saves much developing round time. Yong Dou, Jingfei Jiang, Jinwei Xu |
FPT | 4 |
| 2016 | Coarse-Grained Architecture for Fingerprint MatchingabstractFingerprint matching is a key procedure in fingerprint identification applications. The minutiae-based fingerprint matching algorithm is one of the most typical algorithms achieving a reasonably correct recognition rate. This study proposes a coarse-grained parallel architecture called fingerprint matching core (FMC) to accelerate fingerprint matching. The proposed architecture has a two-level parallel structure (i.e., parallel among groups (PAG) and parallel in group (PIG)). A multirequest controller is added to the PAG structure to obtain a concurrent operation of the multiple processing element group (PEG). The DDR3 controller is used in the PIG structure to read eight minutiae from eight different fingerprints and realize the simultaneous computation of the eight PEs. The whole system is implemented on a Xilinx FPGA board with a Virtex VII XC7VX485T chip. The 16-PEG FMC achieves a throughput of about 9.63 million fingerprint pairs per second, which is larger than that achieved on a Tesla K20c platform. The software execution times are also measured on the 2.93GHz Intel Xeon 5670, 2.3GHz AMD Opteron(tm) Processor 6376, and Tesla K20c platforms. The Intel Xeon 5670 has two processors with 12 cores, and the AMD Opteron(tm) Processor 6376 has two processors with 16 cores. Moreover, the throughput is about 31 times that achieved on a 2.93GHz Intel Xeon 5670 single core. Jinwei Xu, Jingfei Jiang, Yong Dou, Xiaolong Shen |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2015 | Multi-constrained Orientation Field Modeling and Its Application for Fingerprint Indexing
Jinwei Xu, Jiankun Hu |
NSS | 1 |
| 2015 | A Multistaged Automatic Restoration of Noisy Microscopy Cell ImagesabstractAutomated cell segmentation for microscopy cell images has recently become an initial step for further image analysis in cell biology. However, microscopy cell images are easily degraded by noise during the readout procedure via optical-electronic imaging systems. Such noise degradations result in low signal-to-noise ratio (SNR) and poor image quality for cell identification. In order to improve SNR for subsequent segmentation and image-based quantitative analysis, the commonly used state-of-art restoration techniques are applied but few of them are suitable for corrupted microscopy cell images. In this paper, we propose a multistaged method based on a novel integration of trend surface analysis, quantile-quantile plot, bootstrapping, and the Gaussian spatial kernel for the restoration of noisy microscopy cell images. We show this multistaged approach achieves higher performance compared with other state-of-art restoration techniques in terms of peak signal-to-noise ratio and structure similarity in synthetic noise experiments. This paper also reports an experiment on real noisy microscopy data which demonstrated the advantages of the proposed restoration method for improving segmentation performance. Jinwei Xu, Jiankun Hu, Xiuping Jia |
IEEE J. Biomed. Health Informatics | 1 |
| 2012 | FTrack: Infrastructure-free floor localization via mobile phone sensingabstractMobile phone localization plays a key role in the fast-growing Location Based Applications domain. Most of the existing localization schemes rely on infrastructure support such as GSM, WiFi or GPS. In this paper, we present FTrack, a novel floor localization system to identify the floor level in a multi-floor building on which a mobile user is located. FTrack uses the mobile phone's accelerometer only without any infrastructure support. It does not require any prior knowledge of the building such as floor height. By capturing user encounters and analyzing user trails, FTrack finds the mapping from the traveling time (when taking the elevator) or the step counts (when walking on the stairs) between any two floors to the number of floor levels. The mapping can then be used for mobile users to pinpoint their current floor levels. We conduct both simulation and field studies to demonstrate the effectiveness of FTrack. Our field trial in a 10-floor building shows that FTrack achieves an accuracy of over 90% after two hours in our experiment. Haibo Ye, Tao Gu 0001, Jinwei Xu, XianPing Tao, Jian Lu 0001 |
PerCom | 4 |