Wenjin Huang

dblp:149/7880 · DBLP profile ↗
← Back
20ranked-venue papers
2as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 13 · 13 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 2 since 2021Computer networks · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 RRAM: A Reconfigurable Request-Rate-Aware Multi-DNN Accelerator Based on FPGA
abstract
The latency-sensitive image recognition systems in edge data centers are facing the workloads of multiple deep neural networks (Multi-DNN) and multi-user dynamic requests. Most of the existing solutions are based on GPUs or ASICs, whose fixed hardware architecture has limited low-latency capabilities when dealing with varying request rates in random dynamic scenarios. The dynamic reconfigurability of Field-Programmable Gate Arrays (FPGAs) can enable the hardware architecture to adapt in accordance with fluctuating request rates. Therefore, this study proposes a Request-Rate-Aware Multi-DNN accelerator framework (RRAM) based on FPGA, thereby improving latency performance. The RRAM includes a reconfigurable, computation-memory-interleaved multi-core accelerator architecture, a dynamic bandwidth adaptor for stable data transfers, and a performance analytic model for random request scenarios that enables an end-to-end latency evaluation. To efficiently explore the design space, a workload-aware resource boundary search strategy is also introduced, which eliminates a majority of invalid solutions. Implemented on the U250 platform, RRAM achieves better Service-Level Agreement (SLA) satisfaction rates compared to the fixed-architecture baseline accelerators under varying request rates, and improves the Average Normalized Turnaround Time (ANTT) by 1.3x to 6.1x. Furthermore, the average power consumption of RRAM is only 22.8% and 44.9% of the GPUs of RTX3090 and A100, while the latency is improved by up to 18.2x and 4.6x.
Han Jiao 0003, Wenjin Huang, Kailing Zhou, Dinghua Xu, Zhiyong Pang, Yihua Huang 0005
IEEE Internet Things J.2
2026 FPGA-Based Hardware Accelerator of zk-SNARK
abstract
Zero-Knowledge Proof (ZKP) has gained widespread application across various domains, demonstrating remarkable success. Among ZKP algorithms, Zero-Knowledge Succinct Non-Interactive Argument of Knowledge (zk-SNARK) is the most widely used. However, despite its advantages of small proof size and succinct verification, zk-SNARK proof generation faces significant challenges due to high computational demands, limiting its practical application. This paper addresses these challenges by accelerating two computationally intensive operations in zk-SNARK proof generation, Number Theory Transformation (NTT) and Multi-Scalar Multiplication (MSM), using FPGAs. In the implementation of NTT hardware accelerators for zk-SNARK applications, the traditional 4-step algorithm often encounters conflicts between off-chip bandwidth and on-chip memory. To resolve this issue, we propose an innovative approach that enhances accelerator performance by recursively applying the 4-step algorithm to create a more efficient 6-step algorithm. For MSM hardware acceleration on FPGAs, existing works are often constrained by limited on-chip memory, restricting the use of longer slice lengths, which are crucial for higher performance when using the commenly used Pippenger algorithm. To overcome this limitation, we introduce the Batch Method, optimizing off-chip memory consumption, enabling the accelerator to use longer slice lengths and achieve superior performance. Experimental results demonstrate that the proposed NTT design achieves 1.76× higher DSP efficiency than the SAM. Meanwhile, the proposed MSM design demonstrates 1.24× higher performance than the MSMAC with aligned frequency and number of PEs. When benchmarked against the GPU implementation GZKP, our MSM design exhibits 1.16× and 1.46× higher performance than GZKP for BLS12-381 and BN-254, respectively. However, the NTT design remains at a disadvantage due to the bandwidth limitation between our platform, Xilinx Alveo U250, and GZKP’s platforms, Nvidia GTX 1080 Ti and Nvidia Tesla V100.
Baoze Zhao, Conghui Luo, Wenjin Huang, Yihua Huang 0005
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2026 Relationship-Experts Transformer for Image Captioning
abstract
Image captioning is a cross-modal text generation task aimed at understanding the relationships among various objects in an image. Therefore, accurately expressing object–object relations remains a key bottleneck for transformer-based image captioning. Prior methods usually inject semantic and geometric relations once and keep them fixed while only updating visual features, creating a mismatch—evolving visuals vs. frozen relations—that weakens relational guidance and leads to feature entanglement. We propose the Relationship-Experts Transformer (RET), which treats semantic and geometric relations as learnable experts that guide object visual features (students) and co-evolve with them. In RET, we first design the Relationship-Guided Feature Aggregation (RGFA) module, which is analogous to experts-guided student learning, specifically utilizing the relationship kernel (the expert’s knowledge brain) to guide the learning of the object visual features (students). Secondly, we develop the Experts Knowledge Updating (EKU) module, which continuously iterates expert knowledge during training to enhance the expert’s guiding ability over the student. Finally, we design the Student Knowledge Selector (SKS) module to adaptively select object visual features enhanced with different relations under the guidance of semantic and geometric experts to generate descriptive texts embodying semantic and geometric knowledge. Experiments on the MSCOCO dataset demonstrate that our model achieves state-of-the-art performance. All codes are available at https://github.com/songchuanle-1/RET .
Chuanle Song, Wenjin Huang, Han Jiao 0003, Yihua Huang 0005
ACM Trans. Multim. Comput. Commun. Appl.2
2025 An Efficient FPGA-Based Hardware Accelerator of Fully Quantized Mamba-2
abstract
The Mamba-2 model introduces a State Space Duality (SSD) mechanism, based on original State Space Models (SSMs), that accelerates training and improves accuracy. However, efficient hardware acceleration for Mamba-2 faces challenges. Numerous element-wise operations fail to fully utilize GPU tensor cores, diminishing inference efficiency. Furthermore, research on full-quantization strategies for Mamba-2 is lacking. To address this, we propose a hybrid-precision full-quantization strategy, Hfqmamba2, balancing performance and hardware resource usage. Applying this strategy to Mamba-2 models of various sizes shows that accuracy loss remains within an acceptable range. We also propose an efficient FPGA-based hardware accelerator for Mamba-2. Given the distinct data flow characteristics of the two RMSNorm (Root Mean Square Normalization) layers in the hardware implementation, we introduce a reconfigurable hardware architecture based on a segmented quantization strategy, improving efficiency and flexibility by using a segmented lookup table to approximate the inverse square root operation. For the selective SSM layer operations, we design an intra-layer computation pipeline to enhance processing efficiency. Through design space exploration, we configure two versions of the hardware accelerator and evaluate their performance on the Alveo U50 platform. Experimental results show that both configurations achieve 99.63 % bandwidth utilization. Compared to the CPU, the hardware accelerator achieves a 114.05× speedup and a 282.75× improvement in energy efficiency. It also outperforms the PyTorch implementation on the GPU, achieving a 29.81 ×speedup and a 297.87 ×improvement in energy efficiency. Additionally, the hardware implementation shows a 1.94 ×speedup and a 35.89 ×improvement in energy efficiency over the official CUDA-accelerated version.
Kailing Zhou, Han Jiao 0003, Wenjin Huang, Yihua Huang 0005
FCCM3
2025 EViL: An Efficient Vision-LSTM Accelerator Based on FPGA
abstract
Recently, the Vision-LSTM (ViL) model, built upon Extended Long Short-Term Memory (xLSTM) building blocks, has attracted widespread attention due to its excellent performance and its linear complexity. Due to the unified compute architecture of GPUs, efficiently deploying the ViL layer on GPU is challenging because of its complex dataflow, thereby necessitating custom hardware accelerators. Moreover, significant differences in the activation distribution and quantization sensitivity among modules, making most existing quantization methods for CNNs not suitable for ViL layers. To address these challenges, we propose an intra-layer mixed-precision quantization method, termed ClipQuant, for the ViL layer. By introducing variable quantization range parameters$\alpha$and scaling parameters$\beta$, assessing the quantization sensitivity of each module, and imposing suitable quantization parameter constraints, we achieve near-lossless quantization of the ViL layer (accuracy loss$1.10 \times \sim 1.49 \times$performance improvement and a$36.95 \times \sim 45.50 \times$improvement in energy efficiency.
Zexuan Deng, Han Jiao 0003, Wenjin Huang, Yihua Huang 0005
FPL3
2025 An FPGA-based Quantization and Acceleration Framework for Multi-DNN
abstract
Deep neural networks occupy an irreplaceable place in both modern consumer and industrial sectors. The advent of INFerence-as-a-Service (INFasS) by cloud providers facilitates end-users in various domains. Numerous studies have tailored accelerators to boost the efficiency of multi-DNN inference towards these scenarios. However, current research seldom employs quantization techniques to improve the performance of multi-DNN accelerator. FPGAs are characterized by high parallelism and flexible reconfigurability, with their heterogeneous resources enabling support of diverse precision levels. In this context, we propose a framework for quantizing and accelerating multi-DNN on FPGA, including 1. a hardware-aware inter-layer mixed-precision quantization algorithm tailored for multi-DNN inference, 2. an FPGA architecture that facilitates ultra-low bit-width mixed-precision quantization, and 3. a low-overhead scheduling algorithm. Our co-design improves the throughput by up to 2.19x and reduces the response time by up to 28% compared to a unified-precision baseline. In addition, our design achieves similar DSP efficiency and better energy efficiency compared to related work.
Han Jiao 0003, Wenjin Huang, Yihua Huang 0005
ISCAS3
2025 HCG: Streaming DCNN Accelerator With a Hybrid Computational Granularity Scheme on FPGA
abstract
With the growth of field-programmable gate array (FPGA) hardware resources, streaming DCNN accelerators leverage interconvolutional-layer parallelism to enhance throughput. In existing streaming accelerators, convolution nodes typically adopt layer- or column-based tiling methods, where the tiled input feature map (Ifmap) encompasses all input channels. This approach facilitates the comprehensive calculation of the output feature map (Ofmap) and maximizes interlayer parallelism. The computational granularity, defined in this study as the calculated rows or columns of Ofmap based on each tiled Ifmap data, significantly influences on-chip Ifmap storage and off-chip weight bandwidth (BW). The uniform application of computational granularity across all nodes inevitably impacts the memory-BW tradeoff. This article introduces a novel streaming accelerator with a hybrid computational granularity (HCG) scheme. Each node employs an independently optimized computational granularity, enabling a more flexible memory-BW tradeoff and more effective utilization of FPGA resources. However, this hybrid scheme can introduce pipeline bubbles and increase system pipeline complexity and control logic. To address these challenges, this article theoretically analyzes the impact of computational granularity on individual computing nodes and the overall system, aiming to establish a seamless system pipeline without pipeline bubbles and simplify system design. Furthermore, the article develops a hardware overhead model and employs a heuristic algorithm to optimize computational granularity for each computing node, achieving optimal memory-BW tradeoff and higher throughput. Finally, the effectiveness of the proposed design and optimization methodology is validated through the implementation of a 3-TOPS ResNet-18 accelerator on the Alveo U250 development board under BW constraints of 25, 20, and 15 GB/s. Additionally, accelerators for 4-TOPS VGG-16, 4-TOPS ResNet-34, 5-TOPS ResNet-50, 3-TOPS MobileNetV1, 4-TOPS ConvNeXt-T, and 4-TOPS ResNeXt-50 are implemented, surpassing the performance of most existing works.
Wenjin Huang, Conghui Luo, Baoze Zhao, Han Jiao 0003, Yihua Huang 0005
IEEE Trans. Neural Networks Learn. Syst.1
2025 HIN: Hierarchical Interaction Network for Image Captioning
abstract
The purpose of the image captioning task is to understand the content of an image and generate corresponding descriptive text. Traditional approaches to image captioning typically generate descriptive text by extracting different types of visual features from an image and performing feature interactions. However, these methods often fail to fully exploit the interactions between different types of visual features, leading to suboptimal feature integration. To address this limitation, we propose a novel Hierarchical Interaction Network (HIN) , designed to continuously extract and interact with different types of visual features to perform more effective multilevel feature interactions. Our HIN consists of three key modules: firstly, we design the Cross-Type Feature Alignment (CTFA) encoder, which aligns different types of visual features by three global features, so that the subsequent modules can effectively carry out the Hierarchical Interaction (HI) ; secondly, the HI module, which utilizes different types of multilevel features output from the encoder to carry out feature interactions and information mining, so as to generate fully mined multilevel features. The Bottom-up Gated Attention Fusion (BGAF) decoder is finally designed to perform the multilevel decoding of the features mined by our HI module, further enhancing the feature interaction capabilities of our HIN. Moreover, additional experiments on the MS-COCO dataset show that our model achieves new state-of-the-art performance. All codes are available at https://github.com/songchuanle-1/HIN .
Chuanle Song, Wei Zhou 0042, Han Jiao 0003, Wenjin Huang, Yihua Huang 0005
ACM Trans. Multim. Comput. Commun. Appl.4
2024 MRH-GCN: An Efficient GCN Accelerator for Multi-Relation Heterogeneous Graph
abstract
Heterogeneous Graph Convolutional Networks (HGCNs) enhance GCN’ representation learning on diverse graphs by effectively capturing complex relation information. However, the hardware deployment of HGCN remains a challenge. First, the complex network model, massive volume of data, and cross-type message passing of HGCN pose challenges to storage and computational efficiency. Additionally, existing GCN accelerators struggle to handle HGCN for multi-relation modeling and fail to efficiently schedule data. To address this, we proposed a novel hardware accelerator for GCN inference called MRH-GCN for heterogeneous graph data and multirelation HGCN models. Firstly, the proposed MRH-GCN adopted a configurable multiple adjacency matrices fusion strategy to reduce the dimensionality of matrices. Simultaneously, it utilized a sparse-matrix-tile space search compression algorithm to minimize sparse data in the adjacency and feature matrices, reducing data sparsity and enhancing indexing efficiency. Secondly, a regular array pipeline based on data reuse was proposed for scheduling three execution stages in HGCN, matching memory access and computation time, and reducing on-chip memory pressure. Additionally, We designed an SpMM systolic array to enhance computational efficiency, and relation-awareness was introduced in the combination stage to reduce unnecessary computations. Experimental results show that the accelerators produced by our framework achieve significant speedup compared with state-of-the-art DGL implementations on CPU (16.56 × and 30.69 ×), GPU (2.71× and 2.30×) for R-GCN and CompGCN model, respectively.
Wenlu Peng, Wenjin Huang, Yihua Huang 0005
FCCM3
2024 A Novel FPGA Accelerator of R(2+1)D
abstract
As an extension of traditional CNNs, three dimensional convolutional neural networks (3D CNNs) impose significantly higher computational and storage demands compared to their two-dimensional counterparts (2D CNNs). This has led to increased interest in exploring FPGA hardware acceleration methods for 3D CNNs. The R(2+l)D model significantly improves performance compared to conventional 3D CNN architectures. However, due to its incorporation of multiple convolution kernel types and shortcut connections, its hardware design complexity surpasses that of traditional 3D CNNs. This paper presents an efficient computing architecture for (2+l)D convolutions and an FPGA acceleration system for R(2+1)D models. Through systematic exploration of the design space and aligning architectural strategies with FPGA device layout characteristics, a balanced trade-off between parallelism and resource utilization is achieved within the acceleration system. Performance evaluation conducted on the Alveo U50 platform reveals that the accelerator achieves a throughput of 1511 GOP/s. Compared to the Ryzen 3960X CPU, our design exhibits an 8.8× throughput enhancement and a significant 83.5% reduction in latency. Similarly, compared to the RTX 3090 GPU, our design achieves an impressive 12.3×improvement in power efficiency. In contrast to the latest FPGA hardware acceleration design for the R(2+l)D model, our design not only achieves higher throughput but also delivers a 4.7× enhancement in DSP efficiency. When compared to hardware accelerators designed for C3D models, our design demonstrates a roughly equivalent throughput.
Dehao Xiang, Wenjin Huang, Yihua Huang 0005
FCCM3
2024 SpGCN: An FPGA-Based Graph Convolutional Network Accelerator for Sparse Graphs
abstract
Graph convolutional network(GCN) has become one of the most popular graph neural network(GNN) models today. Because the calculation-intensive operations and the memory access-intensive operations during GNN inference, existing GNN accelerators utilize many computing and storage resources but the computing efficiency is still low when the feature matrix or the adjacency matrix is extremely sparse. Therefore, we propose SpGCN, a computationally efficient GCN accelerator suitable for sparse graphs. We design corresponding optimization methods of scheduling, caching and indexing in both combination stage and aggregation stage, aiming at fewer congestion and higher efficiency. Experimental results show that the total computing efficiency is 68.7% on average. Compared with peer work, SpGCN can achieve lower inference delay by using only 18% computing resources and 50 % storage resources.
Xiangzhi Xu, Qi Liu 0061, Wenjin Huang, Wenlu Peng, Yihua Huang 0005
FCCM3
2024 HR-GCN: An Efficient GCN Accelerator for Heterogeneous Graph Data and R-GCN Model
abstract
GCN hardware accelerators currently cater to homogeneous graphs with single connection relationships. Real-world graphs, more diverse in node and connection relationships, benefit from heterogeneous graph models like R-GCN. However, R-GCN's complex architecture and numerous weights result in higher computational and memory demands. Customized R-GCN accelerators are scarce, and running R-GCN on existing GCN accelerators proves highly inefficient. To address this issue, a novel hardware accelerator, HR-GCN, was developed specifically for GCN inference in scenarios involving heterogeneous graph data and the R-GCN Model. First, HR-GCN adopts an adjacency matrix tile-fusing strategy for multiple connection relationships, reducing data dimensions and sparseness in sliced data. Additionally, it utilizes a hybrid inner and outer product tile strategy to significantly enhance the on-chip data reuse rate. Then, to optimize irregular data scheduling and calculation in the aggregation stage, a compressed ring computing array efficiently executes aggregation calculations with a regular ring data flow while compressing the sparse space within the tiles. Finally, for the computation-intensive combination stage, a connection-aware combination processing element is designed to introduce the connection relationship between nodes in advance, reducing the intensive calculation of non-contributing nodes in the tile and reducing computational resource overhead. Compared with the PyG and DGL acceleration frameworks on the 10700K and 1080Ti platforms, HR-GCN achieves an average performance improvement of 56.97× and 1.7× on three public datasets.
Shengjun Xu, Wenlu Peng, Wenjin Huang, Qi Liu 0061, Yihua Huang 0005
FPGA3
2023 A Novel Hardware Accelerator of NeRF Based on Xilinx UltraScale and UltraScale+ FPGA
abstract
Neural Radiance Fields (NeRF) has shown its superiority in various fields, including 3D reconstruction and inverse rendering. However, due to its high computational demands, NeRF typically requires implementation on high-performance GPUs, which is costly and power-intensive, thus limiting its applicability. Compared to GPUs, FPGAs offer a potential solution to implement NeRF at lower cost and power. However, FPGA-based NeRF designs are still rare. To address this issue, a novel hardware NeRF accelerator based on Xilinx UltraScale and UltraScale+ FPGA has been proposed. The proposed design is based on the Multiresolution Hash Encoding NeRF algorithm and comprises Feature Reader, Multilayer Perceptrons (MLP), and Volume Renderer. The Feature Reader and MLP calculate the coordinates and directions of the sampling points on the rays, which are then used to determine the colors and densities of these points. To prevent timing problems and routing congestion caused by high DSPs utilization, we use registers and cascade paths between DSPs to construct MACC matrices in MLP. The Volume Renderer processes these colors and densities to obtain the colors of rays. The proposed accelerator runs at 300MHz on Xilinx Alveo U250, achieving an average 1.60 × performance improvement compared to NVIDIA GTX 1080 Ti. Additionally, the accelerator has reduced energy consumption per image by an average of 5.21x and 4.64x compared to NVIDIA GTX 1080 Ti and NVIDIA RTX 3090, respectively.
Baoze Zhao, Wenjin Huang, Yihua Huang 0005
FPL2
2023 The Learnable Model-Based Genetic Algorithm for the IP Mapping Problem
abstract
The intellectual property (IP) mapping problem is an NP-hard problem in network-on-chip (NoC) designs and is often solved by the genetic Algorithm. In the genetic algorithm (GA), new populations are generated by crossover and mutation operators. However, these operators consider neither the prior knowledge of the IP mapping problem nor the historical experience of the population evolution, easily causing the GA to converge prematurely to the local optima. To solve this problem, a learnable model-based GA (LMGA) is proposed, which introduces a learnable model and a hybridization scheme of the model and GA. 1) the learnable model is implemented by a novel neural network-based probability model for the IP mapping problem, i.e., the message passing attention network (MAN). Moreover, the learnable model is pretrained then updated by learning the features of high-fitness individuals among the population to predict a better probability distribution of the optimal solution in a specific IP mapping problem. 2) The scheme updates the population of certain generations of the GA by sampling the learnable model. Hence, the population evolution is guided by the learnable model to avoid premature convergence. Simulation results show that the MAN achieves a faster training speed and models the large-scale IP mapping problem better than the message passing neural network-pointer network (MPN) in the previous work. The LMGA saves an average of 5.68% and 5.23% in the communication energy and average network delay, respectively, than the state-of-the-art algorithm, i.e., the message passing neural network-pointer network-based genetic algorithm (MPN-GA).
Qingkun Chen, Wenjin Huang, Yihua Huang 0005
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2022 TFR-GCN: A GCN Accelerator with Tile-Fusing Strategy
abstract
Graph convolutional networks(GCN) are widely used because of their superior ability in the graph data processing. However, the irregular data access and the computing intensity in GCN inference bring great challenges to the existing hardware accelerators. Hence, this paper proposes an accelerator using a tiles-fusing rings array (TFR-GCN) in the aggregation phase to achieve regular memory access and improve computing resource utilization. Moreover, this paper proposes a backpressure index calculation unit to optimize sparse-dense high-dimensional vector multiplication in the combination phase. Compared with the performance of software accelerators running on CPU and GPU, this accelerator achieved 208×, 50 × performance improvement on average. Compared with the existing FPGA-based accelerators, this accelerator achieved a performance improvement of (1.1-147)×.
Shengjun Xu, Wenjin Huang, Yihua Huang 0005
FCCM2
2022 FPGA-Based High-Throughput CNN Hardware Accelerator With High Computing Resource Utilization Ratio
abstract
The field-programmable gate array (FPGA)-based CNN hardware accelerator adopting single-computing-engine (CE) architecture or multi-CE architecture has attracted great attention in recent years. The actual throughput of the accelerator is also getting higher and higher but is still far below the theoretical throughput due to the inefficient computing resource mapping mechanism and data supply problem, and so on. To solve these problems, a novel composite hardware CNN accelerator architecture is proposed in this article. To perform the convolution layer (CL) efficiently, a novel multiCE architecture based on a row-level pipelined streaming strategy is proposed. For each CE, an optimized mapping mechanism is proposed to improve its computing resource utilization ratio and an efficient data system with continuous data supply is designed to avoid the idle state of the CE. Besides, to relieve the off-chip bandwidth stress, a weight data allocation strategy is proposed. To perform the fully connected layer (FCL), a single-CE architecture based on a batch-based computing method is proposed. Based on these design methods and strategies, visual geometry group network-16 (VGG-16) and ResNet-101 are both implemented on the XC7VX980T FPGA platform. The VGG-16 accelerator consumed 3395 multipliers and got the throughput of 1 TOPS at 150 MHz, that is, about 98.15% of the theoretical throughput ( 2 ×3395 ×150 MOPS). Similarly, the ResNet-101 accelerator achieved 600 GOPS at 100 MHz, about 96.12% of the theoretical throughput ( 2 ×3121 ×100 MOPS).
Wenjin Huang, Huangtao Wu, Qingkun Chen, Conghui Luo, Shihao Zeng, Tianrui Li 0005, Yihua Huang 0005
IEEE Trans. Neural Networks Learn. Syst.1
2021 A Reinforcement Learning-Based Framework for Solving the IP Mapping Problem
abstract
In network-on-chip (NoC) designs, the intellectual property (IP) mapping problem is a critical issue and is usually solved by heuristic searches. However, heuristic searches suffer from the problem of easily falling into the local optimum. To tackle this problem, this article proposes a reinforcement learning-based framework (RLF), which enhances the performance of heuristic searches through the neural network-based probability model. Within this framework, first, a neural network-based probability model for IP mapping is built and trained by reinforcement learning instead of supervised learning to overcome the difficulty of obtaining a high-quality labeled training set. Second, based on the pretrained probability model, the model-based heuristic uses the probability model to generate the initial population and then employs heuristic searches to find the optimal solution. Two model-based heuristics, i.e., the message passing neural network-pointer network-based genetic algorithm (MPN-GA) and the message passing neural network-pointer network-based PSMAP (MPN-PSMAP), are proposed as specific instances. Simulation results show that the MPN-GA reduces the communication cost by an average of 9.32% than the genetic algorithm (GA). The MPN-PSMAP achieves an average reduction in the communication cost of 8.37% than the PSMAP. Finally, two extensions are given as examples to show the good extensibility of this framework.
Qingkun Chen, Wenjin Huang, Yuze Peng, Yihua Huang 0005
IEEE Trans. Very Large Scale Integr. Syst.2
2021 An IP Core Mapping Algorithm Based on Neural Networks
abstract
The IP core mapping optimization problem is an NP-hard problem in network-on-chip design. Because of the computational complexity of an IP core mapping, the MPNN-Ptr networks composed of the graph networks and the pointer networks are proposed to model the IP core mapping. The neural network IP core mapping model (NNMM) can effectively evaluate the probability of each mapping solution as the optimal solution to an IP core mapping optimization problem. Then, an IP core mapping algorithm, the neural mapping algorithm (NMA), is proposed. In this algorithm, the global search is realized by sampling the mapping solutions with a high probability evaluated by NNMM. The sampled candidate mapping solutions according to probability can effectively decrease the candidate mapping solution space, which can reduce invalid searches and avoid getting stuck at local minima. Then, the 2-opt algorithm is used as a local search to improve the quality of mapping further. Simulation results show that the MPNN-Ptr networks can effectively model the IP core mapping. Compared with the state-of-the-art mapping algorithms, NMA produces better solutions for both specific applications and random applications. NMA achieves a 7.90% communication cost reduction on average than the classic mapping algorithm: NMAP.
Qingkun Chen, Wenjin Huang, Yuanshan Zhang, Yihua Huang 0005
IEEE Trans. Very Large Scale Integr. Syst.2
2020 Toward Constructing a Real-time Social Anxiety Evaluation System: Exploring Effective Heart Rate Features
abstract
Social anxiety is a negative emotion which may impair the health of the heart and social functioning of an individual. This work analyzes the influence of social anxiety on the autonomic nerve control of the heart in two social exposure events: public speaking and thesis defending. In an experiment of public speaking, 59 human subjects were tested, and 11 conventional heartbeat measures and a heartbeat measure named the range of local Hurst exponents (RLHE) were evaluated for their capabilities to reveal the onset of social anxiety. Two-sample t-test between the baseline data and high anxiety data shows that social anxiety significantly reduces the complexity of the heartbeats. In an experiment of thesis defense, heartbeats data were acquired from nine graduate students. With the combination of three conventional features and the RLHE feature, a support vector machine classifier obtained true positive rate and true negative rate of 84.88 and 97.29 percent in the five-fold cross validation process of binary classification between high anxiety status and low anxiety status; the classifier also realized a generalization accuracy of 81.82 percent in detecting the high anxiety status in the thesis defense. A real-time anxiety monitoring system was established based on the above anxiety detecting method.
Wanhui Wen, Guangyuan Liu 0005, Zhi-Hong Mao, Wenjin Huang, Jiemin Yang, Wenyan Jia
IEEE Trans. Affect. Comput.4
2014 Emotion Recognition Based on Multi-Variant Correlation of Physiological Signals
abstract
Emotion recognition based on affective physiological changes is a pattern recognition problem, and selecting specific physiological signals is necessary and helpful to recognize the emotions. Fingertip blood oxygen saturation (OXY), galvanic skin response (GSR) and heart rate (HR) are acquired while amusement, anger, grief and fear of 101 subjects are individually elicited by films. The affective physiological changes in multi-subject GSR, the first derivative of GSR (FD_GSR) and HR are detected by the multi-variant correlation method. The correlation analysis reveals that multi-subject HR, GSR and FD_GSR fluctuations respectively have common intra-class affective patterns. In addition to the conventional features of HR and GSR, the affective HR, GSR and FD_GSR fluctuations are quantified by the local scaling dimension and applied as the affective features. The multi-subject affective database containing 477 cases is classified by a Random Forests classifier. An overall correct rate of 74 percent for quinary classification of amusement, anger, grief, fear and the baseline state are obtained.
Wanhui Wen, Guangyuan Liu 0005, Nanpu Cheng, Pengchao Shangguan, Wenjin Huang
IEEE Trans. Affect. Comput.6