EDBT 2026 Demo / reviewers in the wild / expert
Huapeng Wu
dblp:34/5313
· DBLP profile ↗
50ranked-venue papers
23as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 11 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Security and privacy · 3 · 3 first-authorComputer networks · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A large language model-driven framework for multimodal cognitive workload prediction in industrial human-robot collaboration
Qiwei Xue, Yuchong Zhang, Huapeng Wu, Yuntao Song |
Eng. Appl. Artif. Intell. | 3 |
| 2026 | IMENet: infrared-guided multimodal enhancement network for low-light vision
Zhikai Wei, Huapeng Wu, Chenyang Lu 0010, Zebin Wu 0001, Tianming Zhan |
Multim. Syst. | 3 |
| 2026 | A hybrid spatial and spectral mamba network for hyperspectral image super-resolution
Huapeng Wu, Yimeng Shi, Tianming Zhan |
Multim. Syst. | 4 |
| 2025 | A Kinematics Optimization Framework with Improved Computational Efficiency for Task-Based Optimum Design of Serial Manipulators in Cluttered EnvironmentsabstractIt is challenging to find optimum kinematic designs for non-standard robotic manipulators, e.g., medical, nuclear, and space manipulators, which are demanded to adapt to arbitrary complex tasks in constraints. Such design optimization can be modelled as a multi-dimensional non-convex optimization problem with nonlinear constrained conditions. However, it is non-trivial to ensure the essential reachability condition, i.e., the existence of continuous trajectories between demand positions for serial articulated manipulators, given complex spatial constraints, like obstacles and boundaries. Traditional solutions integrate standard motion planning or inverse kinematics algorithms within a kinematic-design optimization process, resulting in significant demand for time and computing resources. To accelerate design optimization at improved efficiency, we design a novel robust design framework built on a new kinematic design synthesis, which allows for simultaneously optimizing dimension and topology of a serial manipulator's kinematics for arbitrary tasks in constrained environments, using a generalised parametric kinematic model. Significantly, in contrast to standard solutions, we develop a novel computationally effective reachability verification method, which rapidly aborts infeasible motions by exploiting efficient collision checks, based on the Rapidly-exploring Random Tree (RRT) algorithm. The effectiveness of the proposed design framework is verified and evaluated by comparing to baseline benchmarks. Results demonstrate the novel design framework can accelerate kinematic design optimization by an order of magnitude compared to the current state-of-the-art, and optimise link dimension and joint type simultaneously of serial robots for cluttered environments. Nikola Petkov, Ozan Tokatli, Kaiqiang Zhang, Huapeng Wu, Robert Skilton |
ICRA | 4 |
| 2025 | Mastering autonomous assembly in fusion application with learning-by-doing: A peg-in-hole study
Ruochen Yin, Huapeng Wu, Ming Li 0067, Yuntao Song, Hongtao Pan, Heikki Handroos |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | A novel gradient and semantic-aware transformer network for low-light image enhancement
Tianming Zhan, Chenyang Lu 0010, Huapeng Wu, Chenyun Wang |
Multim. Syst. | 3 |
| 2025 | Spatial-Spectral Cross Mamba Network for Hyperspectral and Multispectral Image FusionabstractCurrently, hyperspectral and multispectral image fusion methods based on local and global feature learning (e.g., CNN and Transformer) have achieved promising results. However, as the core part of transformer, the computational cost of the self-attention is quadratic with the image size, which severely limits its practical application. In this paper, we propose a spatial-spectral cross mamba network (SSCM) for hyperspectral and multispectral image fusion. By using the mamba structure, our model is able to obtain long-range spatial-spectral information with less computational complexity in comparison with the transformer structure. Specifically, we introduce a spatial-spectral cross mamba block to facilitate the interaction between hyperspectral and multispectral features, effectively enhancing the spatial-spectral feature representation ability of the network. In addition, a cross-scale spatial-spectral learning module based on the U-shaped structure is proposed to effectively extract the long-range high-frequency feature information at different scales. Extensive experimental results demonstrate that our method achieves comparable performance in comparison with some state-of-the-art image fusion methods. Huapeng Wu, Jiaqiang Qi, Tianming Zhan, Yang Xu 0006, Zhihui Wei |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2025 | KANformer: Dual-Priors-Guided Low-Light Enhancement via KAN and TransformerabstractImages captured under low-light conditions suffer from poor visibility and clarity due to insufficient light. The emergence of deep learning has greatly boosted the development of low-light enhancement techniques and achieved promising results. However, while these low-light enhancement methods have enhanced the perceptual effects of human vision, their results in high-level visual tasks (e.g., object detection and semantic segmentation) are still unstable and even sometimes bring negative effects. Therefore, in this work, we propose a new model, KANformer, which uses a semantic-gradient prior as a guide to recover pixels relevant to the image subject from both high-frequency and low-frequency perspectives. Specifically, our model consists of three key components: Low-Frequency Enhancement (LFE) module, which aims to enhance the restoration of the image subject via the semantic prior obtained from SAM; Low-Frequency-Based High-Frequency Enhancement (LFHE) module, which utilizes the KAN module to obtain information from the low-frequency features conducive to the enhancement of high-frequency features; and Gradient-Based High-Frequency Enhancement (GHE) module, which aims to utilize the original gradient as prior to further enhance the structural information of the image and reduce the effect of noise. In addition, we introduce the discrete wavelet transform as down-sampling method while transforming the spatial domain features to the frequency domain for processing. Experiments on multiple paired and unpaired datasets show that our method achieves better visualization and image fidelity compared to other state-of-the-art methods. In addition, experiments on object detection and segmentation show that our method provides better enhancement in improving low-light high-level vision tasks. Chenyang Lu 0010, Zhikai Wei, Huapeng Wu, Le Sun 0003, Tianming Zhan |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | A multi-level digital twin construction method of assembly line based on hybrid worker digital twin models
Youmin Hu, Huapeng Wu, Ming Li 0067, Heikki Handroos, Bo Wu 0006 |
Adv. Eng. Informatics | 5 |
| 2024 | A hybrid U-shaped and transformer network for change detection in high-resolution remote sensing imagesabstractAbstract Deep convolutional neural networks based remote sensing change detection has recently shown significant performance improvement. However, small region changes and global‐local features in high‐resolution remote sensing images are not fully explored. This paper introduces a hybrid U‐shaped and transformer network for change detection in high‐resolution remote sensing images. Specifically, a UNet++‐based backbone to facilitate feature learning across different scales. In addition, we introduce a transformer‐based feature fusion module for extracting long‐range dependencies, which can enhance the representation ability of the network. Furthermore, the introduced efficient channel attention mechanism can efficiently calibrate the feature representation and concentrate on more important feature information. Thanks to the above designs, the proposed method enjoys a strong ability to extract local and global features for remote sensing change detection. Extensive experimental results on different remote sensing images show that our method can achieve superior performance in comparison with state‐of‐the‐art change detection methods. Huapeng Wu, Mengxue Yuan, Tianming Zhan |
IET Image Process. | 1 |
| 2024 | HCT: a hybrid CNN and transformer network for hyperspectral image super-resolution
Huapeng Wu, Chenyun Wang, Chenyang Lu 0010, Tianming Zhan |
Multim. Syst. | 1 |
| 2024 | A novel spatial and spectral transformer network for hyperspectral image super-resolution
Huapeng Wu, Tianming Zhan |
Multim. Syst. | 1 |
| 2023 | Feedback Pyramid Attention Networks for Single Image Super-ResolutionabstractRecently, convolutional neural network (CNN) based image super-resolution (SR) methods have achieved significant performance improvement. However, most CNN-based methods mainly focus on feed-forward architecture design and neglect to explore the feedback mechanism, which usually exists in the human visual system. In this paper, we propose feedback pyramid attention networks (FPAN) to fully exploit the mutual dependencies of features. Specifically, a novel feedback connection structure is developed to enhance low-level feature expression with high-level information. In our method, the output of each layer in the first stage is also used as the input of the corresponding layer in the next state to re-update the previous low-level filters. Moreover, we introduce a pyramid non-local structure to model global contextual information in different scales and improve the discriminative representation of the network. Extensive experimental results on various datasets demonstrate the superiority of our FPAN in comparison with the state-of-the-art SR methods. Huapeng Wu, Jie Gui, Jun Zhang 0024, James T. Kwok, Zhihui Wei |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Pyramidal dense attention networks for single image super-resolutionabstractAbstract Recently, residual and dense networks have effectively promoted the development of image super‐resolution (SR). However, most dense networks based SR methods do not make full use of dense feature information. To solve this problem, a pyramidal dense attention network for single image super‐resolution is proposed in this paper. In this method, the proposed pyramidal dense learning can gradually increase the width of the densely connected layer inside a pyramidal dense block to extract deep features efficiently. Meanwhile, the adaptive group convolution that the number of groups grows linearly with dense convolutional layers is introduced to relieve the parameter explosion. Besides, a novel joint attention to capture cross‐dimension interaction between the spatial dimensions and channel dimension in an efficient way for providing rich discriminative feature representations is also proposed. Extensive experimental results show that the method achieves comparable performance in comparison with the state‐of‐the‐art SR methods. Huapeng Wu, Jie Gui, Jun Zhang 0024, James T. Kwok, Zhihui Wei |
IET Image Process. | 1 |
| 2022 | Area-Efficient Finite Field Multiplication Using Hybrid SET-MOS TechnologyabstractSingle-electron transistors (SETs) exhibit a unique characteristic of Coulomb oscillation which can find many digital applications with area efficiency. More specifically, both MOS and SET devices can be used to implement XOR gates with almost the same area costs regardless of the number of their inputs, outperforming pure CMOS solutions. As multiple-input XOR gates are abundant in the finite field polynomial multiplication which represents the most frequent computation in elliptic curve cryptosystem, hybrid SET-MOS technology can substantially reduce the area cost for this application. This paper presents polynomial multiplication architectures with hybrid SET-MOS transistors and explores Karatsuba-algorithm based multiplication for further area optimization. Simulations show that the proposed hybrid SET-MOS implementations can typically provide around 37% savings in terms of gate count compared to their traditional CMOS counterparts. Huapeng Wu, Chunhong Chen |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2022 | An Efficient Cross-Modality Self-Calibrated Network for Hyperspectral and Multispectral Image FusionabstractRecently, deep convolutional neural network based hyperspectral and multispectral image fusion methods have shown significant performance. Nevertheless, the rich spatial and spectral details of hyperspectral images (HSIs) have not been fully explored, leaving room for further improve the representation ability of the model. In this paper, we propose an efficient cross-modality self-calibrated network (CMSCN) for hyperspectral and multispectral image fusion. Specifically, we use a cross-modality non-local module to fuse a high-resolution multispectral image (HR-MSI) and a low-resolution hyperspectral image (LR-HSI) to get an enhanced LR-HSI. In addition, a novel cross-scale self-calibrated convolution structure is proposed to explore and exploit multi-scale and hierarchical spatial-spectral features, which can improve the learning ability of the model. The introduced efficient spatial-spectral attention mechanism can calibrate the feature representation at different dimensions, thereby providing more efficient and accurate information for hyperspectral image reconstruction. Extensive experimental results on various hyperspectral images demonstrate the superiority of our method in comparison with the state-of-the-art image fusion methods. Huapeng Wu, Jie Gui, Yang Xu 0006, Zebin Wu 0001, Yuan Yan Tang, Zhihui Wei |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | A Novel Cross-Scale Octave Network for Hyperspectral and Multispectral Image FusionabstractRecently, deep convolutional neural network-based low-resolution hyperspectral image (LR-HSI) and high-resolution multispectral image (HR-MSI) fusion methods have achieved significant performance improvement. However, the rich spatial and spectral information in HSIs is not fully explored. In this article, we propose a novel cross-scale octave network (CSONet) for hyperspectral and multispectral image fusion. Specifically, we adopt a progressive image fusion structure to effectively extract the spatial and spectral information of HR-MSI at multiple resolutions, thereby efficiently complementing LR-HSI’s information. In addition, the proposed cross-scale octave convolution module can extract rich multiscale spatial feature information and concentrate on more important spatial–spectral features at different scales with the multiscale spatial–spectral attention mechanism. Finally, a multisupervised loss function is used to improve the gradient propagation and enhance the representation ability of the network. Ablation analysis on the benchmark datasets shows the effectiveness of each component in the proposed method. Extensive experimental results on different hyperspectral images demonstrate that the proposed CSONet can achieve superior results and strong generalization ability in comparison with some state-of-the-art LR-HSI and HR-MSI fusion methods. Tianming Zhan, Zuolin Bi, Huapeng Wu, Qian Du 0001, Yang Xu 0006, Zebin Wu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | FusionLane: Multi-Sensor Fusion for Lane Marking Semantic Segmentation Using Deep Neural NetworksabstractEffective semantic segmentation of lane marking is crucial for construction of high-precision lane level maps. In recent years, a number of different methods for semantic segmentation of images have been proposed. These methods concentrate mainly on analysis of camera images, due to limitations with the sensor itself, and thus far, the accurate three-dimensional spatial position of the lane marking could not be obtained, which hinders lane level map construction.This article proposes a lane marking semantic segmentation method based on LIDAR and camera image fusion using a deep neural network. In the approach, the object of the semantic segmentation is a bird’s-eye view converted from a LIDAR points cloud instead of an image captured by a camera. First, the DeepLabV3+ network image segmentation method is used to segment the image captured by the camera, and the segmentation result is then merged with the point clouds collected by the LIDAR as the input of the proposed network. A long short-term memory (LSTM) structure is added to the neural network to assist the network in semantic segmentation of lane markings by enabling use of time series information. Experiments on datasets containing more than 14,000 images, which were manually labeled and expanded, showed that the proposed method provides accurate semantic segmentation of the bird’s-eye view LIDAR points cloud. Consequently, automation of high-precision map construction can be significantly improved. Our code is available athttps://github.com/rolandying/FusionLane. Ruochen Yin, Huapeng Wu, Yuntao Song, Biao Yu, Runxin Niu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Multi-Grained Attention Networks for Single Image Super-ResolutionabstractDeep Convolutional Neural Networks (CNN) have drawn great attention in image super-resolution (SR). Recently, visual attention mechanism, which exploits both of the feature importance and contextual cues, has been introduced to image SR and proves to be effective to improve CNN-based SR performance. In this paper, we make a thorough investigation on the attention mechanisms in a SR model and shed light on how simple and effective improvements on these ideas improve the state-of-the-arts. We further propose a unified approach called “multi-grained attention networks (MGAN)” which fully exploits the advantages of multi-scale and attention mechanisms in SR tasks. In our method, the importance of each neuron is computed according to its surrounding regions in a multi-grained fashion and then is used to adaptively re-scale the feature responses. More importantly, the “channel attention” and “spatial attention” strategies in previous methods can be essentially considered as two special cases of our method. We also introduce multi-scale dense connections to extract the image features at multiple scales and capture the features of different layers through dense skip connections. Ablation studies on benchmark datasets demonstrate the effectiveness of our method. In comparison with other state-of-the-art SR methods, our method shows the superiority in terms of both accuracy and model size. Huapeng Wu, Zhengxia Zou, Jie Gui, Wen-Jun Zeng, Jieping Ye, Jun Zhang 0024, Hongyi Liu 0001, Zhihui Wei |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | A combination of CSP-based method with soft margin SVM classifier and generalized RBF kernel for imagery-based brain computer interface applicationsabstractAbstract Several methods utilizing common spatial pattern (CSP) algorithm have been presented for improving the identification of imagery movement patterns for brain computer interface applications. The present study focuses on improving a CSP-based algorithm for detecting the motor imagery movement patterns. A discriminative filter bank of CSP method using a discriminative sensitive learning vector quantization (DFBCSP-DSLVQ) system is implemented. Four algorithms are then combined to form three methods for improving the efficiency of the DFBCSP-DSLVQ method, namely the kernel linear discriminant analysis (KLDA), the kernel principal component analysis (KPCA), the soft margin support vector machine (SSVM) classifier and the generalized radial bases functions (GRBF) kernel. The GRBF is used as a kernel for the KLDA, the KPCA feature selection algorithms and the SSVM classifier. In addition, three types of classifiers, namely K-nearest neighbor (K-NN), neural network (NN) and traditional support vector machine (SVM), are employed to evaluate the efficiency of the classifiers. Results show that the best algorithm is the combination of the DFBCSP-DSLVQ method using the SSVM classifier with GRBF kernel (SSVM-GRBF), in which the best average accuracy, attained are 92.70% and 83.21%, respectively. Results of the Repeated Measures ANOVA shows the statistically significant dominance of this method atp< 0.05. The presented algorithms are then compared with the base algorithm of this study i.e. the DFBCSP-DSLVQ with the SVM-RBF classifier. It is concluded that the algorithms, which are based on the SSVM-GRBF classifier and the KLDA with the SSVM-GRBF classifiers give sufficient accuracy and reliable results. Amin Hekmatmanesh, Huapeng Wu, Fatemeh Jamaloo, Ming Li 0067, Heikki Handroos |
Multim. Tools Appl. | 2 |
| 2019 | Combination of discrete wavelet packet transform with detrended fluctuation analysis using customized mother wavelet with the aim of an imagery-motor control interface for an exoskeletonabstractOne critical issue in brain computer interface (BCI) studies is to extract imaginary movement patterns from electroencephalograph (EEG). In this study, two different techniques —detrended fluctuation analysis (DFA) and discrete wavelet packet transform— are combined (DWPT-DFA) for feature extraction. Both approaches are known as self-similarity quantifier techniques. In wavelet technique, mother wavelets play an important role. Herein, A customized mother wavelet utilizing event related desynchronization (ERD) potential patterns are extracted and updated automatically for individual subjects. Also, three predefined mother wavelets are used, and the results are compared with the customized mother wavelet. The predefined mother wavelets are db4, db8 and coiflet 4. The soft margin support vector machine with the generalized radial basis function (SSVM-GRBF) is employed to classify the DWPT-DFA features. For the efficiency of the method, nine subjects have participated to record EEG based on the imaginary hand movements. The ERDs and features are extracted from FC1 and CP6 channels. Results show that the combination of the DWPT and DFA with the personalize ERD mother wavelet gives the best accuracy of 85.33% with p < 0.001. Based on the results, we conclude that the DWPT-DFA method using the ERD mother wavelets improves significantly the efficiency of the SSVM-GRBF classifier. Amin Hekmatmanesh, Huapeng Wu, Ali Motie Nasrabadi, Ming Li 0067, Heikki Handroos |
Multim. Tools Appl. | 2 |
| 2018 | A Self-adaptive Artificial Bee Colony Algorithm with Guard Stage for Global OptimizationabstractThe artificial bee colony (ABC) algorithm is a heuristic optimization algorithm based on the behavior of honeybee swarms. Inspired by particle swarm optimization (PSO) and differential evolution (DE) algorithms, we propose an improved ABC algorithm, named SAG-ABC, which incorporates a self-adaptive employed bees and guard stage to construct a more efficient algorithm. This algorithm combines the advantages of ABC algorithm, which has good exploration capability and global search ability and ease of implementation with fewer control parameters, DE and PSO algorithm, which exchange information with several individuals and utilize the history best information. The searching strategies in these different swarm intelligent algorithms are presented. The information is exchanged among individuals or elements. For the new SAG-ABC algorithm, the self-adaptive employed bees are guided by the global history best bee to enable search in a wider area. Then the search results are adapted to a smaller area. The guard stage is applied to improve the search performance of the employed bees phase by controlling the frequency with which the employed bees abandon the food source. Comparisons between the PSO algorithm, DE algorithm and ABC algorithm are made based on 16 benchmark functions. The results demonstrate the good performance and searching ability of the proposed algorithm. Bingyam Mao, Zhijiang Xie, Huapeng Wu, Heikki Handroos |
CEC | 4 |
| 2018 | Development and Error Compensation of a Flexible Multi-Joint Manipulator Applied in Nuclear Fusion EnvironmentabstractExperimental Advanced Superconducting Tokamak (EAST) is the world's first fully superconducting tokamak fusion device with non-circular cross-section which was built in China The EAST articulated maintenance arm (EAMA) system is developed for real-time detection and rapid repair operations to damaged internal components during plasma discharges without breaking the EAST ultra-high vacuum (UHV) condition. To achieve the desired objectives, the EAMA system design should guarantee that the robot can stably run in the harsh environments of high temperature (80-120 °C) and high vacuum (~ 10-5Pa). Meanwhile, the errors caused by the deformation of long flexible robot arms should also be predicted and compensated in real-time to obtain high accuracy for maintenance operations. In this paper, the vacuum-available design scheme of the manipulator system was firstly introduced. Secondly, inverse kinematics and obstacle avoidance strategy of the highly redundant EAMA robot was built. Then, flexible errors were predicted utilizing a back-propagation neural network (BPNN) model which was established on the basis of real experimental data. Finally, an integrated control strategy for error prediction and compensation was developed. Shanshuang Shi, Hongtao Pan, Wenlong Zhao 0002, Huapeng Wu |
IROS | 5 |
| 2017 | Low-Power Design for a Digit-Serial Polynomial Basis Finite Field Multiplier Using Factoring TechniqueabstractIn CMOS-based application-specific integrated circuit (ASIC) designs, total power consumption is dominated by dynamic power, where dynamic power consists of two major components, namely, switching power and internal power. In this paper, we present a low-power design for a digit-serial finite field multiplier in GF(2m). In the proposed design, a factoring technique is used to minimize switching power. To the best of our knowledge, factoring method has not been reported in the literature being used in the design of a finite field multiplier at an architectural level. Logic gate substitution is also utilized to reduce internal power. Our proposed design along with several existing similar works have been realized for GF(2233) on ASIC platform, and a comparison is made between them. The synthesis results show that the proposed multiplier design consumes at least 27.8% lower total power than any previous work in comparison. Shoaleh Hashemi Namin, Huapeng Wu, Majid Ahmadi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | Efficient multiplication architecture over truncated polynomial ring for NTRUEncrypt systemabstractTruncated polynomial ring has important applications in cryptography. It was probably first used in NTRU public key cryptosystem which is one of the most well-known post-quantum cryptosystems. Recently it is found that a modification to NTRU supports somewhat fully homomorphic encryption where a slightly different truncated polynomial ring is adopted. In this paper an efficient architecture is proposed for multiplication over truncated polynomial ring with application for NTRUEncrypt system. The proposed multiplier is based on the compact structure of a modified linear feedback shift register (LFSR) which can reduce the latency for small input polynomial. The compact-designed arithmetic unit capable of performing both modular addition and subtraction takes input from either of two registers on the left hand side. FPGA simulation results show that the product of area and latency for the proposed multiplier is at most 84% compared to any existing work in comparison. Huapeng Wu |
ISCAS | 2 |
| 2015 | Efficient radix conversions for classes of radicesabstractRadix conversion for conventional number systems can be computed using one of the four algorithms reviewed and discussed in [1], depending on whether the source or the destination radix arithmetic is used and whether an integer or a fraction is to be converted. Out of the four radix conversion algorithms, expensive division operations are required in two of them. In this short paper, two new radix conversion algorithms requiring no division are proposed for classes of radices. The new algorithms are expected to replace or be alternate methods to the existing ones with division operations for classes of radices. Huapeng Wu |
ISCAS | 1 |
| 2013 | Low complexity LFSR based bit-serial montgomery multiplier in GF(2m)abstractMontgomery multiplication in GF(2m) is defined as ABr-1mod f(α), where f(x) is the irreducible polynomial defining the field and r is a fixed field element. In this paper, a low complexity Montgomery multiplier in GF(2m) is proposed with r=αm-1or αm. Linear feedback shift register (LFSR) is adopted as the main module for the presented architecture. It is shown that the proposed multiplier has lower space complexity than any of the existing similar works we found in the literature. The time complexity of the proposed multiplier is slightly higher than the best result among the existing works. Huapeng Wu |
ISCAS | 1 |
| 2012 | Current mode multiple-valued adder for cryptography processorsabstractThis paper presents the design and implementation of a multiple-valued adder, with application in smart cards and cryptographic processors. The adder is designed based on the principles of truncated Continuous Valued Number System (CVNS), in order to relax the implementation requirements of the circuit topology and designs. The CVNS adder has an almost constant power consumption, independent of the input values. This feature makes this adder immune against side channel attacks, which may use the power consumption pattern to obtain the intermediate values of arithmetic operations. Ashley Novak, Farinoush Saffar, Mitra Mirhassani, Huapeng Wu |
ISCAS | 4 |
| 2012 | High-Speed Architectures for Multiplication Using Reordered Normal BasisabstractNormal basis has been widely used for the representation of binary field elements mainly due to its low-cost squaring operation. Optimal normal basis type II is a special class of normal basis exhibiting very low multiplication complexity and is considered as a safe choice for hardware implementation of cryptographic applications. In this paper, high-speed architectures for binary field multiplication using reordered normal basis are proposed, where reordered normal basis is referred to as a certain permutation of optimal normal basis type II. Complexity comparison shows that the proposed architectures are faster compared to previously presented architectures in the open literature using either an optimal normal basis type II or a reordered normal basis. One advantage of the new word-level architectures is that the critical path delay is a constant (not a function of word size). This enables the multipliers to operate at very high clock rates regardless of the field size or the number of words. Hardware implementation of some practical size multipliers for elliptic curve cryptography is also included. Ashkan Hosseinzadeh Namin, Huapeng Wu, Majid Ahmadi |
IEEE Trans. Computers | 2 |
| 2012 | An Efficient Finite Field Multiplier Using Redundant RepresentationabstractAn efficient word-level finite field multiplier using redundant representation is proposed. The proposed multiplier has a significantly higher speed, compared to previously proposed word-level architectures using either redundant representation or optimal normal basis type I, at the expense of moderately higher area complexity. Furthermore, the new design out-performs other similar proposals when considering the product of area and delay as a measure of performance. ASIC Realization of the proposed design using TSMC’s .18 um CMOS technology for the binary field size of 163 is also presented. Ashkan Hosseinzadeh Namin, Huapeng Wu, Majid Ahmadi |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2011 | A Word-Level Finite Field Multiplier Using Normal BasisabstractHardware implementations of finite field arithmetic using normal basis are advantageous due to the fact that the squaring operation can be done at almost no cost. In this paper, a new word-level finite field multiplier using normal basis is proposed. The proposed architecture takes d clock cycles to compute the product bits, where the value for d, 1\leq d \leq m, can be arbitrarily selected by the designer to set the tradeoff between area and speed. When there exists an optimal normal basis, it is shown that the proposed design has a smaller critical path delay than other word-level normal basis multipliers found in the literature, while its circuit complexities are moderate and comparable to the others. Different word size multipliers were implemented in hardware, and implementation results are also presented. Ashkan Hosseinzadeh Namin, Huapeng Wu, Majid Ahmadi |
IEEE Trans. Computers | 2 |
| 2009 | Efficient Hardware Implementation of the Hyperbolic Tangent Sigmoid FunctionabstractEfficient implementation of the activation function is important in the hardware design of artificial neural networks. Sigmoid, and hyperbolic tangent sigmoid functions are the most widely used activation functions for this purpose. In this paper, we present a simple and efficient architecture for digital hardware implementation of the hyperbolic tangent sigmoid function. The proposed method employs a piecewise linear approximation as a foundation, and further improves the results using a lookup table. Our design proves to be more efficient considering area times delay as a performance metric when compared to similar proposals. VLSI implementation of the proposed design using a 0.18 mum CMOS process is also presented, which shows a 35% improvement over similar recently published architectures. Ashkan Hosseinzadeh Namin, Karl Leboeuf, Roberto Muscedere, Huapeng Wu, Majid Ahmadi |
ISCAS | 4 |
| 2009 | A High-Speed Word Level Finite Field Multiplier in BBF2m Using Redundant RepresentationabstractIn this paper, a high-speed word level finite field multiplier in F2musing redundant representation is proposed. For the class of fields that there exists a type I optimal normal basis, the new architecture has significantly higher speed compared to previously proposed architectures using either normal basis or redundant representation at the expense of moderately higher area complexity. One of the unique features of the proposed multiplier is that the critical path delay is not a function of the field size nor the word size. It is shown that the new multiplier outperforms all the other multipliers in comparison when considering the product of area and delay as a measure of performance. VLSI implementation of the proposed multiplier in a 0.18- mum complimentary metal-oxide-semiconductor (CMOS) process is also presented. Ashkan Hosseinzadeh Namin, Huapeng Wu, Majid Ahmadi |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2008 | A high speed word level finite field multiplier using reordered normal basisabstractReordered normal basis is a certain permutation of a type II optimal normal basis. In this paper, a high speed design of a word level finite field multiplier using reordered normal basis is presented. Proposed architecture has a very regular structure which makes it suitable for VLSI implementation. Architectural complexity comparison shows that the new architecture has smaller critical path delay compared to other word level multipliers available in open literature at the cost of having moderately higher area complexity. The new architecture out performs all other similar proposals considering the product of area and delay as a measure of performance. Ashkan Hosseinzadeh Namin, Huapeng Wu, Majid Ahmadi |
ISCAS | 2 |
| 2008 | A Secure Routing Protocol in Proactive Security Approach for Mobile Ad-Hoc NetworksabstractSecure routing of Mobile Ad-hoc Networks (MANETs) is still a hard problem after years of research. We therefore propose to design a secure routing protocol in a new approach. This protocol starts from a prerequisite secure status and fortifies this status by protecting packets using identity-based cryptography and updating cryptographic keys using threshold cryptography periodically or when necessary. Compared to existing schemes, the main contribution of our proposal is the notion of allowing only legitimate nodes to participate in the bootstrapping process, rather than trying to detect adversary nodes after they are participating in the routing protocol. Besides, the proposal has several improvements in routing setup and maintenance: it does not need any side channel or secret channel; it simplifies secret updates without requiring a node to move around; it does not use flooding to set up initial routing, and does not use multicast to update secrets. Shushan Zhao, Akshai K. Aggarwal, Shuping Liu, Huapeng Wu |
WCNC | 4 |
| 2008 | A New Finite-Field Multiplier Using Redundant RepresentationabstractA novel serial-in parallel-out finite field multiplier using redundant representation is proposed. It is shown that the proposed architecture has either a significantly lower complexity and comparable critical path delay or a significantly smaller critical path delay and comparable complexity in comparison to the previously proposed architectures using the same representation. For the class of fields where there exists a type I optimal normal basis, the proposed multiplier compares favorably to the normal basis multipliers. A digit-level version for the new multiplier is also presented in this paper. Ashkan Hosseinzadeh Namin, Huapeng Wu, Majid Ahmadi |
IEEE Trans. Computers | 2 |
| 2008 | Bit-Parallel Polynomial Basis Multiplier for New Classes of Finite FieldsabstractIn this paper, three small classes of finite fields GF$(2^m)$ are found for which low complexity bit-parallel multipliers are proposed. The proposed multipliers have lower complexities compared to those based on the irreducible pentanomials. It is also shown that there does not always exist an irreducible all-one polynomial, equally-spaced polynomial, or trinomial for the new classes of fields. Huapeng Wu |
IEEE Trans. Computers | 1 |
| 2007 | Comb Architectures for Finite Field Multiplication in F(2^m)abstractTwo high-speed bit-serial word-parallel or comb-style finite field multipliers are proposed in this paper. The first proposal utilizes a redundant representation for any binary field and the other uses a reordered normal basis for the binary field where a type-II optimal normal basis exists. The proposed redundant representation architecture has a smaller critical path delay compared to the previous methods while the complexities remain about the same. The proposed reordered normal basis multiplier has a significantly smaller critical path delay compared to the previous methods using the same basis or normal basis. Field-programmable gate array (FPGA) implementation results of the proposed multipliers are compared to those of the previous methods using the same basis, which confirms that the proposed multipliers allow a much higher clock rate. Ashkan Hosseinzadeh Namin, Huapeng Wu, Majid Ahmadi |
IEEE Trans. Computers | 2 |
| 2002 | Montgomery Multiplier and Squarer for a Class of Finite FieldsabstractMontgomery multiplication in GF(2/sup m/) is defined by a(x)b(x)r/sup -1/(x) mod f(x), where the field is generated by a root of the irreducible polynomial f(x), a(x) and b(x) are two field elements in GF(2/sup m/), and r(x) is a fixed field element in GF(2/sup m/). In this paper, first, a slightly generalized Montgomery multiplication algorithm in GF(2/sup m/) is presented. Then, by choosing r(x) according to f (x), we show that efficient architectures of bit-parallel Montgomery multiplier and squarer can be obtained for the fields generated with an irreducible trinomial. Complexities of the Montgomery multiplier and squarer in terms of gate counts and time delay of the circuits are investigated and found to be as good as or better than that of previous proposals for the same class of fields. Huapeng Wu |
IEEE Trans. Computers | 1 |
| 2002 | Bit-Parallel Finite Field Multiplier and Squarer Using Polynomial BasisabstractBit-parallel finite field multiplication using polynomial basis can be realized in two steps: polynomial multiplication and reduction modulo the irreducible polynomial. In this article, we present an upper complexity bound for the modular polynomial reduction. When the field is generated with an irreducible trinomial, closed form expressions for the coefficients of the product are derived in term of the coefficients of the multiplicands. The complexity of the multiplier architectures and their critical path length are evaluated, and they are comparable to the previous proposals for the same class of fields. An analytical form for bit-parallel squaring operation is also presented. The complexities for bit-parallel squarer are also derived when an irreducible trinomial is used. Consequently, it is argued that to solve multiplicative inverse using polynomial basis can be at least as good as using a normal basis. Huapeng Wu |
IEEE Trans. Computers | 1 |
| 2002 | Finite Field Multiplier Using Redundant RepresentationabstractThis article presents simple and highly regular architectures for finite field multipliers using a redundant representation. The basic idea is to embed a finite field into a cyclotomic ring which is based on the elegant multiplicative structure of a cyclic group. One important feature of our architectures is that they provide area-time trade-offs which enable us to implement the multipliers in a partial-parallel/hybrid fashion. This hybrid architecture has great significance in its VLSI implementation in very large fields. The squaring operation using the redundant representation is simply a permutation of the coordinates. It is shown that, when there is an optimal normal basis, the proposed bit-serial and hybrid multiplier architectures have very low space complexity. Constant multiplication is also considered and is shown to have an advantage in using the redundant representation. Huapeng Wu, M. Anwar Hasan, Ian F. Blake, Shuhong Gao |
IEEE Trans. Computers | 1 |
| 2001 | Efficient exponentiation using weakly dual basisabstractA new architecture for finite field exponentiation using weakly dual bases is presented. An extended bidirectional linear feedback shift register is designed to multiply an arbitrary field element with certain essential multiplicands in weakly dual basis (WDB). Each of these multiplications is done in one single clock cycle. It is shown that a bit parallel implementation of the WDB fourth power has complexities comparable to those of polynomial basis fourth power. The proposed structure can effectively speed up the computation of exponentiation and is expected to reduce the power consumption compared to the conventional square and multiply scheme. Compared to the structure for polynomial basis exponentiation, the new structure is thus advantageous in a system where the WDB is already available. Huapeng Wu, M. Anwar Hasan |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2000 | Montgomery Multiplier and Squarer in GF(2m)
Huapeng Wu |
CHES | 1 |
| 2000 | Utilization of differential evolution in inverse kinematics solution of a parallel redundant manipulatorabstractA novel type of redundant parallel manipulator is presented and studied. The inverse kinematics model of the manipulator is postulated. The static stiffness of the manipulator is discussed. To achieve a minimum deflection in the solution of the inverse kinematics problem the differential evolution method is used. In the inverse kinematics solution also the appropriate link motions to avoid collision and joint limits are selected. Huapeng Wu, Heikki Handroos |
KES | 1 |
| 1999 | Low Complexity Bit-Parallel Finite Field Arithmetic Using Polynomial Basis
Huapeng Wu |
CHES | 1 |
| 1999 | Highly Regular Architectures for Finite Field Computation Using Redundant Basis
Huapeng Wu, M. Anwar Hasan, Ian F. Blake |
CHES | 1 |
| 1999 | Closed-Form Expression for the Average Weight of Signed-Digit RepresentationsabstractIn radix-r number system, the minimal weight signed-digit (SD) representation has minimal number of nonzero signed-digits which belong to the set {/spl plusmn/1, /spl plusmn/2, ..., /spl plusmn/(r-1)}. In this article, we derive closed form expressions for the average number of nonzero digits in the minimal weight SD representation and for the average length of the canonical SD representation, a special case of the minimal weight SD form, of a positive integer whose radix-r form is of length n, n/spl ges/1. Huapeng Wu, M. Anwar Hasan |
IEEE Trans. Computers | 1 |
| 1998 | Low Complexity Bit-Parallel Multipliers for a Class of Finite FieldsabstractNew implementations of bit-parallel multipliers for a class of finite fields are proposed. The class of finite fields is constructed with irreducible AOPs (all one polynomials) and ESPs (equally spaced polynomials). The size and time complexities of our proposed multipliers are lower than or equal to those of the previously proposed multipliers of the same class. Huapeng Wu, M. Anwar Hasan |
IEEE Trans. Computers | 1 |
| 1998 | New Low-Complexity Bit-Parallel Finite Field Multipliers Using Weakly Dual BasesabstractNew structures of bit-parallel weakly dual basis (WDB) multipliers over the binary ground field are proposed. An upper bound on the size complexity of bit-parallel multiplier using an arbitrary generating polynomial is given. When the generating polynomial is an irreducible trinomial x/sup m/+x/sup k/+1, 1/spl les/k/spl les/[m/2], the structure of the proposed bit-parallel multiplier requires only m/sup 2/ two-input AND gates and at most m/sup 2/-1 XOR gates. The time delay is no greater than T/sub A/+([log/sub 2/ m]+2)T/sub x/, where T/sub A/ and T/sub X/ are the time delays of an AND gate and an XOR gate, respectively. Huapeng Wu, M. Anwar Hasan, Ian F. Blake |
IEEE Trans. Computers | 1 |
| 1997 | Efficient Exponentiation of a Primitive Root in GF(2^m)abstractIn this paper, exponentiation of a primitive root in GF(2/sup m/) is considered. Signed digit (SD) number representation is used to efficiently represent the exponent and the corresponding algorithms and structures for exponentiation are developed. For primitive multiplications required in exponentiations, extended bidirectional linear feedback shift registers are proposed and used for the cases where the exponent is represented as a binary or a radix-4 SD number. Comparisons are made with other methods on the bases of space, time, and possible power consumption. Since the proposed structures can effectively reduce power and area when implemented in VLSI, they are especially suitable for battery powered portable devices. Huapeng Wu, M. Anwar Hasan |
IEEE Trans. Computers | 1 |