EDBT 2026 Demo / reviewers in the wild / expert
Jianhao Hu
dblp:07/6821
· DBLP profile ↗
44ranked-venue papers
0as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 23 · 7 since 2021Computer networks · 12 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FD-ICA-R: A novel weak signal extraction algorithm for convolutive mixing models
Xuanbo Zhang, Wenhao Sha, Tianzhu Qin, Jianhao Hu, Kaining Han |
Signal Process. | 5 |
| 2026 | A Highly Efficient Massive MIMO Detector Design With MPNF-IM Scheme
Yubin Zhu, Ruobing Yang, Kaining Han, Jianhao Hu |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2026 | Stochastic Symbol Computing Scheme for Signal Processing Application
Yubin Zhu, Kaining Han, Jianhao Hu |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2025 | Stochastic Multivariate Universal-Radix Finite-State Machine: a Theoretically and Practically Elegant Nonlinear Function ApproximatorabstractNonlinearities are crucial for capturing complex input-output relationships especially in deep neural networks. However, nonlinear functions often incur various hardware and compute overheads. Meanwhile, stochastic computing (SC) has emerged as a promising approach to tackle this challenge by trading output precision for hardware simplicity. To this end, this paper proposes a first-of-its-kind stochastic multivariate universal-radix finite-state machine (SMURF) that harnesses SC for hardware-simplistic multivariate nonlinear function generation at high accuracy. We present the finite-state machine (FSM) architecture for SMURF, as well as analytical derivations of sampling gate coefficients for accurately approximating generic nonlinear functions. Experiments demonstrate the superiority of SMURF, requiring only 16.07% area and 14.45% power consumption of Taylor-series approximation, and merely 2.22% area of look-up table (LUT) schemes. Xincheng Feng, Guodong Shen, Jianhao Hu, Meng Li 0016, Ngai Wong 0001 |
ASP-DAC | 3 |
| 2025 | A Fast Converging and Low Complexity Learning-Based MMSE-IRC Detection Algorithm for Massive MIMOabstractMassive MIMO is a pivotal technology for 5G and beyond, significantly enhancing spectrum efficiency and system capacity by deploying numerous receiving antennas at the base station. However, the increasing number of antennas imposes strong requirements for effective interference suppression in uplink MIMO detection. While traditional MMSE-IRC provides excellent performance in mitigating interference, its computational complexity scales cubically with the number of receiving antennas, limiting its practical applicability in massive MIMO systems. This paper presents a learning-based algorithm employing Mini-batch Gradient Descent (MBGD) to solve the weight matrix of MMSE-IRC, complemented by convergence acceleration strategies that incorporate adaptive initial learning rate determination and noise-plus-interference power regularization. Convergence and robustness of the proposed algorithm are validated across diverse modulation schemes, interference levels, and antenna configurations. Performance results indicate that, under high interference conditions, the iteration count of the proposed method is merely 12.5% that of traditional MBGD, while its BER performance exceeds that of conventional MMSE-IRC by approximately 0.5 dB, achieving a reduction in complexity from$O\left(N_{r x}^{3}\right)$to$O\left(N_{t x} N_{r x}\right)$. Ruobing Yang, Yubin Zhu, Kaining Han, Jianhao Hu |
ICC | 5 |
| 2025 | A Novel Efficient Stochastic Gradient Descent based MIMO Detector with Noise-Free InitializationabstractMultiple Input Multiple Output (MIMO) is a critical component of modern wireless communication systems, widely utilized for its enhancement of system capacity. However, one of the major challenges of MIMO is the high complexity of the detection algorithm. In this paper, we propose a novel low-complexity detection algorithm based on multi-layer perception (MLP) and stochastic gradient descent (SGD). This algorithm significantly accelerates convergence using zero-forcing initialization and momentum optimization method. Additionally, a fixed-point optimized hardware architecture is proposed. The proposed design achieves up to 2.39× higher area efficiency with respect to the existing solutions, making it a promising candidate for practical application. Yubin Zhu, Ruobing Yang, Kaining Han, Jianhao Hu |
ISCAS | 5 |
| 2025 | Single Carrier SCMA Detection Schemes for Frequency Selective ChannelsabstractSparse code multiple access (SCMA) is a candidate technology for future wireless communication systems to enhance network access capacity. Existing multi-carrier techniques, such as orthogonal frequency division multiplexing (OFDM) that are based on SCMA schemes, are well-researched and considered promising for future 6G mobile communication networks. However, the high peak-to-average power ratio (PAPR) issue stemmed from the multi-carrier schemes limits the applications of SCMA in long-distance transmission scenarios and low-cost sensor networks. To address this issue, single-carrier based SCMA transmission scheme (SC-SCMA) with low PAPR has been proposed. However multi-user signal detection remains a significant challenge for SC-SCMA, particularly with frequency-selective channels (FSC). In this paper, we establish the self-interference analysis model of the SC-SCMA system in FSC and construct an extended factor graph (EFG) for multi-user detection. Based on that, we propose the Gaussian approximate interference based belief propagation algorithm processing on EFG (E-GAIBP) to realize efficient detection over FSC with reduced complexity. Meanwhile, an extrinsic information transfer (EXIT) chart is introduced to analyze the detection performance and convergence properties of the E-GAIBP detector. Finally, to further improve performance gains, we propose two joint E-GAIBP detection and LDPC decoding schemes, E-GAIBP-JDD and layered E-GAIBP-JDD. Numerical results and complexity evaluations show that the layered E-GAIBP-JDD excels in detection performance and convergence speed, offering an effective receiver design strategy. Tianzhu Qin, Jia'ao Liang, Kaining Han, Jianhao Hu |
IEEE Trans. Wirel. Commun. | 5 |
| 2024 | A Novel Sample Selection based Detection Algorithm for Asynchronous SCMA SystemsabstractAs a promising contender within non-orthogonal multiple access (NOMA), sparse code multiple access (SCMA) stands out as a potential cornerstone in future-generation wireless communication systems. However, while most SCMA research assumes ideal synchronization among multiple users, the reality of non-ideal synchronization introduces multi-user access inter-ference (MAI), leading to significant performance degradation. In fact, this interference can be severe enough to cause traditional SCMA detection methods to falter in convergence. This paper delves into the theoretical analysis of MAI and introduces an innovative solution: a sample selection-based message passing algorithm (SS-MPA). The aim is to tackle this issue head-on. Through error rate performance simulations and complex-ity evaluations, our results showcase that the proposed SS-MPA algorithm, designed for asynchronous conditions, closely approaches the optimal performance of MPA in synchronous settings, all while maintaining comparable complexity levels. This breakthrough marks a significant step toward resolving MAI-related challenges in SCMA systems. Kaining Han, Yonghang Dai, Xuanbo Zhang, Jianhao Hu |
VTC Spring | 4 |
| 2024 | Receiver Designs for SC-SCMA Systems Over Frequency Selective ChannelsabstractSparse code multiple access (SCMA) is a candidate technology for further wireless communication system to enhance the multiple access capacity. The existing multi carriers technique based SCMA schemes, such as orthogonal frequency division multiplexing (OFDM), have been well studied and considered as a promising candidate for future 6G protocols. However, the high peak-to-average power ratio (PAPR) issue resulted by the multi carriers techniques limits the applications of SCMA in long distance transmission scenarios and low cost sensor networks. To address this issue, single carrier based SCMA transmission scheme (SC-SCMA) has been proposed and demonstrates low PAPR. But the multi user signal detection is still a major chal-lenge for SC-SCMA, especially for frequency selective channels (FSC). In this paper, we establish the interference model of SC-SCMA system in FSC, and construct an extended factor graph (EFG) for multi user detection. Afterward, the Gaussian approx-imate interference based belief propagation algorithm processing on EFG (E-GAIBP) is proposed, realizing efficient detection and complexity reduction. Furthermore, to earn higher performance gain, joint E-GAIBP detection and LDPC decoding scheme (E-GAIBP-JDD) is introduced. Numerical results demonstrate that compared with classical MPA using EFG, E-GAIBP shows great complexity reduction and E-GAIBP-JDD performs great performance gain. Tianzhu Qin, Jia'ao Liang, Kaining Han, Jianhao Hu |
VTC Spring | 5 |
| 2023 | Convergence Optimized Joint Detection and Decoding for LDPC Coded SC-MIMO SystemsabstractWith the increasing requirement for high-speed transmission, multiple input multiple out (MIMO) technologies have been proposed. Beneficial from low peak-to-average power ratio (PAPR), single carrier (SC) MIMO systems are suitable for long-distance transmission. However, one of the significant challenges of SC-MIMO systems is the vast complexity of channel equalization and signal detection, especially for the inter-symbol interference (ISI) caused by the frequency-selective fading channels. Due to acceptable complexity and near maximum likelihood (ML) detection performance, the Gaussian approximate interference (GAI) based belief propagation (BP) algorithm has attracted lots of attention. With the help of extrinsic information feedback from low-density parity check (LDPC) codes, iterative detection and decoding (IDD) scheme has been proposed, which has a near to approximate maximum a posterior (MAP) detection performance. Nevertheless, the IDD scheme still suffers from relatively slow convergence resulting from complex connections on the factor graph. Combined with the joint detection and decoding (JDD) scheme and two optimization methods, namely dynamic message damping and layered message updating, a convergence optimized joint detection and decoding (CO-JDD) algorithm is proposed in this paper. Evaluation results show that the proposed CO-JDD algorithm has 2dB bit error rate (BER) performance gains compared with IDD scheme. Besides, the proposed CO-JDD algorithm also has a 70% iteration reduction compared with the IDD scheme. Xuanbo Zhang, Kaining Han, Jianhao Hu |
GLOBECOM | 3 |
| 2023 | A Novel Stochastic Decoder for Extended BCH Code Based Turbo Product CodesabstractTurbo product codes (TPC) have received recent attention as a competitive forward error correction scheme in high-speed communication protocols and Extended BCH based TPC has been accepted by several optical transmission standards. Decoding complexity is one of the major challenges for TPC. This paper proposes a novel stochastic TPC decoder that processes the decoding in a bit-wise iterative manner to reduce the decoding complexity. Further, a hard-decision decoding temporary frozen strategy is proposed to further speed up the convergence and reduce the decoding complexity. The evaluation results demonstrate significant complexity reduction with error correction performance improvements compared to the existing works, making it a promising candidate for practical application. Kaining Han, Guodong Shen, Jianhao Hu |
ICC | 3 |
| 2023 | High Throughput and Hardware Efficient Hybrid LDPC Decoder Using Bit-Serial Stochastic UpdatingabstractHybrid low-density parity-check (LDPC) decoding combines conventional Belief-Propagation (BP) algorithm with stochastic decoding to achieve high performance and low complexity simultaneously. However, lossy and inefficient stochastic-to-binary (S2B) conversion brings extra performance degradation and decoding latency. In this paper, a bit-serial stochastic updating based hybrid decoding (BSSU-HD) is proposed, which employs fully correlated stochastic (FCS) check nodes (CNs) and probability tracers assisted variable nodes (VNs) to accomplish accurate and efficient S2B conversion. Two strategies, including random source selection and tracing speed switching, are proposed to further improve performance and convergence. A BSSU LDPC decoder for IEEE 802.3an is designed in a 65-nm CMOS process, which occupies 4.6 mm2 silicon area and achieves a throughput of 200.8 Gb/s at$E_{b}/N_{0} = 4.4$dB with 500 MHz clock frequency from a 1.2 V supply voltage. The power and energy efficiency are 2.933 W and 14.61 pJ/bit, respectively. To the best of our known, it achieves the best decoding performance, the highest throughput and hardware efficiency among state-of-the-art IEEE 802.3an LDPC decoders. We also verify that the BSSU-HD can achieve better performance for multi-rate 5th generation (5G) New Ratio (NR) LDPC codes than conventional algorithm, which greatly extends the application of the stochastic decoding. Shuai Hu, Kaining Han, Yubin Zhu, Guodong Shen, Fujie Wang, Jianhao Hu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2022 | Low-complexity linear massive MIMO detection based on the improved BFGS methodabstractAbstract Linear minimum mean square error (MMSE) detection achieves a good trade‐off between performance and complexity for massive multiple‐input multiple‐output (MIMO) systems. To avoid the high‐dimensional matrix inversion involved, MMSE detection can be transformed into an unconstrained optimization problem and then solved by efficient numerical algorithms in an iterative way. Three low‐complexity Broyden‐Fletcher‐Goldfarb‐Shanno (BFGS) quasi‐Newton methods are proposed to iteratively realize massive MIMO MMSE detection without matrix inversion. The complexity can be reduced from to , where K and L denote the number of users and iterations, respectively. Leveraging the special properties of massive MIMO, the authors first explore a simplified BFGS method (named S‐BFGS) to alleviate the computational burden in the search direction. For lower complexity, BFGS method with the unit step size (named U‐BFGS) is presented subsequently. When the base station (BS)‐to‐user‐antenna ratio (BUAR) is large enough, the two proposed BFGS methods can be integrated (named U‐S‐BFGS) to further reduce complexity. In addition, an efficient initialization strategy is devised to accelerate convergence. Simulation results verify that the proposed detection scheme can achieve near‐MMSE performance with a small number of iterations L as low as 2 or 3. Lin Li 0080, Jianhao Hu |
IET Commun. | 2 |
| 2022 | An efficient second-order neural network model for computing the Moore-Penrose inverse of matricesabstractAbstract The computation of the Moore–Penrose inverse is widely encountered in science and engineering. Due to the parallel‐processing nature and strong‐learning ability, the neural network has become a promising approach to solving the Moore–Penrose inverse recently. However, almost all the existing neural networks for matrix inversion are based on the gradient‐descent (GD) method, whose main drawbacks are slow convergence and sensitivity to learning parameters. Moreover, there is no unified neural network to compute the Moore–Penrose inverse for both the full‐rank matrix and rank‐deficient matrix. In this paper, an efficient second‐order neural network model with the improved Newton's method is proposed to obtain the accurate Moore–Penrose inverse of an arbitrary matrix by one epoch without any learning parameter. Compared with the GD‐based neural networks for Moore–Penrose inverse computation, the proposed model converges faster and has lower complexity. Furthermore, through in‐depth derivation, the neural network for computing the Moore–Penrose inverse is well interpretable. Numerical studies and application to the random matrix inversion in multiple‐input multiple‐output detection are provided to validate the efficiency of the proposed model for solving the Moore–Penrose inverse. Lin Li 0080, Jianhao Hu |
IET Signal Process. | 2 |
| 2022 | Hybrid Stochastic LDPC Decoder With Fully Correlated Stochastic ComputationabstractThe ultra-low hardware consumption feature of stochastic decoding has made it a potential candidate for the implementation of low-density parity-check(LDPC) decoders. However, the existing stochastic LDPC decoders still suffer from performance degradation and relatively high decoding cycles caused by the correlation among stochastic bit streams. In this paper, we propose Hybrid Stochastic(HS) decoding, which achieves high performance, high throughput, and high hardware efficiency by jointly using our proposed novel stochastic check node(CN) and Two’s Complement(TCS) variable node(VN) to realize Min-Sum Algorithm(MSA) and its enhancements. Fully correlated stochastic bit streams are used to entirely eliminate the indeterminacy caused by the correlation, which results in high performance and fast convergence and inherits the low complexity of stochastic decoders at the same time. We demonstrate the HS decoding by designing a (2048,1723) decoder in a 65 nm process, which achieves the highest Bit-Error-Ratio(BER) performance, highest throughput, and top hardware efficiency among existing stochastic LDPC decoders. We also demonstrate that HS decoding can achieve excellent decoding performance for different code rates and lengths 5G New Radio(NR) LDPC codes. Thus, HS decoding can be adopted in wide applications. Shuai Hu, Kaining Han, Fujie Wang, Jianhao Hu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2021 | Neural Synaptic Plasticity-Inspired Computing: A High Computing Efficient Deep Convolutional Neural Network AcceleratorabstractDeep convolutional neural networks (DCNNs) have achieved state-of-the-art performance in classification, natural language processing (NLP), and regression tasks. However, there is still a great gap between DCNNs and the human brain in terms of computation efficiency. Inspired by neural synaptic plasticity and stochastic computing (SC), we propose neural synaptic plasticity-inspired computing (NSPC) to simulate the human brain's neural network activity for inference tasks with simple logic gates. The multiplication and accumulation (MAC) is transformed by the wire connectivity in NSPC, which only requires bundles of wires and small width adders. To this end, the NSPC imitates the structure of neural synaptic plasticity from a circuit wires connection perspective. Furthermore, from the principle of NSPC, we use a data mapping method to convert the convolution operations to matrix multiplications. Based on the methodology of NSPC, fully-pipelined and low latency architecture is designed. The proposed NSPC accelerator exhibits high hardware efficiency while maintaining a comparable network accuracy level. The NSPC based DCNN accelerator (NSPC-CNN) processes DCNN at 1.5625M images/s with a power dissipation of 15.42 W and an area of 36.4 mm2. The NSPC based deep neural network (DNN) accelerator (NSPC-DNN) that implements three fully connected layers DNN consumes only 6.6 mm2 area and 2.93 W power, and achieves a throughput of 400M images/s. Compared with conventional fixed-point implementations, the NSPC-CNN achieves 2.77× area efficiency, 2.25× power efficiency; the proposed NSPC-DNN exhibits 2.31× area efficiency and 2.09× power efficiency. Zihan Xia 0002, Jienan Chen, Qiu Huang, Jinting Luo, Jianhao Hu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2021 | An Improved Expectation Propagation Based Detection Scheme for MIMO SystemsabstractMultiple-input Multiple-output (MIMO) technologies play an important role in modern and future wireless communication systems as they can improve the capacity without increasing the bandwidth. MIMO detection is one of the key technologies for MIMO system designs. MIMO detection schemes based on Message Passing (MP) algorithms have attracted extensive attention in recent years. The MIMO detection scheme based on Expectation Propagation (EP), which is also a kind of MP algorithms, has been proved to achieve the Bayes-optimal performance under the conditions of large system limit and compression rate threshold. However, there is improvement space for the detection performance of the EP algorithm when the conditions are not satisfied, which are the common cases in the practical applications. In this paper, when the conditions of the large system limit and compression rate threshold are not satisfied, we analyze the influences of two factors, the initial parameters selection and the moment matching strategy, on the performance of EP based detection scheme. As a result, we propose an improved EP detection scheme based on optimizing the two factors. Simulation results show that the proposed scheme outperforms the original EP detection scheme in different MIMO scenarios. Guoqiang Yao, Jianhao Hu |
IEEE Trans. Commun. | 3 |
| 2020 | Design and implementation of SVM OTPC searching based on Shared Dot Product MatrixabstractIn this paper, we proposed a FPGA implementation architecture for SVM classifier. The architecture is based on the proposed Shared Dot Product Matrix (SDPM) method which computes and stores the dot product of all training data before SVM searching process. We implemented the proposed method by software simulation and hardware implementation. The software simulation of SDPM method achieves twice the speed of LIBSVM, which is one of the most popular SVM implementation libraries. This acceleration mainly results from the reduction of repeat Kernel function calculation. Then the hardware software collaboration architecture for SDPM is also proposed in this paper. Results show that the proposed architecture achieves approximately 30 times faster searching speed compared with LIBSVM. Shang Ma, Shengqiang Jiang, Jianhao Hu, Xiongzhong Xiong |
Integr. | 4 |
| 2019 | Low Complexity Expectation Propagation Detection for SCMA Using Approximate ComputingabstractSparse code multiple access (SCMA) is one of the promising non-orthogonal multiple access (NOMA) techniques for the future wireless communication systems, which can provide more connections and higher spectral efficiency than orthogonal multiple access (OMA). In this paper, we propose a novel expectation propagation algorithm based on approximate computing to achieve low complexity and high performance SCMA detection, which is referred as approximate expectation propagation algorithm (AEPA). Three approximate approaches are provided for variable node update, function node update and log likelihood ratio calculation to reduce the algorithm complexity. The parameter optimization methods are also provided to get the trade-off between the detection performance and the algorithm complexity. Simulation and evaluation results show that AEPA can obtain more than 32% complexity reduction with only 0.1dB performance loss when the codebook size is 4 compared with the traditional expectation propagation algorithm (EPA), and the complexity reduction gain will become more significant when the codebook size is larger. Jianhao Hu, Kaining Han |
GLOBECOM | 2 |
| 2019 | iRAF: A Deep Reinforcement Learning Approach for Collaborative Mobile Edge Computing IoT NetworksabstractRecently, as the development of artificial intelligence (AI), data-driven AI methods have shown amazing performance in solving complex problems to support the Internet of Things (IoT) world with massive resource-consuming and delay-sensitive services. In this paper, we propose an intelligent resource allocation framework (iRAF) to solve the complex resource allocation problem for the collaborative mobile edge computing (CoMEC) network. The core of iRAF is a multitask deep reinforcement learning algorithm for making resource allocation decisions based on network states and task characteristics, such as the computing capability of edge servers and devices, communication channel quality, resource utilization, and latency requirement of the services, etc. The proposed iRAF can automatically learn the network environment and generate resource allocation decision to maximize the performance over latency and power consumption with self-play training. iRAF becomes its own teacher: a deep neural network (DNN) is trained to predict iRAF's resource allocation action in a self-supervised learning manner, where the training data is generated from the searching process of Monte Carlo tree search (MCTS) algorithm. A major advantage of MCTS is that it will simulate trajectories into the future, starting from a root state, to obtain a best action by evaluating the reward value. Numerical results show that our proposed iRAF achieves 59.27% and 51.71% improvement on service latency performance compared with the greedy-search and the deep Q-learning-based methods, respectively. Jienan Chen, Siyu Chen 0018, Qi Wang 0049, Bin Cao 0002, Gang Feng 0004, Jianhao Hu |
IEEE Internet Things J. | 6 |
| 2018 | A Low Complexity Expectation Propagation Detection for Massive MIMO SystemabstractThis paper proposes a low-complexity expectation propagation (EP) algorithm for massive multiple-input multiple-output (MIMO) detections. The original EP detection algorithm shows a near-optimal performance but suffers from the unaffordable computational complexity. In this paper, we use an iterative successive updating scheme to reduce the complexity caused by the exact matrix inversion in each iteration and ameliorate the efficiency and accuracy of messages updating to accelerate the convergence, which leads to a low-complexity high-performance massive MIMO detector. Numerical analysis shows the proposed algorithm can outperform 0.2 dB in bit error rate (BER) with a huge of computational complexity saved compared with the previous EP detector. Guoqiang Yao, Guiwu Yang, Jianhao Hu, Chao Fei 0002 |
GLOBECOM | 3 |
| 2018 | A pseudo-random sequence generation scheme based on RNS and permutation polynomials
Shang Ma, Zeguo Yang, Jianhao Hu |
Sci. China Inf. Sci. | 5 |
| 2018 | Feedback-Based Low-Power Soft-Error-Tolerant Design for Dual-Modular Redundancy
Yufeng Li 0003, Jie Han 0001, Jianhao Hu, Fan Yang 0001, Xuan Zeng 0001, Bruce F. Cockburn, Jie Chen 0002 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2016 | Implementation of efficient parallel discrete cosine transform using stochastic logicabstractThis paper provides a new scheme for the VLSI implementation of a parallel Discrete Cosine Transform (DCT) using stochastic logic. Stochastic computation is a number representation, which can carry out complex computations with very low hardware cost. However, the delay of data output is proportional to the length of serial sequence. We provide a new area-saving parallel DCT design to improve the system throughput by using our proposed stochastic OR-adder and OR-AND-adder. Results show the proposed parallel stochastic DCT can meet the requirement of image processing while maintaining a ± 5% performance difference compared to the traditional DCT implementation. Our synthesized chip design using the TSMC CMOS 130nm technology also shows that the proposed parallel stochastic DCT is at least 10 times more efficient in area and delay than that of the traditional DCT and the serial stochastic DCT. Jianhao Hu, Jie Chen 0002 |
ISCAS | 2 |
| 2016 | Area-efficient partial-clique-energy MRF pair design with ultra-low supply voltageabstractAs the size of CMOS devices continues to scale down, the reliability of circuits becomes one of main challenges in low supply voltage designs. Markov Random Field (MRF) circuits, a probabilistic-based approach, can achieve higher noise immunity compared to traditional designs under conditions of ultra-low supply voltage and low threshold voltage. However, the basic MRF elements have complex structures and become a stringent factor that limits MRF-based VLSI design. In this paper, we provide a partial-clique-energy MRF (PMRF) design method, trading off the noise immunity for area efficiency. We then propose an Enhanced PMRF (EPMRF)-pair for multi-level and multi-function joint PMRF designs. The main idea is to use the joint clique energy of two complementary partial clique energies to make up performance losses. The measurement results show that, the proposed EPMRF pair can operate at 0.25 V with 10-4 dB output noise power with 5.6 dB input signal-noise ratio (SNR). With the 130 nm CMOS technology, the chip of our EPMRF based carry-look-ahead adder achieves 29% area-saving and 55% energy-saving compared to existing ultra-low supply voltage fault tolerant designs. Jianhao Hu, Jie Chen 0002 |
ISCAS | 2 |
| 2016 | Sparse Code Multiple Access Decoding Based on a Monte Carlo Markov Chain MethodabstractNonorthogonal multiple access technology has been proposed for use in 5G communications systems. In particular, the sparse code multiple access (SCMA) scheme is believed to be one of the most promising techniques among the various nonorthogonal approaches that have been investigated. In this letter, we focus on reducing the complexity of SCMA decoding and we propose a Monte Carlo Markov Chain (MCMC) based SCMA decoder. Benefiting from the linearly increasing complexity of the MCMC method, the proposed SCMA decoder has only 10% of the computational load compared to previous state-of-the-art methods when the codebook size is 64. Consequently, the MCMC SCMA decoder has great potential for use in practical system implementations. Jienan Chen, Zhenbing Zhang, Shuaining He, Jianhao Hu, Gerald E. Sobelman |
IEEE Signal Process. Lett. | 4 |
| 2016 | Hardware and Energy-Efficient Stochastic LU Decomposition Scheme for MIMO ReceiversabstractIn this paper, we design a hardware and energy-efficient stochastic lower-upper decomposition (LUD) scheme for multiple-input multiple-output receivers. By employing stochastic computation, the complex arithmetic operations in LUD can be performed with simple logic gates. With proposed dual partition computation method, the stochastic multiplier and divider exhibit high computation accuracy with relative short length stochastic stream. We have designed and synthesized the stochastic LUD with CMOS 130-nm technology. According to the postlayout report, the hardware efficiency of the stochastic LUD is as high as 1.5× compared with the exiting LUD methods, and the energy efficiency is also higher than the state-of-the-art LUD when the matrix dimension is 8 × 8 and larger. Jienan Chen, Jianhao Hu, Jiangyun Zhou |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2015 | Area-sharing cyclic structure MRF cirucits design in ultra-low supply voltageabstractFeedback structure is an efficient topology for noise-reducing in analog circuit while the cyclic circuit has been widely used in digital circuit only to establish sequential circuits due to its data-keeping property. The Markov Random Field (MRF) based circuit is proposed with feedback structure from the point of energy, which has high noise-immune property in ultra-low supply voltage. However the MRF-based elements have complex structures which have almost 20 times area compared to traditional units. In this paper we propose area-sharing feedback NAND path based on the MRF circuit which not only can reduce about 60% of area and complex for MRF element but also has 3dB performance of noise tolerant improvement compared to traditional CMOS NAND. Then we provide area-sharing feedback NOR-NAND path for multifunction calculating with the case study of half adder and 8-bits ripple carry adder, which has about at least 4 dB for output bits performance improvement compared to CMOS half adder and 84.2% energy saving at the 0.25V supply voltage compared to original MRF adder. Jianhao Hu |
ISCAS | 3 |
| 2015 | Hardware Efficient Mixed Radix-25/16/9 FFT for LTE SystemsabstractIn this paper, we propose a hardware-efficient mixed generalized high-radix (GHR) reconfigurable fast Fourier transform (FFT) processor for long-term evolution applications. The GHR processor based on radix-25/16/9 uses a 2-D factorization scheme as the high-radix unit and a 1-D factorization method as the system data routing technology. The 2-D factorization scheme is implemented by an enhanced delay element matrix structure, which supports 25-, 16-, 9-, 8-, 5-, 4-, 3-, and 2-point FFTs. Two different designs were implemented. One design (called discrete Fourier transform core) supports 34 different transform sizes from 12 to 1296 points, while the other design (called FFT core) supports five different power-of-two sizes from 128 to 2048 points. The 1-D factorization method is performed by a coprime accessing technology, which accesses the data in parallel without conflict using a RAM. The GHR combines 2-D and 1-D factorization techniques and improves the throughput by a factor of two to four with comparable hardware cost compared with the previous designs. The speed-area ratio of the proposed scheme is nearly two times better than that of previous FFT processors. Application-specified integrated circuit implementation results based on a 0.18-μm technology are also provided. Jienan Chen, Jianhao Hu, Shuyang Lee, Gerald E. Sobelman |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2014 | High performance MIMO detector based on bidirectional path preserving trellis searchabstractIn this paper we propose a high performance bidirectional path preserving trellis search (PPTS) detector for multiple-input-multiple-output (MIMO) systems. The error analysis between single direction and bidirectional PPTS detector is given. We prove that the bidirectional PPTS detector can minimize the detection error effectively. Moreover, the proposed detector saves 10% hardware cost with a 0.1 dB Frame Error Rate (FER) gain compared with traditional PPTS detectors. Jienan Chen, Lian Huai, Jianhao Hu, Gerald E. Sobelman |
ISCAS | 3 |
| 2014 | Random error analysis and reduction for stochastic computation based on autocorrelation sequenceabstractThis paper proposes the random error analysis method for stochastic computation based on autocorrelation sequence (AS), which is more general than the previous work based on Bernoulli sequence (BS). The analysis results show the use of proper ASs as input streams is able to reduce random error compared to the conventional use of BSs. In order to confirm that conclusion, we apply an AS, referred as Maximal Concentrated Autocorrelation Sequence (MCAS), into the stochastic computation system which implements Bernstein polynomial. Both the theoretical analysis and simulation results reveal that the use of MCAS reduces the random error. Ye Cheng, Jianhao Hu |
ISCAS | 2 |
| 2014 | Extensional design for noise-tolerate MRF standard cells via global mappingabstractAs CMOS devices scaling down according to the Moore' law, the reliability and stability of circuit become the main challenge for the chip designs in low supply voltage. Markov Random field (MRF) based design methodology presents a new approach to establish high noise-immune structure for low-power circuit design from the viewpoint of energy. We use global mapping to synthesize all two-input functions and realize the MRF based circuits, then the final structures our provided with same complexity, which does not depend on the inside logic operations any more, but the number of the input-signal instead, which inspires us to extend the MRF standard cells. Our proposed structures keep the performance of fault-tolerance and achieve higher efficiency for energy, time and area compared with that of traditional MRF-circuits. The extensional basic units also can beneficial for saving area in multi-level MRF-circuit. Jianhao Hu |
ISCAS | 2 |
| 2014 | High performance absolute value calculator based on stochastic computingabstractIn this paper, we propose a Sliding Window Method (SWM) to calculate absolute value based on the stochastic computing. We prove that the absolute value of the stochastic stream can be obtained from the limiting distribution of the Markov chain by establishing a Markov model of SWM. The proposed schemes can be employed to both unipolar and bipolar stochastic streams. The simulation and design reports show that the hardware efficiency of the proposed stochastic absolute value calculator is as four times as existing methods. Jiangyun Zhou, Jianhao Hu, Jienan Chen |
ISCAS | 2 |
| 2013 | A novel low-power filter design via reduced-precision redundancy for voltage overscaling applicationsabstractIn this paper, we apply adaptive-signal processing in the reduced precision redundancy filter design for the voltage scaling (VOS) application to achieve high energy efficiency and high SNR performance. RPR technique can mitigate the soft error caused by VOS in the critical path, but the SNR performance of RPR is limited. Thus, we combined RPR and adaptive signal processing to achieve high SNR and low power consumption performance simultaneously for VOS applications. The adaptive signal processing is used to recover the middle significant bits in the result of the filter. From case studies, we find that the proposed method can obtain up to 64% energy reduction with much less SNR performance degradation losing than the traditional RPR and its optimization schemes. Haoliang Li, Jianhao Hu, Jienan Chen |
GLOBECOM | 2 |
| 2013 | A novel FIR filter based on stochastic logicabstractIn this paper, we proposed a novel Finite Impulse Response Filter based on stochastic logic referred as SFIR. The proposed SFIR only requires the wire selecting scheme without any logic gate resource. We first map the FIR function to stochastic computation domain. The FIR function is implemented by the wire selecting (WS) method which only requires wire interleaving. Hence the SFIR can achieve ultra high throughput and low cost when apply in the stochastic based system. The SFIR with backward conversion module also has lower hardware cost than traditional method under a given quantization width, which can apply in the low SNR required system and the error tolerant system. Jienan Chen, Jianhao Hu |
ISCAS | 2 |
| 2013 | A novel implementation scheme for high area-efficient DCT based on signed stochastic computationabstractIn this paper, we propose a novel implementation scheme of Discrete Cosine Transform (DCT) with ultra high area-efficiency based on signed stochastic computation. Due to the properties of stochastic computation, a new number system, complicated DCT circuits can be realized with very simple logic. According to our analysis, addition operation is the performance and hardware cost bottleneck for stochastic domain. Therefore, our optimization strategies aim at reducing the number of addition operations and designing high precision adder module for signed stochastic computation. A novel high accuracy signed stochastic adder is also provided in this paper. Compared with traditional architecture in Two's Complement System (TCS) domain, the proposed scheme can achieve ultra low hardware cost with less than 10% performance loss. The implementation of stochastic DCT not only meets the performance requirement but also provides low hardware cost and critical path latency. Jianhao Hu |
ISCAS | 2 |
| 2013 | High Throughput Stochastic Log-MAP Turbo-Decoder Based on Low Bits ComputationabstractIn this letter, we propose a high throughput stochastic Low Bits Computation (LBC) turbo decoder. We represent the signal by a 3-bits width stochastic stream, which improves the accuracy of stochastic computation significantly. We have designed and synthesized our design based on CMOS 90 nm technology. The report shows that the proposed decoder can achieve 4.0 Gbps with 7.1 M gate count to decode a 2048-length R=1/3 turbo code, when the bit error rate (BER) is 10-5@ Eb/N0=1.25 dB. Jienan Chen, Jianhao Hu |
IEEE Signal Process. Lett. | 2 |
| 2013 | Energy-Efficient Digital Signal Processing via Voltage-Overscaling-Based Residue Number SystemabstractIn this paper, we apply the voltage overscaling (VOS) technique to the residue-number-system (RNS)-based digital signal processing system for achieving high energy efficiency. To mitigate the soft errors caused by VOS, we propose a new method, called joint RNS-RPR (JRR), which is the combination of RNS and the reduced precision redundancy (RPR) technique. The JRR technology inherits the properties of RNS, including shorter critical path, low complexity, and low power. Moreover, JRR can achieve higher power reduction than RNS for VOS applications. Since the soft errors caused by VOS lead to significant performance degradation of RNS, we use the information from RNS and RPR to achieve a high recovering probability of the soft errors with low hardware complexity. From the case study of finite impulse response (FIR) filter design based on the 0.25- μm 2.5-V CMOS technology, we find that JRR can save 62% more energy compared to the traditional FIR with a less than 2-dB signal noise ratio performance loss. We also find that JRR has lower complexity and better performance than the traditional soft error mitigation methods. Jienan Chen, Jianhao Hu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2012 | Low power digital signal processing scheme via stochastic logic protectionabstractIn this paper, we proposed a low power digital signal processing (DSP) scheme with stochastic logic protection. The reduction of supply voltage will reduce the power consumption effectively. However, the timing violation will be happened when the voltage overscaling (VOS) is applied. Fortunately, the stochastic logic has simple hardware structure, and the critical path is short. Thus, we can use stochastic logic as the error control (EC) module to mitigate the soft error by the VOS. Compared with traditional EC based low power DSP, the proposed method can achieve higher energy efficiency with high performance. According to the case study of a 26-tap FIR filter, the stochastic logic based EC module can achieve 65% power saving within 2.5dB signal-to-noise ratio (SNR) loss, which outperforms 6dB than traditional RPR method. When the soft error for each logic gate is considered, the advantage of SFIR based system is more obvious than traditional method. Jienan Chen, Jianhao Hu |
ISCAS | 2 |
| 2012 | High throughput and hardware efficient FFT architecture for LTE applicationabstractIn this paper, we propose a high throughput and hardware efficient Fast Fourier Transform (FFT) architecture for Long Term Evolution (LTE) application. The proposed enhancement delay element matrix (EDEM) which contains the mixed radix unit supports 25, 16, 9, 8, 5, 4, 3 and 2-point FFTs. The reuse technology is also applied into EDEM to reduce the hardware resource. The EDEM reduces the computation cycles significantly since the high radix decomposition method is applied. Compared with the stated of art technology, the proposed scheme improves 2×~4× throughput rate with comparable hardware cost. For all of the 35 FFT lengths in LTE applications, the computation cycles of proposed method are less than the length of FFT, which can supports the continuous flow of processing data in the same clock domain with I/O data. The speed-area product factor outperforms 2~3 times than the existed FFT processor. The proposed architecture also supports the variable length FFT. Jienan Chen, Jianhao Hu |
WCNC | 2 |
| 2011 | Sliding Window Method for stochastic LDPC decoderabstractThis paper proposes a Sliding Window Method (SWM) for stochastic Low Density Parity Check (LDPC) decoder designing. The SWM is formulated for solving the latch-up problem in the Variable Nodes (VN) information updating. The bit in the latch-up state is evolved from the information bit in the sliding window. Then, an optimized hardware structure is proposed for SWM. Compared with traditional VN structure, the SWM require about 35% less hardware resources to achieve the same BER performance. Jienan Chen, Jianhao Hu |
ISCAS | 2 |
| 2010 | A 2n scaling scheme for signed RNS integers and its VLSI implementation
Shang Ma, Jianhao Hu, Yanlong Ye, Xiang Ling 0002 |
Sci. China Inf. Sci. | 2 |
| 2008 | An efficient RNS parity checker for moduli set {2 n - 1, 2 n + 1, 22 n + 1} and its applications
Shang Ma, Jianhao Hu, Xiang Ling 0002 |
Sci. China Ser. F Inf. Sci. | 2 |
| 2007 | A Novel Multiple-Access Scheme for Chirp UWBabstractChirp UWB has the feature of low power consumption, high processing gain and low cost. However, the multiple-access schemes for chirp UWB are limited. They face many difficulties when considering the implementation, bandwidth efficiency and multipath channel. The novel scheme proposed in this paper offers a new multiple-access solution for chirp UWB. It takes advantages of some unique characters of chirp UWB signals and spread spectrum technology. It works well in the multipath channel. Small multiple-access interference (MAI) is achieved as well as bandwidth efficiency. The authors present the details of our proposed scheme and the simulation results in a 4-user situation. Jianhao Hu |
WCNC | 3 |