EDBT 2026 Demo / reviewers in the wild / expert
Yun Chen 0001
dblp:10/5680-1
· DBLP profile ↗
23ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0002-3736-9456ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 18 · 6 since 2021Computer networks · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Configurable Energy-Efficient Dual-Mode Accelerator for DNN and SNN
Xinyang Wu 0007, Jianghao Wu 0006, Yun Chen 0001 |
ISCAS | 4 |
| 2026 | Adaptive Window-Based CIR Truncation and PSO-Optimized Hybrid Deep Network for UWB LOS/NLOS Classification
Zhouxiang Ye, Jianghao Wu 0006, Xinyang Wu 0007, Yun Chen 0001 |
IEEE Signal Process. Lett. | 6 |
| 2026 | A High Efficiency Dual Mode Buck Converter With a Novel Seamless and Fast Transition SchemeabstractThis paper presents a high-efficiency dual-mode off-time controlled buck converter with a novel seamless and fast mode transition scheme. The proposed seamless fast mode transition (SFMT) technique employs a mode selector consisting of a voltage detector (VD) and ripple-to-digital conversion circuit (RDC), solving the prolonged frequency transition problem in conventional mode-switching methods. Additionally, the introduced dynamic sleep clock management adaptively disables idle circuits based on duty cycle variations, achieving high efficiency across wide load ranges. Implemented in 0.18-$\mu $m technology, the converter occupies$0.5525~mm^{2}$while operating from a 2.5-3.6 V input to deliver 1.2 V output. Measurement results demonstrate$93.7~\%$peak efficiency at 0.4 A load and$80.5~\%$high light-load efficiency at 1 mA. The design achieves$\mu $s slew rate) with 28-$\mu $s recovery time, while maintaining low output ripple. Compared with reported designs, the proposed buck converter achieves an outstanding figure-of-merit ($FoM_{1}$) of$7.26~\mu s/(A^{2}/mm^{2})$. Yun Chen 0001, Jinghu Li 0001, Zhicong Luo |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2025 | An Area and Energy-Efficient Systolic Array Accelerator Architecture for Deep Neural Networks Using Stochastic ComputingabstractDeep neural networks (DNNs) are widely used to handle various intelligent tasks. With the increased model size, the DNNs’ hardware accelerators are challenging the higher area overhead and energy consumption. Stochastic computing (SC) has recently been considered for implementing DNNs and reducing hardware consumption. However, many current SC-based DNN accelerators fail to balance accuracy, performance, and resource overhead. In addition, their limited scalability and flexibility restrict their use in edge devices. In this article, we design an area and energy-efficient DNN accelerator architecture using SC. We propose an SC-binary hybrid processing unit with piecewise shift compensation without significant additional hardware overhead increment to improve the SC accuracy. To balance performance and resource overhead, we conduct a design space exploration (DSE) from an overall architectural perspective. An experimental platform with both software and hardware for SC-based DNNs is established. The software simulation results demonstrate that the best accuracy of the designed SC-DNN on the CIFAR-10 is 91.9%, which is 3.2% higher than that of the previous SC-DNN work. The VLSI implementation of the hardware is synthesized using the TSMC 28-nm CMOS process. Results show that compared to the binary computing counterpart, our design achieves$2.7\times $area efficiency and$3.4\times $energy efficiency. Compared to other SC-DNN accelerator designs, our design can provide$5.3\times $area efficiency and$7.3\times $energy efficiency. Jingguo Wu, Zongru Yang, Yun Chen 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2024 | A Semi-Folded Based High-Power-Efficiency FFT for Frequency Offset EstimateabstractFrequency offset estimate (FOE) is utilized to correct the incremental phase deviation caused by transceiver local oscillator mismatches, and determine the performance of the successive carrier offset compensation module. The FFT-based FOE method offers stability and controlled precision. However, as the FFT scale increases, the algorithm’s complexity also rises, which hinders real-time implementation of high-precision FOE hardware. While several research efforts aim to simplify the algorithm, there is limited focus on hardware implementation. In this paper, we introduce a semi-folded 256-point FFT with 64-way parallelism. By reusing the 64-point FFT module, computations are completed within four clock cycles, resulting in circuit area and power savings. We further leverage the semi-folded FFT structure to implement the FOE circuit. The experiments demonstrate that under the 28nm process technology, the semi-folded FFT achieves an area efficiency of 0.0043/(mm2/GS) and a power efficiency of 2.44mW/GS when working at a frequency of 714MHz. Additionally, when working at 500MHz the semi-foleded based FOE circuit has an area of 0.787mm2and a power consumption of 319.1mW, meets the requirements of a 256Gbps 16QAM communication system. Liyu Lin, Jingguo Wu, Xiaoyang Zeng, Yun Chen 0001 |
ISCAS | 4 |
| 2022 | A Low-latency Carrier Phase Recovery Hardware for Coherent Optical CommunicationabstractCarrier phase recovery (CPR) determines the accuracy of the receiver in modern coherent optical communication. The accurate estimation and tracking of carriers are particularly vital with the increase of throughput for long-distance transmit. It is a challenge to implement a real-time system because the computational complexity increases with fractional bits. Moreover, the conversion between polar coordinates and Cartesian coordinates introduces a high latency. In this paper, we present an FPGA implementation of low latency Viterbi-Viterbi 4thPower Estimation (VV4E) based CPR, which mainly performs the computation in Cartesian coordinates and implements the trigonometric function with a look-up table (LUT). Evaluations on Xilinx ZCU102 show that at a frequency of 370MHz, it introduces a 22-cycle latency to handle the 29.6 GBd QPSK signals, which is the minimum value to our knowledge. Liyu Lin, Kaihui Wang, Yun Chen 0001, Jianjun Yu, Xiaoyang Zeng |
ISCAS | 3 |
| 2022 | Safety Assurance System for Electric Vehicles Based on Infrared LiDARabstractWith the development of electric vehicles, more and more people choose electric vehicles as their means of transportation. Microwave radar and terahertz radar can be used to build safety assurance system for electric vehicles. However, both of the two technologies have the disadvantage of low stability. Infrared LiDAR has the advantages of high resolution, low cost and high reliability, but it has not yet been widely applied in the safety assurance system for electric vehicles. In this paper, we propose a safety assurance system for electric vehicles based on infrared LiDAR. In the design, infrared sensors are used for obstacle detection, and an MCU (microprogrammed control unit) is used for control. The excellent performance of the electric vehicle equipped with the safety assurance system in avoiding obstacles proves the high efficiency and high reliability of the system proposed in this paper. This experiment can deepen students’ understanding of the circuit system and cultivate their ability to construct the system independently, which can be of important meaning in their education. Yiqing Mao, Yun Chen 0001 |
ISCAS | 3 |
| 2017 | Implementation of a pipeline division-free MMSE MIMO detector that support soft-input and soft-output
Ziqiang Li 0004, Liyu Lin, Yun Chen 0001, Xiaoyang Zeng |
APCC | 3 |
| 2016 | Convergence-optimized variable node structure for stochastic LDPC decoderabstractBy using stochastic computation, a fully-parallel low-density parity-check (LDPC) decoder can be implemented using a lower wire complexity. In order to enhance the decoder performance, probability tracers, such as up/down counters, are added at each edge between variable nodes and check nodes, as described in previous literature. However, this causes a large decoding latency and a high number of decoding failures. In this paper, a convergence-optimized structure for variable nodes is proposed that is able to overcome these issues. As a result, the throughput for the proposed decoder is 20.5Gb/s, which is 101% higher than the original counter-based decoder presented in the previous literature. Qichen Zhang, Yun Chen 0001, Di Wu 0016, Xiaoyang Zeng, Yeong-Luh Ueng |
ICASSP | 2 |
| 2015 | Latency-optimized stochastic LDPC decoder for high-throughput applicationsabstractStochastic decoding can be applied to Low-Density Parity-Check codes in order to achieve high throughput with less area. However, most architectures suffer from large decoding latencies, due to the mechanism of stochastic computation. In this paper, three novel strategies, including the LUT-based initialization, the posterior-information-based hard decision and the Bit-Flipping-based post processing, are proposed in order to reduce decoding latency and hence improve throughput. For the standard IEEE 802.3an (2048, 1723) code, simulation indicates 75.7% reduction in average decoding cycles at 4.5 dB with satisfied bit error rate. Moreover, hardware implementation shows that the area of variable node units is reduced significantly in SMIC 65 nm technology. Di Wu 0016, Yun Chen 0001, Qichen Zhang, Lirong Zheng 0001, Xiaoyang Zeng, Yeong-Luh Ueng |
ISCAS | 2 |
| 2014 | A low-complexity LDPC decoder for NAND flash applicationsabstractThis paper presents an efficient min-sum-based decoder for high-rate low-density parity-check (LDPC) codes, where the first minimum and second minimum values are stored in registers. In order to meet a strict cost requirement imposed by NAND flash applications, we provide different upper limits for the first and second minimum values. Furthermore, we use non-uniform quantization for the second minimum value so as to reduce storage complexity. In order to enhance the error-rate performance, the normalization factor is determined based on the difference between the first two minimum values. Using the proposed techniques, a reduction in gate count of 13.36% can be achieved without suffering any degradation in error-rate performance. The implementation results for a rate-0.896 length-18624 layered decoder show that this decoder can achieve a throughput of 765.24 Mb/s at a clock frequency of 166 MHz with a gate count of 620K. Mao-Ruei Li, Hsueh-Chih Chou, Yeong-Luh Ueng, Yun Chen 0001 |
ISCAS | 4 |
| 2014 | Efficient symbol reliability based decoding for QCNB-LDPC codesabstractAs an extension of binary low-density parity-check (LDPC) codes, non-binary LDPC (NB-LDPC) codes show significantly better performance when the code length is moderate or small. Recently, enhanced iterative hard reliability based (EIHRB) decoding algorithm is proposed to reduce the computation complexity. However, the EIHRB algorithm suffers a lot from significant performance degradation when the column weight is small. In this paper, a symbol reliability based (SRB) decoding algorithm, which also performs well when the column weight is low, is proposed for NB-LDPC decoding to improve the decoding performance. With the same maximum iteration number, around 0.38 dB extra coding gain is achieved. Furthermore, the corresponding efficient decoder architecture is proposed. Comparison results have shown that the proposed SRB algorithm can not only achieve good coding gain, but the cost for hardware implementation is reasonable. Leixin Zhou, Jin Sha 0001, Yun Chen 0001, Chuan Zhang 0001, Zhongfeng Wang 0001 |
ISCAS | 3 |
| 2013 | An efficient multi-rate LDPC-CC decoder with layered decoding algorithmabstractAn efficient multi-rate Low-Density Parity-Check Convolutional Code decoder will be present in this paper. We will introduce layered decoding algorithm into LDPC-CC decoding. Simulation results shows that our method can achieve better performance than the original brief propagation algorithm with less processors. Besides a new ASIC architecture which adopt proposed algorithm and can support all code rate (1/2, 2/3, 3/4, 4/5) of the LDPC-CC code in IEEE 1901 is proposed. Based on SMIC 130 nm CMOS process, our decoder attaints a maximum throughput of 333.3 Mb/s at 200 MHz. The core area is 3.55 mm2with 10 processors. The average power consumption is 262 mW at code rate 4/5 and 200 MHz. The VLSI result shows that our decoder is both memory efficient and area efficient. Yun Chen 0001, Changsheng Zhou, Yuebin Huang, Xiaoyang Zeng |
ICC | 1 |
| 2013 | Memory efficient EMS decoding for non-binary LDPC codesabstractNon-binary low-density parity-check (NB-LDPC) codes are an extension of binary LDPC codes with significantly better performance when the code length is moderate. Previously, forward-backward schemes are used to implement check node processing, which need large amount of memory. In this paper, a novel approach-TCL-EMS is proposed for NB-LDPC decoding. Compared to original EMS decoding algorithm, the memory efficiency is improved and the average number of iterations is reduced significantly. Also, the overall decoder architecture is proposed. Leixin Zhou, Jin Sha 0001, Yun Chen 0001, Zhongfeng Wang 0001 |
ISCAS | 3 |
| 2013 | Accurate Sampling Timing Acquisition for Baseband OFDM Power-Line Communication in Non-Gaussian NoiseabstractIn this paper, a novel technique is proposed to address the joint sampling timing acquisition for baseband and broadband power-line communication (BB-PLC) systems using Orthogonal-Frequency-Division-Multiplexing (OFDM), including the sampling phase offset (SPO) and the sampling clock offset (SCO). Under pairwise correlation and joint Gaussian assumption of received signals in frequency domain, an approximated form of the log-likelihood function is derived. Instead of a high complexity two-dimension grid-search on the likelihood function, a five-step method is employed for accurate estimations. Several variants are presented in the same framework with different complexities. Unlike conventional pilot-assisted schemes using the extra phase rotations within one OFDM block, the proposed technique turns to the phase rotations between adjacent OFDM blocks. Analytical expressions of the variances and biases are derived. Extensive simulation results indicate significant performance improvements over conventional schemes. Additionally, effects of several noise models including non-Gaussianity, cyclo-stationarity, and temporal correlation are analyzed and simulated. Robustness of the proposed technique against violation of the joint Gaussian assumption is also verified by simulations. Chen Chen 0011, Yun Chen 0001, Na Ding, Jia-Chin Lin 0001, Xiaoyang Zeng, Defeng Huang |
IEEE Trans. Commun. | 2 |
| 2012 | A single-routing layered LDPC decoder for 10Gbase-T Ethernet in 130nm CMOSabstractA highly-parallel LDPC decoder architecture for 10Gbase-T applications is designed in this paper. Firstly, we reduce the routing complexity and corresponding power consumption by the proposed decoder architecture based on single routing networks. Secondly, the proposed architecture is designed with pipelined layered scheduling and multi-block parallel decoding, which improves operation speed and removes pipeline stalls in conventional highly-parallel layered scheduling. Thirdly, we trade off between hardware cost and throughput by a digit-serial data-path. Fourthly, an efficient early-termination circuit suitable for layered decoding is designed. The decoder is implemented in 130nm 1P8M CMOS process. The core area is 18.4mm2with 14% reduction, and the decoding throughput is 9.48Gbps operating at 278MHz and 5 iterations. The tested power consumption is 774mW at 1.2V and 80MHz. Dan Bao, Xubin Chen, Yuebin Huang, Yun Chen 0001, Xiaoyang Zeng |
ASP-DAC | 5 |
| 2012 | A 60mW baseband SoC for CMMB receiverabstractThis paper describes baseband SoC implementation of China Mobile Multimedia Broadcasting (CMMB) receiver, which integrates analog to digital (ADC), physical layer (PHY) baseband processor and medium access control (MAC) processor in single silicon wafer. MAC functions are fully implemented by firmware on an embedded 32-bit RISC-based processor. In addition, several power management techniques are utilized to reduce the power consumption of baseband SoC. The baseband SoC was successfully fabricated in 0.13µm one-poly six-metal (1P6M) CMOS process. Both analog and digital circuits are integrated on 4.8×4.8 mm2die consuming 60mW total power dissipation under 1.2V and 3.3V supplies. The experiment results reveal the proposed baseband SoC has excellent performance under the multipath channels. Jialin Cao, Dan Bao, Yun Chen 0001, Xiaoyang Zeng |
ASP-DAC | 4 |
| 2012 | An improved coarse synchronization scheme in 3GPP LTE downlink OFDM systemsabstractIn this paper, an improved algorithm of coarse synchronization in the downlink of 3GPP LTE system is presented. This new algorithm can reduce the MSE (mean square error) of detected fractional carrier frequency offset (FCFO) by an order of magnitude and improve the accuracy rate of start point estimation by near 80 percent when SNR equals 0 in contrast with conventional coarse synchronization algorithms, such as ML, MMSE, S&C and MC. The simulation result for the proposed algorithm is presented in comparison with the conventional algorithms. Na Ding, Chen Chen 0011, Wenhua Fan, Yun Chen 0001, Xiaoyang Zeng |
ISCAS | 4 |
| 2011 | An area-Efficient LDPC decoder for multi-standard with conflict resolutionabstractThis paper presents an area efficient decoder architecture that supports both perfectly structured and not perfectly structured LDPC codes. To verify our architecture, an area-efficient LDPC decoder that supports both China Multimedia Mobile Broadcasting (CMMB) and Digital Terrestrial/ Television Multimedia Broadcasting (DTMB) standards is developed. A solution is proposed to avoid memory access conflict problem caused by TDMP algorithm. The main timing schedule is arranged carefully to handle the operations of our solution while avoiding much additional hardware consumption. We also optimize the extrinsic message storing strategy to reduce the memory bits needed. Besides the extrinsic message recover and the accumulate operation are merged together. Based on SMIC 0.13 um standard CMOS process, the core area of the decoder is only 4.75 mm2and the maximum operating clock frequency is 200 MHz. With 5 iterations, the estimated average power consumption is 48.4 mW at 25 MHz for CMMB and 130.9 mW at 50 MHz for DTMB with 1.2V supply. Changsheng Zhou, Yunlong Ge, Xubin Chen, Yun Chen 0001, Xiaoyang Zeng |
ASAP | 4 |
| 2011 | A 4.32 mm2 170mW LDPC decoder in 0.13μm CMOS for WiMax/Wi-Fi applicationsabstractAn energy-efficient programmable LDPC decoder is proposed for WiMax and Wi-Fi applications. The proposed decoder is designed with overlapped processing units, flexible message passing network and medium-grain partitioned memories to achieve flexibility, area reduction, and energy efficiency. The decoder can be programmed by host processor with several special-purpose micro-instructions. Thus, various operation modes can be reconfigured. Fabricated in SMIC 0.13μm 1P8M CMOS process, the chip occupies 4.32 mm2with core area 2.97 mm2, and consumes 170mW with a throughput of 302Mb/s when operating at 145MHz and 1.2V. Dan Bao, Yan Ying, Yun Chen 0001, Xiaoyang Zeng |
ASP-DAC | 4 |
| 2011 | A channel estimation scheme for Chinese DTTB system combating long echo and high doppler shiftabstractA novel channel estimation scheme is presented to combat long echo and high Doppler shift for Chinese digital television terrestrial broadcasting (DTTB) system. With this method, fast fading channel with Doppler shift about 200Hz can be well equalized even when OdB echo at 31.8us exists with head length 420, other modes likes 595 and 945 are also supported. Simulation results are given under Chinese DTTB standard to demonstrate the performance of the algorithm. This method can also be adopted to cyclic padded (CP) or zero padded (ZP) OFDM systems and single carrier ones. Yun Chen 0001, Yunlong Ge, Huxiong Xu, Xiaoyang Zeng |
ISCAS | 2 |
| 2010 | A flexible LDPC decoder architecture supporting two decoding algorithmsabstractIn this paper a programmable and area-efficient decoder architecture supporting two main stream decoding algorithms for any Block-LDPC codes is presented. The novel decoder can be configured to decode in either TPMP or TDMP decoding mode according to different Block-LDPC codes. To verify our proposed architecture, a flexible LDPC decoder which supports IEEE 802.16e is implemented using a 0.13um CMOS process with a total area of 6.3 mm2 and maximum clock frequency of 260 MHz. The estimated comsumption is 270 mW when operates at 125 MHz and 1.2V supply. Shuangqu Huang, Dan Bao, Bo Xiang, Yun Chen 0001, Xiaoyang Zeng |
ISCAS | 4 |
| 2010 | A 128/256-point pipeline FFT/IFFT processor for MIMO OFDM system IEEE 802.16eabstractIn this paper, we present a novel 128/256-point FFT/ IFFT processor for the applications in IEEE 802.16e based on MIMO-OFDM. The pipeline FFT architecture is proposed to efficiently deal with 1-4 multiple data sequences, and increase the throughput. Furthermore, less hardware complexity is needed in our design compared with conventional individual parallel approach. The signal-to-quantization noise ratio (SQNR) is 42.7 dB. The proposed FFT has been designed in 0.13 μm technology with the core size of 1.470 × 1.469 mm2. Huxiong Xu, Wenhua Fan, Yun Chen 0001, Xiaoyang Zeng |
ISCAS | 4 |