Songping Mai

dblp:72/7032 · DBLP profile ↗
← Back
21ranked-venue papers
4as first author
12since 2021 · last 2025
0000-0002-0572-9066ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 16 · 2 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-authorComputer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2025 A 9.13-to-15.31 GHz 205.1 dBc/Hz FOMT Dual-Core Dual-Mode VCO with Forward Body Bias Technique For Sub-7 GHz Applications
abstract
This paper presents a low phase noise (PN) and low power voltage-controlled oscillator (VCO) for Sub-7 GHz applications. We introduce a novel Colpitts VCO with forward body bias technique designed to augment the negative transcon-ductance and to optimize PN. The integration of orthogonal-coupled VCO core is employed for a wide frequency tuning range (FTR) based on the previous frequency expansion developed by the dual-mode, with capacitor banks and varactors. Moreover, a switched varactor unit is employed to realize KVCO linearization and optimize the frequency regulation range. Designed in 65-nm CMOS, the measured results exhibit a low PN of -139.8 dBc/Hz at 10 MHz offset and a figure-of-merit (FOM) of 191.1 dBc/Hz from a 6.108 GHz carrier, while the FTR is 50.6% from 9.13 to 15.31 GHz. The proposed VCO consumes 2.78mW at 0.72-V supply voltage.
Ruiyong Xiang, Xinyi Zhang 0008, Xian Tang, Songping Mai, Haigang Feng
ISCAS5
2025 A Fast Transient FVF LDO With Improved Gate Buffer and Level Triggered Transient Enhancement Circuit in 22 nm process
abstract
This paper presents a fast transient flipped-voltage-follower-based low-dropout regulator (FVF LDO) with improved gate buffer and level-triggered transient enhancement circuit (LTTEC) in a 22 nm process. The fast loop composed of an FVF and the improved gate buffer exhibits ultra-high bandwidth and high slew rate. The proposed LTTEC can supply additional sourcing or sinking current to reduce the undershoot or overshoot during load transitions. The introduced bulk driven technique can effectively improve the power supply rejection (PSR) performance. The proposed LDO achieves 0.25ns response time with simulated 40.0 mV undershoot voltage and 36.7 mV overshoot voltage for the load current transition between 100 µA and 20 mA within 0.1 ns.
Haigang Feng, Songping Mai, Xian Tang
ISCAS4
2025 Activation and Weight Distribution Balancing for Optimal Post-Training Quantization in Learned Image Compression
abstract
Recently, Learned Image Compression (LIC) models have garnered significant attention due to their superior performance in comparison to traditional image codecs. However, the growing complexity of these deep learning-based models results in high memory consumption and computational load, which limits their practical deployment. Quantization has emerged as a promising technique to reduce both the storage requirements and computational overhead. Despite its success in high-level vision tasks like image recognition and object detection, quantization techniques applied to LIC models remain underexplored. In this work, we identify the unique challenges of quantizing LIC models, specifically focusing on the impact of latent distribution ranges in high-bitrate. We observe that the activation layers of high-bitrate models exhibit a wider distribution range, which causes significant performance degradation after quantization. Furthermore, we explore the limitations of existing LIC quantization schemes, such as per-channel quantization for activation layers, which result in poor hardware acceleration performance and increased data storage overhead. To address these challenges, we propose the activation and weight distribution balancing post-training quantization (AWDB-PTQ) method for LIC models, which uses a coarse-to-fine strategy to optimize balancing coefficients. In addition, we employ per-tensor activation quantization and symmetric uniform quantization to better facilitate hardware acceleration. Experimental results demonstrate that our proposed method outperforms existing methods in terms of both compression performance and computational efficiency. Our code and data are available at: https://github.com/jie-yu16/AWDB-PTQ.
Songping Mai, Peng Zhang 0007, Yucheng Jiang
ACM Multimedia2
2025 Mixed-Precision Post-Training Quantization for Learned Image Compression
abstract
With the rapid development of neural networks, learned image compression (LIC) has surpassed traditional methods in the field of image compression. However, current LIC frameworks rely on floating-point operations. On the one hand, the use of conventional floating-point operations results in high computational costs, limiting their deployment in the Internet of Things (IoT) devices. On the other hand, when employing prior models, floating-point errors in prior calculations can cause cross-platform decoding failures. Neural network quantization addresses these issues by using low-bitwidth fixed-point, which reduces computational complexity and memory usage while eliminating inconsistencies from floating-point operations. Although quantizing the entire network to a uniform bitwidth simplifies hardware deployment, it can significantly impair coding efficiency, particularly in LIC networks at high bitrates, where certain layers are more sensitive to quantization noise, and we are first to provide both empirical and theoretical insights into the underlying reasons for these layers’ increased quantization difficulty. To address these challenges, we propose a mixed-precision post-training quantization (MP-PTQ) scheme that requires only a small calibration dataset without retraining, enabling rapid deployment. This method assesses the quantization sensitivity of each layer through rate-distortion loss analysis, thereby allocating the optimal bitwidth and quantization scale for each layer. Our method significantly improves the compression performance of quantized LIC models, achieving a Bjøntegaard-Delta rate (BD-rate) loss of less than 1% compared to their floating-point counterparts. Furthermore, by employing per-tensor activation quantization, we enable hardware-friendly acceleration—achieving approximately 3× speedup over floating-point models.
Songping Mai, Yucheng Jiang
IEEE Internet Things J.2
2025 PEPPR-DWS on FPGA: Elevating Universal Parallelism and Precision Through Pulse-Enhanced Push-Relabel and Diffusion Wave Search
abstract
The push-relabel algorithm is recognized as one of the efficient algorithms in the field of graph cut, finding widespread applications in computer vision. While its pixel-level parallel implementations are prevalent, existing methods predominantly rely on checkerboard scheduling, imposing inherent constraints on neighborhood size, limited to four. This limitation compromises both algorithm precision and efficiency, hindering real-time and high-precision applications. To address these issues, this article introduces a novel approach to accelerate push-relabel algorithm implementation on FPGA in a more universal and efficient manner, supporting variable-sized image block operations. First, by introducing the deferred update strategy, we realize the pulse-enhanced parallel push-relabel (PEPPR) algorithm to address data contention and conflict in parallel processing. Second, the simultaneous weighted push method is proposed, further enhancing parallel operations. Lastly, we introduce the efficient diffusion wave search (DWS) algorithm to expedite algorithm convergence and reduce redundancy. While achieving a modest$1.7\times $acceleration compared to state-of-the-art implementations, the proposed algorithm (PEPPR-DWS) successfully overcomes the inherent limitations of checkerboard scheduling in full pixel-level parallelism. In the test based on Middlebury benchmark V3, the proposed 8-neighborhood implementation exhibits a reduction of error rate by over 1% compared to the typical 4-neighborhood implementation. It provides a versatile and efficient solution for high-precision and real-time applications, holding substantial potential for practical applications.
Yucheng Jiang, Songping Mai
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2024 Deep Learning Based Single-Shot Profilometry by Three-Channel Binary-Defocused Projection
abstract
Fringe projection profilometry (FPP), a widely used 3D reconstruction method, often encounters a dilemma between speed and accuracy for dynamic measurement. This paper proposes a deep-learning based single-shot 3D reconstruction method, which considers both speed and accuracy. We utilize the individual red, green, and blue channels of the projector to successively project three binary patterns in a defocused manner. At the same time, the camera exposures during this process and finally captures a single image. During image processing, we combine the task of the wrapped phase and absolute phase prediction, which enables an end-to-end high-precision estimation of the absolute phase from a single fringe pattern through a single network. Experiments on various scenes, encompassing both static and dynamic objects, substantiate our method’s high-quality 3D reconstruction capability from only a single shot, surpassing the performance of previous approaches.
Tianbo Liu 0007, Songping Mai
ICASSP2
2024 A High Efficiency, Low EMI Non-inverting Buck-Boost Converter in Wireless Power and Data Transfer System for Brain Computer Interface
abstract
This paper presents a high efficiency, fixed- frequency waveform Buck-Boost converter which converts a 2.7V to 4.5V Li-ion battery voltage to the regulated voltage featured high efficiency and low electromagnetic interference (EMI) for brain computer interface applications. A smoothly switching peak-valley current control mode, enabling this converter to switch freely from valley-current-mode control in Buck mode to peak-current-mode control in Boost mode, is adopted to avoid sub-harmonic oscillations. And a fixed- frequency control technique based on cycle alternation in Buck-Boost mode, which fixes the frequency of the inductor current as opposed to random-mode-control, is adopted to obtain high efficiency and low EMI. The proposed Buck-Boost converter is implemented in 180nm BCD process. It can achieve and keep the efficiency of above 92.33%, of which the peak efficiency is 98.97%, over the load current ranging from 100mA to 400mA and the input voltage ranging from 2.7V to 4.5V.
Xiangsheng Xu, Qihang Zhang, Songping Mai
ISCAS4
2023 MispredTable: A Side Branch Predictor to TAGE in Multithreading Processors
abstract
Tagged geometric history length (TAGE) branch predictor shows good prediction accuracy with sufficient storage budget. However, shrinking TAGE's storage budget dramatically reduces its prediction accuracy. In this paper, we propose MT (Misprediction Table)-TAGE, which adds a side predictor to the TAGE branch predictor. This side branch predictor will record the branches with the highest misprediction rate during the running process of the program and modify the prediction result of the branch when the number of mispredictions is higher than dynamic threshold. We also propose a new algorithm that enables MT-TAGE to be implemented in multithreaded processors. This work is written in SystemVerilog and tested on a RISC-V multithreaded processor. The experimental results show that in a quad-thread processor, the misprediction rates of MT-TAGE under the storage budget of 32kbits and 8kbits are 5.03MPKI (misprediction per kilo-instructions) and 5.54MPKI, respectively, which are 12.5% and 15.5% less than TAGE under the same area.
Xincheng Yang, Songping Mai, Rongxin Bao
ISCAS2
2023 A Low-power ASK Demodulator for Wireless Power and Data Transfer Systems Supporting Ultra-low Modulation Depth of 0.03%
abstract
In wireless power and data transfer (WPDT) systems implemented with amplitude shift keying (ASK) data demodulation, the low amplitude modulation depth (MD) is usually preferred as it helps to improve energy harvesting efficiency, transmission range and stability. In this paper, a fully integrated ASK demodulator supporting ultra-low MD is proposed, which comprises a two-stage self-biased shifted limiter (SSL) that provides sufficient conversion gain and operates at low power consumption by introducing an adaptive biasing circuit. This structure is implemented in 0.18$\mu \mathbf{m}$high-voltage Bipolar-CMOS-DMOS technology. The detectable MD is measured as low as 0.03%, while the power consumption is only 52.5$\mu \mathbf{W}$.
Qingbing Zhang, Songping Mai, Ruolin Zhou, Xincheng Yang
ISCAS2
2023 A Low-Power Keyword Spotting System With High-Order Passive Switched-Capacitor Bandpass Filters for Analog-MFCC Feature Extraction
abstract
This paper presents a low-power high-accuracy key-word spotting (KWS) system based on analog passive switched-capacitor (SC) bandpass filters (BPF). The proposed system inno-vatively extracts the Mel-frequency cepstrum coefficient (MFCC) features with all-analog circuits, providing a better spotting performance than the present analog short-time amplitude or energy features under the same condition. At the circuit level, the analog-MFCC extraction includes the low-noise amplifier, BPF, squarer, integrator, and discrete cosine transformer. And thanks to the effective analog-MFCC features, the size of the fully-connected neural network (FCNN) classifier in our KWS system is enormously reduced. A high-order and fully-differential BPF is also proposed, achieving ultra-low power and high dynamic range by combining zero and pole generation stages rather than building stages separately in traditional ways. Fabricated in 0.18-$\mu \text{m}$CMOS, our filterbank of eight BPFs is measured with a power consumption of 83.2 nW, with a 69.7 dB dynamic range at 5% THD, having advantages over other SC BPFs with similar functions. The total power consumption of our feature extractor is 661.7 nW, achieving an accuracy of 96.6% in two-keyword spotting by an FPGA-based, 15k bit parameter FCNN with a 9.6$\mu \text{s}$latency.
Shiying Zhang, Fukun Su, Songping Mai, Kong-Pang Pun, Xian Tang
IEEE Trans. Circuits Syst. I Regul. Pap.4
2022 Narrowband FSK Transceiver Circuit for Wireless Power and Data Transmission in Biomedical Implants
abstract
Biomedical implant is very sensitive to volume, so that single coil for wireless power and data transfer (WPDT) is widely used. Frequency-shift keying (FSK) is a robust data modulation scheme with negligible sensitivity to amplitude noise and movements of the transmitter and receiver coils of an inductive link. However, achieving efficient transmission of power while modulating FSK has always been a difficult task, because high data rate FSK requires a broadband modulator. In this paper, we propose a narrowband FSK transceiver scheme, which enables WPDT link to transmit power efficiently when transmitting data. Using this approach, the power transfer efficiency (PTE) and the power delivered to the load (PDL) were measured to be 51.92% and 92.42 mW, respectively. With carrier frequencies of 3.952 and 4.516 MHz, the WPDT link we have implemented can transmit data at a rate of 564.48 kbps, with a 5.35$-\mu$H transmitter coil and a 0.7$-\mu$H receiver coil. The transmitter SoC is fabricated in a 0.18$-\mu$m CMOS technology and occupies 2 mm2. We also designed the corresponding demodulation circuit, of which power consumption is 8.49 uW. Our approach is expected to provide a new communication solution for biomedical implants.
Songping Mai
ISCAS2
2022 Design of High Efficiency Planar Spiral Coil for Implantable Wireless Power Transfer Systems
abstract
This paper proposes a five-coil array united structure for implantable wireless power transfer system. The coil structure comprises a radiating coil and a driving coil. The driving coil drives radiating coil to radiate power. Both coils resonate at 13.56 MHz. It can generate almost uniform magnetic distribution which would improve the dislocation tolerance of the wireless power transfer system. At 13.56 MHz, the S parameter is used to calculate the system efficiency. The five-coil array united structure is fabricated on planar printed circuit board. The structure is compact and suitable for implants. According to the experimental results, the system efficiency can reach as high as 54.7% when coils align (transfer distance from 10 mm to 20 mm) and keep almost unchanged with a dislocation from 5 mm to 10 mm (transfer distance is 10 mm). And almost all efficiency values are above 50%. Efficient and stable power transmission is realized.
Weicheng Zhao, Songping Mai
ISCAS2
2018 An Energy-Efficient High-Frequency Neuro-Stimulator with Parallel Pulse Generators, Staggered Output and Extended Average Current Range
abstract
This paper presents a high-frequency pulse stimulation (HFPS) output stage of neuro-stimulator with extended average output current range and high power efficiency. The output stage features two parallel buck-boost converters without any filter capacitor at the output node. Compared with HFPS output stage with only one converter, the proposed circuit doubles the maximum average current by staggering the output of two converters. Compared with traditional voltage mode stimulation (VMS), HFPS improves the power efficiency by discharging the inductor current through the tissue load directly, rather than through a filter capacitor with constant voltage. Besides, the control circuit for the proposed HFPS is much simpler than that of traditional VMS converter, which reduces the current consumption significantly and thus improves the efficiency. Test results confirm that with staggered output of two parallel converters, the maximum average current is doubled. The maximum energy efficiency of the proposed HFPS is 76.4%.
Guijie Zhu, Songping Mai, Xian Tang, Chun Zhang 0001, Zhihua Wang 0001, Hong Chen 0002
ISCAS2
2017 A 9.4 pJ/bit 432 MHz 16-QAM/MSK transmitter based on edge-combining power amplifier
abstract
An energy-efficiency 16-QAM/MSK transmitter (TX) working at 432 MHz is presented for the short range communications. By adopting the edge-combining power amplifier (ECPA), only an injection-locked ringing oscillator (ILRO) with very low frequency is required, in comparison to the high frequency local oscillator in conventional TX's, which helps to reduce the power consumption of this TX significantly. A new method to control the amplitude of the edge-combined quadrature signals is proposed to generate the 16-quadrature amplitude modulation (16-QAM) and minimum shift keying (MSK) modulations in this work. A finite impulse response (FIR) filter is introduced in the ECPA for side lobe suppression enhancement. Implemented in a 0.18 μm CMOS process, the proposed TX consumes 468 μW when outputting -15 dBm 50 Mbps 16-QAM/MSK signals in simulation. Under 1 V supply, the TX can achieve an energy efficiency of 9.4 pJ/bit at the maximum data rate of 50 Mbps. The phase noise of the RF output in the injection-locked state is -103 dBc/Hz, at 1MHz offset.
Yanshu Guo, Songping Mai, Zhaoyang Weng, Hanjun Jiang, Zhihua Wang 0001
ISCAS2
2017 An optical tracker based registration method using feedback for robot-assisted insertion surgeries
abstract
Robot-assisted needle insertion technology is now widely used in minimally invasive surgeries. The success of these surgeries highly depends on the accuracy of the position and orientation estimation of the needle tip. In this paper, we proposed an optical tracker based registration method for the marker-robot-needle tip registration in a robot-assisted needle insertion surgery. The method consists of two steps. First, with the guide of the optical tracker, we use the motion vector from the marker's current position to its target point as a feedback control to automatically register the robot-marker and the robot-tracker coordinates respectively. Then, in the procedure of marker-needle tip registration, we use two points on the needle holder to achieve a high accuracy needle orientation registration. We have conducted a series of experiments for the proposed method on a verified platform. Results show that the position estimation error is below 0.5mm and the orientation estimation error is below 0.5 degrees.
Xingtong Liu, Guolin Li, Songping Mai, Zhihua Wang 0001
ISCAS5
2015 A high-voltage, energy-efficient, 4-electrode output stage for implantable neural stimulator
abstract
In order to raise energy efficiency of implantable neural stimulators, an output stage circuit with a self-adaptive stimulating voltage is proposed. The output stage consists of a digital to analog converter (DAC), a single-ended primary inductance converter (SEPIC), a switch array and a current detection module. The SEPIC enables stimulating voltage to vary according to the stimulus current on tissue load or DAC digital input. It improves the output energy efficiency by dynamically cutting down the unnecessary stimulating voltage headroom on the current driver, especially when the stimulus current is small. Simulation showed that the stimulating voltage could change from 1.1V to 18.7V with less than 25mV ripple and the stimulus current could change from 0.7mA to 12.1mA with 0.1mA resolution. The output stage could work smoothly under the voltage supply range between 3.3V and 4.1V. The output efficiency of the proposed work could keep higher than 60% when stimulus current was above 1.4mA and even higher than 70% when stimulus current was above 2.9mA. The output stage is being fabricated on a silicon chip using CSMC 1μm HV technology, occupying 3.46mm×1.94mm.
Jinghui Liu, Songping Mai, Chun Zhang 0001, Zhihua Wang 0001
ISCAS2
2013 A high-performance low-power SoC for mobile one-time password applications
abstract
The presented SoC is an 8-bit MCU based processor with 32-bit encryption algorithm acceleration and special security mechanism dedicated to mobile one-time password (MOTP) applications. The encryption algorithm accelerator can perform several 32-bit key operations for hash function computation, with each in one clock cycle. The SoC also features protection mechanisms which can prevent attackers from stealing the algorithm program code or the secret key, or utilizing under-voltage attacks. The SoC chip, fabricated in 0.25 μm CMOS process, has more than 50% efficiency improvement on hash function computation if compared with those 8-bit general-purpose MCUs and consumes less than 7.2 μA current in average for typical MOTP applications.
Songping Mai, Chunhong Li, Yixin Zhao, Chun Zhang 0001, Zhihua Wang 0001
ISCAS1
2012 The design and implementation of a chipset for the endoscopic Micro-Ball
abstract
A chipset, including one master chip and six same slave chips, is designed and implemented for the endoscopic Micro-Ball, which has six cameras to minimize the blind area while examining the human gastrointestinal (GI) tract. For lowering power consumption, a multi-level clock management for the whole chipset is implemented. A power and area efficient 4 × 4 JPEG image compressor with high image quality is integrated into the slave chip to reduce energy consumption and save FLASH memory space. The master chip and slave chip have been fabricated in 0.18 μm CMOS technology. The experimental results show that the designed chipset for Micro-Ball can support different image frame rates: from 1fps to 6fps with 480×480 image resolution. The power consumption of the chipset is 1.6mW@2fps. The power consumption of the Micro-Ball is estimated about 6mW@2fps. It can work for more than 15 hours, powered by a 1.55V@60mAh battery.
Yingke Gu, Guolin Li, Tianjia Sun, Shouhao Liu, Songping Mai, Zhihua Wang 0001
ISCAS7
2012 A Time-Frequency Aware Cochlear Implant: Algorithm and System
Songping Mai, Yixin Zhao, Chun Zhang 0001, Zhihua Wang 0001
ISNN (2)1
2006 A Cochlear System with Implant DSP
abstract
Cochlear implants have achieved great success in restoring hearing to profoundly deaf people. New generation implants pose more restrict requirements on processing precision and make traditional cochlear systems insufficient to deal with the large amount of data which should be transmitted via the wireless link. A new cochlear system with implant DSP is proposed to address the data rate problem. This system needs to transmit only voice-band signals with a data rate of 100 Kbps and the data rate bottleneck is removed for future development. By optimizing the speech processing algorithm and the inductive wireless link, the power consumption of this system increases less than 10% when compared with the traditional system.
Songping Mai, Chun Zhang 0001, Mian Dong, Zhihua Wang 0001
ICASSP (5)1
2006 An open-source based DSP with enhanced multimedia-processing capacity for embedded applications
abstract
This paper proposes an open-source based 32-bit digital signal processor (DSP) suitable for embedded multimedia processing applications. The DSP is a MIPS-based scalar RISC with Harvard micro architecture, 5-stage integer pipeline and enhanced multimedia-processing capacity. It features a data path including non-aligned data access mechanism, two-dimensional DMA function and separated embedded instruction and data memories. This special architecture greatly promotes execution efficiency of boundary-crossed and block data access operations often found in multimedia processing. The processor achieves 80-MIPS performance with typical power dissipation of 88 mW. The prototype has been fabricated using UMC 0.18mum n-well 1P6M standard CMOS technology
Songping Mai, Wenli Lan, Chun Zhang 0001, Zhihua Wang 0001
ISCAS1