Chun Zhang 0001

dblp:77/2866-1 · DBLP profile ↗
← Back
47ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0001-9791-4500ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 29 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9Artificial intelligence and machine learning · 6Applied, interdisciplinary, general and emerging computing · 3
YearPublicationVenuePosition
2025 HyperGS: Efficient Real-Time 3D Gaussian Rendering Processor Through Hierarchical Sorting
abstract
This paper proposes HyperGS, the first complete hardware accelerator implementation for 3D Gaussian Splatting with algorithmic optimization. Through an in-depth analysis of the Gaussian rendering pipeline characteristics, we designed an efficient and practical hardware architecture with optimized resource utilization. Specifically, we proposed a parallel large-scale sorting unit to enable parallel processing of depth sorting and rasterization, saving 22% of computing time. We also designed a fully pipelined preprocessing calculation unit with low resource usage for model parameter preprocessing, generating key-value pairs. We propose an efficient rasterization unit based on factorization. The rendering pipeline has been improved by reducing preprocessing parameters and decreasing the number of external memory accesses. Experimental results show that HyperGS reduces power consumption by a factor of 84 and improves performance by a factor of 44 compared with existing NVIDIA Jetson GPUs. Using only 19.2GB/s of bandwidth, we achieved rendering speeds ranging from 20.76 to 43.38 frames per second. Our work provides a feasible solution for efficient real-time 3D Gaussian rendering on resource-constrained edge computing devices.
Cheng Nian, Xiaorui Mo, Jiaying Peng, Weiyi Zhang 0002, Fasih Ud Din Farrukh, Chun Zhang 0001
ISCAS7
2025 An 197-μJ/Frame Single-Frame Bundle Adjustment Hardware Accelerator for Mobile Visual Odometry
abstract
This article presents an energy-efficient hardware accelerator for optimized bundle adjustment (BA) for mobile high-frame-rate visual odometry (VO). BA uses graph optimization techniques to optimize poses and landmarks and the applications are robot navigation, virtual reality (VR), and augmented reality (AR). Existing software implementations of BA optimization involve complex computational flows, numerical calculations, Lie group, and Lie algebra conversions. This poses challenges of slow computational speeds and high power consumption. A two-level reuse hardware architecture is proposed and implemented that efficiently updates the Jacobian matrix while reducing the field-programmable gate array (FPGA) hardware resources by 25%. A set of methodologies is proposed to quantify the errors caused by fixed-point systems during optimization. A fully pipelined architecture is implemented to increase computational speed while reducing hardware resources by 29%. This design features a parallel equation solver that improves processing speed by$2\times $compared to conventional approaches. This article employs a single-frame local BA VO on the KITTI dataset and EuRoC dataset, achieving an average translational error of 0.75% and a rotational error of$0.0028~^{\circ } $/m. The proposed hardware achieves a performance ranging from 188 to 345 frames/s in optimizing two main feature extraction methods with a maximum of 512 extracted feature points. Compared to state-of-the-art implementations, the accelerator achieved a minimum energy efficiency ratio of 11.6 mJ and$191~\mu $J on the FPGA platform and application-specific integrated circuits (ASICs) platform, respectively. These improvements underscore the potential of FPGAs to enhance VO systems’ adaptability and efficiency in complex environments.
Cheng Nian, Xiaorui Mo, Weiyi Zhang 0002, Fasih Ud Din Farrukh, Yushi Guo, Chun Zhang 0001
IEEE Trans. Very Large Scale Integr. Syst.7
2024 A 77.79 GOPs/W Retentive Network FPGA Inference Accelerator with Optimized Workload
abstract
Retention Network (RetNet) is a novel neural network model, with a inference complexity of O(1), considered to be the successor of the Transformer model. This work presents an inference accelerator supporting RetNet and DNN algorithms focus on dataflow optimizing, model quantization for improved parallelism and reusability, and addressing K-V cache issues in Transformers. To meet model inference storage bandwidth requirements, we used output block stationary (OBS) dataflow to explore data reusability in time and space. A hardware accelerator for the position encoding module is designed, reducing storage bandwidth and on-chip storage during Extrapolatable Position Embedding (XPOS). Compared to Transformer models of the same size architecture, this work has better inference accuracy, achieving energy efficiency ratio of 77.79 GOPs/W, a 12.5x speedup, and 2.13x energy efficiency compared to GPU implementations, which is 1.33x better than the transformer work baseline. Meanwhile, the inference process of Muti-Scale-Retention(MSR) is completed with the smallest on-chip RAM of 350KB, thereby enabling a more efficient ASIC implementation of hardware accelerators.
Cheng Nian, Weiyi Zhang 0002, Fasih Ud Din Farrukh, Liting Niu, Dapeng Jiang, Chun Zhang 0001
IECON7
2023 Hardware-Software Co-Design of Matrix-Solving for Non-Linear Optimization in SLAM Systems
abstract
Simultaneous Localization and Mapping (SLAM) is one of the most important techniques for autonomous robots that enables the robot aware of its current position and the surrounding environment. There is a significant improvement in the accuracy with the advancement in SLAM algorithms. However, the computation complexity increases accordingly and the embedded processors of autonomous robots struggle to support heavy calculation. Matrix-solving contributes a major portion of calculation time and considering sub-tasks such as bundle adjustment takes over 40% of total time. Therefore, it is significant to optimize the calculations required for matrix-solving. However, previous works for matrix-solving accelerators are generalized and the specific matrix form in SLAM problems is not fully considered. This work concentrates on the dedicated software and hardware codesign of the matrix-solving task in SLAM systems and provides three solutions for different scales of matrix-solving problems in SLAM. The proposed FSFI-Cholesky and FI-Iterative method have achieved up to 120.2x speed improvement over the non-optimized Cholesky algorithm. Moreover, this work also reduces the execution time by more than 7.0x compared to the state-of-the-art design with fewer DSPs used for both dense and sparse matrices.
Liting Niu, Weiyi Zhang 0002, Cheng Nian, Fei Shao, Fasih Ud Din Farrukh, Chun Zhang 0001
IECON6
2023 High Linearity Front-End Circuit for RF Sampling ADCs with Nonlinear Junction Capacitor Cancellation
abstract
This paper presents a high linearity front-end circuit for RF sampling ADCs, including an input buffer and a sampling network. The input buffer uses a two-stage NMOS cascode structure and is powered by a separate LDO to support a larger signal swing input with high power supply rejection (PSR) and linearity. We use bootstrap switch with bulk-switching techniques to ensure sampling linearity, while a feed-through compensation technique with self-cancellation of nonlinear junction capacitor is applied to achieve better performance at high-frequency inputs. The above techniques are validated at a 1GS/s ADC in 65nm process, and the simulation results show that the low-frequency PSR of the input buffer reaches over 120dB, and the SNR, SNDR and SFDR of the overall front-end circuit are 78.74dB, 72.37dB and 75.37dB at 2GHz frequency 1.6Vpp input. The −3dB bandwidth of the front-end circuit achieves 4.4GHz.
Yihang Cheng 0003, Fule Li, Chun Zhang 0001, Zhihua Wang 0001
ISCAS4
2022 A 56-Gbps PAM-4 Wireline Receiver With 4-Tap Direct DFE Employing Dynamic CML Comparators in 65 nm CMOS
abstract
This paper presents a four-level pulse amplitude modulation (PAM-4) receiver that incorporates a continuous time linear equalizer, a variable gain amplifier, a phase interpolator-based clock and data recovery, and a 4-tap direct decision feedback equalizer (DFE) for moderate channel loss applications in wireline communication. A dynamic current-mode logic comparator (DCMLC) is proposed and employed in the DFE. The DCMLC, which adopts dynamic logic, breaks the trade-off between the bandwidth and the clock to Q delay in the traditional current-mode logic comparator (CMLC). Compared with the traditional CMLC, the DCMLC reduces the clock to Q delay by 36%, which allows the implementation of a 4-tap direct DFE. Moreover, the first tap feedback signals are directly tapped from the output of the DCMLC, allowing the first tap feedback current to initiate 0.5UI before the decision clock. The PAM-4 receiver prototype is fabricated in a 65nm CMOS process. At a data rate of 56-Gbps, it can compensate for up to 20.17dB loss and achieve a bit error rate$< 1\text{E}$-10 with a power efficiency of 4.75 pJ/bit.
Dengjie Wang, Jiawei Wang 0004, Zeliang Zhao, Chun Zhang 0001, Zhihua Wang 0001, Hong Chen 0002
IEEE Trans. Circuits Syst. I Regul. Pap.6
2020 A CNN Based Human Bowel Sound Segment Recognition Algorithm with Reduced Computation Complexity for Wearable Healthcare System
abstract
Human bowel sounds (BSs) deliver much useful information about gastrointestinal health status. In recent years, the utilization of advanced bio-acoustic sensors, such as the wired electronic stethoscopes and the wireless wearable sound recording patches, has enabled researchers to record, store and analyze the BSs in digitized manners, i.e., computerized bowel sound analysis. However, for collected BS recordings, how to effectively pick up the segments that contain BS events while ignoring those segments that only contain background noises remains difficult. Moreover, the BS segment recognition algorithms that are meant to be applied in the wearable healthcare scenes are further required to retain a low computation complexity. In this work, a light weighted BS recognizer based on convolutional neural networks (CNNs) is proposed for wearable systems. Specifically, the proposed recognizer firstly converts each one-dimensional segment into a two-dimensional spectrogram by calculating the Mel-frequency cepstrum coefficients (MFCCs) frame by frame and then passes the spectrogram through a CNN to infer the category of the segment. To validate the CNN-based BS recognizer, a 28 minutes of BS dataset that contains 955 BS-present segments and 725 BS-absent segments are made. Experimental results on the dataset show that the recognizer attains the 91.25% and 90.83% mean accuracies for the `not-across-subjects' and `across-subjects' validation, respectively. Moreover, compared with the state-of-the-art LSTM approach, the CNN-based BS recognizer is light-weighted with only 20.35k parameters, which is a quarter of that of the LSTM. Due to the lower model complexity, the CNN-based BS recognizer has the potential to be integrated into the gateways in a wearable system to secure the function that only the BS-present segments are allowed to be relayed to the remote. User privacy can be better protected in this way.
Hanjun Jiang, Chun Zhang 0001, Wen Jia, Zhihua Wang 0001
ISCAS4
2019 A Solution to Optimize Multi-Operand Adders in CNN Architecture on FPGA
abstract
Convolutional Neural Network (CNN) is a very popular method in recent times to solve many computer vision tasks. However, CNN is becoming computationally intensive as time is progressing which requires a dedicated hardware for real time implementation. Graphics Processing Unit (GPU) and Field Programmable Gate Array (FPGA) are two hot choices to execute and accelerate the CNN network. FPGA has an advantage over GPU due to its flexible architecture and it can also provide high performance per unit watt of power. These benefits make FPGA a suitable candidate for CNN acceleration. However, optimization is required for FPGA based accelerator design to accommodate more computations. One of the challenges in accelerator design is to perform addition of intermediate results generated in a process of convolution. Therefore, Multi-Operand Adders (MOAs) are necessary in the accelerator design of CNN on FPGA but consume most of the area. Optimization strategy based on WALLACE tree architecture is proposed in this article to replace the typical binary adder tree in CNN accelerator design. Experimental results show an improvement in terms of area optimization and performance in comparison with the previous implementation.
Fasih Ud Din Farrukh, Tuo Xie, Chun Zhang 0001, Zhihua Wang 0001
ISCAS3
2018 A Model for Detection of Angular Velocity of Image Motion Based on the Temporal Tuning of the Drosophila
Huatian Wang, Paul Baxter 0001, Chun Zhang 0001, Zhihua Wang 0001, Shigang Yue
ICANN (2)4
2018 A Dual-Modal Vision-Based Tactile Sensor for Robotic Hand Grasping
abstract
Humans' fingertips can perceive not only the magnitude and the direction of force but also the texture of object. When we grasp an object, the surface texture sensing of the fingertip helps us recognize the object and the force feeling that is parallel to the skin helps us grasp stably. Focusing on these points, we have developed a dual-modal vision-based tactile sensor that can measure the texture of object and a distribution of force vectors. The tactile sensor consists of a transparent elastomer, a camera, a piece of transparent acrylic board, LEDs and supporting structures. A reflective membrane and markers array are on the surface of the elastomer. An applied force on the elastic body results in movements of the markers, which are acquired by the CCD camera. In addition, the shape and texture of the object's contact surface can be reflected by the membrane deformations. The distribution of force vectors is determined by the BP neural network. The local binary pattern algorithm using captured images calculates the texture information. This paper reports experimental evaluation results concerning accuracy of determination of magnitude, direction of force, and texture recognition rate.
Bin Fang 0003, Fuchun Sun 0001, Chao Yang 0026, Hongxiang Xue, Wendan Chen, Chun Zhang 0001, Di Guo 0002, Huaping Liu 0001
ICRA6
2018 A Bio-inspired Collision Detector for Small Quadcopter
abstract
The sense and avoid capability enables insects to fly versatilely and robustly in dynamic and complex environment. Their biological principles are so practical and efficient that inspired we human imitating them in our flying machines. In this paper, we studied a novel bio-inspired collision detector and its application on a quadcopter. The detector is inspired from Lobula giant movement detector (LGMD) neurons in the locusts, and modeled into an STM32F407 Microcontroller Unit (MCU). Compared to other collision detecting methods applied on quadcopters, we focused on enhancing the collision accuracy in a bio-inspired way that can considerably increase the computing efficiency during an obstacle detecting task even in complex and dynamic environment. We designed the quadcopter's responding operation to imminent collisions and tested this bio-inspired system in an indoor arena. The observed results from the experiments demonstrated that the LGMD collision detector is feasible to work as a vision module for the quadcopter's collision avoidance task.
Jiannan Zhao, Cheng Hu 0006, Chun Zhang 0001, Zhihua Wang 0001, Shigang Yue
IJCNN3
2018 An Energy-Efficient High-Frequency Neuro-Stimulator with Parallel Pulse Generators, Staggered Output and Extended Average Current Range
abstract
This paper presents a high-frequency pulse stimulation (HFPS) output stage of neuro-stimulator with extended average output current range and high power efficiency. The output stage features two parallel buck-boost converters without any filter capacitor at the output node. Compared with HFPS output stage with only one converter, the proposed circuit doubles the maximum average current by staggering the output of two converters. Compared with traditional voltage mode stimulation (VMS), HFPS improves the power efficiency by discharging the inductor current through the tissue load directly, rather than through a filter capacitor with constant voltage. Besides, the control circuit for the proposed HFPS is much simpler than that of traditional VMS converter, which reduces the current consumption significantly and thus improves the efficiency. Test results confirm that with staggered output of two parallel converters, the maximum average current is doubled. The maximum energy efficiency of the proposed HFPS is 76.4%.
Guijie Zhu, Songping Mai, Xian Tang, Chun Zhang 0001, Zhihua Wang 0001, Hong Chen 0002
ISCAS4
2017 Multi-rate polar codes for solid state drives
abstract
As solid state drives (SSDs) are gradually replacing hard disk drives, error correction is critical to SSDs since NAND flash has deteriorating reliability over their life span. Existing error correction codes suffer from limited error correction capability or error floor issues. Polar codes are promising for SSDs since they are theoretically proven optimal codes and have good error floor behavior. In this paper, we first design multi-rate polar codes for SSDs, and then implement encoder and decoder that simultaneously support multiple rates in FPGA. Multi-rate polar codes provide a good tradeoff between reliability and efficiency. Finally, we use an FPGA emulation platform to evaluate the error performance of our polar codes, and examine their error floor behavior.
Chun Zhang 0001, Chenrong Xiong, Zhiyuan Yan 0001
ICASSP2
2016 On the performance of wireless source-location using TDOA measurements under poor geometry
abstract
Booming wireless connectivity has boosted the interest of source-location in various types of network. The existing performance evaluations of source-location algorithms infrequently focus on the geometry issue, which also has considerable impact on the accuracy of source-location. This paper presents analysis and illustrations of the geometry issue for the time-difference-of-arrival (TDOA) method. Several common source-location algorithms are reviewed, followed by the simulation tests of positioning performance under variants of geometry and different measurement errors. Results indicate that positioning algorithms also work under the examined cases of poor geometry, and an RMSE of below 0.3m can be achieved, which implies that the deviation of time-difference measurement error should be no greater than 33 ps.
Chun Zhang 0001, Zhihua Wang 0001
CCNC2
2016 LGMD and DSNs neural networks integration for collision predication
abstract
An ability to predict collisions is essential for current vehicles and autonomous robots. In this paper, an integrated collision predication system is proposed based on neural subsystems inspired from Lobula giant movement detector (LGMD) and directional selective neurons (DSNs) which focus on different part of the visual field separately. The two type of neurons found in the visual pathways of insects respond most strongly to moving objects with preferred motion patterns, i.e., the LGMD prefers looming stimuli and DSNs prefer specific lateral movements. We fuse the extracted information by each type of neurons to make final decision. By dividing the whole field of view into four regions for each subsystem to process, the proposed approaches can detect hazardous situations that had been difficult for single subsystem only. Our experiments show that the integrated system works in most of the hazardous scenarios.
Guopeng Zhang, Chun Zhang 0001, Shigang Yue
IJCNN2
2016 A high efficiency single-inductor dual-output buck converter with adaptive freewheel current and hybrid mode control
abstract
A single-inductor dual-output (SIDO) buck converter with adaptive freewheel current and hybrid mode control is presented in this paper. It operates in pseudo-continuous conduction mode (PCCM) with single freewheel phase to reduce switching loss and minimize cross regulation. Freewheel current adjustment is implemented by detecting and control freewheel time, which can extend the load range and improve efficiency. For unbalanced load, this paper further proposed a hybrid PWM/hysteretic mode control scheme, which enables the two sub-converters to work at their own mode based on the load, independently of each other. The converter was designed with a 0.18μm CMOS process. Simulation results show that the efficiency reaches 87% in PWM mode. At light load, it is still higher than 80%.
Zhaoyang Weng, Hanjun Jiang, Chun Zhang 0001, Zhihua Wang 0001, Qingliang Lin
ISCAS4
2016 High precision intelligent flexible grasping front-end with CMOS interface for robots application
Xu Zhang 0010, Ming Liu 0015, Yuanfang Chen, Peng Li 0012, Zhaolin Yao, Weihua Pei, Chun Zhang 0001, Hongda Chen 0002
Sci. China Inf. Sci.9
2015 A fast AGC method for multimode zero-IF/sliding-IF WPAN/BAN receivers
abstract
This paper presents a fast mixed-signal automatic gain control (AGC) method for zero-IF/sliding-IF receivers used in wireless personal and body-area networks (WPAN/BAN). The preamble defined in Bluetooth low energy (BLE)/802.15.4 /802.15.6 specifications are as low as 1 Byte. It is a tough challenge for the zero-IF/sliding-IF receivers to perform AGC training, frequency synchronization and symbol timing estimation. By detecting the RF input of the quadrature mixer, IF input of analog filter and ADC output, the proposed method can achieve a desired IF amplitude within 1 bit when interference signal is small and 3 bits when carrier-to-interference ratio (CIR) is -27dB.
Jingjing Dong, Hanjun Jiang, Zhaoyang Weng, Jingyi Zheng, Chun Zhang 0001, Zhihua Wang 0001
ISCAS5
2015 A high-voltage, energy-efficient, 4-electrode output stage for implantable neural stimulator
abstract
In order to raise energy efficiency of implantable neural stimulators, an output stage circuit with a self-adaptive stimulating voltage is proposed. The output stage consists of a digital to analog converter (DAC), a single-ended primary inductance converter (SEPIC), a switch array and a current detection module. The SEPIC enables stimulating voltage to vary according to the stimulus current on tissue load or DAC digital input. It improves the output energy efficiency by dynamically cutting down the unnecessary stimulating voltage headroom on the current driver, especially when the stimulus current is small. Simulation showed that the stimulating voltage could change from 1.1V to 18.7V with less than 25mV ripple and the stimulus current could change from 0.7mA to 12.1mA with 0.1mA resolution. The output stage could work smoothly under the voltage supply range between 3.3V and 4.1V. The output efficiency of the proposed work could keep higher than 60% when stimulus current was above 1.4mA and even higher than 70% when stimulus current was above 2.9mA. The output stage is being fabricated on a silicon chip using CSMC 1μm HV technology, occupying 3.46mm×1.94mm.
Jinghui Liu, Songping Mai, Chun Zhang 0001, Zhihua Wang 0001
ISCAS3
2014 A low-power DC offset calibration method independent of IF gain for zero-IF receiver
Jingjing Dong, Hanjun Jiang, Lingwei Zhang 0001, Jianjun Wei, Fule Li, Chun Zhang 0001, Zhihua Wang 0001
Sci. China Inf. Sci.6
2014 A flexible capacitive tactile sensor array with micro structure for robotic application
Xu Zhang 0010, Ming Liu 0015, Yuanfang Chen, Peng Li 0012, Weihua Pei, Chun Zhang 0001, Hongda Chen 0002
Sci. China Inf. Sci.7
2013 Rate distortion Multiple Instance Learning for image classification
abstract
In this paper, we model image classification as a Multiple Instance Learning (MIL) problem, by regarding each image as a bag composed of different regions/patches (i.e., instances). Motivated by the fact that a bag is determined by the most positive instance or the least negative instance, which is also called “witness”, we propose a new algorithm to take advantage of witnesses to improve the performance of MIL for image classification task. In the frame of Rate Distortion (RD), we regard MIL as a source coding of each instance to witness or itself, and then the distortion function is measured by the loss of the discriminant model trained on these encoded instances. Hence compared with the existing algorithms, our proposed RDMIL algorithm has the following advantages. First, the probabilistic approach in source coding well illustrates the generative process of witnesses and measures the different importance of instances. Second, the discriminant model trained in a large-margin approach sufficiently considers the diverse influences from the instances, and thus has a strong discriminative ability. The resulted objective function is decomposed into two convex sub-problems, and we especially design a sequential method to effectively optimize the RD sub-problem. Experimental results on two real-world datasets demonstrate the proposed algorithm is effective and promising.
Chun Zhang 0001, Zhihua Wang 0001
ICIP2
2013 Learning to Detect Frame Synchronization
Chun Zhang 0001, Zhihua Wang 0001
ICONIP (2)2
2013 Live demonstration: A wireless force measurement system for total knee arthroplasty
abstract
This is the demonstration description of a wireless force measurement system for total knee arhtroplasty (TKA), which will be adopted in the operation to help the surgeons to place the implants quickly and accurately and consequently enhance the success ratio of treatment.
Hong Chen 0002, Chun Zhang 0001, Zhihua Wang 0001
ISCAS2
2013 A high-performance low-power SoC for mobile one-time password applications
abstract
The presented SoC is an 8-bit MCU based processor with 32-bit encryption algorithm acceleration and special security mechanism dedicated to mobile one-time password (MOTP) applications. The encryption algorithm accelerator can perform several 32-bit key operations for hash function computation, with each in one clock cycle. The SoC also features protection mechanisms which can prevent attackers from stealing the algorithm program code or the secret key, or utilizing under-voltage attacks. The SoC chip, fabricated in 0.25 μm CMOS process, has more than 50% efficiency improvement on hash function computation if compared with those 8-bit general-purpose MCUs and consumes less than 7.2 μA current in average for typical MOTP applications.
Songping Mai, Chunhong Li, Yixin Zhao, Chun Zhang 0001, Zhihua Wang 0001
ISCAS4
2012 A 10Gbps CDR based on phase interpolator for source synchronous receiver in 65nm CMOS
abstract
In this paper, a 10Gbps PI-based CDR circuit is presented in 65nm CMOS technology. The circuit is composed of a phase selector, a phase interpolator, a sample unit, a synchronize unit, a phase detector, and CDR logic. Half-rate clock is adopted to lessen the problems caused by high speed clocks and reduce power. The simulated worst phase step of phase interpolator is 26.7% larger than the average phase error. The power consumption is 15mW for 1V supply.
Shijie Hu, Ke Huang 0003, Chun Zhang 0001, Xuqiang Zheng, Zhihua Wang 0001
ISCAS4
2012 A 9.6Gb/s 5+1-lane source synchronous transmitter in 65nm CMOS technology
abstract
This paper describes the design of a low-jitter source-synchronous link transmitter macro for data rates of 9.6 Gb/s. The transmitter macro consists of 5 data channels plus 1 forwarded-clock channel. A low jitter PLL with bandwidth linearization is employed to achieve 0.66ps rms jitter. The power supply induced jitter is minimized by employing a hybrid clock distribution network which is proposed for both jitter and power consideration. To minimize the influence of PVT variation, Successive Approximation Register (SAR) sub block is implemented to accurately set the on chip impedance and the signal amplitude. A CML driver with 4 tap feed forward equalizer is implemented to compensate the channel loss. The transmitter is implemented in 65nm CMOS technology, the active chip area is 3.12 mm2.
Ke Huang 0003, Xuqiang Zheng, Ni Xu, Chun Zhang 0001, Woogeun Rhee, Zhihua Wang 0001
ISCAS5
2012 A wireless force measurement system for Total Knee Arthroplasty
abstract
A wireless force measurement system is presented in this paper. The system, which is used during the operation of Total Knee Arthroplasty(TKA), is designed to assist surgeons to determine the proper position of the knee implants and consequently enhance the success ratio of the treatment. It consists of three parts: a device to measure and transmit force data, a receiver and a terminal to display the force data in real time. The transmitter communicates with the receiver by 2.4GHz Radio Frequency (RF) signal. The system consumes no more than 17mA current with 3V voltage supply typically. So it can work with a button battery cell. Experimental results show that the performance of the system meets the requirements.
Hanqing Luo, Ming Liu 0015, Hong Chen 0002, Chun Zhang 0001, Zhihua Wang 0001
ISCAS4
2012 Design of a low-cost low-power baseband-processor for UHF RFID tag with asynchronous design technique
abstract
A low-cost low-power baseband processor for passive UHF RFID Tag based on EPC C1G2 protocol is presented in this paper. In order to minimize the power consumption, a novel digital baseband architecture is proposed and a series of low-power design approaches are adopted, including asynchronous design, clock-gating, low operating frequency, reuse of registers, etc. The baseband processor supports eleven mandatory commands and one optional command (Access) and the C1G2 protocol is completely fulfilled. The whole Tag (including a 1K EEPROM, RF/Analog frontend and the low-power baseband processor) is fabricated in 0.18μm CMOS technology. Real measurements on the final chip indicate that the processor consumes less than 2.7μW at 1V supply voltage and occupies an area of 0.11 mm2.
Dingguo Wei, Chun Zhang 0001, Hong Chen 0002, Zhihua Wang 0001
ISCAS2
2012 A wide dynamic range and fast update rate integrated interface for capacitive sensors array
abstract
A low-power CMOS capacitance-to-digital converter for capacitive sensors array is presented. It consists of a 16-channel MUX, a front-end with wide dynamic range for capacitance-voltage conversion and a novel voltage-pulses convertor. This novel circuit charges the measured capacitor indirectly and converts it into a proportional voltage, and then produces a 16-bits pulses-formed single-line output with multi-steps quantization method. The circuit can operate under 1.2V to 3.6V voltage supply for single battery application. The maximum input is 350pF with 0.75ms/ch update rate and 90μA power consumption when the operating clock is 100KHz. The chip has been design and fabricated in 0.18-μm 1P6M CMOS process with active area of 640×630μm2.
Xu Zhang 0010, Ming Liu 0015, Hong Chen 0002, Chun Zhang 0001, Zhihua Wang 0001
ISCAS4
2012 A current-to-voltage integrator using area-efficient correlated double sampling technique
abstract
This paper provides an area-efficient, low power and high precision current-to-voltage integrator for photocurrent measuring by using correlated double sampling. A novel CDS scheme is devised by using an interstage capacitor instead of the large error store capacitor to reduce kT/C noise. An integrator prototype with the proposed technique is fabricated using 0.8-μm 2P2M CMOS process for demonstration. The measurement results show that the output noise is 95μV and 59μV with the integration capacitor 4pF and 32pF, respectively, both for a 33pF photodiode parasitic capacitor. For a 3V output swing at 5V supply, the resolution is about 15bit.
Xuqiang Zheng, Fule Li, Chun Zhang 0001
ISCAS4
2012 A Time-Frequency Aware Cochlear Implant: Algorithm and System
Songping Mai, Yixin Zhao, Chun Zhang 0001, Zhihua Wang 0001
ISNN (2)3
2011 A 0.13µm CMOS 1.5-to-2.15GHz low power transmitter front-end for SDR applications
abstract
A 0.13 um CMOS 1.5-to-2.15 GHz low power transmitter front-end with direct quadrature voltage modulation for SDR applications is presented. The SDR transmitter front-end consists of PGA, passive RC filter, direct quadrature voltage modulator, 25%-duty-cycle LO generator and programmable PA-Driver with reconfigurable operation frequency bands. An on-chip transformer with switched- capacitor array provides the wide-band reconfigurable frequency band covering and differential-to-single-ended conversion. The passive voltage mixer with 25%-duty-cycle LO implements the frequency up-conversion with low noise and low conversion loss. The SDR transmitter front-end has been implemented in 0.13 um CMOS, the measured results show that this front-end could provide 22 dB gain with a step of 2dB and output higher than OdBm power across 1.5-to-2.15 GHz frequency range with a maximum power consumption of 42 mW. The die area is 1.6 mm* 1.2 mm.
Baoyong Chi, Chun Zhang 0001, Zhihua Wang 0001
ISCAS3
2009 A Novel Demodulator for Low Modulation Index RF Signal in Passive UHF RFID Tag
abstract
This paper presents a novel ASK demodulator for passive UHF RFID (radio frequency identification) tag, which can detect minimum 10% modulation index over a RF power range of 30 dB. Compared with conventional nonlinear voltage and linear current demodulations, the detection sensitivity over a wider range is increased by adopting an innovative divisional linear voltage demodulation with ultra low power consumption. The demodulator is implemented in UMC 0.18 mum mixed-mode CMOS technology. The power consumption is 3.95 muW and it can be reduced to 3.08 muW when the modulation index is large.
Zhongqi Liu, Chun Zhang 0001, Yongming Li 0004, Zhihua Wang 0001
ISCAS2
2009 A Robust Radio Frequency Identification System Enhanced with Spread Spectrum Technique
abstract
A robust passive UHF RFID backscatter system enhanced with spread spectrum technique is presented in this paper. Due to the weak signal energy of the backscatter chain, traditional RFID backscatter communication is easily affected by the noises, interferences, and interceptions from the environment. To solve the problem, spread spectrum technique is introduced into the backscatter link of RFID system. Simulated results show that this approach largely reduces the bit error rate and improves the system's reliability and security. For hardware realization, spectrum spreading operation is implemented in a RFID tag baseband processor, which is finally applied into a complete RFID tag and fabricated using 0.18 um 1P6M CMOS technology. Furthermore, de-spreading operation including the PN acquisition is integrated into a FPGA implementation of RFID reader. Thus, a complete robust RFID backscatter system enhanced with spread spectrum technique is constructed.
Qiuling Zhu, Chun Zhang 0001, Zhongqi Liu, Fule Li, Zhihua Wang 0001
ISCAS2
2008 Bandwidth extension for ultra-wideband CMOS low-noise amplifiers
abstract
Techniques are proposed to enhance the bandwidth of ultra-wideband (UWB) CMOS low-noise amplifiers (LNA). By using multiple-input-branch technique and resistive shunt-feedback technique, LNA could achieve ultra-wideband input impedance matching with small noise figure degradation. The gain bandwidth is enhanced by an L type inter-stage matching network between the input transistors and the cascode transistor which could enhance both the gain and the gain bandwidth. The gain bandwidth could be further enhanced by the resistors inserted between the bulks and the sources of the input transistors since the resistors could reduce the effects of the drain-bulk capacitance of the input transistors. The techniques proposed offer the advantages of ultra-wideband bandwidth, simple circuit topology, and small chip area. The proofs of techniques are demonstrated by designing the UWB LNA in UMC 0.18um CMOS.
Baoyong Chi, Chun Zhang 0001, Zhihua Wang 0001
ISCAS2
2008 An improved method of power control with CMOS class-E power amplifiers
abstract
In this paper, an improved method of power control is introduced to widen the range of output power with high efficiency. Two CMOS class-E power amplifiers (PA) with different output power are adopted in the design. In order to eliminate the phase mismatch of two paths, a continuously adjustable CMOS phase shifter with a predictable phase transfer function is added before each PA. Meanwhile, the drain modulation technique is presented to vary the voltage supply of the amplification unit. With the combination of drain modulation and parallel amplification, a wider range of output power is obtained than the previous one. The amplifiers work in 915MHz with the supply voltage of 1.8V and 2.5V. Simulation results show that the amplification unit could reach a maximum output power of 26dBm and maintain a power-added efficiency (PAE) higher than 40% over the 18dBm–26dBm output power range.
Tongqiang Gao, Chun Zhang 0001, Baoyong Chi, Zhihua Wang 0001
ISCAS2
2007 A Low Power, Fully Pipelined JPEG-LS Encoder for Lossless Image Compression
abstract
By analyzing the features unfit for parallel computation and low power implementation, a VLSI architecture of JPEG-LS encoder for lossless image compression is proposed in this paper. It functionally consists of four parts: Mode decision module, clock controller, three linear parallel pipelines, and a two-tier data packer. Computations are organized in a fully pipelined style in these modules, so that real time data processing can be achieved. The clock management scheme with four interlaced clock domains and a dedicated clock controller is applied to ensure the bottleneck calculation, reduce the clock frequency on non-critical paths, and shut off the working clocks of idle modules, which reduces 15.7% of overall power consumption. The proposed JPEG-LS encoder with the features of low power and high processing speed, has been applied in a wireless endoscopy system.
Xinkai Chen, Guolin Li, Li Zhang 0023, Chun Zhang 0001, Zhihua Wang 0001
ICME6
2007 Power Harvesting With PZT Ceramics
abstract
Piezoelectric materials have been proposed as embedded power source, which are capable of converting mechanical energy into electrical energy. However, power generated from a piezoelectric material usually comes with poor characteristics such as high voltage, low current and high impedance. In order to drive the embedded sensor circuit, piezoelectric power needs to be characterized and regulated. In this paper, we present an analysis on the power generation characteristics of the stiff lead zirconate titanate (PZT) ceramics and its equivalent circuit. It is then verified by simulation and experimental results. Since a single PZT element may not meet the needs of real-world embedded applications, we present experiments and analysis using four identical PZTs. Finally, we outline an application where PZT elements are used for power generation in a Total Knee Replacement (TKR) implant.
Hong Chen 0002, Chun Zhang 0001, Zhihua Wang 0001
ISCAS3
2006 A Cochlear System with Implant DSP
abstract
Cochlear implants have achieved great success in restoring hearing to profoundly deaf people. New generation implants pose more restrict requirements on processing precision and make traditional cochlear systems insufficient to deal with the large amount of data which should be transmitted via the wireless link. A new cochlear system with implant DSP is proposed to address the data rate problem. This system needs to transmit only voice-band signals with a data rate of 100 Kbps and the data rate bottleneck is removed for future development. By optimizing the speech processing algorithm and the inductive wireless link, the power consumption of this system increases less than 10% when compared with the traditional system.
Songping Mai, Chun Zhang 0001, Mian Dong, Zhihua Wang 0001
ICASSP (5)2
2006 A 3V 110µW 3.1 ppm/°C curvature-compensated CMOS bandgap reference
abstract
This paper presents a curvature-compensated bandgap reference (BGR) circuit implemented in 0.35 /spl mu/m CMOS technology. The circuit delivers an output voltage of 1.09 V and achieves a temperature coefficient of 3.1 ppm//spl deg/C over [-20/spl deg/C, 100/spl deg/C] after trimming, a power supply rejection ratio of -80 dB at 1 KHz and an output noise level of 1.43 /spl mu/V/sqrt(Hz) at 1 KHz. The BGR circuit consumes a current of 37 /spl mu/A and the power supply ranges from 1.5 V to 3.3 V.
Xiaokang Guan, Albert Wang 0001, Akira Ishikawa, Satoru Tamura, Kaoru Takasuka, Zhihua Wang 0001, Chun Zhang 0001
ISCAS7
2006 An open-source based DSP with enhanced multimedia-processing capacity for embedded applications
abstract
This paper proposes an open-source based 32-bit digital signal processor (DSP) suitable for embedded multimedia processing applications. The DSP is a MIPS-based scalar RISC with Harvard micro architecture, 5-stage integer pipeline and enhanced multimedia-processing capacity. It features a data path including non-aligned data access mechanism, two-dimensional DMA function and separated embedded instruction and data memories. This special architecture greatly promotes execution efficiency of boundary-crossed and block data access operations often found in multimedia processing. The processor achieves 80-MIPS performance with typical power dissipation of 88 mW. The prototype has been fabricated using UMC 0.18mum n-well 1P6M standard CMOS technology
Songping Mai, Wenli Lan, Chun Zhang 0001, Zhihua Wang 0001
ISCAS4
2005 A New Near-Lossless Image Compression Method in Digital Image Sensors with Bayer Color Filter Arrays
abstract
This paper presents a new method for near-lossless image compression for digital colorful image sensors with Bayer color filter arrays (CFA). In this method, the captured CFA raw data is first smoothed by low-pass filters followed by down-sampling and then compressed directly before full color interpolation which introduces redundancy. Assuring high image quality, the method can provide higher compression ratio and lower complexity than conventional image compression methods and other existing similar methods. The composite peak signal to noise ratio (CPSNR) of the decompressed image is larger than 51 dB.
Guolin Li, Xinkai Chen, Chun Zhang 0001, Zhihua Wang 0001
ICASSP (2)6
2005 A new near-lossless image compression algorithm suitable for hardware design in wireless endoscopy system
abstract
In order to decrease the communication bandwidth and save the transmitting power in the wireless endoscopy capsule, this paper presents a new near-lossless image compression algorithm suitable for hardware design based on the Bayer format image. This algorithm can provide low average compression rate (2.12 bits/pixel) with high image quality (larger than 53.11 dB) for endoscopic images. Especially, it has low complexity hardware overload and supports real-time compressing. In addition, the algorithm can provide lossless compression for the region of interest (ROI) and high quality compression for other regions. The ROI can be selected arbitrarily by varying ROI parameters.
Guolin Li, Chun Zhang 0001, Zhihua Wang 0001
ICIP (1)4
2004 An improved algorithm for rate distortion optimization in JPEG2000 and its integrated circuit implementation
abstract
Rate distortion optimization (RDO) plays an important role in a JPEG2000 encoder. An improved RDO algorithm is presented in this paper. The proposed algorithm is suitable for integrated circuit implementation and can reduce the computational cost. A hardware architecture which includes control unit, memory, divider, and data converter is also given to implement the algorithm. The circuit, based on the improved algorithm, has been integrated in a JPG2000 chip codec core.
Zhihua Wang 0001, Li Zhang 0023, Chun Zhang 0001
ICASSP (5)4
2004 A new approach for near-lossless and lossless image compression with Bayer color filter arrays
abstract
This paper presents a new approach for near-lossless and lossless image compression in digital colorful image sensors with Bayer color filter arrays (CFAs). In this approach, the captured CFA raw data is firstly transformed from quincunx shape to rectangular shape and then smoothed by a low-pass filter. Lastly, the filtered data are compressed directly before full color than conventional interpolation-first image compression methods and other existing similar compression-first methods. Furthermore, through altering the quality control factor, the image quality, PSNR, can be changed from 44.15 dB to infinite with the compression ratio from averagely 3.3 bits/pixel to 6.9 bits/pixel.
Guolin Li, Zhihua Wang 0001, Chun Zhang 0001, Li Zhang 0023
ICIG5
2003 A new variable step size LMS algorithm with application to active noise control
abstract
A new adaptive step size adjustment least mean square (LMS) algorithm is presented. The proposed algorithm modifies the existing LMS using the estimated output error as an important component for the modification of the step size. Experiment results demonstrate that application of the new algorithm leads to a significant gain in SNR (signal-to-noise ratio), thus visibly reduces the output error and greatly improves the convergence speed in the context of adaptive active noise cancellation.
Chun Zhang 0001, Zhihua Wang 0001
ICASSP (5)2