EDBT 2026 Demo / reviewers in the wild / expert
C. Patrick Yue
dblp:18/1066 · also Chik Patrick Yue
· DBLP profile ↗
17ranked-venue papers
2as first author
10since 2021 · last 2026
0000-0002-0211-2394ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FLICKER: A Fine-Grained Contribution-Aware Accelerator for Real-Time 3D Gaussian SplattingabstractRecently, 3D Gaussian Splatting (3DGS) has become a mainstream rendering technique for its photorealistic quality and low latency. However, the need to process massive noncontributing Gaussian points makes it struggle on resource-limited edge computing platforms and limits its use in next-gen AR/VR devices. A contribution-based prior skipping strategy is effective in alleviating this inefficiency, but the associated contribution-testing workload becomes prohibitive when it is further applied to the edge. In this paper, we present FLICKER, a contribution-aware 3DGS accelerator that leverages a hardware–software co-design framework, including adaptive leader pixels, pixel-rectangle grouping, hierarchical Gaussian testing, and mixed-precision architecture, to achieve near-pixel-level, contribution-driven rendering with minimal overhead. Experimental results show that our design achieves up to 1.5× speedup, 2.6× energy efficiency improvement, and 14% area reduction over a state-of-the-art accelerator. Meanwhile, it also achieves 19.8× speedup and 26.7× energy efficiency compared with a common edge GPU. Wenhui Ou, Zhuoyu Wu, Yipu Zhang 0002, Dongjun Wu, Frederick Ziyang Hong, C. Patrick Yue |
DATE | 6 |
| 2025 | PDR-KAN: Pipeline-Driven Reconfigurable Accelerator for Kolmogorov-Arnold Networks with Cross-Mode Sparsity SupportabstractThe commonly used Multi-layer Perceptrons (MLPs) in modern AI applications significantly limit the real-time performance due to their intensive memory access operation. Recently, Kolmogorov–Arnold Networks (KANs) have gained attention for offering a similar structure to MLPs but with a more efficient parameter utilization. However, the lack of customized support in conventional hardware restricts the performance gain from its algorithmic superiority. What’s more, the early-stage development of KAN raises questions about the necessity of incurring additional hardware costs for it. In this work, we present PDR-KAN, a reconfigurable accelerator that features two distinct operating modes, one for KANs and one for MLPs. Apart from its pipeline mode and sparsity encoder (SE) for efficient KAN processing, PDR-KAN can also improve the throughput of MLPs with cross-mode sparsity support. Experiments on a real-world dataset show that PDR-KAN provides a 3.95× acceleration and 13% reduction in accuracy loss by simply replacing MLPs with KANs. For a more accurate KAN model version with 3.33× parameter scaling, the latency overhead on PDR-KAN is only 1.16× compared to the baseline KAN model. Additionally, PDR-KAN achieves a 6.73× speed-up in KAN inferences and 20.89× in energy efficiency, against Quad-core ARM Cortex-A72 CPU. Wenhui Ou, Zhuoyu Wu, Alexandra Geciova, Zheng Wang 0027, C. Patrick Yue |
ISCAS | 6 |
| 2025 | A 25-GHz PLL Achieving 8-ns Phase-Shifting Time With Double-Path Modulation SchemeabstractThis article presents a reference phase-shifting architecture (PSA) based on a phase-locked loop (PLL) and a digital-to-time converter (DTC). The double-path phase modulation scheme (DPMS) is proposed to accelerate the settling time of the reference PSA. Off-chip calibration is added to mitigate the effects of nonlinearity in the DPMS process. Additionally, a DTC with improved retiming is proposed to reduce phase-shifting errors. The reference PSA with the DPMS is designed and fabricated in a commercial 22-nm CMOS technology. It occupies 0.048-mm2 active area and 12.8-mW dc power consumption. It achieves a 360° phase tuning range with a resolution of 1.26° at 24.75 GHz. The rms and peak phase errors are 1.38° and 2.6°, respectively. With the proposed DPMS, the settling time of reference PSA is significantly reduced from more than$1~\mu $s to less than 10 ns. Moreover, the PLL with DTC features a phase noise of −112.1 dBc/Hz at 1-MHz offset from 24.75 GHz and a 79.7-fs jitter integrated from 10 kHz to 30 MHz with 250-MHz reference clock. The figure of merits (FoMs) of jitter versus power for the proposed PLL with and without DTC are −250.9 and −251.4 dB, respectively. Weichen Tao, Yongheng Liu, Xu Yan 0006, C. Patrick Yue, Fujiang Lin |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2024 | A 25Gbps Single-Ended to Differential Low-Noise Transimpedance Amplifier in 45nm SOI for High-Speed Optical TransceiversabstractHigh-speed optical transceivers are advancing towards lower noise and power consumption, necessitating enhanced transimpedance amplifier (TIA) designs. This paper presents a 25Gbps single-ended to differential (S2D) low-noise TIA implemented in 45nm SOI technology. The design utilizes AC-coupling capacitors and a cross-coupled capacitor pair to achieve differential output. To optimize noise and bandwidth performance, a high-gain, low-bandwidth input stage is followed by a continuous-time linear equalizer (CTLE). The CTLE incorporates inductive peaking and negative capacitance techniques, achieving a bandwidth extension ratio (BWER) of 3.9 with less than 0.5dB peaking. Post-simulation results demonstrate the TIA achieves a transimpedance gain of 57.1 dB and a bandwidth of 23.7 GHz, considering a photodiode capacitance of 70fF. The input average noise current spectral density is 10.9$pA/\sqrt{Hz}$, with a power consumption of 38mW at a 1.3V supply voltage. Juncheng Deng, Luyao Yuan, Jinfeng Xie, Ahmed Wahba, Yizhe Hu, Fujiang Lin, C. Patrick Yue, Liheng Lou |
TENCON | 9 |
| 2024 | Real-Time 3D Visual Perception by Cross-Dimensional Refined LearningabstractWe introduce a novel learning method that can effectively perceive both the geometry structure and semantic labels of a 3D scene in real time. Existing real-time 3D scene reconstruction approaches often rely on volumetric schemes to regress a Truncated Signed Distance Function (TSDF) as the 3D representation. However, these volumetric approaches primarily focus on ensuring global coherence in the reconstructed scene, which often results in a lack of local geometric detail. To address this limitation, we propose a solution that leverages the latent geometric knowledge present in 2D image features by explicit depth prediction thereby creating anchored features, which are used to refine the learning of occupancy in the TSDF volume. Furthermore, we discover that this cross-dimensional feature refinement methodology can also be applied to the task of semantic segmentation by utilizing semantic priors. As a result, we propose an end-to-end cross-dimensional refinement neural network (CDRNet) that can extract both the 3D mesh and 3D semantic labeling of a scene in real time. Through experimental evaluation on multiple datasets, we demonstrate that our method achieves state-of-the-art 3D perception capability by boosting over 40% and 18% in 3D semantic segmentation and geometric reconstruction respectively over the prior art. These promising results indicate the significant potential of our approach for various industrial applications. Demo video and code can be found on the project page, https://hafred.github.io/cdrnet/. Frederick Ziyang Hong, C. Patrick Yue |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Adaptive Hybrid Optimization Learning-Based Accurate Motion Planning of Multi-Joint ArmabstractMotion planning is important to the automatic operation of the manipulator. It is difficult for traditional motion planning algorithms to achieve efficient online motion planning in a rapidly changing environment and high-dimensional planning space. The neural motion planning (NMP) algorithm based on reinforcement learning provides a new way to solve the above-mentioned task. Aiming to overcome the difficulty of training the neural network in high-accuracy planning tasks, this article proposes to combine the artificial potential field (APF) method and reinforcement learning. The neural motion planner can avoid obstacles in a wide range; meanwhile, the APF method is exploited to adjust the partial position. Considering that the action space of the manipulator is high-dimensional and continuous, the soft-actor-critic (SAC) algorithm is adopted to train the neural motion planner. By training and testing with different accuracy values in a simulation engine, it is verified that, in the high-accuracy planning tasks, the success rate of the proposed hybrid method is better than using the two algorithms alone. Finally, the feasibility of directly transferring the learned neural network to the real manipulator is verified by a dynamic obstacle-avoidance task. Chengchao Bai, Jiawei Zhang 0014, Jifeng Guo 0004, C. Patrick Yue |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2022 | LiDR: Visible-Light-Communication-Assisted Dead Reckoning for Accurate Indoor LocalizationabstractPedestrian dead reckoning (PDR) is an inertial navigation system that relies on smartphone sensors for estimating a pedestrian’s step movements. However, such systems suffer from poor accuracy due to the drift and inherent noise in sensor readings. In addition, step size variation among pedestrians and device heterogeneity pose further challenges for building a scalable PDR system that can provide uniform performance across various devices and a diverse range of users. Visible light positioning (VLP), which uses LED lights with visible light communication (VLC) capability to provide high-accuracy localization, can achieve precision of a few cm. However, VLP systems suffer from practical limitations due to occasional line-of-sight (LOS) blockage and the sparse density of lighting in large-scale indoor venues. In this work, we propose a light-assisted dead reckoning (LiDR) system, which aims to address the problems of both VLP and PDR. It uses LED lighting as high-accuracy location landmarks to provide regular calibration for the PDR and estimates the individual pedestrians’ step size for increased accuracy. In addition, a light-shape-based heading angle correction algorithm is proposed to reduce the heading angle error and further improve the accuracy. The system is implemented as an Android-based navigation application, with a digital map and cloud-based backend storage for location, device, and user-specific parameters. The real-time performance of the system is evaluated in a 450-$\text{m}^{2}$lab and on a 150-m walking track. The experimental results demonstrate that with a maximum light spacing of 15 m, an overall average accuracy of$\lt ~0.7$m can be achieved for the whole system. Babar Hussain, Yiru Wang 0001, Runzhou Chen, Hoi Chuen Cheng, C. Patrick Yue |
IEEE Internet Things J. | 5 |
| 2022 | Efficient-Grad: Efficient Training Deep Convolutional Neural Networks on Edge Devices with Gradient OptimizationsabstractWith the prospering of mobile devices, the distributed learning approach, enabling model training with decentralized data, has attracted great interest from researchers. However, the lack of training capability for edge devices significantly limits the energy efficiency of distributed learning in real life. This article describes Efficient-Grad, an algorithm-hardware co-design approach for training deep convolutional neural networks, which improves both throughput and energy saving during model training, with negligible validation accuracy loss. The key to Efficient-Grad is its exploitation of two observations. Firstly, the sparsity has potential for not only activation and weight, but gradients and the asymmetry residing in the gradients for the conventional back propagation (BP). Secondly, a dedicated hardware architecture for sparsity utilization and efficient data movement can be optimized to support the Efficient-Grad algorithm in a scalable manner. To the best of our knowledge, Efficient-Grad is the first approach that successfully adopts a feedback-alignment (FA)-based gradient optimization scheme for deep convolutional neural network training, which leads to its superiority in terms of energy efficiency. We present case studies to demonstrate that the Efficient-Grad design outperforms the prior arts by 3.72x in terms of energy efficiency. Frederick Ziyang Hong, C. Patrick Yue |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2021 | Orthogonally Interweaved Data Encryption Method for Screen to Camera CommunicationabstractDue to the broadcast nature of screen to camera communication (SCC) channels, SCC systems are under threat of eavesdropping. In this paper, we propose an orthogonally interweaved data encryption method to enhance the security of SCC systems. In our design, the secret key sequence is arrayed randomly and then used to generate an orthogonal key matrix by interweaving the rows in the identity matrix. By multiplying the key matrix, the user data bits are orthogonally interweaved. Then the key sequence is transmitted together with the user data after color-code modulation. The legitimate receiver's camera will subsequently demodulate the color code and extract the secret key sequence from the captured images. Finally, the orthogonal key matrix can be recovered based on this secret key sequence and its transpose can be directly used to decrypt the data bits. Under this method, even though the eavesdropper can capture the color coded image and demodulate the secret key sequence, it can not retrieve the transmitted data without the key matrix generation rule and the orthogonally interweaved data encryption scheme. Additionally, we analyze the theoretical information leakage of an SCC system encrypted using the proposed method. We perform simulations and illustrate that the proposed encryption method can achieve high security performance. Yiru Wang 0001, C. Patrick Yue |
PCS | 2 |
| 2021 | Sensing and Cancellation Circuits for Mitigating EMI-Related Common Mode Noise in High-Speed PAM-4 TransmitterabstractThe common mode (CM) current in differential circuits generates CM noise which radiates in the environment and causes electromagnetic interference (EMI) with electronic devices in proximity. This noise is generated due to imbalance in the charging and discharging paths of the driver circuit and asymmetric equalization. A systematic study of how asymmetric equalizer increase the CM noise is still lacking. Few active on-chip solutions are proposed for wireline transmitters up to 20-Gbps speed. However, these solution does not suppress CM noise by the equalizer and they could not support four-level pulse amplitude modulation (PAM-4) optical transmitters. This paper analyzes two potential sources of EMI-related CM noise, the common mode logic (CML) driver and feed forward equalization (FFE) circuit, and presents a novel on-chip active circuit technique for automatic CM noise cancellation in high-speed PAM-4 transmitter. This solution provides the benefits of small size and low cost by eliminating the need for discrete components to suppress EMI. The post-layout simulation demonstrates that the CM noise cancellation circuit (CMNC) efficiently mitigate the EMI by suppressing the CM noise up to 90% and consumes 5 mW, with core area of$17 \mu m \times 9 \mu m$. Rehan Azmat, Li Wang 0083, Khawaja Qasim Maqbool, Can Wang 0009, C. Patrick Yue |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2020 | A Low-Power PAM4 Receiver With an Adaptive Variable-Gain Rectifier-Based DecoderabstractThis article presents a low-power 1/4-rate four-level pulse amplitude modulation (PAM4) receiver with an adaptive variable-gain rectifier (AVGR)-based decoder in 28-nm CMOS technology. The PAM4 input signal is preconditioned by a continuous-time linear equalizer (CTLE) then sampled into four branches of decoders by 1/4-rate clocks. The proposed AVGR-based PAM4-to-nonreturn-to-zero (NRZ) decoder performs gain adaptation and amplitude rectification simultaneously for decoding the least significant bit (LSB). The linear sense amplifier in the AVGR is modified from a latch to achieve a high gain and low power. Compared with the full-rate receiver adopting a decoder consisting of three comparators, this design achieves a better power efficiency by employing a 1/4-rate topology and merging a variable-gain function into the decoder. Experimental results demonstrate that the receiver chip can receive and decode a 24-Gb/s 190-mVppPAM4 signal at a BER of 10-11and a bit efficiency of 1.38 pJ/bit. Quan Pan 0002, Li Wang 0083, Xiongshi Luo, C. Patrick Yue |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2017 | EMI common-mode (CM) noise suppression from self-calibration of high-speed SST driver using on-chip process monitoring circuitabstractThis paper presents a CM noise suppression technique from self-calibration of a 20-Gb/s source-series terminated (SST) driver using an on-chip process monitoring circuit designed and simulated in a 65-nm CMOS process. An unbalanced charging and discharging loop results in an asymmetric rise and fall time of the output signal and is an intrinsic source of the CM noise. For an SST driver, the CM noise effect is especially aggravated under process corner variations due to push-pull NMOS configuration. The on-chip sensor circuit monitors and detects all possible process corners over the entire temperature range. The control circuit operates to self-calibrate the output driver to compensate for the process corner variations and results in a 15% higher symmetric rise and fall time of the output signal for the worst case scenario. The simulation results indicate that the proposed technique reduces the peak CM noise by 7.2x (-86%) without power overhead, as it is a background monitoring and self-calibration scheme. Khawaja Qasim Maqbool, Duona Luo, Xingyun Luo, Huichun Yu, C. Patrick Yue |
ISCAS | 6 |
| 2014 | Towards indoor localization using Visible Light Communication for consumer electronic devicesabstractIndoor localization is the fundamental capability for indoor service robots and indoor applications on mobile devices. To realize that, the cost of sensors is of great concern. In order to decode the signal carried out by the LED beacons, we propose two reliable solutions using common sensors available on consumer electronic devices. Firstly, we introduce a dedicated analog sensor, which can be directly connected to the microphone input of a computer or a smart phone. It decodes both the signal pattern and signal strength of a beacon. Secondly, we utilize rolling-shutter cameras to decode the signal pattern, providing potential solutions to the localization of hand-held devices with cameras. In contrast to existing widely-applied indoor localization approaches, like vision-based and laser-based methods, our approach reveals its advantages as low-cost, globally consistent and it retains the potential applications using Visible Light Communication(VLC). We also study the characteristics of the proposed solutions under typical indoor conditions by experiments. Ming Liu 0001, Kejie Qiu, Fengyu Che, Babar Hussain, C. Patrick Yue |
IROS | 7 |
| 2007 | A two-tone test method for continuous-time adaptive equalizers
Dongwoo Hong, Shadi Saberi, Kwang-Ting Cheng, C. Patrick Yue |
DATE | 4 |
| 2003 | Design of a 10GHz clock distribution network using coupled standing-wave oscillatorsabstractIn this paper, a global clock network that incorporates standing waves and coupled oscillators to distribute a high-frequency clock signal with low skew and low jitter is described. The key design issues involved in generating standing waves on a chip are discussed, including minimizing wire loss within an available technology. A standing-wave oscillator, a distributed oscillator that sustains ideal standing waves on lossy wires, is introduced. A clock grid architecture comprised of coupled, standing-wave oscillators and differential, low-swing clock buffers is presented. The measured results for a prototyped standing-wave clock grid operating at 10GHz and fabricated in a 0.18μm 6M CMOS logic process are presented. A technique is proposed for on-chip skew measurements with sub-picosecond precision. Frank O'Mahony, C. Patrick Yue, Mark Horowitz, S. Simon Wong |
DAC | 2 |
| 1999 | Design Strategy of On-Chip Inductors for Highly Integrated RF SystemsabstractThis paper describes a physical model for spiral inductors on silicon which is suitable for circuit simulation and layout optimization. Key issues related to inductor modeling such as skin effect and silicon substrate loss are discussed. An effective ground shield is devised to reduce substrate loss and noise coupling. A practical design methodology based on the trade-off between the series resistance and oxide capacitance of an inductor is presented. This method is applied to optimize inductors in state-of-the-art processes with multilevel interconnects. The impact of interconnect scaling, copper metallization and low-K dielectric on the achievable inductor quality factor is studied. C. Patrick Yue, S. Simon Wong |
DAC | 1 |
| 1993 | Improved universal MOSFET electron mobility degradation models for circuit simulationabstractBased on the physical insights provided by the universal mobility curve, an improved comprehensive universal model for effective electron mobility in inversion layers of n-channel MOSFETs is developed for circuit simulation. This model expresses the effective electron mobility at room temperature as a function of effective vertical field. It exhibits a high degree of accuracy for a wide range of different device characteristics, such as channel doping levels, gate oxide thicknesses, and channel dimensions. In addition, it predicts very well the effective mobility under the effects of substrate biases for gate voltages well above threshold, which is an improvement over earlier models. Moreover, this model has been developed with an emphasis on the functional dependence of mobility on high effective field, and is thus particularly accurate in that range of effective field. This is a significant advantage of the model since today's submicrometer MOSFETs typically operate at high effective fields.> C. Patrick Yue, Victor Martin Agostinelli Jr., Greg Yeric, A. F. Tasch Jr. |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |