Haoyu Bai

dblp:23/10091 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 7 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 3 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 A 49.7-fs, 10.29-to-12.75 GHz Fractional-N ADPLL with a New Quad-Core Class-F3 Oscillator
abstract
This paper presents a fractional-N all digital phase locked loop (ADPLL) based on a new quad-core class-F3oscillator, which reduces the out-of-band phase noise of ADPLL. The LC tank of the proposed oscillator incorporates both parallel and series resonant cavities compactly. It increases an additional impedance peak around the third-harmonic of the fundamental oscillation voltage, thus enforcing a pseudo-square voltage waveform around the LC tank. Consequently, the effective impulse sensitivity function (ISF) decreases, thereby reducing the phase noise. Furthermore, the proposed topology demonstrates low phase noise penalty during multi-core resonance. The ADPLL is implemented in a 40nm CMOS process. The simulated phase noise of the quad-core class-F3oscillator at 1MHz offset varies from -124.5 to -126.1dBc/Hz depending on the output frequency, which ranges from 10.29 to 12.75GHz. The oscillator consumes 24.36mW from the supply and exhibits a figure-of-merit (FoM) of 192.5dBc/Hz. Ultimately, the ADPLL utilizing the proposed oscillator exhibits an excellent RMS jitter of 49.7fs and achieves a FoM of -250.5dB.
Sihao Zhang, Ningyuan Zhang, Chuancheng Wu, Ling Hao, Haoyu Bai, Huailin Liao
ISCAS5
2025 A 4T4R Compact Wide-Band Low Power UWB Beamfoming Transceiver for Efficiency and Sensitivity Improvement
abstract
This paper describes a 4T4R ultra-wideband transceiver (UWB TRX) with compact baseband shifter and digital phase interpolar (DPI) for beamforming to alleviate the contradiction between power consumption and performance. A complementary stacked multi-supply switched-capacitor power amplifier (SCPA) is introduced for efficiency. It allows PA work in both high output and low output mode, which makes the proposed transceiver could work as traditional short-ranging communication and ranging UWB TRX or as a 4 elements beamforming TRX for better system efficiency and sensitivity. A bounding wire based in-chip inductor less matching network is adopted for wide band and farther reducing the chip area. Implemented in 22nm CMOS technology, the proposed TRX works from 3GHz to 10GHz, achieve 3.7–4.2 dB NF and 57.1%@11.8dBm / 48%@-0.23dBm drain efficiency, occupies 8.72mW for RX and 31.9mW / 9.6mW for TX per channel.
Jiazheng Zhou, Haoyu Bai, Ling Hao, Huailin Liao
ISCAS2
2024 A 0.12mm2 K/Ka Band RX Front-end in 40-nm CMOS with Inductor-Less LO Generators
abstract
This paper presents a broadband 18-33GHz receiver (RX) front-end with a core area of only 0.12mm2. In the local oscillator (LO) generators, a compact edge-combining (EC)-based frequency multiplication scheme is proposed to generate wideband and precise in-phase and quadrature (I/Q) LO signals at millimeter-wave frequencies. In the output stage of the LO chain, the switched-capacitor (SC) XOR modules triple the output clock frequency through edge-combining, which provides a 3× frequency reduction for the global clock distribution. The clock tripling is independently performed in I/Q differential LO signals, and the orthogonality of LO signals is not influenced by frequency multiplication. Implemented in TSMC 40nm GP CMOS, the proposed receiver core occupies only 0.46mm × 0.26mm. The RX front-end achieves a conversion gain of 42dB, a Noise Figure of 5.6dB, a baseband bandwidth of 100MHz, and an IQ phase variation < 1.3° (1σ interval, without calibration), while operating in the K/Ka band at 18-33GHz.
Haoyu Bai, Sihao Zhang, Huailin Liao
ISCAS1
2024 A 136μW Over 800m Range Backscatter-Like UHF Band Transceiver
abstract
Backscatter transceivers are often used in ultra-low-power Internet of Things (IoT) applications. However traditional backscatter transceiver (TRX) cannot actively control the transmitted power, which greatly limits the communication distance. This paper introduces a novel backscatter TRX operating in the ultra-high frequency (UHF) band, with a low power consumption of only 136μW and an impressive communication range surpassing 800 meters. Unlike the backscatter, the received continuous wave (CW) is not reflected via impedance modulating at the antenna but is directed into the TRX to be modulated and amplified. This technique significantly extends the communication range at a low power consumption. A coupler serves to prevent the direct entry of transmitted signals into the receiving path, ensuring the received CW remains unaffected to be modulated. Implemented in TSMC 40nm CMOS, the proposed TRX core occupies 0.7 mm2and works under a 0.7V supply voltage with the -54.6dBm input power and -20dBm output power.
Ling Hao, Keer Gao, Haoyu Bai, Chuancheng Wu, Sihao Zhang, Jiazheng Zhou, Huailin Liao
ISCAS3
2024 A Hybrid Deep Learning-Based Framework for Chip Packaging Fault Diagnostics in X-Ray Images
abstract
In the testing of chips, defect diagnostics in X-ray images of packaging chips is mainly performed by humans, which is time-consuming and inefficient. To overcome the abovementioned problems, a novel intelligent defect diagnostics system based on hybrid deep learning for chip X-ray images was proposed. The system consists of four successive stages: image segmentation and normalization, image reconstruction and defect detection, contour matching, and qualification diagnosis. The first stage is used to localize the external contours of the target chip and remove extraneous backgrounds through the improved UNet. Then, considering the variety of defects and the complexity of labeling, an unsupervised learning model is designed to reconstruct defect-free images to detect defects, which requires only normal samples for training. Third, the multicomponent template matching based on structural prior is used to localize the internal contours of the chip. In the final stage, the qualification is diagnosed based on the previous results through the Floyd–Warshall algorithm. The effectiveness and robustness of the proposed methods are verified by experiments on real-world inspection lines. The experimental results demonstrate that the developed system can successfully perform fault diagnostics tasks, achieving a judgment accuracy of 92.5%.
Jie Wang 0142, Gaomin Li, Haoyu Bai, Guixin Yuan, Lijun Zhong
IEEE Trans. Ind. Informatics3
2024 DRPC: Distributed Reinforcement Learning Approach for Scalable Resource Provisioning in Container-Based Clusters
abstract
Microservices have transformed monolithic applications into lightweight, self-contained, and isolated application components, establishing themselves as a dominant paradigm for application development and deployment in public clouds such as Google and Alibaba. Autoscaling emerges as an efficient strategy for managing resources allocated to microservices’ replicas. However, the dynamic and intricate dependencies within microservice chains present challenges to the effective management of scaled microservices. Additionally, the centralized autoscaling approach can encounter scalability issues, especially in the management of large-scale microservice-based clusters. To address these challenges and enhance scalability, we propose an innovative distributed resource provisioning approach for microservices based on the Twin Delayed Deep Deterministic Policy Gradient algorithm. This approach enables effective autoscaling decisions and decentralizes responsibilities from a central node to distributed nodes. Comparative results with state-of-the-art approaches, obtained from a realistic testbed and traces, indicate that our approach reduces the average response time by 15% and the number of failed requests by 24%, validating improved scalability as the number of requests increases.
Haoyu Bai, Minxian Xu, Kejiang Ye, Rajkumar Buyya, Cheng-Zhong Xu 0001
IEEE Trans. Serv. Comput.1
2023 A Single-Ended Digital Transmitter Based on I/Q-Sharing Switched-Capacitor Power Amplifier and Third-Order Harmonic-Rejection
abstract
This paper presents a design of 250MHz single-ended digital transmitter (DTX) with switched-capacitor power amplifier (SCPA) for communication between mobile phones and satellites. The implementation of SCPA enhances linearity and efficiency. For compact die area and high efficiency, 50%-duty-cycled clocks are used in the DTX eliminating second-order harmonic without differential architecture. Aiming at higher efficiency compared to conventional quadrature DTX, the technique of digital-domain in-phase (I) and quadrature phase (Q) cell sharing is applied in this work. Moreover, the combination of three signals with specific amplitude and phase eliminates third-order harmonic, improving the output spectrum significantly. The DTX is designed in 22nm standard CMOS process and reaches a peak output power of 8.87dBm at 250MHz and power-added efficiencies (PAE) of 57.4%. When modulating a 4 MHz 64-QAM signal, the error vector magnitude (EVM) is 0.61%, while the average output power and PAE are 1.93dBm and 32.9%, respectively.
Haoyu Bai, Jiazheng Zhou, Huailin Liao
ISCAS3
2016 Importance Sampling for Online Planning under Uncertainty
Yuanfu Luo, Haoyu Bai, David Hsu, Wee Sun Lee
WAFR2
2015 Intention-aware online POMDP planning for autonomous driving in a crowd
abstract
This paper presents an intention-aware online planning approach for autonomous driving amid many pedestrians. To drive near pedestrians safely, efficiently, and smoothly, autonomous vehicles must estimate unknown pedestrian intentions and hedge against the uncertainty in intention estimates in order to choose actions that are effective and robust. A key feature of our approach is to use the partially observable Markov decision process (POMDP) for systematic, robust decision making under uncertainty. Although there are concerns about the potentially high computational complexity of POMDP planning, experiments show that our POMDP-based planner runs in near real time, at 3 Hz, on a robot golf cart in a complex, dynamic environment. This indicates that POMDP planning is improving fast in computational efficiency and becoming increasingly practical as a tool for robot planning under uncertainty.
Haoyu Bai, Shaojun Cai, David Hsu, Wee Sun Lee
ICRA1
2013 Planning how to learn
abstract
When a robot uses an imperfect system model to plan its actions, a key challenge is the exploration-exploitation trade-off between two sometimes conflicting objectives: (i) learning and improving the model, and (ii) immediate progress towards the goal, according to the current model. To address model uncertainty systematically, we propose to use Bayesian reinforcement learning and cast it as a partially observable Markov decision process (POMDP). We present a simple algorithm for offline POMDP planning in the continuous state space. Offline planning produces a POMDP policy, which can be executed efficiently online as a finite-state controller. This approach seamlessly integrates planning and learning: it incorporates learning objectives in the computed plan, which then enables the robot to learn nearly optimally online and reach the goal. We evaluated the approach in simulations on two distinct tasks, acrobot swing-up and autonomous vehicle navigation amidst pedestrians, and obtained interesting preliminary results.
Haoyu Bai, David Hsu, Wee Sun Lee
ICRA1
2010 Monte Carlo Value Iteration for Continuous-State POMDPs
Haoyu Bai, David Hsu, Wee Sun Lee, Ngo Anh Vien
WAFR1