VLDB 2026 Research / reviewers in the wild / expert
Hongtao Xu
dblp:99/1394
· DBLP profile ↗
21ranked-venue papers
3as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 13 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-authorArtificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ATSGRU: Attention-Sparse Gated Recurrent Unit for Computationally Efficient Wideband Digital Predistortion of Quadrature Digital Power AmplifiersabstractDigital predistortion (DPD) is a widely used technique for enhancing signal quality in modern radio frequency (RF) power amplifiers (PAs). However, the strong performance of deep neural network (DNN)-based DPD models is often offset by their prohibitive computational complexity, which limits their practical deployment in wideband systems. This paper presents an attention-sparse gated recurrent unit (ATSGRU)—a novel neural architecture designed for computationally efficient wideband DPD in quadrature digital PAs (DPAs). The ATSGRU integrates the attention mechanism that evaluates the temporal relevance of input features and prunes redundant components, thereby simplifying the model structure and reducing computational load. The proposed method is validated on a custom 28-nm CMOS DPA chip. Experimental results demonstrate that the proposed ATSGRU achieves superior linearization with a favorable balance between accuracy and complexity compared with the state-of-the-art (SOTA) DPD model, reducing multiply-accumulate (MAC) operations by 54% while maintaining comparable performance. These results highlight its strong potential for efficient and scalable wideband DPD applications. Wending Zhao, Zijian Huang 0017, Yinyin Lin, Yun Yin, Hongtao Xu |
ACM Great Lakes Symposium on VLSI | 6 |
| 2026 | A 196-MHz Bandwidth Transimpedance Amplifier with High Out-of-Band Rejection Employing Transmission Zeros
Hexuan Wu, Wei Li 0038, Yunyou Pu, Hongtao Xu |
ISCAS | 5 |
| 2026 | A Fully-Differential Wideband Configurable Power Combiner and Splitter in 28-nm Bulk CMOS
Feiyang Xu, Hongtao Xu, Yun Yin |
ISCAS | 3 |
| 2025 | Linearization of Quadrature Digital Power Amplifiers by Neural Network of ULR_LSTM: Unsupervised Learning Residual LSTMabstractFor the first time, this paper presents an unsupervised learning residual long short-term memory (ULR_LSTM) neural network to develop a digital predistortion (DPD) method for the linearization of digital power amplifiers (DPAs). Our method eliminates the need for iterative learning control (ILC) to obtain the ideal input of the DPA required by state-of-the-arts (SOTAs), which leads to high computational complexity and extensive training time. We perform behavioral modeling of the DPA using the R_LSTM network. After determining the optimal behavioral model architecture, the corresponding DPD model is obtained through an inverse training process. A 15-bit transformer-based quadrature DPA chip incorporating Class-G and IQ-cell-sharing techniques was implemented in a 28nm CMOS process to validate our proposed method. Experimental results demonstrate outstanding linearization performance comparing to prior arts, achieving an error vector magnitude (EVM) of -40.4dB for the 802.11ax 40MHz 64QAM signal. Luyi Guo, Yicheng Li 0002, Wang Wang, Manni Li, Zijian Huang 0017, Yinyin Lin, Yun Yin, Hongtao Xu |
DATE | 10 |
| 2025 | Digital Predistortion for Quadrature Digital Power Amplifiers Using Deep Neural Network of AT_LSTM: Attention LSTM
Wending Zhao, Yicheng Li 0002, Wang Wang, Manni Li, Zijian Huang 0017, Yinyin Lin, Yun Yin, Hongtao Xu |
ACM Great Lakes Symposium on VLSI | 10 |
| 2025 | Efficient Long Context Fine-tuning with Chunk FlowabstractLong context fine-tuning of large language models(LLMs) involves training on datasets that are predominantly composed of short sequences and a small proportion of longer sequences. However, existing approaches overlook this long-tail distribution and employ training strategies designed specifically for long sequences. Moreover, these approaches also fail to address the challenges posed by variable sequence lengths during distributed training, such as load imbalance in data parallelism and severe pipeline bubbles in pipeline parallelism. These issues lead to suboptimal training performance and poor GPU resource utilization. To tackle these problems, we propose a chunk-centric training method named ChunkFlow. ChunkFlow reorganizes input sequences into uniformly sized chunks by consolidating short sequences and splitting longer ones. This approach achieves optimal computational efficiency and balance among training inputs. Additionally, ChunkFlow incorporates a state-aware chunk scheduling mechanism to ensure that the peak memory usage during training is primarily determined by the chunk size rather than the maximum sequence length in the dataset. Integrating this scheduling mechanism with existing pipeline scheduling algorithms further enhances the performance of distributed training. Experimental results demonstrate that, compared with Megatron-LM, ChunkFlow can be up to 4.53x faster in the long context fine-tuning of LLMs. Furthermore, we believe that ChunkFlow serves as an effective solution for a broader range of scenarios, such as long context continual pre-training, where datasets contain variable-length sequences. Xiulong Yuan, Hongtao Xu, Wenting Shen, Ang Wang, Xiafei Qiu, Jie Zhang 0135, Yuqiong Liu, Bowen Yu 0002, Junyang Lin, Mingzhen Li 0001, Weile Jia, Yong Li 0045, Wei Lin 0016 |
ICML | 2 |
| 2025 | Skrull: Towards Efficient Long Context Fine-tuning through Dynamic Data SchedulingabstractLong-context supervised fine-tuning (Long-SFT) plays a vital role in enhancing the performance of large language models (LLMs) on long-context tasks. To smoothly adapt LLMs to long-context scenarios, this process typically entails training on mixed datasets containing both long and short sequences. However, this heterogeneous sequence length distribution poses significant challenges for existing training systems, as they fail to simultaneously achieve high training efficiency for both long and short sequences, resulting in sub-optimal end-to-end system performance in Long-SFT.
In this paper, we present a novel perspective on data scheduling to address the challenges posed by the heterogeneous data distributions in Long-SFT. We propose Skrull, a dynamic data scheduler specifically designed for efficient long-SFT. Through dynamic data scheduling, Skrull balances the computation requirements of long and short sequences, improving overall training efficiency. Furthermore, we formulate the scheduling process as a joint optimization problem and thoroughly analyze the trade-offs involved. Based on those analysis, Skrull employs a lightweight scheduling algorithm to achieve near-zero cost online scheduling in Long-SFT. Finally, we implement Skrull upon DeepSpeed, a state-of-the-art distributed training system for LLMs. Experimental results demonstrate that Skrull outperforms DeepSpeed by 3.76x on average (up to 7.54x) in real-world long-SFT scenarios. Hongtao Xu, Wenting Shen, Yuanxin Wei, Ang Wang, Guo Runfan, Tianxing Wang 0006, Yong Li 0045, Mingzhen Li 0001, Weile Jia |
NeurIPS | 1 |
| 2025 | Highly efficient Doherty power amplifier with peak/backoff joint matching
Zoufeng Yuan, Yun Yin, Naiqian Zhang, Yi Pei, Hongtao Xu |
Sci. China Inf. Sci. | 6 |
| 2025 | Digital Predistortion for Wide Dynamic Power Range Quadrature Switched-Capacitor Power Amplifiers Using Self-Adaptive Residual LSTM Neural Network
Luyi Guo, Yicheng Li 0002, Yinyin Lin, Yun Yin, Hongtao Xu |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2024 | Nonlinear Analysis of Quadrature Switched-Capacitor Power Amplifier and Digital PredistortionabstractThis paper presents a modified vector combination (MVC) model to analyze the nonlinearity behavior in the quadrature switched-capacitor power amplifier (SCPA), which improves model accuracy while reducing the number of recorded points and iterations compared to the two-dimensional (2-D) LUT-based model. To analyze the nonlinearity caused by efficiency enhancement techniques, high-order Taylor series are applied to meet the model accuracy requirement. Moreover, a two-stage process with MVC and general memory polynomial (GMP) digital predistortion (DPD) is introduced to calibrate the static nonlinearity and memory effect, respectively. In the measurement, a 15-bit transformer-based quadrature SCPA chip with Class-G and IQ-cell-sharing techniques is implemented in 28nm CMOS and employed to verify the effectiveness of the MVC and two-stage DPD methods. For the 802.11ax 40MHz 64QAM signal at 2.4GHz, this chip achieves up to −40.9dB error vector magnitude (EVM) floor and significant EVM improvement even at deep power back-offs after the DPD. Fu Gao, Luyi Guo, Yicheng Li 0002, Yun Yin, Hongtao Xu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2024 | Enhanced-Linearity Wideband Full-Duplex Receiver With Shared Self-Interference CancellerabstractA wideband full-duplex (FD) receiver with enhanced-linearity technique and shared self-interference cancellation (SIC) is implemented in a 40-nm CMOS process. By combining Hilbert-transform-equalization (HTE)-based self-interference (SI) canceller and translational loop, an FD receiver with RF domain cancellation is presented with an extra auxiliary cancellation path by reusing the mixer in the translational loop. By introducing the auxiliary path, the influence of SI circuit to receiver front end is minimized. Meanwhile, a self-loaded linearization technique with acceptable noise degradation and extra power consumption is proposed to be employed in the FD receiver for both receiver and SI canceller. Due to the 2-D regulation, such a technique can achieve a relatively robust linearity improvement and bring flexibility to circuit design. The measurement results show that the proposed FD receiver operates across 0.8–3.5 GHz with a gain of 29.0–31.8 dB and a noise figure of 3.68–5.23 dB. The proposed linearization technique achieves 3.2–4.7-dB linearity improvement for receiver with only 0.45–0.64-dB NF degradation. In addition, the canceller with the proposed linearization method achieves RF domain delays ranging from 1.59 to 4.03 ns while demonstrating more than 6.33-dB linearity improvement. With the implementation of self-loaded technique and shared SIC, a greater than 23.4-dB RF domain SI suppression is measured across 40-MHz bandwidth (BW) with 64-QAM modulated signals in a circulator-based setup for the SIC scheme in this work with RX noise degradation of less than 1.38 dB. Wei Li 0038, Chuangguo Wang, Yunyou Pu, Shijiao Dong, Yun Wang 0008, Hongtao Xu |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2024 | Analysis and Calibration of Bit Weights in SAR and Pipelined SAR ADCs Based on Code DistributionabstractThis article comprehensively analyzes the effects of differential capacitor array mismatches on the decision levels of charge-redistribution-based successive-approximation-register (SAR) analog-to-digital converters (ADCs), reveals the intrinsic relation between the differential CDAC mismatches and the output code distribution of SAR, and proposes a bit-weight calibration (BWC) algorithm for SAR in terms of raw output codes, which does not require any modification to the original design. The proposed BWC is applicable to other architectures that incorporate a SAR as a substage, such as a pipelined SAR (P-SAR). Physical measurement results are presented for a pure SAR and simulation results are presented for a two-stage P-SAR, both of which verified the theoretical analyses and the proposed calibration method. Bingbing Ma, Wei Li 0038, Hongtao Xu |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2023 | A 12.1-16.5GHz Resistance Self-biased Inverse Class-F23 VCO Achieving 20-54kHz 1/f3 Corner FrequencyabstractThis paper presents a resistance self-biased inverse class-F23voltage-controlled oscillator (VCO). The VCO keeps ultra-low amplitude and phase mismatches by inducing a resistance between primary and secondary coils of the transformer. It also induces both 2ndand 3rdharmonic resonances to reduce 1/f3corner frequency. The VCO is designed in 40nm CMOS with frequency range of 12.1GHz to 16.5GHz. The post simulation results show that it achieves − 118dBc/Hz and −109dBc/Hz at 1MHz offset frequency at 12.1GHz and 16.5GHz respectively, whose 1/f3corner frequency is 54kHz and 20kHz. The power consumption of the VCO is 14.3mW at 1.1V supply. The core area is 0.03mm2. Jiayu Jing, Wei Li 0038, Ren Yuan, Hongtao Xu |
ISCAS | 4 |
| 2023 | A Tri-Mode Reconfigurable Receiver for GNSS/NB-IoT/BLE With 68-dB HR3 and 60-dB IMRR in 28-nm CMOSabstractIn this article, a 0.7–2.5-GHz tri-mode reconfigurable receiver (RX) for global navigation satellite system (GNSS)/narrow band-Internet of Things (NB-IoT)/Bluetooth low energy (BLE) applications is realized targeting one chip capability of integrating both wide area network (WAN) and local area network (LAN) protocols as well as positioning characteristics in 28-nm CMOS process with a compact core area of 1.7 mm2. By applying the proposed variable gain wideband active balun-low noise amplifier (LNA), the RX achieves minimum noise figure (NF) of 2.8 dB with the power consumption of 43 mW. The harmonic rejection (HR) method suppresses the signal interference caused by ON-chip coupling. 68-dB 3rd HR ratio (HR3) is achieved with only one-stage HR circuit. The mismatch between$I$channel and$Q$channel may limit the image rejection of the complex bandpass filter (CBPF). Detailed mathematical analysis of the$I/Q$calibration circuit is provided. Owing to the virtual short characteristic of the amplifier, the$I/Q$calibration scheme achieves a high-precision phase calibration. About 60-dB image rejection ratio (IMRR) is obtained after calibration. Yunyou Pu, Wei Li 0038, Chuangguo Wang, Qiaoan Li, Hongtao Xu |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2022 | A Two-Stage Digital Predistortion Method for Quadrature Digital Power AmplifiersabstractIn this paper, a novel two-stage digital predistortion method for quadrature digital power amplifiers (DPAs) is proposed. Digital transmitter architectures have advantages of compact die size owing to advanced CMOS process techniques, reconfigurability, and the usage of highly efficient switching-mode power amplifiers. However, the fluctuation of conductance resistance and interaction between in-phase and quadrature paths in quadrature DPAs complicates the parameters extraction process of the behavior model of power amplifiers, which necessitates the 2-D lookup table (2D-LUT) for compensation. A 2-D LUT is derived from the optimal predistortion input signal based on iterative learning control (ILC), which is aimed to calibrate the static nonlinearities. For wide bandwidth scenarios, the memory effect has to be taken into consideration which is neglected in the conventional 2D-LUT method. A general memory polynomial model is added to the whole system to fix the dynamic effect subsequently. Both the error vector magnitude (EVM) and normalized mean square error (NMSE) are used to evaluate the proposed two-stage predistortion method, which demonstrates EVMs of -32.46dB and -33.05dB for 802.11ax Wi-Fi 20MHz MCS9 and 40MHz MCS9 signals, respectively. Fu Gao, Yun Yin, Hongtao Xu |
ISCAS | 5 |
| 2021 | A Two-Way Power-Combining 60GHz CMOS Power Amplifier with 22.0% PAE and 19.4dBm Psat in 65nm Bulk CMOSabstractIn this paper, the analysis and design of a 57-64 CMOS Power Amplifier is discussed. The power-combining technique and the capacitor neutralization technique are applied to boost the performance of the purposed PA. The PA is designed in 65nm bulk CMOS process to achieve a saturated output power of 19.4dBm and a peak power-added efficiency of 22%. The power amplifier consumes 300mW from a 1.2V power supply at the output-refereed 1dB compression point of 16.1dBm, and the corresponding power-added efficiency is 13.0%. The passive devices, such as the transformers, the power-combiner and the pads, are designed by Electromagnetic Field Simulation on ADS momentum. Junjie Gu, Guixiang Jin, Hongtao Xu, Hao Min, Na Yan 0004 |
ISCAS | 3 |
| 2020 | Transformer-Combining Digital PA with Efficiency Peaking at 0, -6, and -12 dB Backoff in 32nm CMOSabstractA digital switched-capacitor transformer-combining power amplifier (SCPA) that uses load modulation to achieve efficiency peaking at 0, -6, and -12 dB backoff levels is beneficial for signals with high peak-to-average power ratios (PAPR). The PA uses a specific switched-capacitor cell design and turn-on sequence, that, contrary to the prior art, ensures correct-by-construction load modulation at backoff, even in the presence of practical transformer non-idealities. The PA has been implemented in 32nm CMOS and tested as part of a complete WiFi polar transmitter. Parmoon Seddighrad, Yorgos Palaskas, Hongtao Xu, Paolo Madoglio, Kailash Chandrashekar, David J. Allstot |
ISCAS | 3 |
| 2009 | Adaptive Model for Integrating Different Types of Associated Texts for Automated Annotation of Web Images
Hongtao Xu, Tat-Seng Chua |
MMM | 1 |
| 2008 | Web Image Annotation Based on Automatically Obtained Noisy Training Set
Hongtao Xu |
APWeb | 3 |
| 2008 | PictureBook: A Text-and-Image Summary System for Web Search ResultabstractSearch engine technology plays an important role in Web information retrieval. However, with Internet information explosion, traditional searching techniques cannot provide satisfactory result due to problems such as huge number of result Web pages, unintuitive ranking, etc. Therefore, the reorganization and post-processing of Web search results have been extensively studied to help user effectively obtain useful information. Previous studies mainly focused on Web page clustering, document summary, visualization of search results, etc, which are applied separately to either text or image search. In this paper, we propose a demo to illustrate a new Web search result summary system - PictureBook, which combines text and image retrieval using techniques of multiple document summarization and image semantics analysis. Particularly, audience can interactively investigate the effect of the combined text and image summary in Web information searching and knowledge acquisition. We also introduce our new image semantic analysis method based on generalized discriminant analysis (GDA). Hongtao Xu, Guoyu Hao, Wei Wang 0009, Qi Zhang 0025, Baile Shi |
ICDE | 2 |
| 2008 | WISA: a novel web image semantic analysis systemabstractWe present a novel Web Image Semantic Analysis (WISA) system, which explores the problem of adaptively modeling the distributions of the semantic labels of the web image on its surrounding text. To deal with this problem, we employ a new piecewise penalty weighted regression model to learn the weights of the contributions of the different parts of the surrounding text to the semantic labels of images. Experimental results on a real web image data set show that it can improve the performance of web image semantic annotation significantly. Hongtao Xu |
SIGIR | 1 |