VLDB 2026 Research / reviewers in the wild / expert
Zhihui Gao
dblp:142/3312
· DBLP profile ↗
18ranked-venue papers
10as first author
18since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 8 · 5 first-author · 8 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OASIS: Outlier-Aware LUT-Based GEMM with Dual-Side Quantization for LLM Inference AccelerationabstractLarge language models (LLMs) have demonstrated impressive capabilities across a wide range of applications, but demand substantial memory and compute resources during inference. Existing quantization methods expose a trade-off between efficiency and accuracy: weight-only quantization (WOQ) incurs costly dequantization overheads, while integer weight-and-activation quantization (INT-WAQ) reduces precision and degrades model quality. Non-uniform weight-and-activation quantization (NU-WAQ) can better capture the non-uniform distributions of LLM weights and activations, yet remains incompatible with conventional low-precision compute units. This paper presents OASIS, a lookup table (LUT)-based architecture that enables efficient general matrix multiplication (GEMM) between non-uniformly quantized weights and activations without requiring dequantization. OASIS employs pre-computed Cartesian Product LUTs, achieving a 64x reduction in LUT size and enabling a 1024x higher computational parallelism over existing LUT-based GEMM methods. To preserve accuracy under aggressive activation quantization, OASIS introduces an outlier-aware quantization scheme with concurrent LUT-based GEMM and error compensation for outliers. Furthermore, we design Orizuru, an efficient top-k detection engine for real-time activation outlier identification. According to extensive evaluations, OASIS incurs an average accuracy drop of only 1.98% compared to the FP16 baseline, which is 5.18% lower than Atom. On the hardware side, OASIS achieves an average 3.00x speedup and a 1.44x energy efficiency improvement compared to the FIGLUT accelerator. Xueying Wu, Baijun Zhou, Zhihui Gao, Yuzhe Fu, Qilin Zheng, Yintao He, Hai Li 0001 |
ISCA | 3 |
| 2026 | Multivariate time series representation learning with multi-task graph neural network
Zhihui Gao, Baomin Xu, Jidong Yuan |
Eng. Appl. Artif. Intell. | 1 |
| 2026 | FastPFRec: A fast personalized federated recommendation with secure sharing
Zhenxing Yan, Jidong Yuan, Yongqi Sun, Zhihui Gao |
Expert Syst. Appl. | 5 |
| 2025 | MSPTS-Net: Multiscale Parallel Temporal-Spatial Network for Precipitation NowcastingabstractPrecipitation nowcasting plays an important role in people’s daily lives, yet current prediction methods still encounter some challenges. Most radar echo extrapolation methods rely on a single-input-single-output (SISO) structure, which can accumulate errors over time, resulting in a decline in prediction accuracy for longer forecasts. Furthermore, the radar echoes exhibit complex characteristics, making it difficult to capture underlying dynamics solely through the extraction of features at a single scale, especially for high-intensity rainfall events. To overcome these difficulties, this letter proposes a multiscale parallel temporal-spatial network (MSPTS-Net) based on a multi-input-multi-output (MIMO) architecture, which reduces the accumulation error for longer predictions. Specifically, we employ a multiscale temporal feature extraction module to capture both long-term and short-term temporal evolution patterns. In the spatial dimension, we adopt a global-to-local multiscale feature extraction module to model and represent the characteristics of radar echoes across different scales. For better representation of temporal-spatial characteristics, the self-attention mechanism is employed to integrate multiscale features across both temporal and spatial dimensions. During the training process, we use a hybrid loss function by combining mean squared error (mse) and Charbonnier loss to enhance prediction accuracy. Compared with existing methods, the proposed MSPTS-Net structure demonstrates significant advantages in performance. Zhihui Gao |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2024 | Real-time Wideband Software-defined Radio with Python Programmability based on RFSoCabstractNext-generation wireless networks necessitate large signal bandwidth to support the growing demands of high data rates, which poses significant challenges in the design of real-time radio platforms. We demonstrate SPEAR, a realtime wideband software-defined radio (SDR) utilizing the Xilinx RFSoC ZCU216 evaluation board. SPEAR leverages a customized "Streaming Direct Memory Access (DMA)" IP to address the latency issues associated with DMA control, thereby enabling high bandwidth data streaming in real-time. It also features a Python-based hardware configuration tool and signal processing framework incorporating an OFDM-based Physical layer. We showcase a real-time data link using the direct RF radio architecture between two RFSoC ZCU216 boards, achieving an error vector magnitude (EVM) of 3.2% for 256QAM across a bandwidth of 1.25 GHz. Wei Cheng 0006, Zhihui Gao, Tingjun Chen |
MobiCom | 2 |
| 2024 | SPEAR: Software-defined Python-Enhanced RFSoC for Wideband Radio ApplicationsabstractNext-generation wireless systems utilize large signal band-widths to meet the growing data rate demands of emerging applications and to provide enhanced resolution for wireless sensing and imaging. This poses significant challenges in the design of the underlying datapaths that carry and transfer signals across different domains, such as between memory and data converters in various software-defined radio (SDR) platforms. In this paper, we present the design and implementation of SPEAR, which is an SDR platform based on the Xilinx RFSoC ZCU216 evaluation board capable of supporting real-time streaming of signals with a bandwidth of up to 1.25 GHz employing the direct RF radio architecture. SPEAR features hardware-assisted direct memory access (DMA) control for real-time data streaming, and a Python-based hardware configuration tool and signal processing framework. Our experiments show that SPEAR can support a real-time bandwidth of up to 1.25 GHz for 256QAM modulation that satisfies the 3GPP error vector magnitude (EVM) requirement of 3.5%. Wei Cheng 0006, Zhihui Gao, Tingjun Chen |
MobiCom | 2 |
| 2024 | Mambas: Maneuvering Analog Multi-User Beamforming using an Array of Subarrays in mmWave NetworksabstractBeyond-5G and 6G wireless networks exploit the millimeter-wave (mmWave) frequency bands to achieve significantly improved data rates, and existing mmWave systems rely on analog single-user beamforming (SUBF) or hybrid multi-user beamforming (MUBF). In this work, we focus on improving the performance of multi-user communication in mmWave networks by exploring analog MUBF using an array of subarrays (ASA) with reduced system overhead and hardware complexity as it eliminates digital beamforming and the need for estimating the channel state information (CSI). We present Mambas, a novel system that maneuvers analog MUBF using an ASA to support simultaneous communication with multiple users located in close proximity, e.g., within the half-power beamwidth of the ASA. In essence, Mambas effectively decouples the user selection, subarray allocation, and beamforming optimization based on a comprehensive understanding of the multi-user support determined by the ASA. We evaluate Mambas using a 28 GHz software-defined radio testbed and show that, compared to existing methods, Mambas can effectively support users that are 2× more closely spaced while achieving an improved sum rate of up to 2×, using only two subarrays. Large-scale ray tracing-based simulations also show that Mambas can achieve a sum rate gain of 1.92--3.86× and is able to maintain consistent performance with significantly increased user density. Zhihui Gao, Zhenzhou Qi, Tingjun Chen |
MobiCom | 1 |
| 2024 | DeepMon: Wi-Fi Monitoring Using Sub-Nyquist Sampling Rate Receivers with Deep LearningabstractNext-generation Wi-Fi networks employ large signal bandwidth to meet the demands of high data rates, which poses challenges to Wi-Fi monitoring systems that typically rely on a full sampling rate receiver (RX) to capture signals at full bandwidth for demodulation and decoding. Interestingly, preambles of Wi-Fi packets contain unencrypted information that can be decoded to extract Wi-Fi Physical (PHY) layer information such as modulation and coding scheme (MCS), transmission time, and PHY service data unit (PSDU) length. In this paper, we propose DeepMon, which leverages low-cost RXs operating at sub-Nyquist sampling rates and deep learning (DL) to identify the Wi-Fi protocol and decode PHY layer packet properties from the Wi-Fi preamble. To evaluate DeepMon, we use PlutoSDR as the low sampling rate RX to collect a dataset of over 390K real-world 802.11a/n/ac Wi-Fi packets for the DL model training and testing. Our experiments show that for Wi-Fi packets with up to 160 MHz bandwidth, an RX running DeepMon at 3 MHz sampling rate (i.e., a downsampling ratio of >50×) can achieve an average bit decoding accuracy of 96.20% for the legacy signal field, corresponding to a mean absolute error of only 0.077 ms for predicting the packet transmission time. Zhihui Gao, Tingjun Chen |
MobiCom | 1 |
| 2024 | Mining Top-K constrained cross-level high-utility itemsets over data streams
Shujuan Liu, Zhihui Gao, Dongliang Mu |
Knowl. Inf. Syst. | 3 |
| 2024 | Online active learning method for multi-class imbalanced data stream
Dongliang Mu, Zhihui Gao, Shujuan Liu |
Knowl. Inf. Syst. | 4 |
| 2023 | Swirls: Sniffing Wi-Fi Using Radios with Low Sampling RatesabstractNext-generation Wi-Fi systems embrace large signal bandwidth to achieve significantly improved data rates, while requiring efficient methods for network monitoring and spectrum sharing applications. A radio receiver (RX) operating at low sampling rates can largely improve the energy- and cost-efficiency in such systems if it can extract useful network information such as the duration and structure of wireless packets. In this paper, we present the design of Swirls, a novel framework for sniffing Wi-Fi Physical layer information using RXs operating at sampling rates that are (much) smaller than the signal bandwidth. Swirls consists of three modules tailored for low sampling rate RXs: joint packet detection, optimized RX frequency selection, and packet property decoder. We implement Swirls using three software-defined radio platforms and extensively evaluate Swirls in real-world scenarios. The experiments show that for 20/40 MHz 802.11n packets, Swirls with 5 MHz sampling rate can achieve a mean absolute error (MAE) of transmission time and physical service data unit length decoding of 0.06 ms and 1.91 kB, respectively, at only 10 dB signal-to-noise ratio. With the same setting, Swirls simultaneously achieves a classification accuracy for the modulation and coding scheme, number of spatial streams, and bandwidth of 95.3%, 96.1%, and 95.6%, respectively. In an extreme case for 160 MHz 802.11ac/ax packets, Swirls with 2.5 MHz sampling rate (i.e., a downsampling ratio of 64) can still achieve an MAE of transmission time decoding of 0.47/0.67 ms. Zhihui Gao, Yiran Chen 0001, Tingjun Chen |
MobiHoc | 1 |
| 2023 | Open-access millimeter-wave software-defined radios in the PAWR COSMOS testbed: Design, deployment, and experimentation
Tingjun Chen, Prasanthi Maddala, Panagiotis Skrimponis, Jakub Kolodziejski, Abhishek Adhikari, Hang Hu 0008, Zhihui Gao, Arun Paidimarri, Alberto Valdes-Garcia, Myung J. Lee, Sundeep Rangan, Gil Zussman, Ivan Seskar |
Comput. Networks | 7 |
| 2022 | MOM: Microphone based 3D Orientation MeasurementabstractWhile a tremendous amount of effort has been devoted to localization, the orientation of a device, especially in 3D space, is seldom explored. Although many sensor-based methods utilizing gyro-scope, accelerometer, and magnetometer have been proposed to measure 3D orientation, these methods generally suffer from high cumulative errors and performance degradation when the device is moving. In this paper, we present MOM, the first microphone-based system that estimates the 3D orientation of a device. The key idea of MOM is to employ free sound sources in our surrounding environment as anchors. The prior knowledge of these sound sources, including the signal waveform and the locations of the sound sources, is not required to be known. In particular, we propose an angle-of-arrival (AoA) extraction algorithm that compares fine-grained time delays over microphones at a low computational cost. We implement our system on three platforms including a 6-microphone array Seeed Studio ReSpeaker, a commodity earphone Sennheiser AMBEO smart headset and a commodity smartphone Google Pixel 4. Extensive experiments show that MOM can achieve significantly higher accuracy compared with status quo approaches and is robust against cumulative errors. We apply MOM to two real-life applications, i.e., head tracking and 3D reconstruction, to demonstrate the applicability and generality of MOM in practice. Zhihui Gao, Ang Li 0005, Dong Li 0031, Jialin Liu 0004, Jie Xiong 0001, Yu Wang 0002, Bing Li 0017, Yiran Chen 0001 |
IPSN | 1 |
| 2022 | An overview of high utility itemsets mining methods based on intelligent optimization algorithms
Zhihui Gao, Shujuan Liu, Dongliang Mu |
Knowl. Inf. Syst. | 2 |
| 2021 | Hermes: Decentralized Dynamic Spectrum Access System for Massive Devices Deployment in 5G
Zhihui Gao, Ang Li 0005, Yu Wang 0002, Yiran Chen 0001 |
EWSN | 1 |
| 2021 | FedSwap: A Federated Learning based 5G Decentralized Dynamic Spectrum Access SystemabstractThe era of 5G extends the available spectrum from the microwave band to the millimeter-wave band. The thriving Internet of Things (IoT) also enriches the user equipment (UEs) we used in our daily life, such as smart glasses, smart watches, and drones. With such a larger spectrum and massive UEs, existing dynamic spectrum access (DSA) suffers both low spectrum utilization efficiency and unfair spectrum allocation. Thus, a more sophisticated dynamic spectrum access (DSA) system is required in the 5G context. In this paper, we propose a federated learning based system, FedSwap, the first decentralized DSA system that improves both efficiency and fairness simultaneously. In FedSwap, we deploy an improved multi-agent reinforcement learning (iMARL) algorithm on each UE, enabling UEs to share the spectrum coordinately with fewer collisions. Furthermore, we also propose a novel swapping mechanism for aggregating UEs' models periodically so that UEs can fairly share the spectrum resources. Meanwhile, the sensory data of UEs are not transmitted and hence privacy is protected. We evaluate FedSwap's performance in 5G simulations with various settings. Compared to the state-of-the-art decentralized DSA methods, FedSwap can significantly improve the efficiency and fairness of spectrum utilization. Zhihui Gao, Ang Li 0005, Bing Li 0017, Yu Wang 0002, Yiran Chen 0001 |
ICCAD | 1 |
| 2021 | Brief Industry Paper: Tenma: A Real-time LibOS Developed for Industry Embedded SystemsabstractWireless communication base stations have strong requirements in flexibility, performance and real-time responsiveness. Building an application-specific lightweight LibOS on a general OS is a good technique to make a balance among real-time responsiveness, performance and flexibility. Compared to traditional ext-kernel systems, the fast progress of virtualization technology brings a more efficient implementation of LibOS. In this paper, we propose a virtualization assisted real-time LibOS system named Tenma for wireless Baseband Unit (BBU). In Tenma, lightweight process virtualization, enhanced isolation between host OS and guest OS, interrupt direct-pass and application specific lightweight scheduling are used to guarantee the requirements of performance and real-time responsiveness while keeping system flexibility. The Tenma system is already in commercial deployment. Experiments in practice prove that it can meet the requirements of wireless base station applications. Zhihui Gao, Zichang Lin |
RTAS | 1 |
| 2021 | CRISLoc: Reconstructable CSI Fingerprinting for Indoor Smartphone LocalizationabstractChannel-state information (CSI)-based fingerprinting for WIFI indoor localization has attracted lots of attention very recently. The frequency diverse and temporally stable CSI better represents the location-dependent channel characteristics than the coarse received signal strength (RSS). However, the acquisition of CSI requires the cooperation of access points (APs) and involves only data frames, which imposes restrictions on real-world deployment. In this article, we present CRISLoc, the first CSI fingerprinting-based localization prototype system using ubiquitous smartphones. CRISLoc operates in a completely passive mode, overhearing the packets on-the-fly for his own CSI acquisition. The smartphone CSI is sanitized via calibrating the distortion enforced by WiFi amplifier circuits. CRISLoc tackles the challenge of altered APs with a joint clustering and outlier detection method to find them. A novel transfer learning approach is proposed to reconstruct the high-dimensional CSI fingerprint database on the basis of the outdated fingerprints and a few fresh measurements, and an enhanced KNN approach is proposed to pinpoint the location of a smartphone. Our study reveals important properties about the stability and sensitivity of smartphone CSI that has not been reported previously. Experimental results show that CRISLoc can achieve a mean error of around 0.29 m in a 6 m × 8 m research laboratory. The mean error increases by 5.4 and 8.6 cm upon the movement of one and two APs, which validates the robustness of CRISLoc against environmental changes. Zhihui Gao, Sulei Wang, Dan Li 0004, Yuedong Xu 0001 |
IEEE Internet Things J. | 1 |