VLDB 2026 Research / reviewers in the wild / expert
Zhiwei Zheng
dblp:176/8348
· DBLP profile ↗
14ranked-venue papers
3as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 4 · 4 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | WiFi-based Human Pose Estimation via Intermediate Pose Synthesis and Part Grouping
Shigeng Zhang, Xuan Liu 0001, Zhiwei Zheng, Song Guo 0001 |
IWQoS | 4 |
| 2025 | On the Performance Analysis of Momentum Method: A Frequency Domain PerspectiveabstractMomentum-based optimizers are widely adopted for training neural networks. However, the optimal selection of momentum coefficients remains elusive. This uncertainty impedes a clear understanding of the role of momentum in stochastic gradient methods. In this paper, we present a frequency domain analysis framework that interprets the momentum method as a time-variant filter for gradients, where adjustments to momentum coefficients modify the filter characteristics. Our experiments support this perspective and provide a deeper understanding of the mechanism involved. Moreover, our analysis reveals the following significant findings: high-frequency gradient components are undesired in the late stages of training; preserving the original gradient in the early stages, and gradually amplifying low-frequency gradient components during training both enhance performance. Based on these insights, we propose Frequency Stochastic Gradient Descent with Momentum (FSGDM), a heuristic optimizer that dynamically adjusts the momentum filtering characteristic with an empirically effective dynamic magnitude response. Experimental results demonstrate the superiority of FSGDM over conventional momentum optimizers. Xianliang Li, Zhiwei Zheng, Lingkun Wen, Linlong Wu, Sheng Xu 0004 |
ICLR | 3 |
| 2025 | PreFall: Early Detection of Consecutive Fall Events with Commercial Wi-Fi DevicesabstractFalls are significant hazard to the health and safety of elderly individuals. Designing an alarm system capable of detecting fall events is crucial to mitigate these life-threatening risks. Compared to traditional methods, WiFi-based solutions have garnered extensive attention in recent years due to non-contact sensing nature and low-cost deployment advantages. While existing approaches have achieved high fall recognition accuracy, they still encounter inevitable delays in practice. These delays are mainly caused by limitations in signal segmentation, leading to recognition outcomes only after the activity has concluded. Moreover, existing methods overlook the continuity of human behaviors, reducing model robustness and increasing application constraints. To address these issues, this paper proposes PreFall, a WiFi-based fall detection method that achieves ongoing fall recognition before the activity ends. We adopt a deep model with attentional sentence embedding and propose a hybrid loss strategy. They are supposed to minimize the required sample observation for lower latency and guide the model to correctly detect fall events when multiple activities occur consecutively. Experimental results indicate that PreFall advances fall detection time by an average of 1.5 seconds and maintains a recognition accuracy of 88% when user behaves continuously. Zhiwei Zheng, Shigeng Zhang, Yalong Xiao |
MASS | 1 |
| 2025 | RF-Based 3D SLAM Rivaling Vision ApproachesabstractThis paper presents CartoRadar, a novel RF-based SLAM system that delivers high-fidelity 3D mapping with centimeter-level accuracy. CartoRadar builds on top of the advancements in learning-based RF imaging. However, learning-based systems often exhibit variation in prediction accuracy during inference. To address this challenge and enable robust RF sensing, CartoRadar introduces a novel, training-free uncertainty quantification method tailored to RF signals. Additionally, CartoRadar features an efficient SLAM algorithm that incorporates this uncertainty into the mapping process. We deploy CartoRadar on a mobile robot and conduct extensive evaluations across 14 floors in 5 buildings. Results show that CartoRadar achieves a trajectory error of 14.1 cm, outperforming camera-based baselines by 72.1%. For mapping, CartoRadar achieves an accuracy of 7.4 cm and a completion of 8.1 cm, improving over vision methods by 46.2% and 67.6%, respectively. Code, datasets, and demo videos are available on our website. Haowen Lai, Zhiwei Zheng, Mingmin Zhao |
MobiCom | 2 |
| 2025 | A Data-Driven Dynamics Simulation Model for Railway Vehicles Based on Lightweight 3DCNN With Physics-Informed ConstraintsabstractThe dynamics simulation of complex railway vehicles requires a dedicated vehicle model, such as multi-body dynamics model. However, the multi-body model is time-consuming in long-distance simulation due to its computational complexity. This issue can be alleviated by using a data-driven vehicle dynamics model due to its effective generalization and computational speed. Firstly, the construction of the physical model of the vehicle system is carried out to obtain the coupling relationship between the components. Secondly, the coupling relationship between the components is embedded into the loss function of the deep neural network as physics-informed constraints. Further, the network parameters satisfying certain physical laws are obtained by minimizing the loss function. Finally, the proposed lightweight 3D convolutional neural network is used to predict the vibration state of the vehicle system. The dynamic response resulting from both the data-driven simulation model and the multi-body simulation model are investigated and compared. The simulation results show that the data-driven dynamics simulation model can accurately predict the vibration state of the vehicle system. The data-driven simulation model has smaller size and faster operation speed, which can be applied to long-distance prediction research of vehicle systems. Zhiwei Zheng, Cai Yi, Jianhui Lin |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | VeloVox: A Low-Cost and Accurate 4D Object Detector with Single-Frame Point Cloud of Livox LiDARabstractCombining motion prediction in LiDAR-based 3D object detection is an effective method for improving overall accuracy, especially the downstream autonomous driving tasks. The recent development of low-cost LiDARs (e.g. Livox LiDAR) enables us to explore such 4D perception systems with a lower budget and higher performance. In this paper, we propose a 4D object detector, VeloVox, to establish accurate object detection and velocity estimation with a single-frame point cloud of Livox LiDAR. Based on the non-repetitive scanning pattern and point-level temporal nature, we propose a two-stage module to enhance the spatial-temporal point feature interaction along the time dimension. The aggregated feature also benefits a more accurate proposal refinement. To demonstrate the performance, comparison of VeloVox with several SOTA detector based baselines is evaluated on our in-house dataset and synthesized dataset built under Carla simulation. Code will be released at https://github.com/PJLab-ADG/VeloVox. Tao Ma 0002, Zhiwei Zheng, Hongbin Zhou, Xinyu Cai, Xuemeng Yang, Yikang Li 0002, Botian Shi, Hongsheng Li 0001 |
ICRA | 2 |
| 2024 | Unveiling the Dynamics of Video Out-of-Focus Blur: A Paradigm of Realistic Simulation and Novel EvaluationabstractVideo serves as the paramount medium through which humanity perceives the world via ocular faculties, spurring the development of myriad algorithms aimed at replicating the verisimilitude of the tangible realm. Among the numerous visual artifacts, out-of-focus blur remains a ubiquitous challenge, particularly prevalent in handheld camera cinematography. However, despite its common occurrence, the academic discourse on the underlying mechanisms of focal variance in video compositions is notably sparse, with a marked absence of publicly available datasets addressing real-world out-of-focus blur. In this study, we not only elucidate the phenomenon of video out-of-focus blur but also dissect the intrinsic dynamics governing focal fluctuations within authentic video sequences. We propose a novel framework for out-of-focus video simulation that faithfully replicates the nuanced characteristics of real-world scenarios. This framework involves segmenting a multitude of frames into discrete intervals through strategic keyframe selection, uniformly distributing focal planes across intervals, and meticulously rendering frames via a single-frame rendering methodology. Moreover, to evaluate the extent of out-of-focus blur in videos, we introduce an innovative metric that operates independently of ground truth data, thereby offering enhanced reliability compared to existing benchmarks. Zhiwei Zheng |
ISPA | 3 |
| 2024 | Education From Short Video: A Novel Educational Pattern Inspired by Entertainment VideosabstractBlended learning has gained popularity for its flexibility and technological integration. However, many current models simply replicate offline classrooms in virtual environments, lacking engagement for younger students. While gamified education fosters student motivation, it fails to enhance deep knowledge comprehension or assist in automating assessment processes. This paper introduces a novel educational approach, Education From Short Video (EFSV), which uniquely advocates for ‘Short-Video-Izing’ educational content and automating student assessments. We propose a model that transforms content into interactive short videos by blending AI-driven and manual techniques, moving away from traditional definition-based methods. Additionally, a Short Video Platform of Education (SVPE) is developed, where students act as both content creators and viewers. On this platform, student interactions like ‘comments’ and ‘likes’ automate final grading, significantly reducing the workload of teachers. Experimental results demonstrate that this model enhances students’ learning effectiveness and enjoyment. Jiale Yu, Zhiwei Zheng |
ISPA | 4 |
| 2024 | A Communication Protocol Recognition Method Based on AlexNet Deep Learning for Internet of ThingsabstractExisting communication protocol recognition methods of Internet of Things (IoT) suffer from high operational difficulty and low recognition accuracy. Therefore, this paper proposes a Communication Protocol Recognition method based on AlexNet Deep Learning (CPR-ADL) for IoT. The CPR-ADL first converts the collected communication data sequences into time-domain waveform figures, and then uses the AlexNet deep learning model to classify and recognize the communication protocols of these time-domain waveform figures. Experimental results show that CPR-ADL can effectively recognize the communication protocols of IOT communication data with high recognition accuracy, which is beneficial for developing a IoT comprehensive testing instrument with communication protocol recognition capabilities. Zhiwei Zheng, Kui Tie |
MSN | 3 |
| 2024 | Acoustic Volume Rendering for Neural Impulse Response FieldsabstractRealistic audio synthesis that captures accurate acoustic phenomena is essential for creating immersive experiences in virtual and augmented reality. Synthesizing the sound received at any position relies on the estimation of impulse response (IR), which characterizes how sound propagates in one scene along different paths before arriving at the listener position. In this paper, we present Acoustic Volume Rendering (AVR), a novel approach that adapts volume rendering techniques to model acoustic impulse responses. While volume rendering has been successful in modeling radiance fields for images and neural scene representations, IRs present unique challenges as time-series signals. To address these challenges, we introduce frequency-domain volume rendering and use spherical integration to fit the IR measurements. Our method constructs an impulse response field that inherently encodes wave propagation principles and achieves state of-the-art performance in synthesizing impulse responses for novel poses. Experiments show that AVR surpasses current leading methods by a substantial margin. Additionally, we develop an acoustic simulation platform, AcoustiX, which provides more accurate and realistic IR simulations than existing simulators. Code for AVR and AcoustiX are available at https://zitonglan.github.io/avr. Zitong Lan, Chenhao Zheng, Zhiwei Zheng, Mingmin Zhao |
NeurIPS | 3 |
| 2024 | A comparative study of large language model-based zero-shot inference and task-specific supervised classification of breast cancer pathology reportsabstractOBJECTIVE: Although supervised machine learning is popular for information extraction from clinical notes, creating large annotated datasets requires extensive domain expertise and is time-consuming. Meanwhile, large language models (LLMs) have demonstrated promising transfer learning capability. In this study, we explored whether recent LLMs could reduce the need for large-scale data annotations. MATERIALS AND METHODS: We curated a dataset of 769 breast cancer pathology reports, manually labeled with 12 categories, to compare zero-shot classification capability of the following LLMs: GPT-4, GPT-3.5, Starling, and ClinicalCamel, with task-specific supervised classification performance of 3 models: random forests, long short-term memory networks with attention (LSTM-Att), and the UCSF-BERT model. RESULTS: Across all 12 tasks, the GPT-4 model performed either significantly better than or as well as the best supervised model, LSTM-Att (average macro F1-score of 0.86 vs 0.75), with advantage on tasks with high label imbalance. Other LLMs demonstrated poor performance. Frequent GPT-4 error categories included incorrect inferences from multiple samples and from history, and complex task design, and several LSTM-Att errors were related to poor generalization to the test set. DISCUSSION: On tasks where large annotated datasets cannot be easily collected, LLMs can reduce the burden of data labeling. However, if the use of LLMs is prohibitive, the use of simpler models with large annotated datasets can provide comparable results. CONCLUSIONS: GPT-4 demonstrated the potential to speed up the execution of clinical NLP studies by reducing the need for large annotated datasets. This may increase the utilization of NLP-based variables and outcomes in clinical studies. Madhumita Sushil, Travis Zack, Divneet Mandair, Zhiwei Zheng, Ahmed Wali, Yan-Ning Yu, Yuwei Quan, Dmytro Lituiev, Atul J. Butte |
J. Am. Medical Informatics Assoc. | 4 |
| 2023 | FedCML: Federated Clustering Mutual Learning with non-IID Data
Zekai Chen 0010, Fuyi Wang, Shengxing Yu, Ximeng Liu, Zhiwei Zheng |
Euro-Par | 5 |
| 2023 | Fedward: Flexible Federated Backdoor Defense Framework with Non-IID DataabstractFederated learning (FL) enables multiple clients to collaboratively train deep learning models while considering sensitive local datasets’ privacy. However, adversaries can manipulate datasets and upload models by injecting triggers for federated backdoor attacks (FBA). Existing defense strategies against FBA consider specific and limited attacker models, and a sufficient amount of noise to be injected only mitigates rather than eliminates FBA. To address these deficiencies, we introduce a Flexible Federated Backdoor Defense Framework (Fedward) to ensure the elimination of adversarial backdoors. We decompose FBA into various attacks, and design amplified magnitude sparsification (AmGrad) and adaptive OPTICS clustering (AutoOPTICS) to address each attack. Meanwhile, Fedward uses the adaptive clipping method by regarding the number of samples in the benign group as constraints on the boundary. This ensures that Fedward can maintain the performance for the Non-IID scenario. We conduct experimental evaluations over three benchmark datasets and thoroughly compare them to state-of-the-art studies. The results demonstrate the promising defense performance from Fedward, moderately improved by 33% ∼ 75% in clustering defense methods, and 96.98%, 90.74%, and 89.8% for Non-IID to the utmost extent for the average FBA success rate over MNIST, FMNIST, and CIFAR10, respectively. Zekai Chen 0010, Fuyi Wang, Zhiwei Zheng, Ximeng Liu |
ICME | 3 |
| 2023 | Stacked graph bone region U-net with bone representation for hand pose estimation and semi-supervised training
Zhiwei Zheng, Zhongxu Hu, Hui Qin |
Image Vis. Comput. | 1 |