EDBT 2026 Demo / reviewers in the wild / expert
Guangrong Zhao
dblp:250/4466
· DBLP profile ↗
13ranked-venue papers
5as first author
12since 2021 · last 2026
0000-0002-4703-9397ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 6 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dynolayout: robust layout estimation from event-stream for extended reality under dynamic scenarios
Xucheng Guo, Qiang Qu 0004, Guangrong Zhao, Yuanfeng Zhou, Yiran Shen 0001 |
CCF Trans. Pervasive Comput. Interact. | 5 |
| 2025 | RGBE-Gaze: A Large-Scale Event-Based Multimodal Dataset for High Frequency Remote Gaze TrackingabstractHigh-frequency gaze tracking demonstrates significant potential in various critical applications, such as foveatedrendering, gaze-based identity verification, and the diagnosis of mental disorders. However, existing eye-tracking systems based on CCD/CMOS cameras either provide tracking frequencies below 200 Hz or employ high-speedcameras, causing high power consumption and bulky devices. While there have been some high-speed eye-tracking datasets and methods based on event cameras, they are primarily tailored for near-eye camera scenarios. They lackthe advantages associated with remote camera scenarios, such as the absence of the need for direct contact, improved user comfort and head pose freedom. In this work, we present RGBE-Gaze, the first large-scale and multimodal dataset for remote gaze tracking in high-frequency through synchronizing RGB and event cameras. This dataset is collected from 66 participants with diverse genders and age groups. Our setup captures 3.6 million RGB images and 26.3 billion event samples. Additionally, the dataset includes 10.7 million gaze references from the Gazepoint GP3 HD eye tracker and 15,972 sparse points of gaze (PoG) ground truth obtained through manualstimuli clicks by participants. We present dataset characteristics such as head pose, gaze direction, and pupil size. Furthermore, we introduce a hybrid frame-event based gaze estimation method specifically designed for the collected dataset. Moreover, we perform extensive evaluations of different benchmarking methods under variousgaze-related factors. The evaluation results illustrate that introducing event stream as a new modality improves gazetracking frequency and demonstrates greater estimation robustness across diverse gaze-related factors. Guangrong Zhao, Yiran Shen 0001, Zhaoxin Shen, Yuanfeng Zhou, Hongkai Wen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Ui-Ear: On-Face Gesture Recognition Through On-Ear Vibration SensingabstractWith the convenient design and prolific functionalities, wireless earbuds are fast penetrating in our daily life and taking over the place of traditional wired earphones. The sensing capabilities of wireless earbuds have attracted great interests of researchers on exploring them as a new interface for human-computer interactions. However, due to its extremely compact size, the interaction on the body of the earbuds is limited and not convenient. In this paper, we proposeUi-Ear, a new on-face gesture recognition system to enrich interaction maneuvers for wireless earbuds.Ui-Earexploits the sensing capability of Inertial Measurement Units (IMUs) to extend the interaction to the skin of the face near ears. The accelerometer and gyroscope in IMUs perceive dynamic vibration signals induced by on-face touching and moving, which brings rich maneuverability. Since IMUs are provided on most of the budget and high-end wireless earbuds, we believe thatUi-Earhas great potential to be adopted pervasively. To demonstrate the feasibility of the system, we define seven different on-face gestures and design an end-to-end learning approach based on Convolutional Neural Networks (CNNs) for classifying different gestures. To further improve the generalization capability of the system, adversarial learning mechanism is incorporated in the offline training process to suppress the user-specific features while enhancing gesture-related features. We recruit 20 participants and collect a realworld datasets in a common office environment to evaluate the recognition accuracy. The extensive evaluations show that the average recognition accuracy ofUi-Earis over 95% and 82.3% in the user-dependent and user-independent tasks, respectively. Moreover, we also show that the pre-trained model (learned from user-independent task) can be fine-tuned with only few training samples of the target user to achieve relatively high recognition accuracy (up to 95%). At last, we implement the personalization and recognition components ofUi-Earon an off-the-shelf Android smartphone to evaluate its system overhead. The results demonstrateUi-Earcan achieve real-time response while only brings trivial energy consumption on smartphones. Guangrong Zhao, Yiran Shen 0001, Feng Li 0002, Lei Liu 0003, Li-Zhen Cui 0001, Hongkai Wen 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | p-Blend: Privacy- and Utility-Preserving Blendshape Perturbation Against Re-Identification Attacks in Virtual RealityabstractIn this paper, we propose p-Blend, an efficient and effective blendshape perturbation mechanism designed to defend against both intra- and cross-app re-identification attacks in virtual reality. p-Blend provides privacy protection when streaming blendshape data to third-party applications on VR devices. In its design, we consider both privacy and utility. p-Blend not only perturbs blendshape values to resist re-identification attacks but also preserves the smoothness of facial animations and the naturalness of facial expressions, ensuring the continued usability of the data. We validate the effectiveness of p-Blend through extensive empirical evaluations and user studies. Quantitative experiments on a large-scale dataset collected from 45 participants demonstrate that p-Blend significantly reduces re-identification accuracy across a range of machine learning models. While pure-random perturbation fails to prevent attacks that exploit statistical features, p-Blend effectively mitigates these risks in both raw and statistical blendshape data. Additionally, user study results show that facial animations generated from p-Blend-perturbed blendshapes maintain greater smoothness and naturalness compared to those using purely random perturbation. The codes and dataset are available at https://github.com/jingwei1016/p-Blend. Yan Hu 0003, Guangrong Zhao, Qing Yang 0009, Guangdong Bai, Yiran Shen 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2024 | Towards High-Speed Passive Visible Light Communication with Event Cameras and Digital Micro-MirrorsabstractPassive visible light communication (VLC) modulates light propagation or reflection to transmit data without directly modulating the light source. Thus, passive VLC provides an alternative to conventional VLC, enabling communication where the light source cannot be directly controlled. There have been ongoing efforts to explore new methods and devices for modulating light propagation or reflection. The state-of-the-art has broken the 100 kbps data rate barrier for passive VLC by using a digital micro-mirror device (DMD) as the light modulating platform, or transmitter, and a photo-diode as the receiver. We significantly extend this work by proposing a massive spatial data channel framework for DMDs, where individual channels can be decoded in parallel using an event camera at the receiver. For the event camera, we introduce event processing algorithms to detect numerous channels and decode bits from individual channels with high reliability. Our prototype, built with off-the-shelf event cameras and DMDs, can decode up to ~2,000 parallel channels, achieving a data transmission rate of 1.6 Mbps, markedly surpassing current benchmarks by 16x. Yiran Shen 0001, Kenuo Xu, Mahbub Hassan, Guangrong Zhao, Chenren Xu, Wen Hu 0001 |
SenSys | 5 |
| 2024 | KD-Eye: Lightweight Pupil Segmentation for Eye Tracking on VR Headsets via Knowledge Distillation
Yanlin Li 0014, Guangrong Zhao, Yiran Shen 0001 |
WASA (1) | 3 |
| 2024 | EV-Tach: A Handheld Rotational Speed Estimation System With Event CameraabstractRotational speed is one of the important metrics to be measured for calibrating electric motors in manufacturing, monitoring engines during car repairs, detecting faults in electrical appliance and more. However, existing measurement techniques either require prohibitive hardware (e.g., high-speed camera) or are inconvenient to use in real-world application scenarios. In this paper, we propose,EV-Tach, a novel handheld rotational speed estimation system that utilizes emerging imaging sensors known as event cameras or dynamic vision sensors (DVS). The pixels of DVS work independently and trigger an event as soon as a per-pixel intensity change is detected, without global synchronization like conventional RGB cameras. Thus, its unique design features high temporal resolution and generates sparse events, which benefits the high-speed rotation estimation. To achieve accurate and efficient rotational speed estimation, a series of signal processing algorithms are specifically designed for the event streams generated by event cameras on an embedded platform. First, a new cluster-centroids initialization module is proposed to initialize the centroids of the clusters to address the issue that common clustering approaches are easy to fall into a local optimal solution without proper initial centroids. Second, an outlier removal module is designed to suppress the background noise caused by subtle hand movements and host devices vibrations. Third, a coarse-to-fine alignment strategy is proposed with an event stream alignment method to obtain angle of rotation and achieve accurate estimation for rotational speed in a large range. With these bespoke components,EV-Tachis able to extract the rotational speed accurately from the event stream produced by an event camera recording rotary targets. According to our extensive evaluations under controlled and practical experiment settings, the Relative Mean Absolute Error (RMAE) ofEV-Tachis as low as$0.3\%_{0}$, which is comparable to the state-of-the-art laser tachometer under fixed measurement mode. Moreover,EV-Tachis robust to subtle movement of user's hand and dazzling light outdoor, therefore, can be used as a handheld device under challenging lighting condition, where the laser tachometer fails to produce reasonable results. To speed up the processing ofEV-Tachand reduce its resource consumption on embedded devices, event stream is significantly downsampled by merging neighboring events while preserving its formation in spatial-temporal domain. At last, we implementEV-Tachon Raspberry Pi and the evaluation results show that the downsampling process preserves the high measurement accuracy while saving the computation speed and energy consumption by approximately 8 times and 30 times in average. Guangrong Zhao, Yiran Shen 0001, Pengfei Hu 0001, Lei Liu 0003, Hongkai Wen 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2024 | Swift-Eye: Towards Anti-blink Pupil Tracking for Precise and Robust High-Frequency Near-Eye Movement Analysis with Event CamerasabstractEye tracking has shown great promise in many scientific fields and daily applications, ranging from the early detection of mental health disorders to foveated rendering in virtual reality (VR). These applications all call for a robust system for high-frequency near-eye movement sensing and analysis in high precision, which cannot be guaranteed by the existing eye tracking solutions with CCD/CMOS cameras. To bridge the gap, in this paper, we propose Swift-Eye, an offline precise and robust pupil estimation and tracking framework to support high-frequency near-eye movement analysis, especially when the pupil region is partially occluded. Swift-Eye is built upon the emerging event cameras to capture the high-speed movement of eyes in high temporal resolution. Then, a series of bespoke components are designed to generate high-quality near-eye movement video at a high frame rate over kilohertz and deal with the occlusion over the pupil caused by involuntary eye blinks. According to our extensive evaluations on EV-Eye, a large-scale public dataset for eye tracking using event cameras, Swift-Eye shows high robustness against significant occlusion. It can improve the IoU and F1-score of the pupil estimation by 20% and 12.5% respectively, compared with the second-best competing approach, when over 80% of the pupil region is occluded by the eyelid. Lastly, it provides continuous and smooth traces of pupils in extremely high temporal resolution and can support high-frequency eye movement analysis and a number of potential applications, such as mental health diagnosis, behaviour-brain association, etc. The implementation details and source codes can be found at https://github.com/ztysdu/Swift-Eye. Tongyu Zhang, Yiran Shen 0001, Guangrong Zhao, Lin Wang 0025, Xiaoming Chen 0006, Lu Bai 0004, Yuanfeng Zhou |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2023 | EV-Eye: Rethinking High-frequency Eye Tracking through the Lenses of Event CamerasabstractIn this paper, we present EV-Eye, a first-of-its-kind large scale multimodal eye tracking dataset aimed at inspiring research on high-frequency eye/gaze tracking. EV-Eye utilizes an emerging bio-inspired event camera to capture independent pixel-level intensity changes induced by eye movements, achieving sub-microsecond latency. Our dataset was curated over a two-week period and collected from 48 participants encompassing diverse genders and age groups. It comprises over 1.5 million near-eye grayscale images and 2.7 billion event samples generated by two DAVIS346 event cameras. Additionally, the dataset contains 675 thousands scene images and 2.7 million gaze references captured by Tobii Pro Glasses 3 eye tracker for cross-modality validation. Compared with existing event-based high-frequency eye tracking datasets, our dataset is significantly larger in size, and the gaze references involve more natural eye movement patterns, i.e., fixation, saccade and smooth pursuit. Alongside the event data, we also present a hybrid eye tracking method as benchmark, which leverages both the near-eye grayscale images and event data for robust and high-frequency eye tracking. We show that our method achieves higher accuracy for both pupil and gaze estimation tasks compared to the existing solution. Guangrong Zhao, Yurun Yang, Yiran Shen 0001, Hongkai Wen 0001, Guohao Lan |
NeurIPS | 1 |
| 2023 | Demo: EV-DMD: a high-speed VLC systemabstractVisible light communications (VLC) have gained significant attention as a potential solution for the radio spectrum crunch. To achieve high data rates, emerging transmitter devices like 2D digital micro-mirror devices (DMD) have been proposed, offering significantly faster state flipping rates compared to conventional liquid crystalline shutters. However, previous approaches utilizing DMD suffered from a lack of spatial diversity, as they used all micro-mirrors in the same state. This paper introduces EV-DMD, a novel approach that utilizes DMD as a 2D transmitter, working in tandem with an event-based vision (EV) camera. In this method, multiple bit streams are transmitted in parallel through different mirror blocks of the DMD, while an EV camera simultaneously decodes multiple light blocks, enabling a truly 2D high-speed VLC system. To the best of our knowledge, this is the first implementation of a 2D VLC system that achieves an order-of-magnitude improvement in bit rate compared to state-of-the-art solutions. Guangrong Zhao, Kenuo Xu, Yiran Shen 0001, Chenren Xu, Mahbub Hassan, Wen Hu 0001 |
SIGCOMM | 2 |
| 2022 | Event-Stream Representation for Human Gaits Identification Using Deep Neural NetworksabstractDynamic vision sensors (event cameras) have recently been introduced to solve a number of different vision tasks such as object recognition, activities recognition, tracking, etc. Compared with the traditional RGB sensors, the event cameras have many unique advantages such as ultra low resources consumption, high temporal resolution and much larger dynamic range. However, these cameras only produce noisy and asynchronous events of intensity changes, i.e., event-streams rather than frames, where conventional computer vision algorithms can't be directly applied. In our opinion the key challenge for improving the performance of event cameras in vision tasks is finding the appropriate representations of the event-streams so that cutting-edge learning approaches can be applied to fully uncover the spatio-temporal information contained in the event-streams. In this paper, we focus on the event-based human gait identification task and investigate the possible representations of the event-streams when deep neural networks are applied as the classifier. We propose new event-based gait recognition approaches basing on two different representations of the event-stream, i.e., graph and image-like representations, and use graph-based convolutional network (GCN) and convolutional neural networks (CNN) respectively to recognize gait from the event-streams. The two approaches are termed as EV-Gait-3DGraph and EV-Gait-IMG. To evaluate the performance of the proposed approaches, we collect two event-based gait datasets, one from real-world experiments and the other by converting the publicly available RGB gait recognition benchmark CASIA-B. Extensive experiments show that EV-Gait-3DGraph achieves significantly higher recognition accuracy than other competing methods when sufficient training samples are available. However, EV-Gait-IMG converges more quickly than graph-based approaches while training and shows good accuracy with only few number of training samples (less than ten). So image-like presentation is preferable when the amount of training data is limited. Yiran Shen 0001, Bowen Du 0002, Guangrong Zhao, Li-Zhen Cui 0001, Hongkai Wen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2021 | LeaD: Learn to Decode Vibration-based Communication for Intelligent Internet of ThingsabstractIn this article, we propose, LeaD , a new vibration-based communication protocol to Lea rn the unique patterns of vibration to D ecode the short messages transmitted to smart IoT devices. Unlike the existing vibration-based communication protocols that decode the short messages symbol-wise, either in binary or multi-ary, the message recipient in LeaD receives vibration signals corresponding to bits-groups. Each group consists of multiple symbols sent in a burst and the receiver decodes the group of symbols as a whole via machine learning-based approach. The fundamental behind LeaD is different combinations of symbols (1 s or 0 s) in a group will produce unique and reproducible patterns of vibration. Therefore, decoding in vibration-based communication can be modeled as a pattern classification problem. We design and implement a number of different machine learning models as the core engine of the decoding algorithm of LeaD to learn and recognize the vibration patterns. Through the intensive evaluations on large amount of datasets collected, the Convolutional Neural Network (CNN)-based model achieves the highest accuracy of decoding (i.e., lowest error rate), which is up to 97% at relatively high bits rate of 40 bits/s. While its competing vibration-based communication protocols can only achieve transmission rate of 10 bits/s and 20 bits/s with similar decoding accuracy. Furthermore, we evaluate its performance under different challenging practical settings and the results show that LeaD with CNN engine is robust to poses, distances (within valid range), and types of devices, therefore, a CNN model can be generally trained beforehand and widely applicable for different IoT devices under different circumstances. Finally, we implement LeaD on both off-the-shelf smartphone and smart watch to measure the detailed resources consumption on smart devices. The computation time and energy consumption of its different components show that LeaD is lightweight and can run in situ on low-cost smart IoT devices, e.g., smartwatches, without accumulated delay and introduces only marginal system overhead. Guangrong Zhao, Bowen Du 0002, Yiran Shen 0001, Zhenyu Lao, Li-Zhen Cui 0001, Hongkai Wen 0001 |
ACM Trans. Sens. Networks | 1 |
| 2019 | EV-Gait: Event-Based Robust Gait Recognition Using Dynamic Vision SensorsabstractIn this paper, we introduce a new type of sensing modality, the Dynamic Vision Sensors (Event Cameras), for the task of gait recognition. Compared with the traditional RGB sensors, the event cameras have many unique advantages such as ultra low resources consumption, high temporal resolution and much larger dynamic range. However, those cameras only produce noisy and asynchronous events of intensity changes rather than frames, where conventional vision-based gait recognition algorithms can’t be directly applied. To address this, we propose a new Event-based Gait Recognition (EV-Gait) approach, which exploits motion consistency to effectively remove noise, and uses a deep neural network to recognise gait from the event streams. To evaluate the performance of EV-Gait, we collect two event-based gait datasets, one from real-world experiments and the other by converting the publicly available RGB gait recognition benchmark CASIA-B. Extensive experiments show that EV-Gait can get nearly 96% recognition accuracy in the real-world settings, while on the CASIA-B benchmark it achieves comparable performance with state-of-the-art RGB-based gait recognition approaches. Bowen Du 0002, Yiran Shen 0001, Guangrong Zhao, Hongkai Wen 0001 |
CVPR | 5 |