Yiran Shen 0001

dblp:71/11188-1 · DBLP profile ↗
← Back
69ranked-venue papers
13as first author
42since 2021 · last 2026
0000-0003-1385-1480ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 31 · 11 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 18 since 2021Human-computer interaction and ubiquitous computing · 10 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 8 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Security and privacy · 1 · 1 first-author
YearPublicationVenuePosition
2026 Ev-iCRF: Self-supervised Event-guided iCRF Estimation for HDR Image Reconstruction
abstract
In this paper, we present Ev-iCRF, a novel self-supervised pipeline for high dynamic range (HDR) image reconstruction from a single-exposure low dynamic range (LDR) image, guided by asynchronous event streams generated by a bio-inspired event camera. The highlight of Ev-iCRF lies in its formulation of the inverse camera response function (iCRF) based on Event-LDR Correspondence. By leveraging the HDR properties of event data, the method enables direct iCRF estimation, offering a new perspective for event-guided HDR imaging. The pipeline is trained in a self-supervised manner using formulation-driven iCRF estimation loss and refinement loss, without the need for synchronized HDR supervision. Ev-iCRF adopts a two-stage coarse-to-fine reconstruction pipeline, allowing effective fusion of features from both LDR image and event data. The event information is used to optimize the iCRF, enabling accurate HDR reconstruction from LDR inputs. We evaluate Ev-iCRF on real-world datasets, and results show that it outperforms state-of-the-art methods in HDR reconstruction accuracy. Moreover, the reconstructed images demonstrate improved texture fidelity and structural detail.
Xucheng Guo, Lin Wang 0025, Yiran Shen 0001
AAAI4
2026 Towards Event-guided Panoramic HDR Video Reconstruction for Indoor Immersive VR: A Novel Dataset and Approach
abstract
High Dynamic Range (HDR) panoramic video is crucial to enhance immersive experience in Virtual Reality (VR). However, a hurdle is that panoramic cameras often struggle with limited dynamic range and motion blur. Inspired by the event-driven sensing of the human eye, this paper explores the potential of the event cameras to enhance panoramic HDR video reconstruction. As a pioneering research endeavor, we first starts by designing a novel hybrid imaging platform equipped with preprocessing pipelines for event-panorama synchronization, alignment, and HDR ground truth generation. Based on the platform, we then introduce Ev-Pano, the first event-panorama HDR video covering diverse indoor scenes for panoramic HDR video reconstruction. We hope Ev-Pano will establish a foundation to support event-guided panoramic HDR imaging and VR research community. With Ev-Pano, we further propose a novel approach that employs a weighting function-based luminance fusion to enable events to recover missing textures in LDR panoramic videos for panoramic HDR video reconstruction. We conduct extensive experiments to demonstrate the effectiveness of our approach. The results show the best performance of ours than prior arts. Meanwhile, a user study on an HDR-capable head-mounted display (Apple Vision Pro) shows feasible perceptual quality (which is closer to the HDR ground truth) of the reconstructed panoramic HDR videos. The codes and part of the dataset can be accessed via the anonymized link https://anonymous.4open.science/r/Ev-Pano-D2D2/.
Xucheng Guo, Majed Elwardy, Yan Hu 0003, Yuanfeng Zhou, Xiaoming Chen 0006, Yiran Shen 0001
VR10
2026 Dynolayout: robust layout estimation from event-stream for extended reality under dynamic scenarios
Xucheng Guo, Qiang Qu 0004, Guangrong Zhao, Yuanfeng Zhou, Yiran Shen 0001
CCF Trans. Pervasive Comput. Interact.7
2026 NVS-SQA: Exploring Self-Supervised Quality Representation Learning for Neurally Synthesized Scenes Without References
abstract
Neural View Synthesis (NVS), such as NeRF and 3D Gaussian Splatting, effectively creates photorealistic scenes from sparse viewpoints, typically evaluated by quality assessment methods like PSNR, SSIM, and LPIPS. However, these full-reference methods, which compare synthesized views to reference views, may not fully capture the perceptual quality of neurally synthesized scenes (NSS), particularly due to the limited availability of dense reference views. Furthermore, the challenges in acquiring human perceptual labels hinder the creation of extensive labeled datasets, risking model overfitting and reduced generalizability. To address these issues, we propose NVS-SQA, a NSS quality assessment method to learn no-reference quality representations through self-supervision without reliance on human labels. Traditional self-supervised learning predominantly relies on the "same instance, similar representation" assumption and extensive datasets. However, given that these conditions do not apply in NSS quality assessment, we employ heuristic cues and quality scores as learning objectives, along with a specialized contrastive pair preparation process to improve the effectiveness and efficiency of learning. The results show that NVS-SQA outperforms 17 no-reference methods by a large margin (i.e., on average 109.5% in SRCC, 98.6% in PLCC, and 91.5% in KRCC over the second best) and even exceeds 16 full-reference methods across all evaluation metrics (i.e., 22.9% in SRCC, 19.1% in PLCC, and 18.6% in KRCC over the second best).
Qiang Qu 0004, Yiran Shen 0001, Xiaoming Chen 0006, Vera Chung, Tom Weidong Cai, Tongliang Liu
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 Progressive Orthodontic Motion Planning Based on Hierarchical Diffusion Transformer
abstract
Orthodontic motion planning plays a crucial role in digital orthodontics by predicting tooth motion sequences to assist dentists in formulating treatment plans efficiently. Most prior work generates the entire intermediate tooth motion sequence given the initial and target tooth alignments. In practice, only the initial alignment of the patient is obtained. However, no existing method can predict the complete motion sequence using only the initial tooth alignment. To address this gap, we propose OrthoDiff, a novel target-free framework that uses only initial tooth alignment through a progressive generation strategy. This strategy generates tooth motion sequences by decomposing the entire motion sequence into multi-level motions, progressively constraining the inference space and reducing the complexity of target-free planning from coarse to fine. Moreover, we design a hierarchical diffusion transformer as the backbone of OrthoDiff, which treats tooth alignment as a sequence of tooth tokens and fully leverages the topological prior knowledge of the dental model. Through extensive evaluations, we demonstrate that our method significantly outperforms state-of-the-art techniques in target-free tooth motion generation. Ablation studies further confirm the efficacy of key components in our network design. Meanwhile, we also achieve state-of-the-art results in tooth target alignment prediction, benefiting from our framework. The code and data will be publicly available at https://github.com/Intelligent-Orthodontics/OrthoDiff.github.io.
Yeying Fan, Yuanfeng Zhou, Guangshun Wei, Zhiming Cui 0001, Yiran Shen 0001, Yong-Jin Liu 0001, Wenping Wang 0001
IEEE Trans. Medical Imaging6
2025 Te3DFR: Texture-Enabled 3D Face Reconstruction from Monocular Image via Self-supervised Learning
Yishen Bi, Chen Wang 0054, Lei Li 0008, Yiran Shen 0001, Yuanfeng Zhou
CGI (2)4
2025 EEG2Gaussian: Decoding and Visualizing Visual-Evoked EEG for VR Scenes Using 3D Gaussian Splatting
abstract
Decoding and visualizing brain activity evoked by visual stimuli is critical for both understanding neural mechanisms and advancing brain-computer interfaces (BCIs). However, non-invasive signals such as Electroencephalogram (EEG) present significant challenges due to their inherently low signal-to-noise ratios. Although recent deep learning methods have resolved this task, most approaches are confined to 2D visualizations that fail to capture the complexities of real-world 3D perception. In this research, we investigate the relationship between EEG signals and 3D visual stimuli presented in virtual reality (VR) scenes, aiming to extract taskrelevant semantics from the EEG responses elicited by these stimuli. We introduce EEG2Gaussian, a novel framework for decoding and visualizing visual-evoked EEG signals by reconstructing immersive VR scenes using 3D Gaussian Splatting. The framework consists of three stages. The preprocessing stage removes noise and artifacts from raw EEG signals to provide cleaner input for subsequent processing. In the encoding stage, we propose a Neural Temporal-Frequency Encoder (NTF-Encoder) to extract temporal and frequency features using fused channel and band attention mechanisms, and disentangles them into high-level and low-level semantic representations. In the decoding stage, a 3D EEG Decoder takes these multi-level features through separate pathways as conditional inputs to guide the reconstruction of semantically consistent VR scenes. Furthermore, we construct a VR-EEG dataset that pairs real-time EEG recordings with VR scenes, and analyze how different types of scenes affect EEG responses across frequency bands. Our experimental results show that EEG2Gaussian can reconstruct VR scenes that are semantically aligned with the visual stimuli. Ablation studies verify the effectiveness of channel and band attention in EEG feature encoding, and demonstrate that combining high-level and low-level semantic features enhances the consistency and interpretability of the reconstructed scenes.
Qiang Qu 0004, Xiaoming Chen 0006, Longfei Han, Yiran Shen 0001
ISMAR6
2025 MirrorPose: Enabling Full-Body Gestures Interaction for Head-Mounted Devices with a Full-Length Mirror
abstract
Human-computer interaction based on full-body gestures has been successfully adopted in various applications, such as motionsensing games. Typically, full-body gestures are captured using vision-based pose estimation or multiple inertial measurement units (IMUs) attached to the limbs. Gesture-based interactions in virtual and augmented reality environments allow for seamless and intuitive engagement across virtual and real domains. However, due to the design of head-mounted devices, only partial body tracking—such as hand tracking—is typically available for interactions. Capturing full-body pose using head-mounted sensors is inherently challenging due to device placement constraints. Furthermore, the limited computational resources of AR devices (e.g., constrained processing power and memory bandwidth) present significant challenges for the real-time deployment of sophisticated 3D human pose estimation architectures. To address these challenges, we propose MirrorPose, a lightweight framework that integrates a 3D pose estimation network (PoseARNet), optimized for the resource constraints of AR headsets and the dynamic viewpoint changes inherent in mirror-mediated spatial perception. This design enables practical, full-body gesture interaction on AR devices. To demonstrate its practicality, we developed a 3D virtual teaching application on Microsoft HoloLens 2, to enhance students' understanding of human poses. Extensive experiments and evaluations confirm that our system provides users with accurate and timely feedback. The codes and dataset are available at https://github.com/zhchlong/mirror_pose.
Xingwang Xue, Xiyu Sheng, Qiang Qu 0004, Yiran Shen 0001
ISMAR6
2025 WinSpy: Cross-window Side-channel Attacks on Android's Multi-window Mode
abstract
With the development of the Android system and increasing screen size, the use of multi-window mode has become prevalent among users. However, the security and privacy implications associated with this mode have not been thoroughly investigated. This paper uncovers severe and unique security vulnerabilities in Android's multi-window mode, revealing several high-risk side-channels that facilitate diverse cross-window attacks, leading to significant breaches of user privacy. In detail, our research introduces WinSpy, a framework leveraging a newly discovered resource contention side-channel in multi-window mode to fingerprint app launches, web pages, and in-app activities, all without violating Android's permission framework. Our extensive evaluations demonstrate that WinSpy achieves high accuracy (from 70 to 80% detecting website and app launches to over 97% recognizing critical in-app activities). Additionally, we reveal that due to Android's lenient permission management for this mode, window apps can also use Inertial Measurement Unit sensors to launch attacks, such as inferring the user's touch positions outside the window with high precision. Furthermore, we propose systematic mitigations against these vulnerabilities.
Chuan Yan, Liuhuo Wan, Hui Zhuang, Pengfei Hu 0001, Guangdong Bai, Yiran Shen 0001
MobiCom7
2025 VF-Lens: Enhancing Visual Perception of Visually Impaired Users in VR via Adversarial Learning with Visual Field Attention
abstract
This research aims to enhance the image perception of visually impaired users in VR environments. We propose VF-Lens, a model that adaptively compensates for light sensitivity based on the user’s visual field impairment, acting as a virtual lens between the visually impaired users and the VR world. VF-Lens is designed as a tailored generative adversarial learning model with a generator and discriminator, offering applicability to various types of visual impairments while bypassing engineering complexities. The generator creates a "hyperimage" tailored to the user’s visual field impairment, which then undergoes a particular regression process to predict and replicate the real perception of the visually impaired user. The discriminator then evaluates the similarity between the replicated perception and the original image. Through adversarial training, the generator can produce hyperimages that adapt to the user’s visual field parameters, enabling them to perceive the image more similarly to normal-vision users. We further improve VF-Lens by proposing new "visual field attention" mechanisms that prioritize and refine visual information in the user’s visual field. Extensive evaluation, encompassing both visually impaired participants and simulations, has been conducted to demonstrate the effectiveness of VF-Lens in improving visual perception for visually impaired users. Moreover, we establish a standardized evaluation process involving tailored metrics as well as objective and subjective evaluations to promote reusability and comparability for future research in this field.
Xiaoming Chen 0006, Dehao Han, Qiang Qu 0004, Yiran Shen 0001
VR4
2025 Design and evaluation of AR-based adaptive human-computer interaction cognitive training
abstract
As human-computer interaction (HCI) technology advances, the use of augmented reality (AR) in cognitive training is becoming more prevalent. However, traditional training methods often apply a one-size-fits-all approach, failing to accommodate the varied training needs of individuals with different cognitive levels. Additionally, most HCI systems use subjective questionnaires for evaluation, which can be influenced by the subjects' emotional and mental states. To overcome these challenges, this study developed an AR-based adaptive HCI cognitive training system that dynamically adjusts task difficulty based on real-time user performance. We used multi-source data to empirically validate the effectiveness of adaptive HCI in cognitive training. Specifically, we recorded functional Near-Infrared Spectroscopy (fNIRS) data, movement data, task performance, and subjective feedback from 22 elderly participants, dividing them into two groups—low cognitive group and normal cognitive group. The results showed that the system exerted a significant influence on brain functional connectivity (FC) associated with cognition, movement, and vision. Changes in FC may highlight the benefits of adaptive HCI training strategies. Furthermore, participants with normal cognitive abilities significantly outperformed their low cognitive counterparts in task performance. In conclusion, this study designed and evaluated an AR-based adaptive HCI cognitive training system that ensures personalized training. It demonstrated the feasibility of adaptive HCI strategies in cognitive rehabilitation by incorporating physiological and behavioral data, thereby enhancing the precision of quantitative assessments for HCI systems.
Man Chu, Jing Qu 0001, Tan Zou, Qinbiao Li, Lingguo Bu, Yiran Shen 0001
Int. J. Hum. Comput. Stud.6
2025 RGBE-Gaze: A Large-Scale Event-Based Multimodal Dataset for High Frequency Remote Gaze Tracking
abstract
High-frequency gaze tracking demonstrates significant potential in various critical applications, such as foveatedrendering, gaze-based identity verification, and the diagnosis of mental disorders. However, existing eye-tracking systems based on CCD/CMOS cameras either provide tracking frequencies below 200 Hz or employ high-speedcameras, causing high power consumption and bulky devices. While there have been some high-speed eye-tracking datasets and methods based on event cameras, they are primarily tailored for near-eye camera scenarios. They lackthe advantages associated with remote camera scenarios, such as the absence of the need for direct contact, improved user comfort and head pose freedom. In this work, we present RGBE-Gaze, the first large-scale and multimodal dataset for remote gaze tracking in high-frequency through synchronizing RGB and event cameras. This dataset is collected from 66 participants with diverse genders and age groups. Our setup captures 3.6 million RGB images and 26.3 billion event samples. Additionally, the dataset includes 10.7 million gaze references from the Gazepoint GP3 HD eye tracker and 15,972 sparse points of gaze (PoG) ground truth obtained through manualstimuli clicks by participants. We present dataset characteristics such as head pose, gaze direction, and pupil size. Furthermore, we introduce a hybrid frame-event based gaze estimation method specifically designed for the collected dataset. Moreover, we perform extensive evaluations of different benchmarking methods under variousgaze-related factors. The evaluation results illustrate that introducing event stream as a new modality improves gazetracking frequency and demonstrates greater estimation robustness across diverse gaze-related factors.
Guangrong Zhao, Yiran Shen 0001, Zhaoxin Shen, Yuanfeng Zhou, Hongkai Wen 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Ui-Ear: On-Face Gesture Recognition Through On-Ear Vibration Sensing
abstract
With the convenient design and prolific functionalities, wireless earbuds are fast penetrating in our daily life and taking over the place of traditional wired earphones. The sensing capabilities of wireless earbuds have attracted great interests of researchers on exploring them as a new interface for human-computer interactions. However, due to its extremely compact size, the interaction on the body of the earbuds is limited and not convenient. In this paper, we proposeUi-Ear, a new on-face gesture recognition system to enrich interaction maneuvers for wireless earbuds.Ui-Earexploits the sensing capability of Inertial Measurement Units (IMUs) to extend the interaction to the skin of the face near ears. The accelerometer and gyroscope in IMUs perceive dynamic vibration signals induced by on-face touching and moving, which brings rich maneuverability. Since IMUs are provided on most of the budget and high-end wireless earbuds, we believe thatUi-Earhas great potential to be adopted pervasively. To demonstrate the feasibility of the system, we define seven different on-face gestures and design an end-to-end learning approach based on Convolutional Neural Networks (CNNs) for classifying different gestures. To further improve the generalization capability of the system, adversarial learning mechanism is incorporated in the offline training process to suppress the user-specific features while enhancing gesture-related features. We recruit 20 participants and collect a realworld datasets in a common office environment to evaluate the recognition accuracy. The extensive evaluations show that the average recognition accuracy ofUi-Earis over 95% and 82.3% in the user-dependent and user-independent tasks, respectively. Moreover, we also show that the pre-trained model (learned from user-independent task) can be fine-tuned with only few training samples of the target user to achieve relatively high recognition accuracy (up to 95%). At last, we implement the personalization and recognition components ofUi-Earon an off-the-shelf Android smartphone to evaluate its system overhead. The results demonstrateUi-Earcan achieve real-time response while only brings trivial energy consumption on smartphones.
Guangrong Zhao, Yiran Shen 0001, Feng Li 0002, Lei Liu 0003, Li-Zhen Cui 0001, Hongkai Wen 0001
IEEE Trans. Mob. Comput.2
2025 Beyond Subspace Isolation: Many-to-Many Transformer for Light Field Image Super-Resolution
abstract
The effective extraction of spatial-angular features plays a crucial role in light field image super-resolution (LFSR) tasks, and the introduction of convolution and Transformers leads to significant improvement in this area. Nevertheless, due to the large 4D data volume of light field images, many existing methods opted to decompose the data into a number of lower-dimensional subspaces and perform Transformers in each sub-space individually. As a side effect, these methods inadvertently restrict the self-attention mechanisms to a One-to-One scheme accessing only a limited subset of LF data, explicitly preventing comprehensive optimization on all spatial and angular cues. In this paper, we identify this limitation as subspace isolation and introduce a novel Many-to-Many Transformer (M2MT) to address it. M2MT aggregates angular information in the spatial subspace before performing the self-attention mechanism. It enables complete access to all information across all sub-aperture images (SAIs) in a light field image. Consequently, M2MT is enabled to comprehensively capture long-range correlation dependencies. With M2MT as the foundational component, we develop a simple yet effective M2MT network for LFSR. Our experimental results demonstrate that M2MT achieves state-of-the-art performance across various public datasets, and it offers a favorable balance between model performance and efficiency, yielding higher-quality LFSR results with substantially lower demand for memory and computation. We further conduct in-depth analysis using local attribution maps (LAM) to obtain visual interpretability, and the results validate that M2MT is empowered with a truly non-local context in both spatial and angular subspaces to mitigate subspace isolation and acquire effective spatial-angular representation.
Zexi Hu, Xiaoming Chen 0006, Vera Chung, Yiran Shen 0001
IEEE Trans. Multim.4
2025 EX-Gaze: High-Frequency and Low-Latency Gaze Tracking with Hybrid Event-Frame Cameras for On-Device Extended Reality
abstract
The integration of gaze/eye tracking into virtual and augmented reality devices has unlocked new possibilities, offering a novel human-computer interaction (HCI) modality for on-device extended reality (XR). Emerging applications in XR, such as low-effort user authentication, mental health diagnosis, and foveated rendering, demand real-time eye tracking at high frequencies, a capability that current solutions struggle to deliver. To address this challenge, we present EX-Gaze, an event-based real-time eye tracking system designed for on-device extended reality. EX-Gaze achieves a high tracking frequency of 2KHz, providing decent accuracy and low tracking latency. The exceptional tracking frequency of EX-Gaze is achieved through the use of event cameras, cutting-edge, bio-inspired vision hardware that delivers event-stream output at high temporal resolution. We have developed a lightweight tracking framework that enables real-time pupil region localization and tracking on mobile devices. To effectively leverage the sparse nature of event-streams, we introduce the sparse event-patch representation and the corresponding sparse event patches transformer as key components to reduce computational time. Implemented on Jetson Orin Nano, a low-cost, small-sized mobile device with hybrid GPU and CPU components capable of parallel processing of multiple deep neural networks, EX-Gaze maximizes the computation power of Jetson Orin Nano through sophisticated computation scheduling and offloading between GPUs and CPUs. This enables EX-Gaze to achieve real-time tracking at 2KHz without accumulating latency. Evaluation on public datasets demonstrates that EX-Gaze outperforms other event-based eye tracking methods by striking the best balance between accuracy and efficiency on mobile devices. These results highlight EX-Gaze's potential as a groundbreaking technology to support XR applications that require high-frequency and real-time eye tracking. The code is available at https://github.com/Ningreka/EX-Gaze.
Yiran Shen 0001, Tongyu Zhang, Yanni Yang 0003, Hongkai Wen 0001
IEEE Trans. Vis. Comput. Graph.2
2025 p-Blend: Privacy- and Utility-Preserving Blendshape Perturbation Against Re-Identification Attacks in Virtual Reality
abstract
In this paper, we propose p-Blend, an efficient and effective blendshape perturbation mechanism designed to defend against both intra- and cross-app re-identification attacks in virtual reality. p-Blend provides privacy protection when streaming blendshape data to third-party applications on VR devices. In its design, we consider both privacy and utility. p-Blend not only perturbs blendshape values to resist re-identification attacks but also preserves the smoothness of facial animations and the naturalness of facial expressions, ensuring the continued usability of the data. We validate the effectiveness of p-Blend through extensive empirical evaluations and user studies. Quantitative experiments on a large-scale dataset collected from 45 participants demonstrate that p-Blend significantly reduces re-identification accuracy across a range of machine learning models. While pure-random perturbation fails to prevent attacks that exploit statistical features, p-Blend effectively mitigates these risks in both raw and statistical blendshape data. Additionally, user study results show that facial animations generated from p-Blend-perturbed blendshapes maintain greater smoothness and naturalness compared to those using purely random perturbation. The codes and dataset are available at https://github.com/jingwei1016/p-Blend.
Yan Hu 0003, Guangrong Zhao, Qing Yang 0009, Guangdong Bai, Yiran Shen 0001
IEEE Trans. Vis. Comput. Graph.7
2025 Self-Supervised Learning of Event-Guided Video Frame Interpolation for Rolling Shutter Frames
abstract
Most consumer cameras use rolling shutter (RS) exposure, the captured videos often suffer from distortions (e.g., skew and jelly effect). Also, these videos are impeded by the limited bandwidth and frame rate, which inevitably affect the video streaming experience. In this paper, we excavate the potential of event cameras as they enjoy high temporal resolution. Accordingly, we propose a framework to recover the global shutter (GS) high frame rate (i.e., slow motion) video without RS distortion from an RS camera and event camera. One challenge is the lack of real-world datasets for supervised training. Therefore, we explore self-supervised learning with the key idea of estimating the displacement field-a non-linear and dense 3D spatiotemporal representation of all pixels during the exposure time. This allows for a mutual reconstruction between RS and GS frames and facilitates slow-motion video recovery. We then combine the input RS frames with the DF to map them to the GS frames (RS-to-GS). Given the under-constrained nature of this mapping, we integrate it with the inverse mapping (GS-to-RS) and RS frame warping (RS-to-RS) for self-supervision. We evaluate our framework via objective analysis (i.e., quantitative and qualitative comparisons on four datasets) and subjective studies (i.e., user study). The results show that our framework can recover slow-motion videos without distortion, with much lower bandwidth (94% drop) and higher inference speed ($ 16\; {\rm ms}/{\rm frame}$16 ms / frame ) under $32 \times$32× frame interpolation.
Yunfan Lu, Guoqiang Liang 0003, Yiran Shen 0001, Lin Wang 0025
IEEE Trans. Vis. Comput. Graph.3
2025 AirtypeLogger: How Short Keystrokes in Virtual Space Can Expose Your Semantic Input to Nearby Cameras
abstract
Considering the issue of privacy leakage and motivating more sophisticated protection methods for air-typing with XR devices, in this paper, we propose AirtypeLogger, a new approach towards practical video-based attacks on the air-typing activities of XR users in virtual space. Different from the existing approaches, AirtypeLogger considers a scenario in which the users are typing a short text fragment with semantic meaning occasionally under the spy of video cameras. It detects and localizes the air-typing events in video streams and proposes the spatial-temporal representation to encode the keystrokes' relative positions and temporal order. Then, high-precision inference can be achieved by applying a Transformer-based network to the spatial and temporal encodings of the keystroke sequences. Finally, according to our extensive real-world experiments, AirtypeLogger can achieve a Character Error Rate (CER) of less than 0.1 as long as 7 air-typing events are observed, which is impossible for previous approaches that require long-term observation of the typing activities online before launching inference attacks. The implementation details and source codes can be found at https://github.com/ztysdu/AirtypeLogger.
Tongyu Zhang, Yiran Shen 0001, Yuanfeng Zhou
IEEE Trans. Vis. Comput. Graph.2
2024 E2HQV: High-Quality Video Generation from Event Camera via Theory-Inspired Model-Aided Deep Learning
abstract
The bio-inspired event cameras or dynamic vision sensors are capable of asynchronously capturing per-pixel brightness changes (called event-streams) in high temporal resolution and high dynamic range. However, the non-structural spatial-temporal event-streams make it challenging for providing intuitive visualization with rich semantic information for human vision. It calls for events-to-video (E2V) solutions which take event-streams as input and generate high quality video frames for intuitive visualization. However, current solutions are predominantly data-driven without considering the prior knowledge of the underlying statistics relating event-streams and video frames. It highly relies on the non-linearity and generalization capability of the deep neural networks, thus, is struggling on reconstructing detailed textures when the scenes are complex. In this work, we propose E2HQV, a novel E2V paradigm designed to produce high-quality video frames from events. This approach leverages a model-aided deep learning framework, underpinned by a theory-inspired E2V model, which is meticulously derived from the fundamental imaging principles of event cameras. To deal with the issue of state-reset in the recurrent components of E2HQV, we also design a temporal shift embedding module to further improve the quality of the video frames. Comprehensive evaluations on the real world event camera datasets validate our approach, with E2HQV, notably outperforming state-of-the-art approaches, e.g., surpassing the second best by over 40% for some evaluation metrics.
Qiang Qu 0004, Yiran Shen 0001, Xiaoming Chen 0006, Vera Chung, Tongliang Liu
AAAI2
2024 Towards High-Speed Passive Visible Light Communication with Event Cameras and Digital Micro-Mirrors
abstract
Passive visible light communication (VLC) modulates light propagation or reflection to transmit data without directly modulating the light source. Thus, passive VLC provides an alternative to conventional VLC, enabling communication where the light source cannot be directly controlled. There have been ongoing efforts to explore new methods and devices for modulating light propagation or reflection. The state-of-the-art has broken the 100 kbps data rate barrier for passive VLC by using a digital micro-mirror device (DMD) as the light modulating platform, or transmitter, and a photo-diode as the receiver. We significantly extend this work by proposing a massive spatial data channel framework for DMDs, where individual channels can be decoded in parallel using an event camera at the receiver. For the event camera, we introduce event processing algorithms to detect numerous channels and decode bits from individual channels with high reliability. Our prototype, built with off-the-shelf event cameras and DMDs, can decode up to ~2,000 parallel channels, achieving a data transmission rate of 1.6 Mbps, markedly surpassing current benchmarks by 16x.
Yiran Shen 0001, Kenuo Xu, Mahbub Hassan, Guangrong Zhao, Chenren Xu, Wen Hu 0001
SenSys2
2024 Understanding the Impact of Longitudinal VR Training on Users with Mild Cognitive Impairment Using fNIRS and Behavioral Data
abstract
With the growing needs on rehabilitation of the mild cognitive impairment (MCI) users group and the advantages of virtual reality (VR) technologies in cognitive training, the development of VR-based rehabilitation training methods has become a hot spot recently. However, the challenges in accurately measuring users’ needs and quantifying training system efficacy are still not well resolved, especially for longitudinal tracking. In this study, a VR-based cognitive training and evaluation system is designed and implemented, targeting at fulfilling the rehabilitation needs of MCI users. It evaluates the impact of longitudinal VR-based training on MCI users with a number of feedback methodologies including brain activation indicators, brain network connectivity indicators, behavioral indicators and the Montreal Cognitive Assessment (MoCA) scale scores, extracted from multi-modal data collected while training. A two-month longitudinal tracking ergonomics experiment was conducted to validate the usability of the feedback methodologies and to explore the influence of the training duration on the rehabilitation efficacy. The results showed that our proposed VR-based cognitive training and evaluation system had a positively significant impact on the rehabilitation of the MCI group. Meanwhile, the multi-source feedbacks can also help the updates and iterations of VR-based rehabilitation training systems. Finally, this study provides guidance for the selection of rehabilitation cycles and emphasizes the importance of quantitative studies with longitudinal follow-up in assessing rehabilitation efficacy.
Jing Qu 0001, Shantong Zhu, Yiran Shen 0001, Lingguo Bu
VR3
2024 KD-Eye: Lightweight Pupil Segmentation for Eye Tracking on VR Headsets via Knowledge Distillation
Yanlin Li 0014, Guangrong Zhao, Yiran Shen 0001
WASA (1)4
2024 Profinder: Towards Professionals Recognition on Mobile Devices for Users with Cognitive Decline
Yong Wang 0020, Yiran Shen 0001
WASA (3)4
2024 Ef-kpress: joint event-frame compression and video generation with event cameras for low-bandwidth VR streaming
Guangyong Hao, Yiran Shen 0001, Feng Li 0002
CCF Trans. Pervasive Comput. Interact.2
2024 BudsAuth: Toward Gesture-Wise Continuous User Authentication Through Earbuds Vibration Sensing
abstract
The surge in popularity of wireless headphones, particularly wireless earbuds, as smart wearables, has been notable in recent years. These devices, empowered by artificial intelligence (AI), are broadening their utility in areas such as speech recognition, augmented reality, pose recognition, and health care monitoring, thereby enriching user experiences through novel interactive interfaces driven by embedded sensors. However, the widespread adoption of wireless earbuds has spurred concerns regarding security and privacy, necessitating robust bespoke security measures. Despite the miniaturization of mobile chips enabling the integration of sophisticated algorithms into smart wearables, the research and industrial communities have yet to accord adequate attention to earbud security. This paper focuses on empowering wireless earbuds to authenticate their legitimate users, tackling the challenges associated with conventional authentication methods. Instead of relying on input interface authentication methods like PIN or lock patterns, this research delves into leveraging Inertial Measurement Unit (IMU) data collected during interactions with devices to extract novel biometric features, presenting an alternative approach that nonetheless confronts challenges related to signal capture and interference. Consequently, we propose and design BudsAuth, an implicit user authentication framework that harnesses built-in IMU sensors in smart earbuds to capture vibration signals induced by on-face touching interactions with the earbuds. These vibrations are utilized to deliver continuous and implicit user authentication with high precision and compatibility across various earbud models. Extensive evaluation demonstrates BudsAuth’s capability to achieve an Equal Error Rate (EER) of 0.0003, representing an approximate 99.97% accuracy with seven consecutive samples of interactive gestures for implicit authentication.
Yong Wang 0020, Feng Li 0002, Pengfei Hu 0001, Yiran Shen 0001
IEEE Internet Things J.6
2024 EV-Perturb: event-stream perturbation for privacy-preserving classification with dynamic vision sensors
Yong Wang 0020, Qing Yang 0009, Yiran Shen 0001, Hongkai Wen 0001
Multim. Tools Appl.4
2024 EvRepSL: Event-Stream Representation via Self-Supervised Learning for Event-Based Vision
abstract
Event-stream representation is the first step for many computer vision tasks using event cameras. It converts the asynchronous event-streams into a formatted structure so that conventional machine learning models can be applied easily. However, most of the state-of-the-art event-stream representations are manually designed and the quality of these representations cannot be guaranteed due to the noisy nature of event-streams. In this paper, we introduce a data-driven approach aiming at enhancing the quality of event-stream representations. Our approach commences with the introduction of a new event-stream representation based on spatial-temporal statistics, denoted as EvRep. Subsequently, we theoretically derive the intrinsic relationship between asynchronous event-streams and synchronous video frames. Building upon this theoretical relationship, we train a representation generator, RepGen, in a self-supervised learning manner accepting EvRep as input. Finally, the event-streams are converted to high-quality representations, termed as EvRepSL, by going through the learned RepGen (without the need of fine-tuning or retraining). Our methodology is rigorously validated through extensive evaluations on a variety of mainstream event-based classification and optical flow datasets (captured with various types of event cameras). The experimental results highlight not only our approach's superior performance over existing event-stream representations but also its versatility, being agnostic to different event cameras and tasks.
Qiang Qu 0004, Xiaoming Chen 0006, Vera Chung, Yiran Shen 0001
IEEE Trans. Image Process.4
2024 EV-Tach: A Handheld Rotational Speed Estimation System With Event Camera
abstract
Rotational speed is one of the important metrics to be measured for calibrating electric motors in manufacturing, monitoring engines during car repairs, detecting faults in electrical appliance and more. However, existing measurement techniques either require prohibitive hardware (e.g., high-speed camera) or are inconvenient to use in real-world application scenarios. In this paper, we propose,EV-Tach, a novel handheld rotational speed estimation system that utilizes emerging imaging sensors known as event cameras or dynamic vision sensors (DVS). The pixels of DVS work independently and trigger an event as soon as a per-pixel intensity change is detected, without global synchronization like conventional RGB cameras. Thus, its unique design features high temporal resolution and generates sparse events, which benefits the high-speed rotation estimation. To achieve accurate and efficient rotational speed estimation, a series of signal processing algorithms are specifically designed for the event streams generated by event cameras on an embedded platform. First, a new cluster-centroids initialization module is proposed to initialize the centroids of the clusters to address the issue that common clustering approaches are easy to fall into a local optimal solution without proper initial centroids. Second, an outlier removal module is designed to suppress the background noise caused by subtle hand movements and host devices vibrations. Third, a coarse-to-fine alignment strategy is proposed with an event stream alignment method to obtain angle of rotation and achieve accurate estimation for rotational speed in a large range. With these bespoke components,EV-Tachis able to extract the rotational speed accurately from the event stream produced by an event camera recording rotary targets. According to our extensive evaluations under controlled and practical experiment settings, the Relative Mean Absolute Error (RMAE) ofEV-Tachis as low as$0.3\%_{0}$, which is comparable to the state-of-the-art laser tachometer under fixed measurement mode. Moreover,EV-Tachis robust to subtle movement of user's hand and dazzling light outdoor, therefore, can be used as a handheld device under challenging lighting condition, where the laser tachometer fails to produce reasonable results. To speed up the processing ofEV-Tachand reduce its resource consumption on embedded devices, event stream is significantly downsampled by merging neighboring events while preserving its formation in spatial-temporal domain. At last, we implementEV-Tachon Raspberry Pi and the evaluation results show that the downsampling process preserves the high measurement accuracy while saving the computation speed and energy consumption by approximately 8 times and 30 times in average.
Guangrong Zhao, Yiran Shen 0001, Pengfei Hu 0001, Lei Liu 0003, Hongkai Wen 0001
IEEE Trans. Mob. Comput.2
2024 VibHead: An Authentication Scheme for Smart Headsets through Vibration
abstract
Recent years have witnessed the fast penetration of Virtual Reality (VR) and Augmented Reality (AR) systems into our daily life, the security and privacy issues of the VR/AR applications have been attracting considerable attention. Most VR/AR systems adopt head-mounted devices (i.e., smart headsets) to interact with users and the devices usually store the users’ private data. Hence, authentication schemes are desired for the head-mounted devices. Traditional knowledge-based authentication schemes for general personal devices have been proved vulnerable to shoulder-surfing attacks, especially considering the headsets may block the sight of the users. Although the robustness of the knowledge-based authentication can be improved by designing complicated secret codes in virtual space, this approach induces a compromise of usability. Another choice is to leverage the users’ biometrics; however, it either relies on highly advanced equipments which may not always be available in commercial headsets or introduce heavy cognitive load to users. In this paper, we propose a vibration-based authentication scheme, VibHead, for smart headsets. Since the propagation of vibration signals through human heads presents unique patterns for different individuals, VibHead employs a CNN-based model to classify registered legitimate users based the features extracted from the vibration signals. We also design a two-step authentication scheme where the above user classifiers are utilized to distinguish the legitimate user from illegitimate ones. We implement VibHead on a Microsoft HoloLens equipped with a linear motor and an IMU sensor which are commonly used in off-the-shelf personal smart devices. According to the results of our extensive experiments, with short vibration signals (≤ 1s ), VibHead has an outstanding authentication accuracy; both FAR and FRR are around 5%.
Feng Li 0002, Huan Yang 0001, Dongxiao Yu, Yuanfeng Zhou, Yiran Shen 0001
ACM Trans. Sens. Networks6
2024 PDSR: A Privacy-Preserving Diversified Service Recommendation Method on Distributed Data
abstract
The last decade has witnessed a tremendous growth of service computing, while efficient service recommendation methods are desired to recommend high-quality services to users. It is well known that collaborative filtering is one of the most popular methods for service recommendation based on QoS, and many existing proposals focus on improving recommendation accuracy, i.e., recommending high-quality redundant services. Nevertheless, users may have different requirements on QoS, and hence diversified recommendation has been attracting increasing attention in recent years to fulfill users’ diverse demands and to explore potential services. Unfortunately, the recommendation performances relies on a large volume of data (e.g., QoS data), whereas the data may be distributed across multiple platforms. Therefore, to enable data sharing across the different platforms for diversified service recommendation, we propose aPrivacy-preserving Diversified Service Recommendation(PDSR) method. Specifically, we innovate in leveraging the Locality-Sensitive Hashing (LSH) mechanism such that privacy-preserved data sharing across different platforms is enabled to construct a service similarity graph. Based on the similarity graph, we propose a novel accuracy-diversity metric and design a 2-approximation algorithm to select$K$services to recommend by maximizing the accuracy-diversity measure. Extensive experiments on real datasets are conducted to verify the efficacy of our PDSR method.
Huan Yang 0001, Yiran Shen 0001, Chao Liu 0008, Lianyong Qi, Xiuzhen Cheng, Feng Li 0002
IEEE Trans. Serv. Comput.3
2024 NeRF-NQA: No-Reference Quality Assessment for Scenes Generated by NeRF and Neural View Synthesis Methods
abstract
Neural View Synthesis (NVS) has demonstrated efficacy in generating high-fidelity dense viewpoint videos using a image set with sparse views. However, existing quality assessment methods like PSNR, SSIM, and LPIPS are not tailored for the scenes with dense viewpoints synthesized by NVS and NeRF variants, thus, they often fall short in capturing the perceptual quality, including spatial and angular aspects of NVS-synthesized scenes. Furthermore, the lack of dense ground truth views makes the full reference quality assessment on NVS-synthesized scenes challenging. For instance, datasets such as LLFF provide only sparse images, insufficient for complete full-reference assessments. To address the issues above, we propose NeRF-NQA, the first no-reference quality assessment method for densely-observed scenes synthesized from the NVS and NeRF variants. NeRF-NQA employs a joint quality assessment strategy, integrating both viewwise and pointwise approaches, to evaluate the quality of NVS-generated scenes. The viewwise approach assesses the spatial quality of each individual synthesized view and the overall inter-views consistency, while the pointwise approach focuses on the angular qualities of scene surface points and their compound inter-point quality. Extensive evaluations are conducted to compare NeRF-NQA with 23 mainstream visual quality assessment methods (from fields of image, video, and light-field assessment). The results demonstrate NeRF-NQA outperforms the existing assessment methods significantly and it shows substantial superiority on assessing NVS-synthesized scenes without references. An implementation of this paper are available at https://github.com/VincentQQu/NeRF-NQA.
Qiang Qu 0004, Hanxue Liang, Xiaoming Chen 0006, Vera Chung, Yiran Shen 0001
IEEE Trans. Vis. Comput. Graph.5
2024 Swift-Eye: Towards Anti-blink Pupil Tracking for Precise and Robust High-Frequency Near-Eye Movement Analysis with Event Cameras
abstract
Eye tracking has shown great promise in many scientific fields and daily applications, ranging from the early detection of mental health disorders to foveated rendering in virtual reality (VR). These applications all call for a robust system for high-frequency near-eye movement sensing and analysis in high precision, which cannot be guaranteed by the existing eye tracking solutions with CCD/CMOS cameras. To bridge the gap, in this paper, we propose Swift-Eye, an offline precise and robust pupil estimation and tracking framework to support high-frequency near-eye movement analysis, especially when the pupil region is partially occluded. Swift-Eye is built upon the emerging event cameras to capture the high-speed movement of eyes in high temporal resolution. Then, a series of bespoke components are designed to generate high-quality near-eye movement video at a high frame rate over kilohertz and deal with the occlusion over the pupil caused by involuntary eye blinks. According to our extensive evaluations on EV-Eye, a large-scale public dataset for eye tracking using event cameras, Swift-Eye shows high robustness against significant occlusion. It can improve the IoU and F1-score of the pupil estimation by 20% and 12.5% respectively, compared with the second-best competing approach, when over 80% of the pupil region is occluded by the eyelid. Lastly, it provides continuous and smooth traces of pupils in extremely high temporal resolution and can support high-frequency eye movement analysis and a number of potential applications, such as mental health diagnosis, behaviour-brain association, etc. The implementation details and source codes can be found at https://github.com/ztysdu/Swift-Eye.
Tongyu Zhang, Yiran Shen 0001, Guangrong Zhao, Lin Wang 0025, Xiaoming Chen 0006, Lu Bai 0004, Yuanfeng Zhou
IEEE Trans. Vis. Comput. Graph.2
2023 EV-Eye: Rethinking High-frequency Eye Tracking through the Lenses of Event Cameras
abstract
In this paper, we present EV-Eye, a first-of-its-kind large scale multimodal eye tracking dataset aimed at inspiring research on high-frequency eye/gaze tracking. EV-Eye utilizes an emerging bio-inspired event camera to capture independent pixel-level intensity changes induced by eye movements, achieving sub-microsecond latency. Our dataset was curated over a two-week period and collected from 48 participants encompassing diverse genders and age groups. It comprises over 1.5 million near-eye grayscale images and 2.7 billion event samples generated by two DAVIS346 event cameras. Additionally, the dataset contains 675 thousands scene images and 2.7 million gaze references captured by Tobii Pro Glasses 3 eye tracker for cross-modality validation. Compared with existing event-based high-frequency eye tracking datasets, our dataset is significantly larger in size, and the gaze references involve more natural eye movement patterns, i.e., fixation, saccade and smooth pursuit. Alongside the event data, we also present a hybrid eye tracking method as benchmark, which leverages both the near-eye grayscale images and event data for robust and high-frequency eye tracking. We show that our method achieves higher accuracy for both pupil and gaze estimation tasks compared to the existing solution.
Guangrong Zhao, Yurun Yang, Yiran Shen 0001, Hongkai Wen 0001, Guohao Lan
NeurIPS5
2023 Demo: EV-DMD: a high-speed VLC system
abstract
Visible light communications (VLC) have gained significant attention as a potential solution for the radio spectrum crunch. To achieve high data rates, emerging transmitter devices like 2D digital micro-mirror devices (DMD) have been proposed, offering significantly faster state flipping rates compared to conventional liquid crystalline shutters. However, previous approaches utilizing DMD suffered from a lack of spatial diversity, as they used all micro-mirrors in the same state. This paper introduces EV-DMD, a novel approach that utilizes DMD as a 2D transmitter, working in tandem with an event-based vision (EV) camera. In this method, multiple bit streams are transmitted in parallel through different mirror blocks of the DMD, while an EV camera simultaneously decodes multiple light blocks, enabling a truly 2D high-speed VLC system. To the best of our knowledge, this is the first implementation of a 2D VLC system that achieves an order-of-magnitude improvement in bit rate compared to state-of-the-art solutions.
Guangrong Zhao, Kenuo Xu, Yiran Shen 0001, Chenren Xu, Mahbub Hassan, Wen Hu 0001
SIGCOMM4
2023 VasLine: Realize online detection and augmented NIR using deep learning
Zhongxin Chen, Yiran Shen 0001, Panling Huang, Hengchang Zang, Yongxia Guan
Eng. Appl. Artif. Intell.2
2023 In-Situ Fish Heart-Rate Estimation and Feeding Event Detection Using an Implantable Biologger
abstract
Monitoring of physiology and behavior of marine animals living undisturbed in their natural habitats can provide valuable information about their well-being and response to environmental stressors. We focus on detecting the feeding behavior in predatory fish using implantable biologgers that record and analyze electrocardiogram (ECG) signals. We propose a novel processing pipeline for resource-constrained embedded systems that can infer higher-level information, such as heart-rate and feeding events, from the ECG signals in situ. Our main contributions are in proposing efficient event detection algorithms that can reliably detect fish feeding events from noisy heart-rate data based on the unique statistical properties of feeding-induced changes in the heart-rate. We evaluate our approaches using an in-house biologger that we surgically implant in twelve coral trout fish and use to collect data during an experiment for a period of ten weeks and show that our signal processing pipeline performs well with noisy ECG signals overall. Specifically, our heart-rate estimation algorithm achieves errors of less than one beat per minute even in scenarios where popular algorithms used by domain specialists perform poorly. Furthermore, our feeding detection algorithms offer improved accuracy compared with the state-of-the-art algorithms while requiring significantly reduced computational and energy resources. We implement the proposed heart-rate estimation and feeding detection algorithms on the biologger and evaluate the associated system overhead. The results show that our proposed heart-rate estimation and feeding detection algorithms can run in-situ on the biologger as they demand rather small computational and energy resources that can conveniently be provisioned. This work is an important first step towards developing effective tools for long-term monitoring of high-level parameters pertaining to the health and behavior of marine animals in the wild.
Yiran Shen 0001, Reza Arablouei, Frank de Hoog, Jacques Malan, James Sharp, Sara Shoouri, Timothy D. Clark, Carine Lefevre, Frederieke Kroon, Andrea Severati, Branislav Kusy
IEEE Trans. Mob. Comput.1
2023 EV-LFV: Synthesizing Light Field Event Streams from an Event Camera and Multiple RGB Cameras
abstract
Light field videos captured in RGB frames (RGB-LFV) can provide users with a 6 degree-of-freedom immersive video experience by capturing dense multi-subview video. Despite its potential benefits, the processing of dense multi-subview video is extremely resource-intensive, which currently limits the frame rate of RGB-LFV (i.e., lower than 30 fps) and results in blurred frames when capturing fast motion. To address this issue, we propose leveraging event cameras, which provide high temporal resolution for capturing fast motion. However, the cost of current event camera models makes it prohibitive to use multiple event cameras for RGB-LFV platforms. Therefore, we propose EV-LFV, an event synthesis framework that generates full multi-subview event-based RGB-LFV with only one event camera and multiple traditional RGB cameras. EV-LFV utilizes spatial-angular convolution, ConvLSTM, and Transformer to model RGB-LFV's angular features, temporal features, and long-range dependency, respectively, to effectively synthesize event streams for RGB-LFV. To train EV-LFV, we construct the first event-to-LFV dataset consisting of 200 RGB-LFV sequences with ground-truth event streams. Experimental results demonstrate that EV-LFV outperforms state-of-the-art event synthesis methods for generating event-based RGB-LFV, effectively alleviating motion blur in the reconstructed RGB-LFV.
Zhicheng Lu, Xiaoming Chen 0006, Vera Chung, Tom Weidong Cai, Yiran Shen 0001
IEEE Trans. Vis. Comput. Graph.5
2022 A differential privacy-based classification system for edge computing in IoT
Wanli Xue, Yiran Shen 0001, Chengwen Luo 0001, Weitao Xu, Wen Hu 0001, Aruna Seneviratne
Comput. Commun.2
2022 Event-Stream Representation for Human Gaits Identification Using Deep Neural Networks
abstract
Dynamic vision sensors (event cameras) have recently been introduced to solve a number of different vision tasks such as object recognition, activities recognition, tracking, etc. Compared with the traditional RGB sensors, the event cameras have many unique advantages such as ultra low resources consumption, high temporal resolution and much larger dynamic range. However, these cameras only produce noisy and asynchronous events of intensity changes, i.e., event-streams rather than frames, where conventional computer vision algorithms can't be directly applied. In our opinion the key challenge for improving the performance of event cameras in vision tasks is finding the appropriate representations of the event-streams so that cutting-edge learning approaches can be applied to fully uncover the spatio-temporal information contained in the event-streams. In this paper, we focus on the event-based human gait identification task and investigate the possible representations of the event-streams when deep neural networks are applied as the classifier. We propose new event-based gait recognition approaches basing on two different representations of the event-stream, i.e., graph and image-like representations, and use graph-based convolutional network (GCN) and convolutional neural networks (CNN) respectively to recognize gait from the event-streams. The two approaches are termed as EV-Gait-3DGraph and EV-Gait-IMG. To evaluate the performance of the proposed approaches, we collect two event-based gait datasets, one from real-world experiments and the other by converting the publicly available RGB gait recognition benchmark CASIA-B. Extensive experiments show that EV-Gait-3DGraph achieves significantly higher recognition accuracy than other competing methods when sufficient training samples are available. However, EV-Gait-IMG converges more quickly than graph-based approaches while training and shows good accuracy with only few number of training samples (less than ten). So image-like presentation is preferable when the amount of training data is limited.
Yiran Shen 0001, Bowen Du 0002, Guangrong Zhao, Li-Zhen Cui 0001, Hongkai Wen 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 Event-Based American Sign Language Recognition Using Dynamic Vision Sensor
Yong Wang 0020, Chanying Huang, Yiran Shen 0001
WASA (3)6
2021 Towards a Compressive-Sensing-Based Lightweight Encryption Scheme for the Internet of Things
abstract
Internet of Things (IoT) is flourishing and has penetrated deeply into people's daily life. With the seamless connection to the physical world, IoT provides tremendous opportunities to a wide range of applications. However, potential risks exist when the IoT system collects sensor data and uploads it to the Cloud. The leakage of private data can be severe with curious database administrator or malicious hackers who compromise the Cloud. In this work, we propose Kryptein, a compressive-sensing-based lightweight encryption scheme for Cloud-enabled IoT systems to secure the interaction between the IoT devices and the Cloud. Kryptein supports random compressed encryption, statistical computation over cipher, and accurate raw data decryption. According to our evaluation based on two real datasets, Kryptein provides strong protection to the data. It is 250 times faster than other state-of-the-art systems and incurs 120 times less energy consumption. The performance of Kryptein is also measured on off-the-shelf IoT devices, and the result shows Kryptein can run efficiently on IoT devices. After comparing with other state-of-the-art lightweight ciphers on IoT (Simon and Speck), IoT system with Kryptein is expected to have a much more longevity with about 35 percent extended lifetime. Further, experiments illustrated IoT data variance will not affect Kryptein's accuracy in a long term usage, and Krpytein is also able to support basic analytics tasks like machine learning (e.g., classification).
Wanli Xue, Chengwen Luo 0001, Yiran Shen 0001, Rajib Rana, Guohao Lan, Sanjay K. Jha, Aruna Seneviratne, Wen Hu 0001
IEEE Trans. Mob. Comput.3
2021 LeaD: Learn to Decode Vibration-based Communication for Intelligent Internet of Things
abstract
In this article, we propose, LeaD , a new vibration-based communication protocol to Lea rn the unique patterns of vibration to D ecode the short messages transmitted to smart IoT devices. Unlike the existing vibration-based communication protocols that decode the short messages symbol-wise, either in binary or multi-ary, the message recipient in LeaD receives vibration signals corresponding to bits-groups. Each group consists of multiple symbols sent in a burst and the receiver decodes the group of symbols as a whole via machine learning-based approach. The fundamental behind LeaD is different combinations of symbols (1 s or 0 s) in a group will produce unique and reproducible patterns of vibration. Therefore, decoding in vibration-based communication can be modeled as a pattern classification problem. We design and implement a number of different machine learning models as the core engine of the decoding algorithm of LeaD to learn and recognize the vibration patterns. Through the intensive evaluations on large amount of datasets collected, the Convolutional Neural Network (CNN)-based model achieves the highest accuracy of decoding (i.e., lowest error rate), which is up to 97% at relatively high bits rate of 40 bits/s. While its competing vibration-based communication protocols can only achieve transmission rate of 10 bits/s and 20 bits/s with similar decoding accuracy. Furthermore, we evaluate its performance under different challenging practical settings and the results show that LeaD with CNN engine is robust to poses, distances (within valid range), and types of devices, therefore, a CNN model can be generally trained beforehand and widely applicable for different IoT devices under different circumstances. Finally, we implement LeaD on both off-the-shelf smartphone and smart watch to measure the detailed resources consumption on smart devices. The computation time and energy consumption of its different components show that LeaD is lightweight and can run in situ on low-cost smart IoT devices, e.g., smartwatches, without accumulated delay and introduces only marginal system overhead.
Guangrong Zhao, Bowen Du 0002, Yiran Shen 0001, Zhenyu Lao, Li-Zhen Cui 0001, Hongkai Wen 0001
ACM Trans. Sens. Networks3
2020 Estimating Heart Rate and Detecting Feeding Events of Fish Using an Implantable Biologger
abstract
Monitoring of physiology and behavior of marine animals living undisturbed in their natural habitats can provide valuable data on their well-being and response to environmental stressors. We focus on detection of feeding of predatory fish using implantable biologgers that record electrocardiogram (ECG) signals. We propose a novel processing pipeline for resource-constrained embedded systems that can infer higher-level information, such as heart-rate and feeding events, from the ECG signals. Our main contribution is a lightweight change-detection algorithm, that can reliably detect fish feeding in noisy heart-rate data based on unique statistical properties of feeding-induced changes in heart-rate. We evaluate our approach using an in-house biologger that we surgically implant in twelve coral trouts over a period of ten weeks. We show that our signal processing pipeline performs well with noisy ECG signals overall. Specifically, our heart-rate estimation algorithm achieves errors of less than one beat per minute even in scenarios where popular algorithms used by domain scientists perform poorly. Furthermore, our feeding detection algorithm achieves good accuracy and matches the performance of state-of-the-art algorithms while requiring significantly less memory and computational resources. This work is an important first step towards long-term monitoring of high-level condition and health of marine animals in the wild.
Yiran Shen 0001, Reza Arablouei, Frank de Hoog, Jacques Malan, James Sharp, Sara Shoouri, Timothy D. Clark, Carine Lefevre, Frederieke Kroon, Andrea Severati, Branislav Kusy
IPSN1
2020 Gait-Watch: A Gait-based context-aware authentication system for smart watch via sparse coding
Weitao Xu, Yiran Shen 0001, Chengwen Luo 0001, Jianqiang Li 0001, Wei Li 0058, Albert Y. Zomaya
Ad Hoc Networks2
2020 Design and Implementation of Secret Key Agreement for Platoon-based Vehicular Cyber-physical Systems
abstract
In a platoon-based vehicular cyber-physical system (PVCPS), a lead vehicle that is responsible for managing the platoon’s moving directions and velocity periodically disseminates control messages to the vehicles that follow. Securing wireless transmissions of the messages between the vehicles is critical for privacy and confidentiality of the platoon’s driving pattern. However, due to the broadcast nature of radio channels, the transmissions are vulnerable to eavesdropping. In this article, we propose a cooperative secret key agreement (CoopKey) scheme for encrypting/decrypting the control messages, where the vehicles in PVCPS generate a unified secret key based on the quantized fading channel randomness. Channel quantization intervals are optimized by dynamic programming to minimize the mismatch of keys. A platooning testbed is built with autonomous robotic vehicles, where a TelosB wireless node is used for onboard data processing and multi-hop dissemination. Extensive real-world experiments demonstrate that CoopKey achieves significantly low secret bit mismatch rate in a variety of settings. Moreover, the standard NIST test suite is employed to verify randomness of the generated keys, where the p-values of our CoopKey pass all the randomness tests. We also evaluate CoopKey with an extended platoon size via simulations to investigate the effect of system scalability on performance.
Kai Li 0002, Wei Ni 0001, Yousef Emami, Yiran Shen 0001, Ricardo Severino, David Pereira, Eduardo Tovar
ACM Trans. Cyber Phys. Syst.4
2020 Securing Cyber-Physical Social Interactions on Wrist-Worn Devices
abstract
Since ancient Greece, handshaking has been commonly practiced between two people as a friendly gesture to express trust and respect, or form a mutual agreement. In this article, we show that such physical contact can be used to bootstrap secure cyber contact between the smart devices worn by users. The key observation is that during handshaking, although belonged to two different users, the two hands involved in the shaking events are often rigidly connected, and therefore exhibit very similar motion patterns. We propose a novel key generation system, which harvests motion data during user handshaking from the wrist-worn smart devices such as smartwatches or fitness bands, and exploits the matching motion patterns to generate symmetric keys on both parties. The generated keys can be then used to establish a secure communication channel for exchanging data between devices. This provides a much more natural and user-friendly alternative for many applications, e.g., exchanging/sharing contact details, friending on social networks, or even making payments, since it doesn’t involve extra bespoke hardware, nor require the users to perform pre-defined gestures. We implement the proposed key generation system on off-the-shelf smartwatches, and extensive evaluation shows that it can reliably generate 128-bit symmetric keys just after around 1s of handshaking (with success rate >99%), and is resilient to different types of attacks including impersonate mimicking attacks, impersonate passive attacks, or eavesdropping attacks. Specifically, for real-time impersonate mimicking attacks, in our experiments, the Equal Error Rate (EER) is only 1.6% on average. We also show that the proposed key generation system can be extremely lightweight and is able to run in-situ on the resource-constrained smartwatches without incurring excessive resource consumption.
Yiran Shen 0001, Bowen Du 0002, Weitao Xu, Chengwen Luo 0001, Bo Wei 0003, Li-Zhen Cui 0001, Hongkai Wen 0001
ACM Trans. Sens. Networks1
2019 EV-Gait: Event-Based Robust Gait Recognition Using Dynamic Vision Sensors
abstract
In this paper, we introduce a new type of sensing modality, the Dynamic Vision Sensors (Event Cameras), for the task of gait recognition. Compared with the traditional RGB sensors, the event cameras have many unique advantages such as ultra low resources consumption, high temporal resolution and much larger dynamic range. However, those cameras only produce noisy and asynchronous events of intensity changes rather than frames, where conventional vision-based gait recognition algorithms can’t be directly applied. To address this, we propose a new Event-based Gait Recognition (EV-Gait) approach, which exploits motion consistency to effectively remove noise, and uses a deep neural network to recognise gait from the event streams. To evaluate the performance of EV-Gait, we collect two event-based gait datasets, one from real-world experiments and the other by converting the publicly available RGB gait recognition benchmark CASIA-B. Extensive experiments show that EV-Gait can get nearly 96% recognition accuracy in the real-world settings, while on the CASIA-B benchmark it achieves comparable performance with state-of-the-art RGB-based gait recognition approaches.
Bowen Du 0002, Yiran Shen 0001, Guangrong Zhao, Hongkai Wen 0001
CVPR3
2019 Weakly Supervised Brain Lesion Segmentation via Attentional Representation Learning
Bowen Du 0002, Man Luo 0001, Hongkai Wen 0001, Yiran Shen 0001, Jianfeng Feng
MICCAI (3)5
2019 Labelling issue reports in mobile apps
abstract
Millions of mobile apps have been released to the market. Developers need to maintain these apps so that they can continue to benefit end users, who usually submit issue reports to describe the bugs, the feature requests, and other changes appearing in apps. The labels (e.g. bug, feature request) are important resources to indicate which issue reports should be resolved first or next. According to the investigation, 35.6% of issue reports in top‐17 popular mobile apps are not labelled. Developers have to spend additional time to manually verify each unlabelled issue report so that they can decide to resolve the most important issues. In order to help developers to reduce the workload, in this study, the authors propose a novel approach to automatically tag the unlabelled issue reports. This approach not only computes the similarity between each unlabelled issue report and user reviews related to bugs and features but also calculates the textual similarity scores between each unlabelled issue report and labelled ones. As a result, among all textual similarity measures, this approach using cosine similarity with MCG shows the best performance. Moreover, this approach performs better than the method proposed in the authors' previous study.
Tao Zhang 0001, Haoming Li 0005, Zhou Xu 0003, Rubing Huang, Yiran Shen 0001
IET Softw.6
2019 GaitLock: Protect Virtual and Augmented Reality Headsets Using Gait
abstract
With the fast penetration of commercial Virtual Reality (VR) and Augmented Reality (AR) systems into our daily life, the security issues of those devices have attracted significant interests from both academia and industry. Modern VR/AR systems typically use head-mounted devices (i.e., headsets) to interact with users, and often store private user data, e.g., social network accounts, online transactions or even payment information. This poses significant security threats, since in practice the headset can be potentially obtained and accessed by unauthenticated parties, e.g., identity thieves, and thus cause catastrophic breach. In this paper, we propose a novel GaitLock system, which can reliably authenticate users using their gait signatures. Our system doesn't require extra hardware, e.g., fingerprint sensors or retina scanners, but only uses the on-board inertial measurement units (IMUs) equipped in almost all mainstream VR/AR headsets to authenticate the legitimate users from intruders, by simply asking them to walk a few steps. To achieve that, we propose a new gait recognition model Dynamic-SRC, which combines the strength of Dynamic Time Warping (DTW) and Sparse Representation Classifier (SRC), to extract unique gait patterns from the inertial signals during walking. We implement GaitLock on Google Glass (a typical AR headset), and extensive experiments show that GaitLock outperforms the state-of-the-art systems significantly in recognition accuracy (> 98 percent success in 5 steps), and is able to run in-situ on the resource-constrained VR/AR headsets without incurring high energy cost.
Yiran Shen 0001, Hongkai Wen 0001, Chengwen Luo 0001, Weitao Xu, Tao Zhang 0001, Wen Hu 0001, Daniela Rus
IEEE Trans. Dependable Secur. Comput.1
2019 Predictable Privacy-Preserving Mobile Crowd Sensing: A Tale of Two Roles
abstract
The rise of mobile crowd sensing has brought privacy issues into a sharp view. In this paper, our goal is to achieve the predictable privacy-preserving mobile crowd sensing, which we envision to have the capability to quantify the privacy protections, and simultaneously allowing application users to predict the utility loss at the same time. TheSalusalgorithm is first proposed to protect the private data against the data reconstruction attacks. To understand privacy protection, we quantify the privacy risks in terms of private data leakage under reconstruction attacks. To predict the utility, we provide accurate utility predictions for various crowd sensing applications using Salus. The risk assessments can be generally applied to different type of sensors on the mobile platform, and the utility prediction can also be used to support various applications that use data aggregators such as average, histogram, and classifiers. Finally, we propose and implement the$P^{3}$application framework. Both measurement results using online datasets and real-world case studies show that the$P^{3}$provides accurate risk assessments and utility estimations, which makes it a promising framework to support future privacy-preserving mobilecrowd sensing applications.
Chengwen Luo 0001, Wanli Xue, Yiran Shen 0001, Jianqiang Li 0001, Wen Hu 0001, Alex X. Liu
IEEE/ACM Trans. Netw.4
2018 Deepauth: in-situ authentication for smartwatches via deeply learned behavioural biometrics
abstract
This paper proposes DeepAuth, an in-situ authentication framework that leverages the unique motion patterns when users entering passwords as behavioural biometrics. It uses a deep recurrent neural network to capture the subtle motion signatures during password input, and employs a novel loss function to learn deep feature representations that are robust to noise, unseen passwords, and malicious imposters even with limited training data. DeepAuth is by design optimised for resource constrained platforms, and uses a novel split-RNN architecture to slim inference down to run in real-time on off-the-shelf smartwatches. Extensive experiments with real-world data show that DeepAuth outperforms the state-of-the-art significantly in both authentication performance and cost, offering real-time authentication on a variety of smartwatches.
Xiaoxuan Lu 0001, Bowen Du 0002, Peijun Zhao, Hongkai Wen 0001, Yiran Shen 0001, Andrew Markham, Agathoniki Trigoni
UbiComp5
2018 Shake-n-Shack: Enabling Secure Data Exchange Between Smart Wearables via Handshakes
abstract
Since ancient Greece, handshaking has been commonly practiced between two people as a friendly gesture to express trust and respect, or form a mutual agreement. In this paper, we show that such physical contact can be used to bootstrap secure cyber contact between the smart devices worn by users. The key observation is that during handshaking, although belonged to two different users, the two hands involved in the shaking events are often rigidly connected, and therefore exhibit very similar motion patterns. We propose a novel Shake-n-Shack system, which harvests motion data during user handshaking from the wrist worn smart devices such as smartwatches or fitness bands, and exploits the matching motion patterns to generate symmetric keys on both parties. The generated keys can be then used to establish a secure communication channel for exchanging data between devices. This provides a much more natural and user-friendly alternative for many applications, e.g., exchanging/sharing contact details, friending on social networks, or even making payments, since it doesn't involve extra bespoke hardware, nor require the users to perform pre-defined gestures. We implement the proposed Shake-n-Shack1system on off-the-shelf smartwatches, and extensive evaluation shows that it can reliably generate 128-bit symmetric keys just after around 1s of handshaking (with success rate >99%), and is resilient to real-time mimicking attacks: in our experiments the Equal Error Rate (EER) is only 1.6% on average. We also show that the proposed Shake-n-Shack system can be extremely lightweight, and is able to run in-situ on the resource-constrained smartwatches without incurring excessive resource consumption.
Yiran Shen 0001, Fengyuan Yang 0001, Bowen Du 0002, Weitao Xu, Chengwen Luo 0001, Hongkai Wen 0001
PerCom1
2018 Privacy-preserving sparse representation classification in cloud-enabled mobile applications
Yiran Shen 0001, Chengwen Luo 0001, Dan Yin, Hongkai Wen 0001, Daniela Rus, Wen Hu 0001
Comput. Networks1
2018 HealCam: Energy-efficient and privacy-preserving human vital cycles monitoring on camera-enabled smart devices
Qing Yang 0009, Yiran Shen 0001, Fengyuan Yang 0001, Jianpei Zhang, Wanli Xue, Hongkai Wen 0001
Comput. Networks2
2018 Generate domain-specific sentiment lexicon for review sentiment analysis
Hongyu Han, Jianpei Zhang, Jing Yang 0010, Yiran Shen 0001, Yongshi Zhang
Multim. Tools Appl.4
2018 Sensor-Assisted Multi-View Face Recognition System on Smart Glass
abstract
Face recognition is a hot research topic with a variety of application possibilities, including video surveillance and mobile payment. It has been well researched in traditional computer vision community. However, new research issues arise when it comes to resource constrained devices, such as smart glasses, due to the overwhelming computation and energy requirements of the accurate face recognition methods. In this paper, we propose a robust and efficient sensor-assisted face recognition system on smart glasses by exploring the power of multimodal sensors including the camera and Inertial Measurement Unit (IMU) sensors. The system is based on a novel face recognition algorithm, namely Multi-view Sparse Representation Classification (MVSRC), by exploiting the prolific information among multi-view face images. To improve the efficiency of MVSRC on smart glasses, we propose two novel sampling optimization strategies using the less expensive inertial sensors. Our evaluations on public and private datasets show that the proposed method is up to 10 percent more accurate than the state-of-the-art multi-view face recognition methods while its computation cost is the same order as an efficient benchmark method (e.g., Eigenfaces). Finally, extensive real-world experiments show that our proposed system improves recognition accuracy by up to 15 percent while achieving the same level of system overhead compared to the existing face recognition system (OpenCV algorithms) on smart glasses.
Weitao Xu, Yiran Shen 0001, Neil W. Bergmann, Wen Hu 0001
IEEE Trans. Mob. Comput.2
2017 Normal direction local binary pattern for fragment reconstruction
abstract
Fragment reconstruction aims to restore broken images and documents via matching spatial adjacent fragments. As the existing solutions in the literature still remain problematic, we present a novel feature descriptor, Normal Direction Local Binary Pattern (termed as ND-LBP), for document/image fragment matching. ND-LBP is based on the conventional LBP descriptor, however, it outstands LBP by introducing new features derived from shapes and contents of fragments to promote its discrimination. With normal direction operation, ND-LBP is rotation-invariant, and thus could effectively and efficiently match fragments of arbitrary orientation. According to our extensive evaluations on real world datasets, the fragment reconstruction approach with ND-LBP feature has high precision, robustness and efficiency, and outperforms existing features.
Yiran Shen 0001
ICME4
2017 Learn to Recognise: Exploring Priors of Sparse Face Recognition on Smartphones
abstract
Face recognition is one of the important components of many smart devices apps, e.g., face unlocking, people tagging and games on smart phones, tablets, or smart glasses. Sparse Representation Classification (SRC) is a state-of-the-art face recognition algorithm, which has been shown to outperform many classical face recognition algorithms in OpenCV, e.g., Eigenface algorithm. The success of SRC is due to its use of 21 optimization, which makes SRC robust to noise and occlusions. Since 21 optimization is computationally intensive, SRC uses random projection matrices to reduce the dimension of the 21 problem. However, random projection matrices do not give consistent classification accuracy as they ignored the prior knowledge of the training set. In this paper, we propose to exploit the prior knowlege of the training set to improve the recognition accuracy. It first learns the optimized projection matrix from the training set to produce consistent recognition performance then applies 21-based classification based on the group sparsity structure of SRC to further improve the recognition accuracy. Our evaluations, based on publicly available databases and real experiment, show that face recognition using optimized projection matrix is 8-17 percent more accurate than its random counterpart and Eigenface algorithm, and the recognition accuracy can be further improved by up to 5 percent by exploiting group sparsity structure. Furthermore, the optimized projection matrix does not have to be re-calculated even if new faces are added to the training set. We implement the SRC with optimized projection matrix on Android smartphones and find that the computation of residuals in SRC is a severe bottleneck, taking up 85-90 percent of the computation time. To address this problem, we propose a method to compute the residuals approximately, which is 50 times faster with little sacrificing recognition accuracy. Lastly, we demonstrate the feasibility of our new algorithm by the implementation and evaluation of a new face unlocking app and show its robustness to variation of poses, facial expressions, lighting changes, and occlusions.
Yiran Shen 0001, Mingrui Yang, Bo Wei 0003, Chun Tung Chou, Wen Hu 0001
IEEE Trans. Mob. Comput.1
2016 Sensor-Assisted Face Recognition System on Smart Glass via Multi-View Sparse Representation Classification
abstract
Face recognition is one of the most popular research problems on various platforms. New research issues arise when it comes to resource constrained devices, such as smart glasses, due to the overwhelming computation and energy requirements of the accurate face recognition methods. In this paper, we propose a robust and efficient sensor-assisted face recognition system on smart glasses by exploring the power of multimodal sensors including the camera and Inertial Measurement Unit (IMU) sensors. The system is based on a novel face recognition algorithm, namely Multi-view Sparse Representation Classification (MVSRC), by exploiting the prolific information among multi-view face images. To improve the efficiency of MVSRC on smart glasses, we propose a novel sampling optimization strategy using the less expensive inertial sensors. Our evaluations on public and private datasets show that the proposed method is up to 10% more accurate than the state-of-the-art multi-view face recognition methods while its computation cost is in the same order as an efficient benchmark method (e.g., Eigenfaces). Finally, extensive real-world experiments show that our proposed system improves recognition accuracy by up to 15% while achieving the same level of system overhead compared to the existing face recognition system (OpenCV algorithms) on smart glasses.
Weitao Xu, Yiran Shen 0001, Neil W. Bergmann, Wen Hu 0001
IPSN2
2016 Real-Time and Robust Compressive Background Subtraction for Embedded Camera Networks
abstract
Real-time target tracking is an important service provided by embedded camera networks. The first step in target tracking is to extract the moving targets from the video frames, which can be realised by using background subtraction. For a background subtraction method to be useful in embedded camera networks, it must be both accurate and computationally efficient because of the resource constraints on embedded platforms. This makes many traditional background subtraction algorithms unsuitable for embedded platforms because they use complex statistical models to handle subtle illumination changes. These models make them accurate but the computational requirement of these complex models is often too high for embedded platforms. In this paper, we propose a new background subtraction method which is both accurate and computationally efficient. We propose a baseline version which uses luminance only and then extend it to use colour information. The key idea is to use random projection matrics to reduce the dimensionality of the data while retaining most of the information. By using multiple datasets, we show that the accuracy of our proposed background subtraction method is comparable to that of the traditional background subtraction methods. Moreover, to show the computational efficiency of our methods is not platform specific, we implement it on various platforms. The real implementation shows that our proposed method is consistently better and is up to six times faster, and consume significantly less resources than the conventional approaches. Finally, we demonstrated the feasibility of the proposed method by the implementation and evaluation of an end-to-end real-time embedded camera network target tracking application.
Yiran Shen 0001, Wen Hu 0001, Mingrui Yang, Junbin Liu, Bo Wei 0003, Simon Lucey, Chun Tung Chou
IEEE Trans. Mob. Comput.1
2015 Opportunistic Radio Assisted Navigation for Autonomous Ground Vehicles
abstract
Navigating autonomous ground vehicles with visual sensors has many advantages - it does not rely on global maps, yet is accurate and reliable even in GPS-denied environments. However, due to the limitation of the camera field of view, one typically has to record a large number of visual experiences for practical navigation. In this paper, we explore new avenues in linking together visual experiences, by opportunistically harvesting and sharing a variety of radio signals emitted by surrounding stationary access points and mobile devices. We propose a novel navigation approach, which exploits side-channel information of co-location to thread up visually-separated experiences with short exploration phases. The proposed approach empowers users to trade travel time for manual navigation effort, allowing them to choose the itinerary that best serves their needs. We evaluate the proposed approach with data collected from a typical urban area, and show that it achieves much better navigation performance in both reach ability and cost, comparing with the state of the arts that only use visual information.
Hongkai Wen 0001, Yiran Shen 0001, Savvas Papaioannou, Winston Churchill, Agathoniki Trigoni, Paul Newman 0001
DCOSS2
2015 Poster: An Online Approach for Gait Recognition on Smart Glasses
abstract
With the fast development and increasing population of the wearable devices involves in our daily life, the security of the privacy information on those devices is attracting significant attentions. One of the possible solution is to enable the devices to recognise the real owner with authentication system. Biometrics recognition is popular used for authentication systems. The biometrics used including faces, fingerprints, gait cycles and etc. Using gait cycles as the criteria for identities recognition is superior than other biometrics as the gait information can be collected by the IMU sensors which are most popular embedded on portable devices and they cannot be reproduced by the invaders. We propose, Securitas, the continuous authentication system exploits the information from IMU sensors on the smart glasses to distinguish different wearers.
Yiran Shen 0001, Chengwen Luo 0001, Weitao Xu, Wen Hu 0001
SenSys1
2015 Poster: Robust and Efficient Sensor-assisted Face Recognition System on Smart Glass
abstract
Face recognition is one of the most popular research problems on various platforms. New research issues arise when it comes to resource constrained devices, such as smart glasses, due to the overwhelming computation and energy requirements of the accurate face recognition methods. In this paper, we have prototyped a robust and efficient sensor-assisted face recognition system on smart glasses by exploring the power of multimodal sensors including the camera and Inertial Measurement Unit (IMU) sensors. Evaluation shows that the prototyped system is up to 10% more accurate than the state-of-the-art face recognition methods while its computational cost is in the same order as an efficient benchmark method (e.g., Eigenface).
Weitao Xu, Yiran Shen 0001, Neil W. Bergmann, Wen Hu 0001
SenSys2
2014 Face recognition on smartphones via optimised sparse representation classification
Yiran Shen 0001, Wen Hu 0001, Mingrui Yang, Bo Wei 0003, Simon Lucey, Chun Tung Chou
IPSN1
2013 Projection matrix optimisation for compressive sensing based applications in embedded systems
abstract
The information-preserving sampling properties of compressive sensing have found a number of successful applications, such as sensor scheduling, localisation and tracking to deal with the resource constraints of the embedded systems. In this paper, we investigate an approach to improve the performance of compressive sensing applications through a novel strategy for optimising the projection matrix. We formulate the projection matrix optimisation problem and apply greedy algorithm to solve the optimisation problem efficiently. We evaluate the proposed approach by an emerging background subtraction method designed specifically for the embedded systems and show the proposed approach outperforms existing approaches significantly with little overhead.
Yiran Shen 0001, Wen Hu 0001, Mingrui Yang, Bo Wei 0003, Chun Tung Chou
SenSys1
2013 Real-time classification via sparse representation in acoustic sensor networks
abstract
Acoustic Sensor Networks (ASNs) have a wide range of applications in natural and urban environment monitoring, as well as indoor activity monitoring. In-network classification is critically important in ASNs because wireless transmission costs several orders of magnitude more energy than computation. The main challenges of in-network classification in ASNs include effective feature selection, intensive computation requirement and high noise levels. To address these challenges, we propose a sparse representation based feature-less, low computational cost, and noise resilient framework for in-network classification in ASNs. The key component of Sparse Approximation based Classification (SAC), ℓ1 minimization, is a convex optimization problem, and is known to be computationally expensive. Furthermore, SAC algorithms assumes that the test samples are a linear combination of a few training samples in the training sets. For acoustic applications, this results in a very large training dictionary, making the computation infeasible to be performed on resource constrained ASN platforms. Therefore, we propose several techniques to reduce the size of the problem, so as to fit SAC for in-network classification in ASNs. Our extensive evaluation using two real-life datasets (consisting of calls from 14 frog species and 20 cricket species respectively) shows that the proposed SAC framework outperforms conventional approaches such as Support Vector Machines (SVMs) and k-Nearest Neighbor (kNN) in terms of classification accuracy and robustness. Moreover, our SAC approach can deal with multi-label classification which is common in ASNs. Finally, we explore the system design spaces and demonstrate the real-time feasibility of the proposed framework by the implementation and evaluation of an acoustic classification application on an embedded ASN testbed.
Bo Wei 0003, Mingrui Yang, Yiran Shen 0001, Rajib Rana, Chun Tung Chou, Wen Hu 0001
SenSys3
2012 Efficient background subtraction for tracking in embedded camera networks
abstract
Background subtraction is often the first step in many computer vision applications such as object localisation and tracking. It aims to segment out moving parts of a scene that represent object of interests. In the field of computer vision, researchers have dedicated their efforts to improve the robustness and accuracy of such segmentations but most of their methods are computationally intensive, making them non-viable options for our targeted embedded camera platform whose energy and processing power is significantly more constrained. To address this problem as well as maintain an acceptable level of performance, we introduce Compressive Sensing (CS) to the widely used Mixture of Gaussian to create a new background subtraction method. The results show that our method not only can decrease the computation significantly (a factor of 7 in a DSP setting) but remains comparably accurate.
Yiran Shen 0001, Wen Hu 0001, Mingrui Yang, Junbin Liu, Chun Tung Chou
IPSN1
2012 Efficient background subtraction for real-time tracking in embedded camera networks
abstract
Background subtraction is often the first step of many computer vision applications. For a background subtraction method to be useful in embedded camera networks, it must be both accurate and computationally efficient because of the resource constraints on embedded platforms. This makes many traditional background subtraction algorithms unsuitable for embedded platforms because they use complex statistical models to handle subtle illumination changes. These models make them accurate but the computational requirement of these complex models is often too high for embedded platforms. In this paper, we propose a new background subtraction method which is both accurate and computational efficient. The key idea is to use compressive sensing to reduce the dimensionality of the data while retaining most of the information. By using multiple datasets, we show that the accuracy of our proposed background subtraction method is comparable to that of the traditional background subtraction methods. Moreover, real implementation on an embedded camera platform shows that our proposed method is at least 5 times faster, and consumes significantly less energy and memory resources than the conventional approaches. Finally, we demonstrated the feasibility of the proposed method by the implementation and evaluation of an end-to-end real-time embedded camera network target tracking application.
Yiran Shen 0001, Wen Hu 0001, Junbin Liu, Mingrui Yang, Bo Wei 0003, Chun Tung Chou
SenSys1