EDBT 2026 Demo / reviewers in the wild / expert
Hongkai Wen 0001
dblp:55/7703
· DBLP profile ↗
70ranked-venue papers
6as first author
31since 2021 · last 2026
0000-0003-1159-090XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 17 since 2021Computer networks · 21 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 6 since 2021Databases, data management, data science and information retrieval · 6 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Systems, architecture and hardware · 3Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TRACE: Trajectory Recovery with State Propagation Diffusion for Urban MobilityabstractHigh-quality GPS trajectories are essential for location-based web services and smart city applications, including navigation, ride-sharing and delivery. However, due to low sampling rates and limited infrastructure coverage during data collection, real-world trajectories are often sparse and feature unevenly distributed location points. Recovering these trajectories into dense and continuous forms is essential but challenging, given their complex and irregular spatio-temporal patterns. In this paper, we introduce a novel diffusion model for TRA jectory rEC overy named TRACE, which reconstruct dense and continuous trajectories from sparse and incomplete inputs. At the core of TRACE, we propose a State Propagation Diffusion Model (SPDM), which integrates a novel memory mechanism, so that during the denoising process, TRACE can retain and leverage intermediate results from previous steps to effectively reconstruct those hard-to-recover trajectory segments. Extensive experiments on multiple real-world datasets show that TRACE outperforms the state-of-the-art, offering >26% accuracy improvement without significant inference overhead. Our work strengthens the foundation for mobile and web-connected location services, advancing the quality and fairness of data-driven urban applications. Code is available at:~ https://github.com/JinmingWang/TRACE Hai Wang 0019, Hongkai Wen 0001, Geyong Min, Man Luo 0001 |
WWW | 3 |
| 2026 | CityVerse: A Unified Data Framework for Evaluating Large Language Models in Urban ComputingabstractLarge Language Models (LLMs) show remarkable potential for urban computing, from spatial reasoning to predictive analytics. However, evaluating LLMs across diverse urban tasks faces two critical challenges: lack of unified platforms for consistent multi-source data access, and fragmented task definitions that hinder fair comparison. To address these challenges, we present CityVerse, the first unified platform integrating multi-source urban data, capability-based task taxonomy, and dynamic simulation for systematic LLM evaluation in urban contexts. CityVerse provides: (i) coordinate-based Data APIs unifying ten categories of urban data—including spatial features, temporal dynamics, demographics, and multi-modal imagery-with over 38 million curated records; (ii) Task APIs organizing 43 urban computing tasks into a four-level cognitive hierarchy: Perception, Spatial Understanding, Reasoning Prediction, and Decision Interaction, enabling standardized evaluation across capability levels; (iii) an interactive visualization frontend supporting real-time data retrieval, multi-layer display, and simulation replay for intuitive exploration and validation. We validate the platform's effectiveness through evaluations on mainstream LLMs across representative tasks, demonstrating its capability to support reproducible and systematic assessment. CityVerse provides a reusable foundation for advancing LLMs and multi-task approaches in the urban computing domain. The framework is publicly available at https://julietjobs.github.io/City_Verse/. Yaqiao Zhu 0001, Hongkai Wen 0001, Mark H. Birkin, Man Luo 0001 |
WWW | 2 |
| 2026 | HALO: Hierarchical Reinforcement Learning for Large-Scale Adaptive Traffic Signal ControlabstractAdaptive traffic signal control (ATSC) is essential for mitigating urban congestion in modern smart cities, where traffic infrastructure is evolving into interconnected Web-of-Things (WoT) environments with thousands of sensing-and-control nodes. However, existing methods face a critical scalability-coordination tradeoff: centralized approaches optimize global objectives but become computationally intractable at city scale, while decentralized multi-agent methods scale efficiently yet lack network-level coherence, resulting in suboptimal performance. In this paper, we present HALO, a hierarchical reinforcement learning framework that addresses this tradeoff for large-scale ATSC. HALO decouples decision-making into two levels: a high-level global guidance policy employs Transformer-LSTM encoders to model spatio-temporal dependencies across the entire network and broadcast compact guidance signals, while low-level local intersection policies execute decentralized control conditioned on both local observations and global context. To ensure better alignment of global-local objectives, we introduce an adversarial goal-setting mechanism where the global policy proposes challenging-yet-feasible network-level targets that local policies are trained to surpass, fostering robust coordination. We evaluate HALO extensively on multiple standard benchmarks, and a newly constructed large-scale Manhattan-like network with 2,668 intersections under real-world traffic patterns, including peak transitions, adverse weather and holiday surges. Results demonstrate HALO shows competitive performance and becomes increasingly dominant as network complexity grows across small-scale benchmarks, while delivering the strongest performance in all large-scale regimes, offering up to 6.8% lower average travel time and 5.0% lower average delay than the best state-of-the-art. Yaqiao Zhu 0001, Hongkai Wen 0001, Geyong Min, Man Luo 0001 |
WWW | 2 |
| 2026 | Road Network Generation From Noisy Trajectories With Graph Latent Diffusion ModelsabstractModern navigation systems and logistics services rely heavily on digital road networks, which can be accurately constructed from crowd-sourced GPS trajectories. Yet, most existing trajectory-based road network generators project GPS coordinates to 2D grid, followed by semantic segmentation to obtain road heatmap, and further processed to road network graph via heuristic processing. The data through this pipeline involve multiple cross-modal heuristic conversions, which cause serious information loss and harm overall robustness. In this paper, we propose a novel end-to-end learnable deep learning approach, named GraphWalker, to generate accurate road network graphs directly from raw trajectories with minimal information loss. In particular, GraphWalker is a latent diffusion model (LDM) consisting of two key components: theTrajectory-to-Walks Diffusion Transformer(T2W-DiT) generateswalkswith trajectories as conditions, and theWalks-to-Graph Variational Auto-Encoder(W2G-VAE) reconstructs road network graphs from walks. GraphWalker keeps the data in geographical space throughout the pipeline to retain all information within raw trajectories. Additionally, the end-to-end learnable design removes the heuristic processing and enhances the robustness and generalizability. Extensive experiments on multiple datasets demonstrate that the proposed GraphWalker can effectively generate high-quality road networks from noisy and sparse trajectories, showcasing significant improvements of up to 55.43% over state-of-the-art. Hongkai Wen 0001, Geyong Min, Man Luo 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2025 | DPaI: Differentiable Pruning at Initialization with Node-Path Balance PrincipleabstractPruning at Initialization (PaI) is a technique in neural network optimization characterized by the proactive elimination of weights before the network's training on designated tasks. This innovative strategy potentially reduces the costs for training and inference, significantly advancing computational efficiency. A key factor leading to PaI's effectiveness is that it considers the saliency of weights in an untrained network, and prioritizes the trainability and optimization potential of the pruned subnetworks. Recent methods can effectively prevent the formation of hard-to-optimize networks, e.g. through iterative adjustments at each network layer. However, this way often results in large-scale discrete optimization problems, which could make PaI further challenging. This paper introduces a novel method, called DPaI, that involves a differentiable optimization of the pruning mask. DPaI adopts a dynamic and adaptable pruning process, allowing easier optimization processes and better solutions. More importantly, our differentiable formulation enables readily use of the existing rich body of efficient gradient-based methods for PaI. Our empirical results demonstrate that DPaI significantly outperforms current state-of-the-art PaI methods on various architectures, such as Convolutional Neural Networks and Vision-Transformers. Code is available at https://github.com/QuanNguyen-Tri/DPaI.git Lichuan Xiang, Quan Nguyen-Tri, Lan-Cuong Nguyen, Khoat Than, Long Tran-Thanh, Hongkai Wen 0001 |
ICLR | 7 |
| 2025 | FlexControl: Computation-Aware Conditional Control with Differentiable Router for Text-to-Image GenerationabstractSpatial conditioning control offers a powerful way to guide diffusion-based generative models. Yet, most implementations (e.g., ControlNet) rely on ad-hoc heuristics to choose which network blocks to control — an approach that varies unpredictably with different tasks. To address this gap, we propose FlexControl, a novel framework that equips all diffusion blocks with control signals during training and employs a trainable gating mechanism to dynamically select which control signal to activate at each denoising step. By introducing a computation-aware loss, we can encourage the control signal to activate only when it benefits the generation quality. By eliminating manual control unit selection, FlexControl enhances adaptability across diverse tasks and streamlines the design pipeline with computation-aware training loss in an end-to-end training manner. Through comprehensive experiments on both UNet and DiT architectures on different control methods, we show that our method can upgrade existing controllable generative models in certain key aspects of interest. As evidenced by both quantitative and qualitative evaluations, FlexControl preserves or enhances image fidelity while also reducing computational overhead by selectively activating the most relevant blocks to control. These results underscore the potential of a flexible, data-driven approach for controlled diffusion and open new avenues for efficient generative model design. The code will soon be available at https://github.com/Daryu-Fan/FlexControl. Lichuan Xiang, Kaicheng Zhou, Hongkai Wen 0001 |
ICML | 5 |
| 2025 | High-Fidelity Road Network Generation with Latent Diffusion ModelsabstractRoad networks are the vein of modern cities. Yet, maintaining up-to-date and accurate road network information is a persistent challenge, especially in areas with rapid urban changes or limited surveying resources. Crowdsourced trajectories, e.g., from GPS records collected by mobile devices and vehicles, have emerged as a powerful data source for continuously mapping the urban areas. However, the inherent noise, irregular and often sparse sampling rates, and the vast variability in movement patterns make the problem of road network generation from trajectories a non-trivial task. Existing methods often approach this from an appearance-based perspective: they typically render trajectories as 2D density maps and then employ heuristic algorithms to extract road networks - leading to inevitable information loss and thus poor performance especially when trajectories are sparse or ambiguities present, e.g. flyovers. In this paper, we propose a novel approach, called GraphWalker, to generate high-fidelity road network graphs from raw trajectories in an end-to-end manner. We achieve this by designing a bespoke latent diffusion transformer T2W-DiT, which treats input trajectories as generation conditions, and gradually denoises samples from a latent space to obtain the corresponding walks on the underlying road network graph - then assemble them together as the final road network. Extensive experiments on multiple datasets demonstrate the proposed GraphWalker can effectively generate high quality road networks from noisy and sparse trajectories, showcasing significant improvements over state-of-the-art. Hongkai Wen 0001, Geyong Min, Man Luo 0001 |
IJCAI | 2 |
| 2025 | BGM: Demand Prediction for Expanding Bike-Sharing Systems with Dynamic Graph ModelingabstractAccurate demand prediction is crucial for the equitable and sustainable expansion of bike-sharing systems, which help reduce urban congestion, promote low-carbon mobility, and improve transportation access in underserved areas. However, expanding these systems presents societal challenges, particularly in ensuring fair resource distribution and operational efficiency. A major hurdle is the difficulty of demand prediction at new stations, which lack historical usage data and are heavily influenced by the existing network. Additionally, new stations dynamically reshape demand patterns across time and space, complicating efforts to balance supply and accessibility in evolving urban environments. Existing methods model relationships between new and existing stations but often assume static patterns, overlooking how new stations reshape demand dynamics over time and space. To tackle these challenges, we propose a novel demand prediction framework for expanding bike-sharing systems, namely BGM, which leverages dynamic graph modeling to capture the evolving inter-station correlations while accounting for spatial and temporal heterogeneity. Specifically, we develop a knowledge transfer approach that studies the embeddings transformation across existing and new stations through a learnable orthogonal mapping matrix. We further design a gated selecting vector-based feature fusion mechanism to integrate the transferred embeddings and the intrinsic features of stations for precise predictions. Experiments on real-world bike-sharing data demonstrate that BGM outperforms existing methods. Hongkai Wen 0001, Man Luo 0001 |
IJCAI | 2 |
| 2025 | Shortcutting Pre-trained Flow Matching Diffusion Models is Almost Free LunchabstractWe present an ultra-efficient post-training method for shortcutting large-scale pre-trained flow matching diffusion models into efficient few-step samplers, enabled by novel velocity field self-distillation. While shortcutting in flow matching, originally introduced by shortcut models, offers flexible trajectory-skipping capabilities, it requires a specialized step-size embedding incompatible with existing models unless retraining from scratch—a process nearly as costly as pretraining itself.
Our key contribution is thus imparting a more aggressive shortcut mechanism to standard flow matching models (e.g., Flux), leveraging a unique distillation principle that obviates the need for step-size embedding. Working on the velocity field rather than sample space and learning rapidly from self-guided distillation in an online manner, our approach trains efficiently, e.g., producing a 3-step Flux <1 A100 day. Beyond distillation, our method can be incorporated into the pretraining stage itself, yielding models that inherently learn efficient, few-step flows without compromising quality. This capability also enables, to our knowledge, the first few-shot distillation method (e.g., 10 text-image pairs) for dozen-billion-parameter diffusion models, delivering state-of-the-art performance at almost free cost. Qianli Chen, Lichuan Xiang, Hongkai Wen 0001 |
NeurIPS | 6 |
| 2025 | MixSignGraph: A Sign Sequence is Worth Mixed Graphs of NodesabstractRecent advances in sign language research have benefited from CNN-based backbones, which are primarily transferred from traditional computer vision tasks (\eg object detection, image recognition). However, these CNN-based backbones usually excel at extracting features like contours and texture, but may struggle with capturing sign-related features.
To capture such sign-related features, SignGraph model extracts the cross-region sign features by building the Local Sign Graph (LSG) module and the Temporal Sign Graph (TSG) module. However, we emphasize that although capturing cross-region dependencies can improve sign language performance, it may degrade the representation quality of local regions.
To mitigate this, we introduce MixSignGraph, which represents sign sequences as a group of mixed graphs for feature extraction.
Specifically, besides the LSG module and TSG module that model the intra-frame and inter-frame cross-regions features, we design a simple yet effective Hierarchical Sign Graph (HSG) module, which enhances local region representations following the extraction of cross-region features, by aggregating the same-region features from different-granularity feature maps of a frame, \ie to boost discriminative local features.
In addition, to further improve the performance of gloss-free sign language task, we propose a simple yet counter-intuitive Text-based CTC Pre-training (TCTC) method, which generates pseudo gloss labels from text sequences for model pre-training. Extensive experiments conducted on the current five sign language datasets demonstrate that MixSignGraph surpasses the most current models on multiple sign language tasks across several datasets, without relying on any additional cues. Code and models are available at: \href{https://github.com/gswycf/SignLanguage}{\textcolor{blue}{https://github.com/gswycf/SignLanguage}}. Shiwei Gan, Yafeng Yin 0002, Zhiwei Jiang 0001, Lei Xie 0004, Sanglu Lu, Hongkai Wen 0001 |
NeurIPS | 6 |
| 2025 | DCA: Graph-Guided Deep Embedding Clustering for Brain AtlasesabstractBrain atlases are essential for reducing the dimensionality of neuroimaging data and enabling interpretable analysis. However, most existing atlases are predefined, group-level templates with limited flexibility and resolution. We present Deep Cluster Atlas (DCA), a graph-guided deep embedding clustering framework for generating individualized, voxel-wise brain parcellations. DCA combines a pretrained voxel-level fMRI autoencoder with spatially regularized deep clustering to produce functionally coherent and spatially contiguous regions. Our method supports flexible control over resolution and anatomical scope, and generalizes to arbitrary brain structures. We further introduce a standardized benchmarking platform for atlas evaluation, using multiple large-scale fMRI datasets. Across multiple datasets and scales, DCA outperforms state-of-the-art atlases, improving functional homogeneity by 98.8% and silhouette coefficient by 29%, and achieves superior performance in downstream tasks. Furthermore, atlases demonstrate heterogeneous performance across various tasks and an atlas derived from a fine-tuned model yields superior results for its specific application. Codes are available at https://github.com/ncclab-sustech/DCA. Kaining Peng, Jingsheng Tang, Hongkai Wen 0001, Quanying Liu |
NeurIPS | 4 |
| 2025 | Robust spatio-temporal demand prediction for bike-sharing systems with dynamic hierarchical structureabstractAbstract Bike-sharing systems have gained popularity as a solution for short-distance urban travel, highlighting the need for precise demand prediction. Most existing methods predict bike-sharing demand across all regions without adequately accounting for spatial heterogeneity, meaning they overlook that different regions may have distinct and skewed demand distributions. Moreover, these models frequently miss the temporal heterogeneity introduced by varying demand patterns, as they tend to model temporal correlations with a uniform parameter set across all time periods, failing to account for the dynamic fluctuations in demand over time. To address these shortcomings, we introduce a novel Robust Spatio-Temporal Demand Prediction (RST) model that enhances bike-sharing demand prediction. The proposed approach employs a convolutional spatio-temporal graph, which integrates the features from multiple sources, including Points of Interest (POIs), road networks, and weather. Upon that, it proposes a scheme to classify the stations based on their traffic irregularity. The proposed scheme captures the traffic dependencies between stations by generating an embedding sequence from the historical data and ignoring correlations to the low-relevant stations, ensuring accurate predictions. Furthermore, the model introduces a dynamic hierarchical structure, where an additional retraining layer is activated only for stations with underperforming predictions, leveraging data from similar stations to refine and improve accuracy. This two-layered hierarchical structure ensures robustness and adaptability in complex scenarios such as different weather conditions or events. Our methodology sets a new standard in demand prediction for bike-sharing, demonstrating significant improvements in prediction accuracy on real-world datasets from New York City’s Citi-Bike and Beijing’s BJ-Bike. The results validate our model’s superior performance, offering a scalable and effective strategy for enhancing urban mobility. Bowen Du 0002, Man Luo 0001, Hongkai Wen 0001 |
CCF Trans. Pervasive Comput. Interact. | 4 |
| 2025 | Fast Sampling Through The Reuse Of Attention Maps In Diffusion ModelsabstractAbstract Text-to-image diffusion models have demonstrated unprecedented capabilities for flexible and realistic image synthesis. Nevertheless, these models rely on a time-consuming sampling procedure, which has motivated attempts to reduce their latency. When improving efficiency, researchers often use the original diffusion model to train an additional network designed specifically for fast image generation. In contrast, our approach seeks to reduce latency directly, without any retraining, fine-tuning, or knowledge distillation. In particular, we find the repeated calculation of attention maps to be costly yet redundant, and instead suggest reusing them during sampling. Our specific reuse strategies are based on ODE theory, which implies that the later a map is reused, the smaller the distortion in the final image. We empirically compare our reuse strategies with few-step sampling procedures of comparable latency, finding that reuse generates images that are closer to those produced by the original high-latency diffusion model. Rosco Hunter, Lukasz Dudziak, Mohamed S. Abdelfattah, Abhinav Mehrotra, Sourav Bhattacharya, Hongkai Wen 0001 |
Int. J. Comput. Vis. | 6 |
| 2025 | RGBE-Gaze: A Large-Scale Event-Based Multimodal Dataset for High Frequency Remote Gaze TrackingabstractHigh-frequency gaze tracking demonstrates significant potential in various critical applications, such as foveatedrendering, gaze-based identity verification, and the diagnosis of mental disorders. However, existing eye-tracking systems based on CCD/CMOS cameras either provide tracking frequencies below 200 Hz or employ high-speedcameras, causing high power consumption and bulky devices. While there have been some high-speed eye-tracking datasets and methods based on event cameras, they are primarily tailored for near-eye camera scenarios. They lackthe advantages associated with remote camera scenarios, such as the absence of the need for direct contact, improved user comfort and head pose freedom. In this work, we present RGBE-Gaze, the first large-scale and multimodal dataset for remote gaze tracking in high-frequency through synchronizing RGB and event cameras. This dataset is collected from 66 participants with diverse genders and age groups. Our setup captures 3.6 million RGB images and 26.3 billion event samples. Additionally, the dataset includes 10.7 million gaze references from the Gazepoint GP3 HD eye tracker and 15,972 sparse points of gaze (PoG) ground truth obtained through manualstimuli clicks by participants. We present dataset characteristics such as head pose, gaze direction, and pupil size. Furthermore, we introduce a hybrid frame-event based gaze estimation method specifically designed for the collected dataset. Moreover, we perform extensive evaluations of different benchmarking methods under variousgaze-related factors. The evaluation results illustrate that introducing event stream as a new modality improves gazetracking frequency and demonstrates greater estimation robustness across diverse gaze-related factors. Guangrong Zhao, Yiran Shen 0001, Zhaoxin Shen, Yuanfeng Zhou, Hongkai Wen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | AdaFlowLite: Scalable and Non-Blocking Inference on Asynchronous Mobile DataabstractThe rise of mobile devices equipped with numerous sensors, such as LiDAR and cameras, has driven the adoption of multi-modal deep intelligence for distributed sensing tasks, such as smart cabins and driving assistance. However, the arrival time of mobile sensory data vary due to modality size and network dynamics, which can lead to delays (if waiting for slow data) or accuracy decline (if inference proceeds without waiting). Moreover, the diversity and dynamic nature of mobile systems exacerbate this challenge. In response, we present a shift toopportunisticinference for asynchronous distributed multi-modal data, enabling inference as soon as partial data arrives. While existing methods focus on optimizing modality consistency and complementarity, known as modal affinity, they lack acomputationalapproach to control this affinity in open-world mobile environments.${\sf AdaFlowLite}$pioneers the formulation of structured cross-modality affinity in mobile contexts using a hierarchical analysis-based normalized matrix. This approach accommodates the diversity and dynamics of modalities, generalizing across different types and numbers of inputs. Employing an multi-modal lightweight Swin Transformer (MMLST),${\sf AdaFlowLite}$facilitates real-time and flexible data imputation, adapting to various modalities and downstream tasks without retraining. Experiments show that${\sf AdaFlowLite}$significantly reduces inference latency by up to 80.4% and enhances accuracy by up to 62.1%, while achieving nearly a 50% reduction in energy consumption, outperforming status quo approaches. Also, this method can enhance LLM performance to preprocess asynchronous data. Sicong Liu 0005, Fengmin Wu, Bin Guo 0001, Zimu Zhou, Hongkai Wen 0001, Zhiwen Yu 0001 |
IEEE Trans. Mob. Comput. | 6 |
| 2025 | Ui-Ear: On-Face Gesture Recognition Through On-Ear Vibration SensingabstractWith the convenient design and prolific functionalities, wireless earbuds are fast penetrating in our daily life and taking over the place of traditional wired earphones. The sensing capabilities of wireless earbuds have attracted great interests of researchers on exploring them as a new interface for human-computer interactions. However, due to its extremely compact size, the interaction on the body of the earbuds is limited and not convenient. In this paper, we proposeUi-Ear, a new on-face gesture recognition system to enrich interaction maneuvers for wireless earbuds.Ui-Earexploits the sensing capability of Inertial Measurement Units (IMUs) to extend the interaction to the skin of the face near ears. The accelerometer and gyroscope in IMUs perceive dynamic vibration signals induced by on-face touching and moving, which brings rich maneuverability. Since IMUs are provided on most of the budget and high-end wireless earbuds, we believe thatUi-Earhas great potential to be adopted pervasively. To demonstrate the feasibility of the system, we define seven different on-face gestures and design an end-to-end learning approach based on Convolutional Neural Networks (CNNs) for classifying different gestures. To further improve the generalization capability of the system, adversarial learning mechanism is incorporated in the offline training process to suppress the user-specific features while enhancing gesture-related features. We recruit 20 participants and collect a realworld datasets in a common office environment to evaluate the recognition accuracy. The extensive evaluations show that the average recognition accuracy ofUi-Earis over 95% and 82.3% in the user-dependent and user-independent tasks, respectively. Moreover, we also show that the pre-trained model (learned from user-independent task) can be fine-tuned with only few training samples of the target user to achieve relatively high recognition accuracy (up to 95%). At last, we implement the personalization and recognition components ofUi-Earon an off-the-shelf Android smartphone to evaluate its system overhead. The results demonstrateUi-Earcan achieve real-time response while only brings trivial energy consumption on smartphones. Guangrong Zhao, Yiran Shen 0001, Feng Li 0002, Lei Liu 0003, Li-Zhen Cui 0001, Hongkai Wen 0001 |
IEEE Trans. Mob. Comput. | 6 |
| 2025 | EX-Gaze: High-Frequency and Low-Latency Gaze Tracking with Hybrid Event-Frame Cameras for On-Device Extended RealityabstractThe integration of gaze/eye tracking into virtual and augmented reality devices has unlocked new possibilities, offering a novel human-computer interaction (HCI) modality for on-device extended reality (XR). Emerging applications in XR, such as low-effort user authentication, mental health diagnosis, and foveated rendering, demand real-time eye tracking at high frequencies, a capability that current solutions struggle to deliver. To address this challenge, we present EX-Gaze, an event-based real-time eye tracking system designed for on-device extended reality. EX-Gaze achieves a high tracking frequency of 2KHz, providing decent accuracy and low tracking latency. The exceptional tracking frequency of EX-Gaze is achieved through the use of event cameras, cutting-edge, bio-inspired vision hardware that delivers event-stream output at high temporal resolution. We have developed a lightweight tracking framework that enables real-time pupil region localization and tracking on mobile devices. To effectively leverage the sparse nature of event-streams, we introduce the sparse event-patch representation and the corresponding sparse event patches transformer as key components to reduce computational time. Implemented on Jetson Orin Nano, a low-cost, small-sized mobile device with hybrid GPU and CPU components capable of parallel processing of multiple deep neural networks, EX-Gaze maximizes the computation power of Jetson Orin Nano through sophisticated computation scheduling and offloading between GPUs and CPUs. This enables EX-Gaze to achieve real-time tracking at 2KHz without accumulating latency. Evaluation on public datasets demonstrates that EX-Gaze outperforms other event-based eye tracking methods by striking the best balance between accuracy and efficiency on mobile devices. These results highlight EX-Gaze's potential as a groundbreaking technology to support XR applications that require high-frequency and real-time eye tracking. The code is available at https://github.com/Ningreka/EX-Gaze. Yiran Shen 0001, Tongyu Zhang, Yanni Yang 0003, Hongkai Wen 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | Behavior-aware Sparse Trajectory Recovery in Last-mile Delivery with Multi-scale Attention FusionabstractTrajectory data is a valuable asset for service management and spatio-temporal mining in transportation and logistics systems. However, due to equipment failure, network delay, and energy constraints, some trajectory point may be missed, which makes it difficult for trajectory-based management. Some researchers have focused on recovering sparse trajectories from road networks and historical trajectory data, but these methods are ineffective when the road network is incomplete. Recent research works have explored learning-based methods to recover trajectories in free space but lack user movement behavior modeling and efficient feature extraction on sparse long-range trajectories. Our work exploits the periodic behavior of couriers and fine-grained Area of Interest (AOI) data for sparse trajectory recovery in last-mile delivery. However, we face challenges with AOI access sequence deviations due to GPS inaccuracies and abnormal courier behaviors, as well as the complex, dynamic relationships within and between courier routes due to uncertain pick-up demands. To address these challenges, we design a graph-based multi-task learning framework, focusing on multi-scale attention fusion for end-to-end free space trajectory recovery. Our approach starts with a behavior-aware graph network that generates detailed spatial features. Following this, we propose a multi-scale attention fusion mechanism to extract intra- and inter-trajectory features. Finally, we design a multi-task learning module that predicts both coarse-grained spatial access sequences and fine-grained trajectory points. We evaluate the model with six-month data involved with more than 360,000 trajectory segments and more than 7.2 million waybills collected from one of the largest logistic companies in China. Extensive experiments on real-world datasets demonstrate that our method outperforms state-of-the-arts in multiple metrics. Hai Wang 0019, Shuai Wang 0008, Li Lin 0011, Yu Yang 0010, Shuai Wang 0021, Hongkai Wen 0001 |
CIKM | 6 |
| 2024 | SignGraph: A Sign Sequence is Worth Graphs of NodesabstractDespite the recent success of sign language research, the widely adopted CNN-based backbones are mainly migrated from other computer vision tasks, in which the contours and texture of objects are crucial for identifying objects. They usually treat sign frames as grids and may fail to capture effective cross-region features. In fact, sign language tasks need to focus on the correlation of different regions in one frame and the interaction of different regions among adjacent frames for identifying a sign sequence. In this paper, we propose to represent a sign sequence as graphs and introduce a simple yet effective graph-based sign language processing architecture named SignGraph, to extract crossregion features at the graph level. SignGraph consists of two basic modules: Local Sign Graph (LSG) module for learning the correlation of intra-frame cross-region features in one frame and Temporal Sign Graph (TSG) module for tracking the interaction of inter-frame cross-region features among adjacent frames. With LSG and TSG, we build our model in a multiscale manner to ensure that the representation of nodes can capture cross-region features at different granularities. Extensive experiments on current public sign language datasets demonstrate the superiority of our SignGraph model. Our model achieves very competitive performances with the SOTA model, while not using any extra cues. Code and models are available at: https://github.com/gswycf/SignGraph. Shiwei Gan, Yafeng Yin 0002, Zhiwei Jiang 0001, Hongkai Wen 0001, Lei Xie 0004, Sanglu Lu |
CVPR | 4 |
| 2024 | Towards Neural Architecture Search through Hierarchical Generative ModelingabstractNeural Architecture Search (NAS) aims to automate deep neural network design across various applications, while a good search space design is core to NAS performance. A too-narrow search space may fail to cover diverse task requirements, whereas a too-broad one can escalate computational expenses and reduce efficiency. %We propose automatically generating the search space to tailor it to specific task conditions, optimizing search costs and producing viable architectures. In this work, we aim to address this challenge by leaning on the recent advances in generative modelling -- we propose a novel method that can navigate through an extremely large, general-purpose initial search space efficiently by training a two-level generative model hierarchy. The first level uses Conditional Continuous Normalizing Flow (CCNF) for micro-cell design, while the second employs a transformer-based sequence generator to craft macro architectures aligned with task needs and architectural constraints. To ensure computational feasibility, we pretrain the generative models in a task-agnostic manner using a metric space of graph and zero-cost (ZC) similarities between architectures. We show our approach can achieve state-of-the-art performance among other low-cost NAS methods across different tasks on CIFAR-10/100, ImageNet and NAS-Bench-360. Lichuan Xiang, Lukasz Dudziak, Mohamed S. Abdelfattah, Abhinav Mehrotra, Nicholas D. Lane, Hongkai Wen 0001 |
ICML | 6 |
| 2024 | AdaFlow: Opportunistic Inference on Asynchronous Mobile Data with Generalized Affinity ControlabstractThe rise of mobile devices equipped with numerous sensors, such as LiDAR and cameras, has spurred the adoption of multi-modal deep intelligence for distributed sensing tasks, such as smart cabins and driving assistance. However, the arrival times of mobile sensory data vary due to modality size and network dynamics, which can lead to delays (if waiting for slower data) or accuracy decline (if inference proceeds without waiting). Moreover, the diversity and dynamic nature of mobile systems exacerbate this challenge. In response, we present a shift to opportunistic inference for asynchronous distributed multi-modal data, enabling inference as soon as partial data arrives. While existing methods focus on optimizing modality consistency and complementarity, known as modal affinity, they lack a computational approach to control this affinity in open-world mobile environments. AdaFlow pioneers the formulation of structured cross-modality affinity in mobile contexts using a hierarchical analysis-based normalized matrix. This approach accommodates the diversity and dynamics of modalities, generalizing across different types and numbers of inputs. Employing an affinity attention-based conditional GAN (ACGAN), AdaFlow facilitates flexible data imputation, adapting to various modalities and downstream tasks without retraining. Experiments show that AdaFlow significantly reduces inference latency by up to 79.9% and enhances accuracy by up to 61.9%, outperforming status quo approaches. Also, this method can enhance LLM performance to preprocess asynchronous data. Fengmin Wu, Sicong Liu 0005, Kehao Zhu, Bin Guo 0001, Zhiwen Yu 0001, Hongkai Wen 0001, Xiangrui Xu 0005, Lehao Wang |
SenSys | 7 |
| 2024 | EV-Perturb: event-stream perturbation for privacy-preserving classification with dynamic vision sensors
Yong Wang 0020, Qing Yang 0009, Yiran Shen 0001, Hongkai Wen 0001 |
Multim. Tools Appl. | 5 |
| 2024 | EV-Tach: A Handheld Rotational Speed Estimation System With Event CameraabstractRotational speed is one of the important metrics to be measured for calibrating electric motors in manufacturing, monitoring engines during car repairs, detecting faults in electrical appliance and more. However, existing measurement techniques either require prohibitive hardware (e.g., high-speed camera) or are inconvenient to use in real-world application scenarios. In this paper, we propose,EV-Tach, a novel handheld rotational speed estimation system that utilizes emerging imaging sensors known as event cameras or dynamic vision sensors (DVS). The pixels of DVS work independently and trigger an event as soon as a per-pixel intensity change is detected, without global synchronization like conventional RGB cameras. Thus, its unique design features high temporal resolution and generates sparse events, which benefits the high-speed rotation estimation. To achieve accurate and efficient rotational speed estimation, a series of signal processing algorithms are specifically designed for the event streams generated by event cameras on an embedded platform. First, a new cluster-centroids initialization module is proposed to initialize the centroids of the clusters to address the issue that common clustering approaches are easy to fall into a local optimal solution without proper initial centroids. Second, an outlier removal module is designed to suppress the background noise caused by subtle hand movements and host devices vibrations. Third, a coarse-to-fine alignment strategy is proposed with an event stream alignment method to obtain angle of rotation and achieve accurate estimation for rotational speed in a large range. With these bespoke components,EV-Tachis able to extract the rotational speed accurately from the event stream produced by an event camera recording rotary targets. According to our extensive evaluations under controlled and practical experiment settings, the Relative Mean Absolute Error (RMAE) ofEV-Tachis as low as$0.3\%_{0}$, which is comparable to the state-of-the-art laser tachometer under fixed measurement mode. Moreover,EV-Tachis robust to subtle movement of user's hand and dazzling light outdoor, therefore, can be used as a handheld device under challenging lighting condition, where the laser tachometer fails to produce reasonable results. To speed up the processing ofEV-Tachand reduce its resource consumption on embedded devices, event stream is significantly downsampled by merging neighboring events while preserving its formation in spatial-temporal domain. At last, we implementEV-Tachon Raspberry Pi and the evaluation results show that the downsampling process preserves the high measurement accuracy while saving the computation speed and energy consumption by approximately 8 times and 30 times in average. Guangrong Zhao, Yiran Shen 0001, Pengfei Hu 0001, Lei Liu 0003, Hongkai Wen 0001 |
IEEE Trans. Mob. Comput. | 6 |
| 2023 | Zero-Cost Operation Scoring in Differentiable Architecture SearchabstractWe formalize and analyze a fundamental component of dif- ferentiable neural architecture search (NAS): local “opera- tion scoring” at each operation choice. We view existing operation scoring functions as inexact proxies for accuracy, and we find that they perform poorly when analyzed empir- ically on NAS benchmarks. From this perspective, we intro- duce a novel perturbation-based zero-cost operation scor- ing (Zero-Cost-PT) approach, which utilizes zero-cost prox- ies that were recently studied in multi-trial NAS but de- grade significantly on larger search spaces, typical for dif- ferentiable NAS. We conduct a thorough empirical evalu- ation on a number of NAS benchmarks and large search spaces, from NAS-Bench-201, NAS-Bench-1Shot1, NAS- Bench-Macro, to DARTS-like and MobileNet-like spaces, showing significant improvements in both search time and accuracy. On the ImageNet classification task on the DARTS search space, our approach improved accuracy compared to the best current training-free methods (TE-NAS) while be- ing over 10× faster (total searching time 25 minutes on a single GPU), and observed significantly better transferabil- ity on architectures searched on the CIFAR-10 dataset with an accuracy increase of 1.8 pp. Our code is available at: https://github.com/zerocostptnas/zerocost operation score. Lichuan Xiang, Lukasz Dudziak, Mohamed S. Abdelfattah, Thomas C. P. Chau, Nicholas D. Lane, Hongkai Wen 0001 |
AAAI | 6 |
| 2023 | Towards Data-Agnostic Pruning At Initialization: What Makes a Good Sparse Mask?abstractPruning at initialization (PaI) aims to remove weights of neural networks before training in pursuit of training efficiency besides the inference. While off-the-shelf PaI methods manage to find trainable subnetworks that outperform random pruning, their performance in terms of both accuracy and computational reduction is far from satisfactory compared to post-training pruning and the understanding of PaI is missing. For instance, recent studies show that existing PaI methods only able to find good layerwise sparsities not weights, as the discovered subnetworks are surprisingly resilient against layerwise random mask shuffling and weight re-initialization.
In this paper, we study PaI from a brand-new perspective -- the topology of subnetworks. In particular, we propose a principled framework for analyzing the performance of Pruning and Initialization (PaI) methods with two quantities, namely, the number of effective paths and effective nodes. These quantities allow for a more comprehensive understanding of PaI methods, giving us an accurate assessment of different subnetworks at initialization. We systematically analyze the behavior of various PaI methods through our framework and observe a guiding principle for constructing effective subnetworks: *at a specific sparsity, the top-performing subnetwork always presents a good balance between the number of effective nodes and the number of effective paths.*
Inspired by this observation, we present a novel data-agnostic pruning method by solving a multi-objective optimization problem. By conducting extensive experiments across different architectures and datasets, our results demonstrate that our approach outperforms state-of-the-art PaI methods while it is able to discover subnetworks that have much lower inference FLOPs (up to 3.4$\times$). Code will be fully released. The-Anh Ta, Shiwei Liu 0003, Lichuan Xiang, Dung Le, Hongkai Wen 0001, Long Tran-Thanh |
NeurIPS | 6 |
| 2023 | EV-Eye: Rethinking High-frequency Eye Tracking through the Lenses of Event CamerasabstractIn this paper, we present EV-Eye, a first-of-its-kind large scale multimodal eye tracking dataset aimed at inspiring research on high-frequency eye/gaze tracking. EV-Eye utilizes an emerging bio-inspired event camera to capture independent pixel-level intensity changes induced by eye movements, achieving sub-microsecond latency. Our dataset was curated over a two-week period and collected from 48 participants encompassing diverse genders and age groups. It comprises over 1.5 million near-eye grayscale images and 2.7 billion event samples generated by two DAVIS346 event cameras. Additionally, the dataset contains 675 thousands scene images and 2.7 million gaze references captured by Tobii Pro Glasses 3 eye tracker for cross-modality validation. Compared with existing event-based high-frequency eye tracking datasets, our dataset is significantly larger in size, and the gaze references involve more natural eye movement patterns, i.e., fixation, saccade and smooth pursuit. Alongside the event data, we also present a hybrid eye tracking method as benchmark, which leverages both the near-eye grayscale images and event data for robust and high-frequency eye tracking. We show that our method achieves higher accuracy for both pupil and gaze estimation tasks compared to the existing solution. Guangrong Zhao, Yurun Yang, Yiran Shen 0001, Hongkai Wen 0001, Guohao Lan |
NeurIPS | 6 |
| 2023 | Fleet Rebalancing for Expanding Shared e-Mobility Systems: A Multi-Agent Deep Reinforcement Learning ApproachabstractThe electrification of shared mobility has become popular across the globe. Many cities have their new shared e-mobility systems deployed, with continuously expanding coverage from central areas to the city edges. A key challenge in the operation of these systems is fleet rebalancing, i.e., how EVs should be repositioned to better satisfy future demand. This is particularly challenging in the context of expanding systems, because i) the range of the EVs is limited while charging time is typically long, which constrain the viable rebalancing operations; and ii) the EV stations in the system are dynamically changing, i.e., the legitimate targets for rebalancing operations can vary over time. We tackle these challenges by first investigating rich sets of data collected from a real-world shared e-mobility system for one year, analyzing the operation model, usage patterns and expansion dynamics of this new mobility mode. With the learned knowledge we design a high-fidelity simulator, which is able to abstract key operation details of EV sharing at fine granularity. Then we model the rebalancing task for shared e-mobility systems under continuous expansion as a Multi-Agent Reinforcement Learning (MARL) problem, which directly takes the range and charging properties of the EVs into account. We further propose a novel policy optimization approach with action cascading, which is able to cope with the expansion dynamics and solve the formulated MARL. We evaluate the proposed approach extensively, and experimental results show that our approach outperforms the state-of-the-art, offering significant performance gain in both satisfied demand and net revenue. Man Luo 0001, Bowen Du 0002, Tianyou Song, Kun Li 0030, Hongming Zhu, Mark H. Birkin, Hongkai Wen 0001 |
IEEE Trans. Intell. Transp. Syst. | 8 |
| 2022 | BLOX: Macro Neural Architecture Search Benchmark and AlgorithmsabstractNeural architecture search (NAS) has been successfully used to design numerous high-performance neural networks. However, NAS is typically compute-intensive, so most existing approaches restrict the search to decide the operations and topological structure of a single block only, then the same block is stacked repeatedly to form an end-to-end model. Although such an approach reduces the size of search space, recent studies show that a macro search space, which allows blocks in a model to be different, can lead to better performance. To provide a systematic study of the performance of NAS algorithms on a macro search space, we release Blox – a benchmark that consists of 91k unique models trained on the CIFAR-100 dataset. The dataset also includes runtime measurements of all the models on a diverse set of hardware platforms. We perform extensive experiments to compare existing algorithms that are well studied on cell-based search spaces, with the emerging blockwise approaches that aim to make NAS scalable to much larger macro search spaces. The Blox benchmark and code are available at https://github.com/SamsungLabs/blox. Thomas C. P. Chau, Lukasz Dudziak, Hongkai Wen 0001, Nicholas D. Lane, Mohamed S. Abdelfattah |
NeurIPS | 3 |
| 2022 | Event-Stream Representation for Human Gaits Identification Using Deep Neural NetworksabstractDynamic vision sensors (event cameras) have recently been introduced to solve a number of different vision tasks such as object recognition, activities recognition, tracking, etc. Compared with the traditional RGB sensors, the event cameras have many unique advantages such as ultra low resources consumption, high temporal resolution and much larger dynamic range. However, these cameras only produce noisy and asynchronous events of intensity changes, i.e., event-streams rather than frames, where conventional computer vision algorithms can't be directly applied. In our opinion the key challenge for improving the performance of event cameras in vision tasks is finding the appropriate representations of the event-streams so that cutting-edge learning approaches can be applied to fully uncover the spatio-temporal information contained in the event-streams. In this paper, we focus on the event-based human gait identification task and investigate the possible representations of the event-streams when deep neural networks are applied as the classifier. We propose new event-based gait recognition approaches basing on two different representations of the event-stream, i.e., graph and image-like representations, and use graph-based convolutional network (GCN) and convolutional neural networks (CNN) respectively to recognize gait from the event-streams. The two approaches are termed as EV-Gait-3DGraph and EV-Gait-IMG. To evaluate the performance of the proposed approaches, we collect two event-based gait datasets, one from real-world experiments and the other by converting the publicly available RGB gait recognition benchmark CASIA-B. Extensive experiments show that EV-Gait-3DGraph achieves significantly higher recognition accuracy than other competing methods when sufficient training samples are available. However, EV-Gait-IMG converges more quickly than graph-based approaches while training and shows good accuracy with only few number of training samples (less than ten). So image-like presentation is preferable when the amount of training data is limited. Yiran Shen 0001, Bowen Du 0002, Guangrong Zhao, Li-Zhen Cui 0001, Hongkai Wen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2022 | Deployment Optimization for Shared e-Mobility Systems With Multi-Agent Deep Neural SearchabstractShared e-mobility services have been widely tested and piloted in cities across the globe, and already woven into the fabric of modern urban planning. This paper studies a practical yet important problem in those systems: how to deploy and manage their infrastructure across space and time, so that the services areubiquitousto the users whilesustainablein profitability. However, in real-world systems evaluating the performance of different deployment strategies and then finding the optimal plan is prohibitively expensive, as it is often infeasible to conduct many iterations of trial-and-error. We tackle this by designing a high-fidelity simulation environment, which abstracts the key operation details of the shared e-mobility systems at fine-granularity, and is calibrated using data collected from the real-world. This allows us to try out arbitrary deployment plans to learn the optimal given specific context, before actually implementing any in the real-world systems. In particular, we propose a novel multi-agent neural search approach, in which we design a hierarchical controller to produce tentative deployment plans. The generated deployment plans are then tested using a multi-simulation paradigm, i.e., evaluated in parallel, where the results are used to train the controller with deep reinforcement learning. With this closed loop, the controller can be steered to have higher probability of generating better deployment plans in future iterations. The proposed approach has been evaluated extensively in our simulation environment, and experimental results show that it outperforms baselines e.g., human knowledge, and state-of-the-art heuristic-based optimization approaches in both service coverage and net revenue. Man Luo 0001, Bowen Du 0002, Konstantin Klemmer, Hongming Zhu, Hongkai Wen 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2021 | LeaD: Learn to Decode Vibration-based Communication for Intelligent Internet of ThingsabstractIn this article, we propose, LeaD , a new vibration-based communication protocol to Lea rn the unique patterns of vibration to D ecode the short messages transmitted to smart IoT devices. Unlike the existing vibration-based communication protocols that decode the short messages symbol-wise, either in binary or multi-ary, the message recipient in LeaD receives vibration signals corresponding to bits-groups. Each group consists of multiple symbols sent in a burst and the receiver decodes the group of symbols as a whole via machine learning-based approach. The fundamental behind LeaD is different combinations of symbols (1 s or 0 s) in a group will produce unique and reproducible patterns of vibration. Therefore, decoding in vibration-based communication can be modeled as a pattern classification problem. We design and implement a number of different machine learning models as the core engine of the decoding algorithm of LeaD to learn and recognize the vibration patterns. Through the intensive evaluations on large amount of datasets collected, the Convolutional Neural Network (CNN)-based model achieves the highest accuracy of decoding (i.e., lowest error rate), which is up to 97% at relatively high bits rate of 40 bits/s. While its competing vibration-based communication protocols can only achieve transmission rate of 10 bits/s and 20 bits/s with similar decoding accuracy. Furthermore, we evaluate its performance under different challenging practical settings and the results show that LeaD with CNN engine is robust to poses, distances (within valid range), and types of devices, therefore, a CNN model can be generally trained beforehand and widely applicable for different IoT devices under different circumstances. Finally, we implement LeaD on both off-the-shelf smartphone and smart watch to measure the detailed resources consumption on smart devices. The computation time and energy consumption of its different components show that LeaD is lightweight and can run in situ on low-cost smart IoT devices, e.g., smartwatches, without accumulated delay and introduces only marginal system overhead. Guangrong Zhao, Bowen Du 0002, Yiran Shen 0001, Zhenyu Lao, Li-Zhen Cui 0001, Hongkai Wen 0001 |
ACM Trans. Sens. Networks | 6 |
| 2020 | Journey Towards Tiny Perceptual Super-Resolution
Royson Lee, Lukasz Dudziak, Mohamed S. Abdelfattah, Stylianos I. Venieris, Hyeji Kim, Hongkai Wen 0001, Nicholas D. Lane |
ECCV (26) | 6 |
| 2020 | 3DCFS: Fast and Robust Joint 3D Semantic-Instance Segmentation via Coupled Feature SelectionabstractWe propose a novel fast and robust 3D point clouds segmentation framework via coupled feature selection, named 3DCFS, that jointly performs semantic and instance segmentation. Inspired by the human scene perception process, we design a novel coupled feature selection module, named CFSM, that adaptively selects and fuses the reciprocal semantic and instance features from two tasks in a coupled manner. To further boost the performance of the instance segmentation task in our 3DCFS, we investigate a loss function that helps the model learn to balance the magnitudes of the output embedding dimensions during training, which makes calculating the Euclidean distance more reliable and enhances the generalizability of the model. Extensive experiments demonstrate that our 3DCFS outperforms state-of-the-art methods on benchmark datasets in terms of accuracy, speed and computational cost. Codes are available at: https://github.com/Biotan/3DCFS. Liang Du 0004, Jingang Tan, Xiangyang Xue 0001, Hongkai Wen 0001, Jianfeng Feng, Jiamao Li |
ICRA | 5 |
| 2020 | Rebalancing Expanding EV Sharing Systems with Deep Reinforcement LearningabstractElectric Vehicle (EV) sharing systems have recently experienced unprecedented growth across the world. One of the key challenges in their operation is vehicle rebalancing, i.e., repositioning the EVs across stations to better satisfy future user demand. This is particularly challenging in the shared EV context, because i) the range of EVs is limited while charging time is substantial, which constrains the rebalancing options; and ii) as a new mobility trend, most of the current EV sharing systems are still continuously expanding their station networks, i.e., the targets for rebalancing can change over time. To tackle these challenges, in this paper we model the rebalancing task as a Multi-Agent Reinforcement Learning (MARL) problem, which directly takes the range and charging properties of the EVs into account. We propose a novel approach of policy optimization with action cascading, which isolates the non-stationarity locally, and use two connected networks to solve the formulated MARL. We evaluate the proposed approach using a simulator calibrated with 1-year operation data from a real EV sharing system. Results show that our approach significantly outperforms the state-of-the-art, offering up to 14% gain in order satisfied rate and 12% increase in net revenue. Man Luo 0001, Tianyou Song, Kun Li 0030, Hongming Zhu, Bowen Du 0002, Hongkai Wen 0001 |
IJCAI | 7 |
| 2020 | Securing Cyber-Physical Social Interactions on Wrist-Worn DevicesabstractSince ancient Greece, handshaking has been commonly practiced between two people as a friendly gesture to express trust and respect, or form a mutual agreement. In this article, we show that such physical contact can be used to bootstrap secure cyber contact between the smart devices worn by users. The key observation is that during handshaking, although belonged to two different users, the two hands involved in the shaking events are often rigidly connected, and therefore exhibit very similar motion patterns. We propose a novel key generation system, which harvests motion data during user handshaking from the wrist-worn smart devices such as smartwatches or fitness bands, and exploits the matching motion patterns to generate symmetric keys on both parties. The generated keys can be then used to establish a secure communication channel for exchanging data between devices. This provides a much more natural and user-friendly alternative for many applications, e.g., exchanging/sharing contact details, friending on social networks, or even making payments, since it doesn’t involve extra bespoke hardware, nor require the users to perform pre-defined gestures. We implement the proposed key generation system on off-the-shelf smartwatches, and extensive evaluation shows that it can reliably generate 128-bit symmetric keys just after around 1s of handshaking (with success rate >99%), and is resilient to different types of attacks including impersonate mimicking attacks, impersonate passive attacks, or eavesdropping attacks. Specifically, for real-time impersonate mimicking attacks, in our experiments, the Equal Error Rate (EER) is only 1.6% on average. We also show that the proposed key generation system can be extremely lightweight and is able to run in-situ on the resource-constrained smartwatches without incurring excessive resource consumption. Yiran Shen 0001, Bowen Du 0002, Weitao Xu, Chengwen Luo 0001, Bo Wei 0003, Li-Zhen Cui 0001, Hongkai Wen 0001 |
ACM Trans. Sens. Networks | 7 |
| 2019 | EV-Gait: Event-Based Robust Gait Recognition Using Dynamic Vision SensorsabstractIn this paper, we introduce a new type of sensing modality, the Dynamic Vision Sensors (Event Cameras), for the task of gait recognition. Compared with the traditional RGB sensors, the event cameras have many unique advantages such as ultra low resources consumption, high temporal resolution and much larger dynamic range. However, those cameras only produce noisy and asynchronous events of intensity changes rather than frames, where conventional vision-based gait recognition algorithms can’t be directly applied. To address this, we propose a new Event-based Gait Recognition (EV-Gait) approach, which exploits motion consistency to effectively remove noise, and uses a deep neural network to recognise gait from the event streams. To evaluate the performance of EV-Gait, we collect two event-based gait datasets, one from real-world experiments and the other by converting the publicly available RGB gait recognition benchmark CASIA-B. Extensive experiments show that EV-Gait can get nearly 96% recognition accuracy in the real-world settings, while on the CASIA-B benchmark it achieves comparable performance with state-of-the-art RGB-based gait recognition approaches. Bowen Du 0002, Yiran Shen 0001, Guangrong Zhao, Hongkai Wen 0001 |
CVPR | 7 |
| 2019 | Weakly Supervised Brain Lesion Segmentation via Attentional Representation Learning
Bowen Du 0002, Man Luo 0001, Hongkai Wen 0001, Yiran Shen 0001, Jianfeng Feng |
MICCAI (3) | 4 |
| 2019 | Autonomous Learning for Face Recognition in the Wild via Ambient Wireless CuesabstractFacial recognition is a key enabling component for emerging Internet of Things (IoT) services such as smart homes or responsive offices. Through the use of deep neural networks, facial recognition has achieved excellent performance. However, this is only possibly when trained with hundreds of images of each user in different viewing and lighting conditions. Clearly, this level of effort in enrolment and labelling is impossible for wide-spread deployment and adoption. Inspired by the fact that most people carry smart wireless devices with them, e.g. smartphones, we propose to use this wireless identifier as a supervisory label. This allows us to curate a dataset of facial images that are unique to a certain domain e.g. a set of people in a particular office. This custom corpus can then be used to finetune existing pre-trained models e.g. FaceNet. However, due to the vagaries of wireless propagation in buildings, the supervisory labels are noisy and weak. We propose a novel technique, AutoTune, which learns and refines the association between a face and wireless identifier over time, by increasing the inter-cluster separation and minimizing the intra-cluster distance. Through extensive experiments with multiple users on two sites, we demonstrate the ability of AutoTune to design an environment-specific, continually evolving facial recognition system with entirely no user effort. Xiaoxuan Lu 0001, Xuan Kan, Bowen Du 0002, Changhao Chen, Hongkai Wen 0001, Andrew Markham, Agathoniki Trigoni, John A. Stankovic |
WWW | 5 |
| 2019 | Dense 3D Object Reconstruction from a Single Depth ViewabstractIn this paper, we propose a novel approach, 3D-RecGAN++, which reconstructs the complete 3D structure of a given object from a single arbitrary depth view using generative adversarial networks. Unlike existing work which typically requires multiple views of the same object or class labels to recover the full 3D geometry, the proposed 3D-RecGAN++ only takes the voxel grid representation of a depth view of the object as input, and is able to generate the complete 3D occupancy grid with a high resolution of 2563by recovering the occluded/missing regions. The key idea is to combine the generative capabilities of 3D encoder-decoder and the conditional adversarial networks framework, to infer accurate and fine-grained 3D structures of objects in high-dimensional voxel space. Extensive experiments on large synthetic datasets and real-world Kinect datasets show that the proposed 3D-RecGAN++ significantly outperforms the state of the art in single view 3D object reconstruction, and is able to reconstruct unseen types of objects. Bo Yang 0027, Stefano Rosa, Andrew Markham, Agathoniki Trigoni, Hongkai Wen 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2019 | GaitLock: Protect Virtual and Augmented Reality Headsets Using GaitabstractWith the fast penetration of commercial Virtual Reality (VR) and Augmented Reality (AR) systems into our daily life, the security issues of those devices have attracted significant interests from both academia and industry. Modern VR/AR systems typically use head-mounted devices (i.e., headsets) to interact with users, and often store private user data, e.g., social network accounts, online transactions or even payment information. This poses significant security threats, since in practice the headset can be potentially obtained and accessed by unauthenticated parties, e.g., identity thieves, and thus cause catastrophic breach. In this paper, we propose a novel GaitLock system, which can reliably authenticate users using their gait signatures. Our system doesn't require extra hardware, e.g., fingerprint sensors or retina scanners, but only uses the on-board inertial measurement units (IMUs) equipped in almost all mainstream VR/AR headsets to authenticate the legitimate users from intruders, by simply asking them to walk a few steps. To achieve that, we propose a new gait recognition model Dynamic-SRC, which combines the strength of Dynamic Time Warping (DTW) and Sparse Representation Classifier (SRC), to extract unique gait patterns from the inertial signals during walking. We implement GaitLock on Google Glass (a typical AR headset), and extensive experiments show that GaitLock outperforms the state-of-the-art systems significantly in recognition accuracy (> 98 percent success in 5 steps), and is able to run in-situ on the resource-constrained VR/AR headsets without incurring high energy cost. Yiran Shen 0001, Hongkai Wen 0001, Chengwen Luo 0001, Weitao Xu, Tao Zhang 0001, Wen Hu 0001, Daniela Rus |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2019 | Efficient Indoor Positioning with Visual Experiences via Lifelong LearningabstractPositioning with visual sensors in indoor environments has many advantages: it doesn’t require infrastructure or accurate maps, and is more robust and accurate than other modalities such as WiFi. However, one of the biggest hurdles that prevents its practical application on mobile devices is the time-consuming visual processing pipeline. To overcome this problem, this paper proposes a novel lifelong learning approach to enable efficient and real-time visual positioning. We explore the fact that when following a previous visual experience for multiple times, one could gradually discover clues on how to traverse it with much less effort, e.g., which parts of the scene are more informative, and what kind of visual elements we should expect. Such second-order information is recorded as parameters, which provide key insights of the context and empower our system to dynamically optimise itself to stay localised with minimum cost. We implement the proposed approach on an array of mobile and wearable devices, and evaluate its performance in two indoor settings. Experimental results show our approach can reduce the visual processing time up to two orders of magnitude, while achieving sub-metre positioning accuracy. Hongkai Wen 0001, Ronald Clark, Sen Wang 0002, Xiaoxuan Lu 0001, Bowen Du 0002, Wen Hu 0001, Agathoniki Trigoni |
IEEE Trans. Mob. Comput. | 1 |
| 2018 | Deepauth: in-situ authentication for smartwatches via deeply learned behavioural biometricsabstractThis paper proposes DeepAuth, an in-situ authentication framework that leverages the unique motion patterns when users entering passwords as behavioural biometrics. It uses a deep recurrent neural network to capture the subtle motion signatures during password input, and employs a novel loss function to learn deep feature representations that are robust to noise, unseen passwords, and malicious imposters even with limited training data. DeepAuth is by design optimised for resource constrained platforms, and uses a novel split-RNN architecture to slim inference down to run in real-time on off-the-shelf smartwatches. Extensive experiments with real-world data show that DeepAuth outperforms the state-of-the-art significantly in both authentication performance and cost, offering real-time authentication on a variety of smartwatches. Xiaoxuan Lu 0001, Bowen Du 0002, Peijun Zhao, Hongkai Wen 0001, Yiran Shen 0001, Andrew Markham, Agathoniki Trigoni |
UbiComp | 4 |
| 2018 | Simultaneous Localization and Mapping with Power Network Electromagnetic FieldabstractVarious sensing modalities have been exploited for indoor location sensing, each of which has well understood limitations, however. This paper presents a first systematic study on using the electromagnetic field (EMF) induced by a building's electric power network for simultaneous localization and mapping (SLAM). A basis of this work is a measurement study showing that the power network EMF sensed by either a customized sensor or smartphone's microphone as a side-channel sensor is spatially distinct and temporally stable. Based on this, we design a SLAM approach that can reliably detect loop closures based on EMF sensing results. With the EMF feature map constructed by SLAM, we also design an efficient online localization scheme for resource-constrained mobiles. Evaluation in three indoor spaces shows that the power network EMF is a promising modality for location sensing on mobile devices, which is able to run in real time and achieve sub-meter accuracy. Xiaoxuan Lu 0001, Yang Li 0147, Peijun Zhao, Changhao Chen, Linhai Xie, Hongkai Wen 0001, Rui Tan 0001, Agathoniki Trigoni |
MobiCom | 6 |
| 2018 | Shake-n-Shack: Enabling Secure Data Exchange Between Smart Wearables via HandshakesabstractSince ancient Greece, handshaking has been commonly practiced between two people as a friendly gesture to express trust and respect, or form a mutual agreement. In this paper, we show that such physical contact can be used to bootstrap secure cyber contact between the smart devices worn by users. The key observation is that during handshaking, although belonged to two different users, the two hands involved in the shaking events are often rigidly connected, and therefore exhibit very similar motion patterns. We propose a novel Shake-n-Shack system, which harvests motion data during user handshaking from the wrist worn smart devices such as smartwatches or fitness bands, and exploits the matching motion patterns to generate symmetric keys on both parties. The generated keys can be then used to establish a secure communication channel for exchanging data between devices. This provides a much more natural and user-friendly alternative for many applications, e.g., exchanging/sharing contact details, friending on social networks, or even making payments, since it doesn't involve extra bespoke hardware, nor require the users to perform pre-defined gestures. We implement the proposed Shake-n-Shack1system on off-the-shelf smartwatches, and extensive evaluation shows that it can reliably generate 128-bit symmetric keys just after around 1s of handshaking (with success rate >99%), and is resilient to real-time mimicking attacks: in our experiments the Equal Error Rate (EER) is only 1.6% on average. We also show that the proposed Shake-n-Shack system can be extremely lightweight, and is able to run in-situ on the resource-constrained smartwatches without incurring excessive resource consumption. Yiran Shen 0001, Fengyuan Yang 0001, Bowen Du 0002, Weitao Xu, Chengwen Luo 0001, Hongkai Wen 0001 |
PerCom | 6 |
| 2018 | Automatic Face Recognition Adaptation via Ambient Wireless IdentifiersabstractFace recognition is a key enabling service for smart-spaces, allowing building management agents to easily monitor 'who is where', anticipating user needs and tailoring their local environment and experiences. Although facial recognition, especially through the use of deep neural networks, has achieved stellar performance over large datasets, the majority of approaches require supervised learning, that is, to be trained with tens or hundreds of images of users in different poses and lighting conditions. In this paper, we motivate that this enrollment effort is unnecessary if the smart-space has access to a wireless identifier e.g., through a smart-phone's MAC address. By learning and refining the noisy and weak association between a user's smart-phone and facial images, AutoTune can fine-tune a deep neural network to tailor it to the environment, users and conditions of a particular camera or set of cameras. Xiaoxuan Lu 0001, Peijun Zhao, Bowen Du 0002, Hongkai Wen 0001, Andrew Markham, Stefano Rosa, Agathoniki Trigoni |
SenSys | 4 |
| 2018 | Privacy-preserving sparse representation classification in cloud-enabled mobile applications
Yiran Shen 0001, Chengwen Luo 0001, Dan Yin, Hongkai Wen 0001, Daniela Rus, Wen Hu 0001 |
Comput. Networks | 4 |
| 2018 | HealCam: Energy-efficient and privacy-preserving human vital cycles monitoring on camera-enabled smart devices
Qing Yang 0009, Yiran Shen 0001, Fengyuan Yang 0001, Jianpei Zhang, Wanli Xue, Hongkai Wen 0001 |
Comput. Networks | 6 |
| 2017 | VINet: Visual-Inertial Odometry as a Sequence-to-Sequence Learning ProblemabstractIn this paper we present an on-manifold sequence-to-sequence learning approach to motion estimation using visual and inertial sensors. It is to the best of our knowledge the first end-to-end trainable method for visual-inertial odometry which performs fusion of the data at an intermediate feature-representation level. Our method has numerous advantages over traditional approaches. Specifically, it eliminates the need for tedious manual synchronization of the camera and IMU as well as eliminating the need for manual calibration between the IMU and camera. A further advantage is that our model naturally and elegantly incorporates domain specific information which significantly mitigates drift. We show that our approach is competitive with state-of-the-art traditional methods when accurate calibration data is available and can be trained to outperform them in the presence of calibration and synchronization errors. Ronald Clark, Sen Wang 0002, Hongkai Wen 0001, Andrew Markham, Agathoniki Trigoni |
AAAI | 3 |
| 2017 | VidLoc: A Deep Spatio-Temporal Model for 6-DoF Video-Clip RelocalizationabstractMachine learning techniques, namely convolutional neural networks (CNN) and regression forests, have recently shown great promise in performing 6-DoF localization of monocular images. However, in most cases image-sequences, rather only single images, are readily available. To this extent, none of the proposed learning-based approaches exploit the valuable constraint of temporal smoothness, often leading to situations where the per-frame error is larger than the camera motion. In this paper we propose a recurrent model for performing 6-DoF localization of video-clips. We find that, even by considering only short sequences (20 frames), the pose estimates are smoothed and the localization error can be drastically reduced. Finally, we consider means of obtaining probabilistic pose estimates from our model. We evaluate our method on openly-available real-world autonomous driving and indoor localization datasets. Ronald Clark, Sen Wang 0002, Andrew Markham, Agathoniki Trigoni, Hongkai Wen 0001 |
CVPR | 5 |
| 2017 | DeepVO: Towards end-to-end visual odometry with deep Recurrent Convolutional Neural NetworksabstractThis paper studies monocular visual odometry (VO) problem. Most of existing VO algorithms are developed under a standard pipeline including feature extraction, feature matching, motion estimation, local optimisation, etc. Although some of them have demonstrated superior performance, they usually need to be carefully designed and specifically fine-tuned to work well in different environments. Some prior knowledge is also required to recover an absolute scale for monocular VO. This paper presents a novel end-to-end framework for monocular VO by using deep Recurrent Convolutional Neural Networks (RCNNs). Since it is trained and deployed in an end-to-end manner, it infers poses directly from a sequence of raw RGB images (videos) without adopting any module in the conventional VO pipeline. Based on the RCNNs, it not only automatically learns effective feature representation for the VO problem through Convolutional Neural Networks, but also implicitly models sequential dynamics and relations using deep Recurrent Neural Networks. Extensive experiments on the KITTI VO dataset show competitive performance to state-of-the-art methods, verifying that the end-to-end Deep Learning technique can be a viable complement to the traditional VO systems. Sen Wang 0002, Ronald Clark, Hongkai Wen 0001, Agathoniki Trigoni |
ICRA | 3 |
| 2017 | SCAN: learning speaker identity from noisy sensor dataabstractSensor data acquired from multiple sensors simultaneously is featuring increasingly in our evermore pervasive world. Buildings can be made smarter and more efficient, spaces more responsive to users. A fundamental building block towards smart spaces is the ability to understand who is present in a certain area. A ubiquitous way of detecting this is to exploit the unique vocal features as people interact with one another. As an example, consider audio features sampled during a meeting, yielding a noisy set of possible voiceprints. With a number of meetings and knowledge of participation (e.g. through a calendar or MAC address), can we learn to associate a specific identity with a particular voiceprint? Obviously enrolling users into a biometric database is time-consuming and not robust to vocal deviations over time. To address this problem, the standard approach is to perform a clustering step (e.g. of audio data) followed by a data association step, when identity-rich sensor data is available. In this paper we show that this approach is not robust to noise in either type of sensor stream; to tackle this issue we propose a novel algorithm that jointly optimises the clustering and association process yielding up to three times higher identification precision than approaches that execute these steps sequentially. We demonstrate the performance benefits of our approach in two case studies, one with acoustic and MAC datasets that we collected from meetings in a non-residential building, and another from an online dataset from recorded radio interviews. Xiaoxuan Lu 0001, Hongkai Wen 0001, Sen Wang 0002, Andrew Markham, Agathoniki Trigoni |
IPSN | 2 |
| 2017 | Towards Self-supervised Face Labeling via Cross-modality AssociationabstractFace recognition has become the de facto authentication solution in a broad spectrum of applications, from smart buildings, to industrial monitoring and security services. However, in many of those real-world scenarios, tracking or identifying people with facial recognition is extremely challenging due to the variations in the environment such as lighting conditions, camera viewing angles and subject motion. For most of the state-of-the-art face recognition systems, they need to be trained on a large dataset containing a good variety of labelled face images to work well. However, collecting and manually labelling such datasets is difficult and time consuming, probably more so than developing the algorithms. In this paper, we propose a novel framework to automatically label user identities with their face images in smart spaces, exploiting the fact that the users tend to carry their smart devices while seen by the surveillance cameras. We evaluate our method on 10 users in a smart building setting, and the experimental results show that our method can achieve > 0.9 f1 score on average. Xiaoxuan Lu 0001, Xuan Kan, Stefano Rosa, Bowen Du 0002, Hongkai Wen 0001, Andrew Markham, Agathoniki Trigoni |
SenSys | 5 |
| 2016 | NaviGlass: Indoor Localisation Using Smart Glasses
Yongtuo Zhang, Wen Hu 0001, Weitao Xu, Hongkai Wen 0001, Chun Tung Chou |
EWSN | 4 |
| 2016 | Poster Abstract: Efficient Visual Positioning with Adaptive Parameter LearningabstractPositioning with vision sensors is gaining its popularity, since it is more accurate, and requires much less bootstrapping and training effort. However, one of the major limitations of the existing solutions is the expensive visual processing pipeline: on resource-constrained mobile devices, it could take up to tens of seconds to process one frame. To address this, we propose a novel learning algorithm, which adaptively discovers the place dependent parameters for visual processing, such as which parts of the scene are more informative, and what kind of visual elements one would expect, as it is employed more and more by the users in a particular setting. With such meta- information, our positioning system dynamically adjust its behaviour, to localise the users with minimum effort. Preliminary results show that the proposed algorithm can reduce the cost on visual processing significantly, and achieve sub-metre positioning accuracy. Hongkai Wen 0001, Sen Wang 0002, Ronnie Clark, Savvas Papaioannou, Agathoniki Trigoni |
IPSN | 1 |
| 2016 | Keyframe based large-scale indoor localisation using geomagnetic field and motion patternabstractThis paper studies indoor localisation problem by using low-cost and pervasive sensors. Most of existing indoor localisation algorithms rely on camera, laser scanner, floor plan or other pre-installed infrastructure to achieve sub-meter or sub-centimetre localisation accuracy. However, in some circumstances these required devices or information may be unavailable or too expensive in terms of cost or deployment. This paper presents a novel keyframe based Pose Graph Simultaneous Localisation and Mapping (SLAM) method, which correlates ambient geomagnetic field with motion pattern and employs low-cost sensors commonly equipped in mobile devices, to provide positioning in both unknown and known environments. Extensive experiments are conducted in large-scale indoor environments to verify that the proposed method can achieve high localisation accuracy similar to state-of-the-arts, such as vision based Google Project Tango. Sen Wang 0002, Hongkai Wen 0001, Ronald Clark, Agathoniki Trigoni |
IROS | 2 |
| 2016 | Robust occupancy inference with commodity WiFiabstractAccurate occupancy information of indoor environments is one of the key prerequisites for many pervasive and context-aware services, e.g. smart building/home systems. Some of the existing occupancy inference systems can achieve impressive accuracy, but they either require labour-intensive calibration phases, or need to install bespoke hardware such as CCTV cameras, which are privacy-intrusive by default. In this paper, we present the design and implementation of a practical end-to-end occupancy inference system, which requires minimum user effort, and is able to infer room-level occupancy accurately with commodity WiFi infrastructure. Depending on the needs of different occupancy information subscribers, our system is flexible enough to switch between snapshot estimation mode and continuous inference mode, to trade estimation accuracy for delay and communication cost. We evaluate the system on a hardware testbed deployed in a 600m2workspace with 25 occupants for 6 weeks. Experimental results show that the proposed system significantly outperforms competing systems in both inference accuracy and robustness. Xiaoxuan Lu 0001, Hongkai Wen 0001, Han Zou, Hao Jiang 0008, Lihua Xie 0001, Agathoniki Trigoni |
WiMob | 2 |
| 2015 | Opportunistic Radio Assisted Navigation for Autonomous Ground VehiclesabstractNavigating autonomous ground vehicles with visual sensors has many advantages - it does not rely on global maps, yet is accurate and reliable even in GPS-denied environments. However, due to the limitation of the camera field of view, one typically has to record a large number of visual experiences for practical navigation. In this paper, we explore new avenues in linking together visual experiences, by opportunistically harvesting and sharing a variety of radio signals emitted by surrounding stationary access points and mobile devices. We propose a novel navigation approach, which exploits side-channel information of co-location to thread up visually-separated experiences with short exploration phases. The proposed approach empowers users to trade travel time for manual navigation effort, allowing them to choose the itinerary that best serves their needs. We evaluate the proposed approach with data collected from a typical urban area, and show that it achieves much better navigation performance in both reach ability and cost, comparing with the state of the arts that only use visual information. Hongkai Wen 0001, Yiran Shen 0001, Savvas Papaioannou, Winston Churchill, Agathoniki Trigoni, Paul Newman 0001 |
DCOSS | 1 |
| 2015 | Accurate Positioning via Cross-Modality TrainingabstractIn this paper we propose a novel algorithm for tracking people in highly dynamic industrial settings, such as construction sites. We observed both short term and long term changes in the environment; people were allowed to walk in different parts of the site on different days, the field of view of fixed cameras changed over time with the addition of walls, whereas radio and magnetic maps proved unstable with the movement of large structures. To make things worse, the uniforms and helmets that people wear for safety make them very hard to distinguish visually, necessitating the use of additional sensor modalities. In order to address these challenges, we designed a positioning system that uses both anonymous and id-linked sensor measurements and explores the use of cross-modality training to deal with environment dynamics. The system is evaluated in a real construction site and is shown to outperform state of the art multi-target tracking algorithms designed to operate in relatively stable environments. Savvas Papaioannou, Hongkai Wen 0001, Zhuoling Xiao, Andrew Markham, Agathoniki Trigoni |
SenSys | 2 |
| 2015 | Robust Indoor Positioning With Lifelong LearningabstractInertial tracking and navigation systems have been playing an increasingly important role in indoor tracking and navigation. They have the competitive advantage of leveraging not requiring expensive infrastructure-only existing smart mobile devices with embedded inertial measurement units. When aided with other sources of information, such as radio data from existing WiFi/BLE infrastructure, and environment constraints from floor plans or radio maps, they often report great performance of 0.5-2 m. Given the promising results, what is it that prevents the widespread adoption of this tracking solution? We argue that pedestrian dead reckoning (PDR) techniques are often evaluated in a specific context and are not mature enough to handle variations in user motion, device type, device placement, or environment. They typically use a number of parameters that require careful context-specific tuning, which is labor intensive and requires expert knowledge. In this paper, we propose two novel approaches to address these problems. Our first contribution is a robust PDR algorithm, which is based on general physics principles that underpin human motion and is by design robust to context changes. The second contribution is a novel way of interaction between the PDR and map matching layers based on the principle of lifelong learning. Unlike traditional approaches where information flows unidirectionally from the PDR to the map matching layer, we introduce a feedback loop that can be used to automatically tune the parameters of the PDR algorithm. This is not dissimilar to the way that people improve their navigation skills when they repeatedly visit the same environment. Extensive experiments in multiple sites, with a variety of users, devices, and device placements, show that the combination of a robust PDR with a lifelong learning tracker can achieve submeter accuracy with no user effort for parameter tuning. Zhuoling Xiao, Hongkai Wen 0001, Andrew Markham, Agathoniki Trigoni |
IEEE J. Sel. Areas Commun. | 2 |
| 2015 | Accuracy Estimation for Sensor SystemsabstractIn most sensing applications, the measurements generated by sensor networks are noisy and usually annotated with some measure of uncertainty. The question that we address in this paper is how to estimate the accuracy of these uncertain sensor measurements. Existing studies on estimating the accuracy of uncertain measurements in real sensing applications are limited in three ways. First, they tend to be application-specific. Second, they typically employ learning techniques to estimate the parameters of sensor noise models, and ignore alternative state estimation approaches without learning. Third, they do not explore whether exploiting the dynamics of the monitored state can yield significant benefits. We address the above limitations as follows: we define the accuracy estimation problem in a general manner that applies to a broad spectrum of application scenarios. We present a general framework to address this problem, and show that the proposed framework can be implemented in a number of different ways. We evaluate and compare the different implementations in the context of two real sensing scenarios, and discuss how they trade accuracy for computation cost, and how this trade-off largely depends on the user's knowledge of the application scenario. Hongkai Wen 0001, Zhuoling Xiao, Andrew Markham, Agathoniki Trigoni |
IEEE Trans. Mob. Comput. | 1 |
| 2015 | Indoor Tracking Using Undirected Graphical ModelsabstractIndoor tracking and navigation is a fundamental need for pervasive and context-aware smartphone applications. Although indoor maps are becoming increasingly available, there is no practical and reliable indoor map matching solution available at present. We present MapCraft, a novel, robust and responsive technique that is extremely computationally efficient (running in under 10 ms on an Android smartphone), does not require training in different sites, and tracks well even when presented with very noisy sensor data. Key to our approach is expressing the tracking problem as a conditional random field (CRF), a technique which has had great success in areas such as natural language processing. Unlike directed graphical models like Hidden Markov Models, CRFs capture arbitrary constraints that express how well observations support state transitions, given map constraints. In addition, we show how to further improve tracking accuracy, by tuning the parameters of the motion sensing model using an unsupervised EM-style optimization scheme. Extensive experiments in multiple sites show how MapCraft outperforms state-of-the art approaches, demonstrating excellent tracking error and accurate reconstruction of tortuous trajectories with zero training effort. As proof of its robustness, we also demonstrate how it is able to accurately track the position of a user from accelerometer and magnetometer measurements only (i.e., gyroand Wi-Fi-free). We believe that such an energy-efficient approach will enable always-on background localisation, enabling a new era of location-aware applications to be developed. Zhuoling Xiao, Hongkai Wen 0001, Andrew Markham, Agathoniki Trigoni |
IEEE Trans. Mob. Comput. | 2 |
| 2015 | Non-Line-of-Sight Identification and Mitigation Using Received Signal StrengthabstractIndoor wireless systems often operate under non-line-of-sight (NLOS) conditions that can cause ranging errors for location-based applications. As such, these applications could benefit greatly from NLOS identification and mitigation techniques. These techniques have been primarily investigated for ultra-wide band (UWB) systems, but little attention has been paid to WiFi systems, which are far more prevalent in practice. In this study, we address the NLOS identification and mitigation problems using multiple received signal strength (RSS) measurements from WiFi signals. Key to our approach is exploiting several statistical features of the RSS time series, which are shown to be particularly effective. We develop and compare two algorithms based on machine learning and a third based on hypothesis testing to separate LOS/NLOS measurements. Extensive experiments in various indoor environments show that our techniques can distinguish between LOS/NLOS conditions with an accuracy of around 95%. Furthermore, the presented techniques improve distance estimation accuracy by 60% as compared to state-of-the-art NLOS mitigation techniques. Finally, improvements in distance estimation accuracy of 50% are achieved even without environment-specific training data, demonstrating the practicality of our approach to real world implementations. Zhuoling Xiao, Hongkai Wen 0001, Andrew Markham, Agathoniki Trigoni, Phil Blunsom, Jeff Frolik |
IEEE Trans. Wirel. Commun. | 2 |
| 2014 | Robust pedestrian dead reckoning (R-PDR) for arbitrary mobile device placementabstractPedestrian dead reckoning, especially on smart-phones, is likely to play an increasingly important role in indoor tracking and navigation, due to its low cost and ability to work without any additional infrastructure. A challenge however, is that positioning, both in terms of step detection and heading estimation, must be accurate and reliable, even when the use of the device is so varied in terms of placement (e.g. handheld or in a pocket) or orientation (e.g holding the device in either portrait or landscape mode). Furthermore, the placement can vary over time as a user performs different tasks, such as making a call or carrying the device in a bag. A second challenge is to be able to distinguish between a true step and other periodic motion such as swinging an arm or tapping when the placement and orientation of the device is unknown. If this is not done correctly, then the PDR system typically overestimates the number of steps taken, leading to a significant long term error. We present a fresh approach, robust PDR (R-PDR), based on exploiting how bipedal motion impacts acquired sensor waveforms. Rather than attempting to recognize different placements through sensor data, we instead simply determine whether the motion of one or both legs impact the measurements. In addition, we formulate a set of techniques to accurately estimate the device orientation, which allows us to very accurately (typically over 99%) reject false positives. We demonstrate that regardless of device placement, we are able to detect the number of steps taken with >99.4% accuracy. R-PDR thus addresses the two main limitations facing existing PDR techniques. Zhuoling Xiao, Hongkai Wen 0001, Andrew Markham, Agathoniki Trigoni |
IPIN | 2 |
| 2014 | Lightweight map matching for indoor localisation using conditional random fields
Zhuoling Xiao, Hongkai Wen 0001, Andrew Markham, Agathoniki Trigoni |
IPSN | 2 |
| 2014 | Fusion of Radio and Camera Sensor Data for Accurate Indoor PositioningabstractIndoor positioning systems have received a lot of attention recently due to their importance for many location-based services, e.g. indoor navigation and smart buildings. Lightweight solutions based on WiFi and inertial sensing have gained popularity, but are not fit for demanding applications, such as expert museum guides and industrial settings, which typically require sub-meter location information. In this paper, we propose a novel positioning system, RAVEL (Radio And Vision Enhanced Localization), which fuses anonymous visual detections captured by widely available camera infrastructure, with radio readings (e.g. WiFi radio data). Although visual trackers can provide excellent positioning accuracy, they are plagued by issues such as occlusions and people entering/exiting the scene, preventing their use as a robust tracking solution. By incorporating radio measurements, visually ambiguous or missing data can be resolved through multi-hypothesis tracking. We evaluate our system in a complex museum environment with dim lighting and multiple people moving around in a space cluttered with exhibit stands. Our experiments show that although the WiFi measurements are not by themselves sufficiently accurate, when they are fused with camera data, they become a catalyst for pulling together ambiguous, fragmented, and anonymous visual tracklets into accurate and continuous paths, yielding typical errors below 1 meter. Savvas Papaioannou, Hongkai Wen 0001, Andrew Markham, Agathoniki Trigoni |
MASS | 2 |
| 2013 | Comparison of Accuracy Estimation Approaches for Sensor NetworksabstractWith sensor technology gaining maturity and becoming ubiquitous, we are experiencing an unprecedented wealth of sensor data. In most sensing applications, users receive sensor measurements, which are prone to error. As a result, they are often annotated with some measure of uncertainty, such as the distribution variance or a confidence interval, and will be hereafter referred to as probabilistic measurements. The question that we address in this paper is how to estimate the accuracy of these probabilistic measurements, that is, how far they lie from the ground truth of the measured attribute. Existing studies on estimating the accuracy of probabilistic measurements in real sensing applications are limited in three ways. First, they tend to be application-specific. Second, they typically employ learning techniques to estimate the parameters of sensor noise models, and ignore alternative approaches that rely on simple state estimation without learning. Third, they do not explore whether exploiting the dynamics of the monitored state can yield significant benefits in terms of accuracy estimation. In this paper, we address the above limitations as follows: We define the problem of accuracy estimation in a general way that applies to a wide spectrum of application scenarios. We then propose a taxonomy of accuracy estimation techniques, which include both state estimation and parameter learning. These techniques are further subdivided into static and dynamic, depending on whether they exploit knowledge of system dynamics. All different approaches in the taxonomy are then applied and compared with each other in the context of two real sensing applications. We discuss how they trade accuracy for computation cost, and how this tradeoff largely depends on the user's knowledge of the application scenario. Hongkai Wen 0001, Zhuoling Xiao, Andrew Colquhoun Symington, Andrew Markham, Agathoniki Trigoni |
DCOSS | 1 |
| 2013 | On Assessing the Accuracy of Positioning Systems in Indoor Environments
Hongkai Wen 0001, Zhuoling Xiao, Agathoniki Trigoni, Phil Blunsom |
EWSN | 1 |
| 2013 | Identification and mitigation of non-line-of-sight conditions using received signal strengthabstractVarious applications, such as localisation of persons and objects could benefit greatly from non-line-of-sight (NLOS) identification and mitigation techniques. However, such techniques have been primarily investigated for ultra-wide band (UWB) signals, leaving the area of WiFi signals untouched. In this study, we propose two accurate approaches using only received signal strength (RSS) measurements from WiFi signals to identify NLOS conditions and mitigate the effects. We first explore several features from the RSS which are later demonstrated as very effective in identifying and mitigating NLOS conditions. After that, we develop and compare two major optimization problems based on a machine learning technique and hypothesis testing according to different user requirements and information available. Extensive experiments in various indoor environments have shown that our techniques can not only accurately distinguish between LOS/NLOS conditions, but also mitigate the impact of NLOS conditions as well. Zhuoling Xiao, Hongkai Wen 0001, Andrew Markham, Agathoniki Trigoni, Phil Blunsom, Jeff Frolik |
WiMob | 2 |
| 2012 | Operational transformation for orthogonal conflict resolution in real-time collaborative 2d editing systemsabstractOperational Transformation (OT) is commonly used for conflict resolution in real-time collaborative applications, but none of existing OT techniques is able to solve a special type of conflict - orthogonal conflict, which may occur when concurrent operations are inserting/deleting an arbitrary number of objects in different dimensions of a two-dimensional (2D) workspace, such as spreadsheet documents. This paper is the first to identify and solve the orthogonal conflict problem by extending OT with a new capability of resolving 2D conflicts. Extending OT from one- to two-dimensional conflict resolution is fundamental to the theory and application of OT, and technically challenging as well because 2D orthogonal conflict is different from but intimately related to the one-dimensional positional shifting conflict and necessitates new and integral solutions for multi-dimensional conflicts. In this paper, we present formal definitions of orthogonal conflict, pseudo-code description, design rationale analysis, and correctness verification and complexity analysis of the 2DOT solution. Chengzheng Sun, Hongkai Wen 0001, Hongfei Fan |
CSCW | 2 |
| 2012 | Ranking Query Answers in Probabilistic Databases: Complexity and Efficient AlgorithmsabstractIn many applications of probabilistic databases, the probabilities are mere degrees of uncertainty in the data and are not otherwise meaningful to the user. Often, users care only about the ranking of answers in decreasing order of their probabilities or about a few most likely answers. In this paper, we investigate the problem of ranking query answers in probabilistic databases. We give a dichotomy for ranking in case of conjunctive queries without repeating relation symbols: it is either in polynomial time or NP-hard. Surprisingly, our syntactic characterisation of tractable queries is not the same as for probability computation. The key observation is that there are queries for which probability computation is \#P-hard, yet ranking can be computed in polynomial time. This is possible whenever probability computation for distinct answers has a common factor that is hard to compute but irrelevant for ranking. We complement this tractability analysis with an effective ranking technique for conjunctive queries. Given a query, we construct a share plan, which exposes sub queries whose probability computation can be shared or ignored across query answers. Our technique combines share plans with incremental approximate probability computation of sub queries. We implemented our technique in the SPROUT query engine and report on performance gains of orders of magnitude over Monte Carlo simulation using FPRAS and exact probability computation based on knowledge compilation. Dan Olteanu, Hongkai Wen 0001 |
ICDE | 2 |