Shubham Jain 0003

dblp:132/6759-3 · DBLP profile ↗
← Back
26ranked-venue papers
4as first author
19since 2021 · last 2025
0000-0002-4864-6420ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 16 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2025 A Landmark-Aware Visual Navigation Dataset for Map Representation Learning
abstract
Map representations learned by expert demonstrations have shown promising research value. However, the field of visual navigation still faces challenges due to the lack of real-world human-navigation datasets that can support efficient, supervised, representation learning of environments. We present a Landmark-Aware Visual Navigation (LAVN) dataset to allow for supervised learning of human-centric exploration policies and map building. We collect RGBD observation and human point-click pairs as a human annotator explores virtual and real-world environments with the goal of full coverage exploration of the space. The human annotators also provide distinct landmark examples along each trajectory, which we intuit will simplify the task of map or graph building and localization. These human point-clicks serve as direct supervision for waypoint prediction when learning to explore in environments. Our dataset covers a wide spectrum of scenes, including rooms in indoor environments, as well as walkways outdoors. We releaseour dataset with detailed documentation at https://huggingface.co/datasets/visnavdataset/lavn (DOI: l0.57967/hf/2386) and a plan for long-term preservation.
Faith M. Johnson, Kristin J. Dana, Bryan Bo Cao, Shubham Jain 0003, Ashwin Ashok
HRI4
2025 Few-Class Arena: A Benchmark for Efficient Selection of Vision Models and Dataset Difficulty Measurement
abstract
We propose Few-Class Arena (FCA), as a unified benchmark with focus on testing efficient image classification models for few classes. A wide variety of benchmark datasets with many classes (80-1000) have been created to assist Computer Vision architectural evolution. An increasing number of vision models are evaluated with these many-class datasets. However, real-world applications often involve substantially fewer classes of interest (2-10). This gap between many and few classes makes it difficult to predict performance of the few-class applications using models trained on the available many-class datasets. To date, little has been offered to evaluate models in this Few-Class Regime. We conduct a systematic evaluation of the ResNet family trained on ImageNet subsets from 2 to 1000 classes, and test a wide spectrum of Convolutional Neural Networks and Transformer architectures over ten datasets by using our newly proposed FCA tool. Furthermore, to aid an up-front assessment of dataset difficulty and a more efficient selection of models, we incorporate a difficulty measure as a function of class similarity. FCA offers a new tool for efficient machine learning in the Few-Class Regime, with goals ranging from a new efficient class similarity proposal, to lightweight model architecture design, to a new scaling law. FCA is user-friendly and can be easily extended to new models and datasets, facilitating future research work. Our benchmark is available at https://github.com/bryanbocao/fca.
Bryan Bo Cao, Lawrence O'Gorman, Michael Coss, Shubham Jain 0003
ICLR4
2024 Hand Gesture Recognition for Blind Users by Tracking 3D Gesture Trajectory
abstract
Hand gestures provide an alternate interaction modality for blind users and can be supported using commodity smartwatches without requiring specialized sensors. The enabling technology is an accurate gesture recognition algorithm, but almost all algorithms are designed for sighted users. Our study shows that blind user gestures are considerably diferent from sighted users, rendering current recognition algorithms unsuitable. Blind user gestures have high inter-user variance, making learning gesture patterns difcult without large-scale training data. Instead, we design a gesture recognition algorithm that works on a 3D representation of the gesture trajectory, capturing motion in free space. Our insight is to extract a micro-movement in the gesture that is user-invariant and use this micro-movement for gesture classifcation. To this end, we develop an ensemble classifer that combines image classifcation with geometric properties of the gesture. Our evaluation demonstrates a 92% classifcation accuracy, surpassing the next best state-of-the-art which has an accuracy of 82%.
Prerna Khanna, I. V. Ramakrishnan, Shubham Jain 0003, Xiaojun Bi 0001, Aruna Balasubramanian
CHI3
2024 A Lightweight Measure of Classification Difficulty from Application Dataset Characteristics
Bryan Bo Cao, Lawrence O'Gorman, Michael Coss, Shubham Jain 0003
ICPR (5)5
2024 VideoJam: Self-Balancing Architecture for Live Video Analytics
abstract
Edge-based live video analytics are a promising approach to reduce bandwidth overheads caused by the transmission of raw video streams to the cloud. However, the limited resources available on edge devices make it challenging to successfully process video streams in real-time. This gets further exacerbated when attempting to process video streams from mobile cameras. While mobile cameras are a desirable source of information, thanks to them being in the right place at the right time, they are inherently dynamic and unpredictable. To address these challenges, we propose VideoJam, a decentralized load balancing solution for live video analytics. VideoJam uses a set of load balancers to balance incoming video traffic across replicas without the need of centralized coordination. Exploiting the inherent load dynamicity generated by different video sources, VideoJam predicts the incoming load for each processing component and offloads excessive traffic to less-loaded neighbors. Further, VideoJam operates independently of deployed configurations and cameras present in the system, dynamically adapting to handle load changes and balance video traffic across available resources. Our evaluation shows that VideoJam can adapt to different mixes of mobile and fixed cameras, as well as quickly adapting to configuration changes occurring at runtime. Compared to state-of-the-art solutions, VideoJam achieves 2.91× lower response time, while reducing video data loss by more than 4.64× and generating lower bandwidth overheads.
Youssouph Faye, Francescomaria Faticanti, Shubham Jain 0003, Francesco Bronzino
SEC3
2024 OVIDA: Orchestrator for Video Analytics on Disaggregated Architecture
abstract
Millions of video cameras are deployed globally across major cities for learning-based video analytic (VA) applications, such as object detection. Video streams from the cameras are either sent over the wide-area network to be processed by the cloud or are (at least partially) processed in a local edge workstation, incurring significant latency and elevated financial costs. In this paper, to minimize reliance on the cloud and overcome the unavailability of high-compute workstations on edge, we investigate the use of heterogeneous and distributed embedded devices as edge nodes shared by multiple cameras to fully serve the video processing needs of a VA application (without requiring cloud support). We present OVIDA, an edge-only orchestrator to deploy VA application(s) on a distributed edge environment to maximize accuracy. Given the resource-constrained nature of edge nodes, OVIDA disaggregates the VA application pipeline into multiple modules. OVIDA's core functionality and contributions are: (i) optimizing the placement and replication of the VA application modules across the edge nodes to maximize the throughput, and in turn, accuracy; and (ii) an adaptive model selection algorithm for VA modules based on accuracy-throughput tradeoff to maximize accuracy in response to varying load conditions. To further improve performance, OVIDA employs a central-queue-based design (instead of the usual push-based design), which also obviates the need for complex load balancing algorithms. We implement OVIDA on top of Kubernetes and evaluate its performance for three VA applications, supported over a heterogeneous edge cluster under varying network conditions. When compared against several baselines in our evaluation, we achieve throughput and accuracy gains of at least 51% and 28%.
Manavjeet Singh, Sri Pramodh Rachuri, Bryan Bo Cao, Venkata Bhumireddy, Francesco Bronzino, Samir Ranjan Das, Anshul Gandhi, Shubham Jain 0003
SEC9
2024 Representation Similarity: A Better Guidance of DNN Layer Sharing for Edge Computing without Training
abstract
Edge computing has emerged as an alternative to reduce transmission and processing delay and preserve privacy of the video streams. However, the ever-increasing complexity of Deep Neural Networks (DNNs) used in video-based applications (e.g. object detection) exerts pressure on memory-constrained edge devices. Model merging is proposed to reduce the DNNs' memory footprint by keeping only one copy of merged layers' weights in memory. In existing model merging techniques, (i) only architecturally identical layers can be shared; (ii) requires computationally expensive retraining in the cloud; (iii) assumes the availability of ground truth for retraining. The re-evaluation of a merged model's performance, however, requires a validation dataset with ground truth, typically runs at the cloud. Common metrics to guide the selection of shared layers include the size or computational cost of shared layers or representation size. We propose a new model merging scheme by sharing representations (i.e., outputs of layers) at the edge, guided by representation similarity S. We show that S is extremely highly correlated with merged model's accuracy with Pearson Correlation Coefficient |r| > 0.94 than other metrics, demonstrating that representation similarity can serve as a strong validation accuracy indicator without ground truth. We present our preliminary results of the newly proposed model merging scheme with identified challenges, demonstrating a promising research future direction.
Bryan Bo Cao, Manavjeet Singh, Anshul Gandhi, Samir Ranjan Das, Shubham Jain 0003
MobiCom6
2024 Enabling Accessible and Ubiquitous Interaction in Next-Generation Wearables: An Unvoiced Speech Approach
abstract
As wearable devices increase, there's a growing need for intuitive, private, and accessible interaction methods. This position paper builds on the research on unvoiced speech interaction and authentication to propose a vision for interaction in next-generation wearables. This paper draws upon our previous work on unvoiced speech interfaces that leverage jaw movements and facial vibrations for command recognition and user authentication. We argue that unvoiced speech interaction can provide a robust, privacy-preserving, and noise-resistant alternative to traditional interfaces, enhancing accessibility and offering discrete interaction in public spaces. We discuss the potential integration of these systems into commercial devices and explore gesture-based interactions as an alternative to touch. Additionally, we discuss the future direction of unvoiced speech interfaces. This paper sets the stage for implementing unvoiced speech and gesture-based interaction in mainstream wearables in our daily interactions with technology.
Tanmay Srivastava, Prerna Khanna, Shijia Pan, V. P. Nguyen, Shubham Jain 0003
MobiCom5
2024 Unvoiced: Designing an LLM-assisted Unvoiced User Interface using Earables
abstract
We present Unvoiced, a novel unvoiced user interface that leverages jaw motion to enable users to silently interact with their devices using earables. The core idea is to translate low-frequency jaw motion signals into high-frequency information-rich mel spectrograms. Our proposed cross-modal translation incorporates phonetic, contextual, and syntactic information, while the specialized loss function optimizes for these linguistic features. This ensures that the generated spectrograms capture nuanced speech characteristics. Evaluated for 19 users across four tasks, Unvoiced demonstrates >94% task completion rate and <9% word error rate for over 90% of phrases. Further, Unvoiced maintains >90% task completion rate in noisy conditions.
Tanmay Srivastava, Prerna Khanna, Shijia Pan, Phuc Nguyen 0002, Shubham Jain 0003
SenSys5
2024 Poster Unvoiced: Designing an Unvoiced User Interface using Earables and LLMs
abstract
This poster presents the design and implementation of Unvoiced, a silent speech interaction system. Unvoiced transforms subtle jaw movements into rich speech spectrograms, enabling seamless and private device interaction. Our system captures low-frequency jaw motion signals using ear-worn IMUs and translates them into high-fidelity mel-spectrograms through cross-modal translation techniques. By incorporating phonetic, contextual, and syntactic information, Unvoiced generates high-fidelity spectrograms that existing speech recognition systems can process. In our evaluation with 19 users across four common tasks, Unvoiced achieved a remarkable >94% task completion rate and <9% Word Error Rate (WER) for over 90% of phrases, maintaining robust performance even in noisy conditions.
Tanmay Srivastava, Prerna Khanna, Shijia Pan, Phuc Nguyen 0002, Shubham Jain 0003
SenSys5
2023 AccessWear: Making Smartphone Applications Accessible to Blind Users
abstract
In this paper, we present AccessWear, a system that improves the accessibility of smartphone touchscreen interactions for blind users using smartwatch gestures. Our system design is human-centered, namely, it incorporates the design goals that were learned from a formative user study with 9 blind participants. The formative study showed that blind users liked the idea of using smartwatch gestures as an alternative: 4 participants liked that when using smart-watch gestures, they did not have to bring their expensive phones out in public and 6 participants liked that smart-watch gestures can be performed with one-hand, as the other hand is usually occupied in holding a cane or a guide dog. Even though there are several advantages to smartwatch gestures, our study also shows that gestures performed by blind users have different patterns compared to sighted users, making gesture recognition more challenging. To this end, AccessWear makes two contributions. The first is a gesture recognition system that works specifically for blind users that is lightweight and does not require per-person training. The second is a near-zero-effort gesture replacement system that does not require any changes to the original application. AccessWear uses input virtualization techniques so that a given gesture can replace the touchscreen input seamlessly. We implement AccessWear on an Android smartphone and Android watch. We perform a quantitative and qualitative study with 8 blind participants. Our study shows that AccessWear can recognize gestures with a 92% accuracy and the end-to-end latency when using an alternate gesture was 53 msec on average. The qualitative study shows that when participants perform a task, consisting of a series of gestures, the system is robust, does not have perceived delays, and does not add physical or mental load on the users.
Prerna Khanna, Shirin Feiz, Jian Xu 0013, I. V. Ramakrishnan, Shubham Jain 0003, Xiaojun Bi 0001, Aruna Balasubramanian
MobiCom5
2023 Jawthenticate: Microphone-free Speech-based Authentication using Jaw Motion and Facial Vibrations
abstract
In this paper, we present Jawthenticate, an earable system that authenticates a user using audible or inaudible speech without using a microphone. This system can overcome the shortcomings of traditional voice-based authentication systems like unreliability in noisy conditions and spoofing using microphone-based replay attacks. Jawthenticate derives distinctive speech-related features from the jaw motion and associated facial vibrations. This combination of features makes Jawthenticate resilient to vocal imitations as well as camera-based spoofing. We use these features to train a two-class SVM classifier for each user. Our system is invariant to the content and language of speech. In a study conducted with 41 subjects, who speak different native languages, Jawthenticate achieves a Balanced Accuracy (BAC) of 97.07%, True Positive Rate (TPR) of 97.75%, and True Negative Rate (TNR) of 96.4% with just 3 seconds of speech data.
Tanmay Srivastava, Shijia Pan, Phuc Nguyen 0002, Shubham Jain 0003
SenSys4
2022 RadioTransformer: A Cascaded Global-Focal Transformer for Visual Attention-Guided Disease Classification
Moinak Bhattacharya, Shubham Jain 0003, Prateek Prasanna
ECCV (21)2
2022 Vi-Fi: Associating Moving Subjects across Vision and Wireless Sensors
abstract
In this paper, we present Vi-Fi, a multi-modal system that leverages a user's smartphone WiFi Fine Timing Measurements (FTM) and inertial measurement unit (IMU) sensor data to associate the user detected on a camera footage with their corresponding smartphone identifier (e.g. WiFi MAC address). Our approach uses a recurrent multi-modal deep neural network that exploits FTM and IMU measurements along with distance between user and camera (depth information) to learn affinity matrices. As a baseline method for comparison, we also present a traditional non deep learning approach that uses bipartite graph matching. To facilitate evaluation, we collected a multi-modal dataset that comprises camera videos with depth information (RGB-D), WiFi FTM and IMU measurements for multiple participants at diverse real-world settings. Using association accuracy as the key metric for evaluating the fidelity of Vi-Fi in associating human users on camera feed with their phone IDs, we show that Vi-Fi achieves between 81% (real-time) to 91% (offline) association accuracy.
Hansi Liu, Abrar Alali, Mohamed Ibrahim Ahmed 0001, Bryan Bo Cao, Nicholas Meegan, Marco Gruteser, Shubham Jain 0003, Kristin J. Dana, Ashwin Ashok, Bin Cheng 0002, Hongsheng Lu
IPSN8
2022 GazeRadar: A Gaze and Radiomics-Guided Disease Localization Framework
Moinak Bhattacharya, Shubham Jain 0003, Prateek Prasanna
MICCAI (3)2
2022 Leveraging earables for unvoiced command recognition
abstract
We demonstrate an ear-worn technology that recognizes unvoiced human commands by tracking jaw motion. The ear-worn system is designed to achieve continual unvoiced command recognition for robust human-computer interaction (HCI) applications. First, the system reliably extracts the jaw motion signals buried under the noise caused by head motion, walking, and other motion artifacts to track single secondary voice articulator (i.e., word). Then, learning from linguistics and human speech anatomy, we design a novel algorithm that localizes the phonemes in the command, and reconstructs the word. We evaluate the proposed system in real-world experiments with 15 volunteers. Our preliminary results show that the proposed system obtains a word recognition accuracy of 95.6% in noise-free conditions and 93.2% and 91.6%, while head nodding and walking.
Tanmay Srivastava, Prerna Khanna, Shijia Pan, Phuc Nguyen 0002, Shubham Jain 0003
MobiSys5
2022 ViTag: Online WiFi Fine Time Measurements Aided Vision-Motion Identity Association in Multi-person Environments
abstract
In this paper, we present ViTag to associate user identities across multimodal data, particularly those obtained from cameras and smartphones. ViTag associates a sequence of vision tracker generated bounding boxes with Inertial Mea-surement Unit (IMU) data and Wi-Fi Fine Time Measurements (FTM) from smartphones. We formulate the problem as association by sequence to sequence (seq2seq) translation. In this two-step process, our system first performs cross-modal translation using a multimodal LSTM encoder-decoder network (X-Translator) that translates one modality to another, e.g. recon-structing IMU and FTM readings purely from camera bounding boxes. Second, an association module finds identity matches between camera and phone domains, where the translated modality is then matched with the observed data from the same modality. In contrast to existing works, our proposed approach can associate identities in multi-person scenarios where all users may be performing the same activity. Extensive experiments in real-world indoor and outdoor environments demonstrate that online association on camera and phone data (IMU and FTM) achieves an average Identity Precision Accuracy (IDP) of 88.39% on a 1 to 3 seconds window, outperforming the state-of-the-art Vi-Fi (82.93%). Further study on modalities within the phone domain shows the FTM can improve association performance by 12.56% on average. Finally, results from our sensitivity experiments demonstrate the robustness of ViTag under different noise and environment variations.
Bryan Bo Cao, Abrar Alali, Hansi Liu, Nicholas Meegan, Marco Gruteser, Kristin J. Dana, Ashwin Ashok, Shubham Jain 0003
SECON8
2022 A Survey of Parking Solutions for Smart Cities
abstract
Existing surveys look at parking solutions from the perspective of sensors, communication protocols, and the hardware-software interface. While this is a worthwhile approach, it suffers from three obvious shortcomings, namely that present-day sensors are likely to become obsolete in a few years, communication protocols get discontinued, and present-day software will almost certainly not run on tomorrow’s platforms. Consequently, these approaches are not promising for the Smart Cities of the near future. Unlike previous surveys, we look at parking in Smart Cities through the lens of market-based allocation of goods and services. In competitive markets prices act as signals used to allocate goods to those who value them most. In the case of parking spots, some drivers are willing to pay higher prices for the use of those parking spots that offer them the highest utility. What makes our survey unique is that we are looking at the recent literature with an eye for market-oriented solutions including pricing as an instrument for shaping traffic and for incentivizing socially-desirable driver behavior. We believe that one of the important contributions of any survey paper, over and above being a compendium of known art, is to suggest new lines of research. With this in mind, we have peppered the manuscript with slightly unorthodox perspectives. These perspectives are intended to be thought-provoking and to open new avenues for possible investigations.
Meshari Aljohani, Stephan Olariu, Abrar Alali, Shubham Jain 0003
IEEE Trans. Intell. Transp. Syst.4
2021 Lost and Found!: associating target persons in camera surveillance footage with smartphone identifiers
abstract
We demonstrate an application of finding target persons on a surveillance video. Each visually detected participant is tagged with a smartphone ID and the target person with the query ID is highlighted. This work is motivated by the fact that establishing associations between subjects observed in camera images and messages transmitted from their wireless devices can enable fast and reliable tagging. This is particularly helpful when target pedestrians need to be found on public surveillance footage, without the reliance on facial recognition. The underlying system uses a multi-modal approach that leverages WiFi Fine Timing Measurements (FTM) and inertial sensor (IMU) data to associate each visually detected individual with a corresponding smartphone identifier. These smartphone measurements are combined strategically with RGB-D information from the camera, to learn affinity matrices using a multi-modal deep learning network.
Hansi Liu, Abrar Alali, Mohamed Ibrahim Ahmed 0001, Marco Gruteser, Shubham Jain 0003, Kristin J. Dana, Ashwin Ashok, Bin Cheng 0002, Hongsheng Lu
MobiSys6
2019 FusionEye: Perception Sharing for Connected Vehicles and its Bandwidth-Accuracy Trade-offs
abstract
Automated driving and advanced driver assistance systems benefit from complete understandings of traffic scenes around vehicles. Existing systems gather such data through cameras and other sensors in vehicles but scene understanding can be limited due to the sensing range of sensors or occlusion from other objects. To gather information beyond the view of one vehicle, we propose and explore FusionEye - a connected vehicle system that allows multiple vehicles to share perception data over vehicle-to-vehicle communications and collaboratively merge this data into a more complete traffic scene. FusionEye uses a self-adaptive topology merging algorithm based on bipartite graph. We explore its network bandwidth requirements and the trade-off with merging accuracy. Experimental results show that FusionEye creates more complete scenes and achieves a merging accuracy of 88% with 5% packet drop rate and transmission latency around 200ms. We show that richer vehicle descriptors offer only marginal accuracy improvements compared to lower communication overhead options.
Hansi Liu, Shubham Jain 0003, Mohannad Murad, Marco Gruteser, Fan Bai 0002
SECON3
2019 Recognizing Textures with Mobile Cameras for Pedestrian Safety Applications
abstract
As smartphone rooted distractions become commonplace, the lack of compelling safety measures has led to a rise in the number of injuries to distracted walkers. Various solutions address this problem by sensing a pedestrian's walking environment. Existing camera-based approaches have been largely limited to obstacle detection and other forms of object detection. Instead, we present TerraFirma, an approach that performs material recognition on the pedestrian's walking surface. We explore, first, how well commercial off-the-shelf smartphone cameras can learn texture to distinguish among paving materials in uncontrolled outdoor urban settings. Second, we aim at identifying when a distracted user is about to enter the street, which can be used to support safety functions such as warning the user to be cautious. To this end, we gather a unique dataset of street/sidewalk imagery from a pedestrian's perspective, that spans major cities like New York, Paris, and London. We demonstrate that modern phone cameras can be enabled to distinguish materials of walking surfaces in urban areas with more than 90 percent accuracy, and accurately identify when pedestrians transition from sidewalk to street.
Shubham Jain 0003, Marco Gruteser
IEEE Trans. Mob. Comput.1
2017 Panoptes: servicing multiple applications simultaneously using steerable cameras
abstract
Steerable surveillance cameras offer a unique opportunity to support multiple vision applications simultaneously. However, state-of-art camera systems do not support this as they are often limited to one application per camera. We believe that we should break the one-to-one binding between the steerable camera and the application. By doing this we can quickly move the camera to a new view needed to support a different vision application. When done well, the scheduling algorithm can support a larger number of applications over an existing network of surveillance cameras. With this in mind we developed Panoptes, a technique that virtualizes a camera view and presents a different fixed view to different applications. A scheduler uses camera controls to move the camera appropriately providing the expected view for each application in a timely manner, minimizing the impact on application performance. Experiments with a live camera setup demonstrate that Panoptes can support multiple applications, capturing up to 80% more events of interest in a wide scene, compared to a fixed view camera.
Shubham Jain 0003, Viet Nguyen, Marco Gruteser, Paramvir Bahl
IPSN1
2015 LookUp: Enabling Pedestrian Safety Services via Shoe Sensing
abstract
Motivated by safety challenges resulting from distracted pedestrians, this paper presents a sensing technology for fine-grained location classification in an urban environment. It seeks to detect the transitions from sidewalk locations to in-street locations, to enable applications such as alerting texting pedestrians when they step into the street. In this work, we use shoe-mounted inertial sensors for location classification based on surface gradient profile and step patterns. This approach is different from existing shoe sensing solutions that focus on dead reckoning and inertial navigation. The shoe sensors relay inertial sensor measurements to a smartphone, which extracts the step pattern and the inclination of the ground a pedestrian is walking on. This allows detecting transitions such as stepping over a curb or walking down sidewalk ramps that lead into the street. We carried out walking trials in metropolitan environments in United States (Manhattan) and Europe (Turin). The results from these experiments show that we can accurately determine transitions between sidewalk and street locations to identify pedestrian risk.
Shubham Jain 0003, Carlo Borgiattino, Yanzhi Ren, Marco Gruteser, Yingying Chen 0001, Carla Fabiana Chiasserini
MobiSys1
2015 Video: LookUp!: Enabling Pedestrian Safety Services via Shoe Sensing
abstract
This video is a demonstration of the work discussed in our full paper available in the MobiSys'15 proceedings. The video illustrates a sensing technology for fine-grained location classification in an urban environment, for enhancing pedestrian safety. Our system seeks to detect the transitions from sidewalk locations to in-street locations, to enable applications such as alerting texting pedestrians when they step into the street. Existing positioning technologies are not sufficiently precise to allow distinguishing a position on the sidewalk from a position in the street, as explored in our previous work. To this end, we use shoe-mounted inertial sensors for location classification based on surface gradient profile and step patterns. This approach is different from existing shoe sensing solutions that focus on dead reckoning and inertial navigation. The shoe sensors relay inertial sensor measurements to a smartphone, which extracts the step pattern and the inclination of the ground a pedestrian is walking on. This allows detecting transitions such as stepping over a curb or walking down sidewalk ramps that lead into the street. We carried out walking trials in metropolitan environments in United States (Manhattan) and Europe (Turin). The results from these experiments show that we can accurately determine transitions between sidewalk and street locations to identify pedestrian risk.
Shubham Jain 0003, Carlo Borgiattino, Yanzhi Ren, Marco Gruteser, Yingying Chen 0001, Carla Fabiana Chiasserini
MobiSys1
2014 Towards City-Scale Smartphone Sensing of Potentially Unsafe Pedestrian Movements
abstract
This paper proposes large scale collection of pedestrian movement data to promote pedestrian safety in our rapidly developing urban environments. As a first step, we develop and test algorithms for sensing unsafe pedestrian movements. With distracted pedestrian fatalities on the rise, and larger than ever use of smart devices, we propose to use the smartphone to protect pedestrians by leveraging the in-built inertial sensors on the smartphone. We discuss how to use these sensors for recognizing user movements that could be potentially risky when walking on the street, while also accounting for different phone orientations. We introduce a simple path prediction technique and use this to compute potential street crossings. In order to evaluate our algorithms, we conducted walking trials and collected data from all relevant sensors. Initial tests indicate a 90.5% success rate in predicting that a pedestrians trajectory will cross a road.
Trisha Datta, Shubham Jain 0003, Marco Gruteser
MASS2
2014 Capacity of pervasive camera based communication under perspective distortions
abstract
Cameras are ubiquitous and increasingly being used not just for capturing images but also for communicating information. For example, the pervasive QR codes can be viewed as communicating a short code to camera-equipped sensors and recent research has explored using screen-to-camera communications for larger data transfers. Such communications could be particularly attractive in pervasive camera based applications, where such camera communications can reuse the existing camera hardware and also leverage from the large pixel array structure for high data-rate communication. While several prototypes have been constructed, the fundamental capacity limits of this novel communication channel in all but the simplest scenarios remains unknown. The visual medium differs from RF in that the information capacity of this channel largely depends on the perspective distortions while multipath becomes negligible. In this paper, we create a model of this communication system to allow predicting the capacity based on receiver perspective (distance and angle to the transmitter). We calibrate and validate this model through lab experiments wherein information is transmitted from a screen and received with a tablet camera. Our capacity estimates indicate that tens of Mbps is possible using a smartphone camera even when the short code on the screen images onto only 15% of the camera frame. Our estimates also indicate that there is room for at least 2.5x improvement in throughput of existing screen - camera communication prototypes.
Ashwin Ashok, Shubham Jain 0003, Marco Gruteser, Narayan B. Mandayam, Wenjia Yuan, Kristin J. Dana
PerCom2