John-Ross Rizzo

dblp:144/5277 · also JohnRoss Rizzo · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
8since 2021 · last 2024
0000-0002-4084-0085ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 since 2021Systems, architecture and hardware · 5 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2024 NYC-Indoor-VPR: A Long-Term Indoor Visual Place Recognition Dataset with Semi-Automatic Annotation
abstract
Visual Place Recognition (VPR) in indoor environments is beneficial to humans and robots for better localization and navigation. It is challenging due to appearance changes at various frequencies, and difficulties of obtaining ground truth metric trajectories for training and evaluation. This paper introduces the NYC-Indoor-VPR dataset, a unique and rich collection of over 36,000 images compiled from 13 distinct crowded scenes in New York City taken under varying lighting conditions with appearance changes. Each scene has multiple revisits across a year. To establish the ground truth for VPR, we propose a semiautomatic annotation approach that computes the positional information of each image. Our method specifically takes pairs of videos as input and yields matched pairs of images along with their estimated relative locations. The accuracy of this matching is refined by human annotators, who utilize our annotation software to correlate the selected keyframes. Finally, we present a benchmark evaluation of several state-of-the-art VPR algorithms using our annotated dataset, revealing its challenge and thus value for VPR research.
Diwei Sheng, Anbang Yang, John-Ross Rizzo, Chen Feng 0002
ICRA3
2024 5G Edge Vision: Wearable Assistive Technology for People with Blindness and Low Vision
abstract
In an increasingly visual world, people with blindness and low vision (pBLV) face substantial challenges in navigating their surroundings and interpreting visual information. From our previous work, VIS4ION is a smart wearable that helps pBLV in their daily challenges. It enables multiple microservices based on artificial intelligence (AI), such as visual scene processing, navigation, and vision-language inference. These microservices require powerful computational resources and, in some cases, stringent inference times, hence the need to offload computation to edge servers. This paper introduces a novel video streaming platform that improves the capabilities of VIS4ION by providing real-time support of the microservices at the network edge. When video is offloaded wirelessly to the edge, the time-varying nature of the wireless network requires adaptation strategies for a seamless video service. We demonstrate the performance of our adaptive real-time video streaming platform through experimentation with an open-source 5G deployment based on open air interface (OAI). The experiments demonstrate the ability to provide microservices robustly in time-varying network conditions.
Tommy Azzino, Marco Mezzavilla, Sundeep Rangan, Yao Wang 0001, John-Ross Rizzo
WCNC5
2023 Exploring Roundabout Navigation Training with 3D-Printed Tactile Maps
abstract
Navigating new traffic patterns can be challenging for everyone, but it poses particular difficulties for individuals who are blind or visually impaired (BVI). With the recent introduction of roundabouts in intersection design in the United States, many BVI individuals tend to be unfamiliar with them. Orientation and Mobility (O&M) specialists have been teaching their clients to cross traditional intersections for decades leveraging rectilinear geometry and the predictable rhythm of vehicular and pedestrian traffic to ensure safe street crossing. Conventional training methods for crossing intersections do not directly translate to roundabout navigation since they present unique challenges due to the continuous flow of traffic, complex layouts, and the absence of predictable auditory cues. We propose the development and integration of 3D-printed tactile maps into roundabout navigation training tools to address these challenges and provide tactile educational materials that are portable, durable, lightweight, and informative. Through iterative modifications based on feedback from an O&M specialist, we refined the tactile maps to improve usability and conceptual understanding of roundabout intersections. Our work highlights the potential of 3D-printed tactile maps to empower BVI individuals to confidently and independently navigate roundabout intersections.
Gaurav Seth, Vera Liqian Zhong, Lukas Franck, Anita Perr, John-Ross Rizzo, Amy Hurst
ASSETS5
2023 Understanding the Impact of Image Quality and Distance of Objects to Object Detection Performance
abstract
Object detection is a fundamental task for autonomous driving, which aim to identify and localize objects within an image. Deep learning has made great strides for object detection, with popular models including Faster R-CNN, YOLO, and SSD. The detection accuracy and computational cost of object detection depend on the spatial resolution of an image, which may be constrained by both the camera and storage considerations. Furthermore, original images are often compressed and uploaded to a remote server for object detection. Compression is often achieved by reducing either spatial or amplitude resolution or, at times, both, both of which have well-known effects on performance. Detection accuracy also depends on the distance of the object of interest from the camera. Our work examines the impact of spatial and amplitude resolution, as well as object distance, on object detection accuracy and computational cost. As existing models are optimized for uncompressed (or lightly compressed) images over a narrow range of spatial resolution, we develop a resolution-adaptive variant of YOLOv5 (RA-YOLO), which varies the number of scales in the feature pyramid and detection head based on the spatial resolution of the input image. To train and evaluate this new method, we created a dataset of images with diverse spatial and amplitude resolutions by combining images from the TJU and Eurocity datasets and generating different resolutions by applying spatial resizing and compression. We first show that RA-YOLO achieves a good trade-off between detection accuracy and inference time over a large range of spatial resolutions. We then evaluate the impact of spatial and amplitude resolutions on object detection accuracy using the proposed RA-YOLO model. We demonstrate that the optimal spatial resolution that leads to the highest detection accuracy depends on the ‘tolerated’ image size (constrained by the available bandwidth or storage). We further assess the impact of the distance of an object to the camera on the detection accuracy and show that higher spatial resolution enables a greater detection range. These results provide important guidelines for choosing the image spatial resolution and compression settings predicated on available bandwidth, storage, desired inference time, and/or desired detection range, in practical applications.
Haoyang Pei, Yixuan Lyu, Zhongzheng Yuan, John-Ross Rizzo, Yao Wang 0001, Yi Fang 0006
IROS5
2022 'Are They Doing Better In The Clinic Or At Home?': Understanding Clinicians' Needs When Visualizing Wearable Sensor Data Used In Remote Gait Assessments For People With Multiple Sclerosis
abstract
Walking impairment is a debilitating symptom of Multiple Sclerosis (MS), a disease affecting 2.8 million people worldwide. While clinicians’ in-person observational gait assessments are important, research suggests that data from wearable sensors can indicate early onset of gait impairment, track patients’ responses to treatment, and support remote and longitudinal assessment. We present an inquiry into supporting the transition from research to clinical practice. Co-design by HCI, biomedical, neurology and rehabilitation researchers resulted in a data-rich interface prototype for augmented gait analysis based on visualized sensor data. We used this as a prompt in interviews with ten experienced clinicians from a range of MS rehabilitation roles. We find that clinicians value quantitative sensor data within a whole patient narrative, to help track specific rehabilitation goals, but identify a tension between grasping critical information quickly and more detailed understanding. Based on the findings we make design recommendations for data-rich remote rehabilitation interfaces.
Ayanna Seals, Giuseppina Pilloni, Raul Sanchez, John-Ross Rizzo, Leigh Charvet, Oded Nov, Graham Dove
CHI5
2022 Deep Augmentation for Electrode Shift Compensation in Transient High-density sEMG: Towards Application in Neurorobotics
abstract
Going beyond the traditional sparse multi-channel peripheral human-machine interface that has been used widely in neurorobotics, high-density surface electromyography (HD-sEMG) has shown significant potential for decoding upper-limb motor control. We have recently proposed heterogeneous temporal dilation of LSTM in a deep neural network architecture for a large number of gestures (>60), securing spatial resolution and fast convergence. However, several fundamental questions remain unanswered. One problem targeted explicitly in this paper is the issue of “electrode shift,” which can happen specifically for high-density systems and during doffing and donning the sensor grid. Another real-world problem is the question of transient versus plateau classification, which connects to the temporal resolution of neural interfaces and seamless control. In this paper, for the first time, we implement gesture prediction on the transient phase of HD-sEMG data while robustifying the human-machine interface decoder to electrode shift. For this, we propose the concept of deep data augmentation for transient HD-sEMG. We show that without using the proposed augmentation, a slight shift of 10mm may drop the decoder's performance to as low as 20%. Combining the proposed data augmentation with a 3D Convolutional Neural Network (CNN), we recovered the performance to 84.6% while securing a high spatiotemporal resolution, robustifying to the electrode shift, and getting closer to large-scale adoption by the end-users, enhancing resiliency.
Tianyun Sun, Jacqueline Libby, John-Ross Rizzo, Seyed Farokh Atashzar
IROS3
2022 Digital Technologies in Orientation and Mobility Instruction for People Who are Blind or Have Low Vision
abstract
This paper investigates the tools and practices used by Orientation and Mobility (O&M) specialists in instructing people who are blind or have low vision in concepts, skills, and techniques for safe and independent travel. Based on interviews with experienced instructors who practice in different O&M settings we find that a shortage of qualified specialists and restrictions on in-person activities during COVID-19 has accelerated interest in remote instruction and assessment, while widespread adoption of smartphones with accessibility support has driven interest in assistive apps. This presents both opportunities and challenges for a practice that is traditionally conducted in-person and assessed through qualitative observations. In response we identify multiple opportunities for HCI research in service of O&M, including: supporting a 'physician's assistant' model of remote O&M instruction and assessment, matching O&M instructors' clients with guide dogs, highlighting clients' progress towards O&M goals, and collaboratively planning routes and monitoring clients' independent travel progress.
Graham Dove, Adelle Fernando, Kim Hertz, John-Ross Rizzo, William H. Seiple, Oded Nov
Proc. ACM Hum. Comput. Interact.5
2021 NYU-VPR: Long-Term Visual Place Recognition Benchmark with View Direction and Data Anonymization Influences
abstract
Visual place recognition (VPR) is critical in not only localization and mapping for autonomous driving vehicles, but also assistive navigation for the visually impaired population. To enable a long-term VPR system on a large scale, several challenges need to be addressed. First, different applications could require different image view directions, such as front views for self-driving cars while side views for the low vision people. Second, VPR in metropolitan scenes can often cause privacy concerns due to the imaging of pedestrian and vehicle identity information, calling for the need for data anonymization before VPR queries and database construction. Both factors could lead to VPR performance variations that are not well understood yet. To study their influences, we present the NYU-VPR dataset that contains more than 200,000 images over a 2km×2km area near the New York University campus, taken within the whole year of 2016. We present benchmark results on several popular VPR algorithms showing that side views are significantly more challenging for current VPR methods while the influence of data anonymization is almost negligible, together with our hypothetical explanations and in-depth analysis.
Diwei Sheng, Yuxiang Chai, Chen Feng 0002, Jianzhe Lin, Cláudio T. Silva, John-Ross Rizzo
IROS7
2020 A Low-Vision Navigation Platform for Economies in Transition Countries
abstract
An ability to move freely, when wanted, is an essential activity for healthy living. Visually impaired and completely blinded persons encounter many disadvantages in their day-to-day activities, including performing work-related tasks. They are at risk of mobility losses, illness, debility, social isolation, and premature mortality. A novel wearable device and computing platform called VIS4ION is reducing the disadvantage gaps and raising living standards for the visually challenged. It provides personal mobility navigational services that serves as a customizable, human-in-the-loop, sensing-to-feedback platform to deliver functional assistance. The platform is configured as a wearable that provides on-board microcomputers, human-machine interfaces, and sensory augmentation. Mobile edge computing enhances functionality as more services are unleashed with the computational gains. The meta-level goal is to support spatial cognition, personal freedom, and activities, and to promoting health and wellbeing. VIS4ION can be conceptualized as the dovetailing of two thrusts: an on-person navigational and computing device and a multimodal functional aid providing microservices through the cloud. The device has on-board wireless capabilities connected through Wi-Fi or 4/5G. The cloud-based microservices reduce hardware and power requirements while allowing existing and new services to be enhanced and added such as loading new map and real-time communication via haptic or audio signals. This technology can be made available and affordable in the economies of transition countries.
John-Ross Rizzo, Chen Feng 0002, Wachara Riewpaiboon, Pattanasak Mongkolwat
SERVICES1
2019 An assistive low-vision platform that augments spatial cognition through proprioceptive guidance: Point-to-Tell-and-Touch
abstract
Spatial cognition, as gained through the sense of vision, is one of the most important capabilities of human beings. However, for the visually impaired (VI), lack of this perceptual capability poses great challenges in their life. Therefore, we have designed Point-to-Tell-and-Touch, a wearable system with an ergonomic human-machine interface, for assisting the VI with active environmental exploration, with a particular focus on spatial intelligence and navigation to objects of interest in an alien environment. Our key idea is to link visual signals, as decoded synthetically, to the VI's proprioception for more intelligible guidance, in addition to vision-to-audio assistance, i.e., finger pose, as indicated by pointing, is used as “proprioceptive laser pointer” to target an object in that line of sight. The whole system consists of two features, Point-to-Tell and Point-to-Touch, both of which can work independently or cooperatively. The Point-to-Tell feature contains a camera with a novel one-stage neural network tailored for blind-centered object detection and recognition, and a headphone telling the VI the semantic label and distance from the pointed object. the Point-to-Touch, the second feature, leverages a vibrating wrist band to create a haptic feedback tool that supplements the initial vectorial guidance provided by the first stage (hand pose being direction and the distance being the extent, offered through audio cues). Both platform features utilize proprioception or joint position sense. Through hand pose, the VI end user knows where he or she is pointing relative to their egocentric coordinate system and we are able to use this foundation to build spatial intelligence. Our successful indoor experiments demonstrate the proposed system to be effective and reliable in helping the VI gain spatial cognition and explore the world in a more intuitive way.
Wenjun Gui, Shuaihang Yuan, John-Ross Rizzo, Lakshay Sharma, Chen Feng 0002, Anthony Tzes, Yi Fang 0006
IROS4
2019 Learning domain-invariant feature for robust depth-image-based 3D shape retrieval
Jing Zhu 0002, John-Ross Rizzo, Yi Fang 0006
Pattern Recognit. Lett.2
2014 Mechanical force redistribution: enabling seamless, large-format, high-accuracy surface interaction
abstract
We present Mechanical Force Redistribution (MFR): a method of sensing which creates an anti-aliased image of forces applied to a surface. This technique mechanically focuses the force from a surface onto adjacent discrete forcels (force sensing cells) by way of protrusions (small bumps or pegs), allowing for high-accuracy interpolation between adjacent discrete forcels. MFR works with any force transducing technique or material, including force variable resistive inks, piezoelectric materials and capacitive force plates. MFR sensors can be tiled such that the signal is continuous across contiguous tiles. By minimizing active materials and computational complexity, MFR makes large-format interactive walls, collaborative tabletops and high-resolution floor tiles possible and economically feasible.
Alex M. Grau, Charles Hendee, John-Ross Rizzo, Ken Perlin
CHI3