Juan P. Wachs

dblp:89/1212 · also Juan Pablo Wachs · DBLP profile ↗
← Back
82ranked-venue papers
6as first author
21since 2021 · last 2026
0000-0002-6425-5745ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 52 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 3 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 29 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 1 first-author · 6 since 2021Systems, architecture and hardware · 11 · 5 since 2021
YearPublicationVenuePosition
2026 Automated Unified Reasoning with Vision-Language Models for Multi-modal Burn Assessment
abstract
In emerging clinical applications such as ultrasound-based burn assessment, the lack of domain-specific data presents a significant challenge for developing robust AI systems. Vision-language models (VLMs) have shown strong performance in general computer vision tasks, yet their application to medical imaging remains limited, particularly due to insufficient reasoning capabilities and the scarcity of high-quality training data. We introduce AURA (Automated Unified Reasoning for Burn Assessment), a multi-modal approach that integrates pre-trained VLMs with symbolic first-order logic (FOL) reasoning to improve diagnostic accuracy and interpretability in this data-limited setting. For this study, we collected real-patient data over a one-year period at a U.S. burn center, performing all experiments in a real clinical setting to ensure practical relevance. The dataset includes both conventional B-Mode ultrasound and Tissue Doppler Imaging (TDI), with TDI introduced here for the first time in burn assessment, underscoring the emerging nature of this work. Beyond burn severity classification, we assess the system’s ability to produce expert-level surgical insight directly from imaging data. On the retrospective dataset, it achieves up to 93% accuracy in surgical classification and 87% in fine-grained burn depth prediction, comparable to expert-informed predictions and substantially exceeding the 70% accuracy of traditional visual inspection by human experts. These results, obtained from a novel multi-modal dataset collected in a real clinical burn center setting, highlight the potential of this approach to improve decision-making in burn care. To further support future deployment, we demonstrate a prototype integration with an Electronic Medical Record (EMR) system that aligns with clinical workflows and supports scalable, real-world implementation.
Md. Masudur Rahman 0001, Mohamed El Masry, Gayle Gordillo, Juan P. Wachs
AAAI4
2026 Trauma THOMPSON: A Dataset and Realistic Generative Framework for AI Copilots in Emergency Care
abstract
We introduce Trauma THOMPSON, a dataset and suite of benchmarks designed to accelerate the development of AI-powered copilots for real-time decision-making in emergency and resource-limited medical settings. This work proposes a method to address a critical bottleneck for future deployment: models trained on simulations may not perform well in the real world. The dataset features 3,717 unscripted, first-person video clips of five emergency procedures, uniquely including "just-in-time" (JIT) interventions that mirror the improvisational nature of field medicine. To obtain realistic patient data without ethical issues and identity concerns that medical data often encounter, we also propose TraumaGen, a novel framework for generating photorealistic patient and wound images from manikins while preserving clinical context. We establish benchmarks for action recognition, anticipation, and visual question answering (VQA), evaluating state-of-the-art models to demonstrate the challenges and potential of our dataset. By focusing on realism and improvisation, Trauma THOMPSON provides a crucial resource and a clear path toward developing and validating robust AI assistants for future deployment in real-world emergency care.
Yupeng Zhuo, Eddie Zhang, Xiangchen Yu, Aditya Pachpande, Andrew W. Kirkpatrick, Jessica L. McKee, Juan P. Wachs
AAAI7
2025 Context-Aware Vision Language Model for Action Recognition
Eddie Zhang, Yupeng Zhuo, Juan P. Wachs
ACIVS3
2025 A Magnetic-Actuated Vision-Based Whisker Array for Contact Perception and Grasping
abstract
Tactile sensing and the manipulation of delicate objects are critical challenges in robotics. This study presents a vision-based magnetic-actuated whisker array sensor that integrates these functions. The sensor features eight whiskers arranged circularly, supported by an elastomer membrane and actuated by electromagnets and permanent magnets. A camera tracks whisker movements, enabling high-resolution tactile feedback. The sensor's performance was evaluated through object classification and grasping experiments. In the classification experiment, the sensor approached objects from four directions and accurately identified five distinct objects with a classification accuracy of 99.17% using a Multi-Layer Perceptron model. In the grasping experiment, the sensor tested configurations of eight, four, and two whiskers, achieving the highest success rate of 87% with eight whiskers. These results highlight the sensor's potential for precise tactile sensing and reliable manipulation.
Zhixian Hu, Juan P. Wachs, Yu She
ICRA2
2023 Robotic Sonographer: Autonomous Robotic Ultrasound using Domain Expertise in Bayesian Optimization
abstract
Ultrasound is a vital imaging modality utilized for a variety of diagnostic and interventional procedures. However, an expert sonographer is required to make accurate maneuvers of the probe over the human body while making sense of the ultrasound images for diagnostic purposes. This procedure requires a substantial amount of training and up to a few years of experience. In this paper, we propose an autonomous robotic ultrasound system that uses Bayesian Optimization (BO) in combination with the domain expertise to predict and effectively scan the regions where diagnostic quality ultrasound images can be acquired. The quality map, which is a distribution of image quality in a scanning region, is estimated using Gaussian process in BO. This relies on a prior quality map modeled using expert's demonstration of the high-quality probing maneuvers. The ultrasound image quality feedback is provided to BO, which is estimated using a deep convolution neural network model. This model was previously trained on database of images labelled for diagnostic quality by expert radiologists. Experiments on three different urinary bladder phantoms validated that the proposed autonomous ultrasound system can acquire ultrasound images for diagnostic purposes with a probing position and force accuracy of 98.7% and 97.8%, respectively.
Deepak Raina, Chandrashekhara SH, Richard M. Voyles, Juan P. Wachs, Subir Kumar Saha
ICRA4
2023 COMPlacent: A Compliant Whisker Manipulator for Object Tactile Exploration
abstract
Handling fragile objects requires minimally invasive interaction skills in order to avoid any permanent deformation, alternation or damages. Such need is often required in tactile exploration tasks. In this paper, we propose an innovative whisker manipulator (COMPlacent), which is designed to accomplish tactile exploration with minimum intrusiveness. The design is inspired by biological whiskers observed in animals, where whiskers are used as means of tactile exploration in an analogous way as fingers are. Artificial whiskers are compliant but robust, which mitigates contact forces by bending or conforming to the object surface. The intrusiveness is further reduced by reactive control, which is implemented based on tactile sensors and actuators installed on each whisker. This allows the whisker to be retracted from the object surface, so that the energy transferred by contacts is minimized. The tactile sensor is designed to be ultrasensitive, which allows it to gather contact information with high fidelity. By modeling contact pressure as a time-series signal, a machine learning framework is leveraged to discriminate object properties including shape and texture. Evaluation experiments were conducted on real objects, which successfully demonstrates object classification at an accuracy of 97.3%, and texture discrimination accuracy of 92.1%.
Chenxi Xiao, Juan P. Wachs
IROS2
2023 Touchless Interfaces in the Operating Room: A Study in Gesture Preferences
abstract
Touchless interfaces allow surgeons to control medical imaging systems autonomously while maintaining total asepsis in the Operating Room. This is specially relevant as it applies to the recent outbreak of COVID-19 disease. The choice of best gestures/commands for such interfaces is a critical step that determines the overall efficiency of surgeon-computer interaction. In this regard, usability metrics such as task completion time, memorability and error rate have a long-standing as potential entities in determining the best gestures. In addition, previous works concerned with this problem utilized qualitative measures to identify the best gestures. In this work, we hypothesize that there is a correlation between gestures’ qualitative properties and their usability metrics. In this regard, we conducted a user experiment with language experts to quantify gestures’ properties (v). Next, we developed a gesture-based system that facilitates surgeons to control the medical imaging software in a touchless manner. Next, a usability study was conducted with neurosurgeons and the standard usability metrics (u) were measured in a systematic manner. Lastly, multi-variate correlation analysis was used to find the relations between u and v. Statistical analysis showed that the v scores were significantly correlated with the usability metrics with an R2≈0.40 and p < 0.05. Once the correlation is established, we can utilize either gestures’ qualitative properties or usability metrics to identify the best set of gestures.
Naveen Madapana, Daniela Chanci, Glebys T. Gonzalez, Lingsong Zhang, Juan P. Wachs
Int. J. Hum. Comput. Interact.5
2023 Tactile and Chemical Sensing With Haptic Feedback for a Telepresence Explosive Ordnance Disposal Robot
abstract
Robots can be used to mitigate risks in unsafe and austere settings. In recent years, explosive ordnance disposal robots have reduced the technician's time-on-target, and thus, reduce the direct risk of exposure. This article focuses on the study and development of innovative techniques as the foundational work for a new robot platform. The proposed system includes an organic electrochemical transistor device to detect the existence of explosive residues, and lead to decisions for safe-removal progress. Taurus' surgical gripper facilitates object tactile exploration, and manipulation with control precision to the millimeter range. The highly sensitive triboelectric tactile sensor could reduce intrusiveness during contact, and mitigate the risk of detonation. Haptic devices and visual displays are used to convey important signals, in order to improve the situational awareness of the teleoperator. A machine learning classifier can be used to assist the user to identify objects from tactile sampling. The integration of these methodologies allows for a sensitive approach to concealed objects that are only accessible through tactile sensing.
Chenxi Xiao, Aaron Benjamin Woeppel, Gina M. Clepper, Shengjie Gao, Shujia Xu, Johannes F. Rueschen, Daniel Kruse, Wenzhuo Wu, Hong Z. Tan, Thomas Low, Stephen P. Beaudoin, Bryan W. Boudouris, William G. Haris, Juan P. Wachs
IEEE Trans. Robotics14
2022 JSE: Joint Semantic Encoder for zero-shot gesture learning
Naveen Madapana, Juan P. Wachs
Pattern Anal. Appl.2
2022 Active Multiobject Exploration and Recognition via Tactile Whiskers
abstract
Robotic exploration under uncertain environments is challenging when optical information is not available. In this article, we propose an autonomous solution of exploring an unknown task space based on tactile sensing alone. We first designed a whisker sensor based on MEMS barometer devices. This sensor can acquire contact information by interacting with the environment nonintrusively. This sensor is accompanied by a planning technique to generate exploration trajectories by using mere tactile perception. This technique relies on a hybrid policy for tactile exploration, which includes a proactive informative path planner for object searching, and a reactive Hopf oscillator for contour tracing. Results indicate that the hybrid exploration policy can increase the efficiency of object discovery. Last, scene understanding was facilitated by segmenting objects and classification. A classifier was developed to recognize the object categories based on the geometric features collected by the whisker sensor. Such an approach demonstrates the whisker sensor, together with the tactile intelligence, can provide sufficiently discriminative features to distinguish objects.
Chenxi Xiao, Shujia Xu, Wenzhuo Wu, Juan P. Wachs
IEEE Trans. Robotics4
2021 ZF-SSE: A Unified Sequential Semantic Encoder for Zero-Few-Shot Learning
abstract
Humans can inherently recognize various objects and gestures either from very few examples or from their descriptions. However, supervised gesture classification methods require hundreds of examples to learn to classify, and cannot rapidly generalize from either few examples or from their high-level descriptions. This gap can be bridged using few-shot (FSL) and zero-shot (ZSL) learning methods (jointly referred to as ZFSL). Previous approaches studied the problems of ZSL and FSL in isolation, and there are few works concerned with ZFSL for temporal problems such as unfamiliar gesture recognition. In this context, this paper represents categories as semantic descriptions in a high-level attribute space and proposes a unified framework to facilitate ZFSL. This work further introduces an approach referred to as Unified Sequential Semantic Encoder (ZF -SSE) to explore temporal patterns and to predict the semantic information through simultaneously optimizing for both classification and semantic tasks. The proposed framework is validated in the domain of hand gesture recognition and the results show that the ZF -SSE approach significantly outperforms existing approaches by at least 4-10% in zero-shot, few-shot, and open-set experimental conditions.
Naveen Madapana, Juan P. Wachs
FG2
2021 Learning Multimodal Contact-Rich Skills from Demonstrations Without Reward Engineering
abstract
Everyday contact-rich tasks, such as peeling, cleaning, and writing, demand multimodal perception for effective and precise task execution. However, these present a novel challenge to robots as they lack the ability to combine these multimodal stimuli for performing contact-rich tasks. Learning-based methods have attempted to model multi-modal contact-rich tasks, but they often require extensive training examples and task-specific reward functions which limits their practicality and scope. Hence, we propose a generalizable model-free learning-from-demonstration framework for robots to learn contact-rich skills without explicit reward engineering. We present a novel multi-modal sensor data representation which improves the learning performance for contact-rich skills. We performed training and experiments using the real-life Sawyer robot for three everyday contact-rich skills – cleaning, writing, and peeling. Notably, the framework achieves a success rate of 100% for the peeling and writing skill, and 80% for the cleaning skill. Hence, this skill learning framework can be extended for learning other physical manipulation skills.
Mythra V. Balakuntala, Upinder Kaur, Xin Ma 0008, Juan P. Wachs, Richard M. Voyles
ICRA4
2021 DESERTS: DElay-tolerant SEmi-autonomous Robot Teleoperation for Surgery
abstract
Telesurgery can be hindered by high-latency and low-bandwidth communication networks, often found in austere settings. Even delays of less than one second are known to negatively impact surgeries. To tackle the effects of connectivity associated with telerobotic surgeries, we propose the DESERTS framework. DESERTS provides a novel simulator interface where the surgeon can operate directly on a virtualized reality simulation and the activities are mirrored in a remote robot, almost simultaneously. Thus, the surgeon can perform the surgery uninterrupted, while high-level commands are extracted from his motions and are sent to a remote robotic agent. The simulated setup mirrors the remote environment, including an alpha-blended view of the remote scene. The framework abstracts the actions into atomic surgical maneuvers (surgemes) which eliminate the need to transmit compressed video information. This system uses a deep learning based architecture to perform live recognition of the surgemes executed by the operator. The robot then executes the received surgemes, thereby achieving semi-autonomy. The framework’s performance was tested on a peg transfer task. We evaluated the accuracy of the recognition and execution module independently as well as during live execution. Furthermore, we assessed the framework’s performance in the presence of increasing delays. Notably, the system maintained a task success rate of 87% from no-delays to 5 seconds of delay.
Glebys T. Gonzalez, Mridul Agarwal, Mythra V. Balakuntala, Md. Masudur Rahman 0001, Upinder Kaur, Richard M. Voyles, Vaneet Aggarwal, Yexiang Xue, Juan P. Wachs
ICRA9
2021 Dexterous Skill Transfer between Surgical Procedures for Teleoperated Robotic Surgery
abstract
In austere environments, teleoperated surgical robots could save the lives of critically injured patients if they can perform complex surgical maneuvers under limited communication bandwidth. The bandwidth requirement is reduced by transferring atomic surgical actions (referred to as “surgemes”) instead of the low-level kinematic information. While such a policy reduces the bandwidth requirement, it requires accurate recognition of the surgemes. In this paper, we demonstrate that transfer learning across surgical tasks can boost the performance of surgeme recognition. This is demonstrated by using a network pre-trained with peg-transfer data from Yumi robot to learn classification on debridement on data from Taurus robot. Using a pre-trained network improves the classification accuracy achieves a classification accuracy of 76% with only 8 sequences in target domain, which is 22.5% better than no-transfer scenario. Additionally, ablations on transfer learning indicate that transfer learning requires 40% less data compared to no-transfer to achieve same classification accuracy. Further, the convergence rate of the transfer learning setup is significantly higher than the no-transfer setup trained only on the target domain.
Mridul Agarwal, Glebys T. Gonzalez, Mythra V. Balakuntala, Md. Masudur Rahman 0001, Vaneet Aggarwal, Richard M. Voyles, Yexiang Xue, Juan P. Wachs
RO-MAN8
2021 ICONS: Imitation CONStraints for Robot Collaboration
abstract
Skill imitation has been an important ability in human-robot collaboration since it allows expeditious robot teaching of new tasks never seen before. To mimic the human pose, inverse kinematic solvers have been used to endow kinematic structures with human-like motion. Nevertheless, these solutions tend to be formulated for a specific robot or task. To address this generalization issue, this work presents the ICONS framework for imitation constraints, which proposes a general formulation for pose imitation, paired with a computationally efficient solver, inspired by the FABRIK algorithm. Three versions of the solver were developed to optimize the presented constraints. To assess the performance of ICONS, two tasks were evaluated, an incision task, and an assembly task. Fifty demonstrations were collected for each task. We compared the performance of our method, using pose accuracy and occlusion, against the numerical solver baseline (FABRIK). Notably, the ICONS framework improved the pose accuracy by 58% and reduced the environment occlusion by 38%. Moreover, the computational efficiency of the ICONS framework was assessed. Results show that the proposed algorithm maintains the efficiency of the baseline, finding the target solution under 10 iterations.
Glebys T. Gonzalez, Juan P. Wachs
RO-MAN2
2021 SACHETS: Semi-Autonomous Cognitive Hybrid Emergency Teleoperated Suction
abstract
Blood suction and irrigation are among the most critical support tasks in robotic-assisted minimally invasive surgery (RMIS). Usually, suction/irrigation tools are controlled by a surgical assistant to maintain a clear view of the surgical field. Thus, the assistant’s contribution to other emergency support tasks is limited. Similarly, when the surgical assistant is not available to perform the blood suction, the leading surgeon must take over this task, which in a complex surgical procedure can result in an unnecessary increment in the cognitive load. To alleviate this problem, we have developed a semi-autonomous robotic suction assistant, which was integrated with a Da Vinci Research Kit (DVRK). At the heart of the algorithm, there is an autonomous control based on a deep learning model to segment and identify the location of blood accumulations. This system provides automatic suction allowing the leading surgeon to focus exclusively on the main task through the control of key instruments of the robot. We conducted a user study to evaluate the user’s workload demands and performance while doing a surgical task under two modalities: (1) autonomous suction action and (2) a surgeon-controlled-suction. Our results indicate that users working with the autonomous system completed the task 161 seconds faster than in the surgeon-controlled-suction modality. Furthermore, the autonomous modality led to a lower percentage of bleeding in the surgical field and workload demands on the users (p-value<0.05). These results show how leveraging state-of-the-art AI algorithms can reduce cognitive demands and enhance performance.
Juan Barragan Noguera, Daniela Chanci, Denny Yu, Juan P. Wachs
RO-MAN4
2021 Sequential Prediction with Logic Constraints for Surgical Robotic Activity Recognition
abstract
Many real-world time-sensitive and high-stake applications (e.g., surgical, rescue, and recovery robotics) exhibit sequential nature; thus, applying Recurrent Neural Network (RNN)-based sequential models is an attractive approach to detect robotic activity. One limitation of such approaches is data scarcity. As a result, limited training samples may lead to over-fitting, producing incorrect predictions during deployment. Nevertheless, abundant domain knowledge may still be available, which may help formulate logic constraints. In this paper, we propose a novel way to integrate domain knowledge into RNN-based sequential prediction. We build a Markov Logic Network (MLN)-based classifier that automatically learns constraint weights from data. We propose two methods to incorporate this MLN-based prediction: (i) PriorLayer, in which the values of the hidden layer of the RNN are combined with weights learned from logic constraints in an additional neural network layer, and (ii) Conflation, in which class probabilities from RNN predictions and constraint weights are combined based on the conflation of class probabilities. We evaluate robotic activity classification methods on a simulated OpenAI Gym environment and a real-world DESK dataset for surgical robotics. We observe that our proposed MLN-based approaches boost the performance of LSTM-based networks. In particular, MLN boosts the accuracy of LSTM from 71% to 84% on the Gym dataset and from 68% to 72% on the Taurus robot dataset. Furthermore, MLN (i.e., PriorLayer) shows regularization capability where it improves accuracy in initial LSTM training while avoiding over-fitting early, thus improves the final classification accuracy on unseen data. The code is available at https://github.com/masud99r/prediction-with-logic-constraints.
Md. Masudur Rahman 0001, Richard M. Voyles, Juan P. Wachs, Yexiang Xue
RO-MAN3
2021 One-Shot Image Recognition Using Prototypical Encoders with Reduced Hubness
abstract
Humans have the innate ability to recognize new objects just by looking at sketches of them (also referred as to proto-type images). Similarly, prototypical images can be used as an effective visual representations of unseen classes to tackle few-shot learning (FSL) tasks. Our main goal is to recognize unseen hand signs (gestures) traffic-signs, and corporate-logos, by having their iconographic images or prototypes. Previous works proposed to utilize variational prototypical-encoders (VPE) to address FSL problems. While VPE learns an image-to-image translation task efficiently, we discovered that its performance is significantly hampered by the so-called hubness problem and it fails to regulate the representations in the latent space. Hence, we propose a new model (VPE++) that inherently reduces hubness and incorporates contrastive and multi-task losses to increase the discriminative ability of FSL models. Results show that the VPE++ approach can generalize better to the unseen classes and can achieve superior accuracies on logos, traffic signs, and hand gestures datasets as compared to the state-of-the-art.
Chenxi Xiao, Naveen Madapana, Juan P. Wachs
WACV3
2021 Triangle-Net: Towards Robustness in Point Cloud Learning
abstract
Three dimensional (3D) object recognition is becoming a key desired capability for many computer vision systems such as autonomous vehicles, service robots and surveillance drones to operate more effectively in unstructured environments. These real-time systems require effective classification methods that are robust to various sampling resolutions, noisy measurements, and unconstrained pose configurations. Previous research has shown that points' sparsity, rotation and positional inherent variance can lead to a significant drop in the performance of point cloud based classification techniques. However, neither of them is sufficiently robust to multifactorial variance and significant sparsity. In this regard, we propose a novel approach for 3D classification that can simultaneously achieve invariance towards rotation, positional shift, scaling, and is robust to point sparsity. To this end, we introduce a new feature that utilizes graph structure of point clouds, which can be learned end-to-end with our proposed neural network to acquire a robust latent representation of the 3D object. We show that such latent representations can significantly improve the performance of object classification and retrieval tasks when points are sparse. Further, we show that our approach outperforms PointNet and 3DmFV by 35.0% and 28.1% respectively in ModelNet 40 classification tasks using sparse point clouds of only 16 points under arbitrary SO(3) rotation.
Chenxi Xiao, Juan P. Wachs
WACV2
2021 Assessing task understanding in remote ultrasound diagnosis via gesture analysis
Edgar Rojas-Muñoz, Juan P. Wachs
Pattern Anal. Appl.2
2021 Assessing Collaborative Physical Tasks Via Gestural Analysis
abstract
Recent studies have shown that gestures are useful indicators of understanding, learning, and memory retention. However, and specially in collaborative settings, current metrics that estimate task understanding often neglect the information expressed through gestures. This work introduces the physical instruction assimilation (PIA) metric, a novel approach to estimate task understanding by analyzing the way in which collaborators use gestures to convey, assimilate, and execute physical instructions. PIA estimates task understanding by inspecting the number of necessary gestures required to complete a shared task. PIA is calculated based on the multiagent gestural instruction comparer (MAGIC) architecture, a previously proposed framework to represent, assess, and compare gestures. To evaluate our metric, we collected gestures from collaborators remotely completing the following three tasks: block assembly, origami, and ultrasound training. The PIA scores of these individuals are compared against two other metrics used to estimate task understanding: number of errors and amount of idle time during the task. Statistically significant correlations between PIA and these metrics are found. Additionally, a Taguchi design is used to evaluate PIA's sensitivity to changes in the MAGIC architecture. The factors evaluated the effect of changes in time, order, and motion trajectories of the collaborators' gestures. PIA is shown to be robust to these changes, having an average mean change of 0.45. These results hint that gestures, in the form of the assimilation of physical instructions, can reveal insights of task understanding and complement other commonly used metrics.
Edgar Rojas-Muñoz, Juan P. Wachs
IEEE Trans. Hum. Mach. Syst.2
2020 Gesture Agreement Assessment Using Description Vectors
abstract
Participatory design is a popular design technique that involves the end-users in the early stages of the design process to obtain user-friendly gestural interfaces. Guessability studies followed by agreement analyses are often used to elicit and comprehend the preferences (gestures/proposals) of the participants. Previous approaches to assess agreement, grouped the gestures into equivalence classes and ignored the integral properties that are shared between them. In this work, we represent the gestures using binary description vectors to allow them to be partially similar. In this context, we introduce a new metric referred to as a soft agreement rate (SAR) to quantify the level of consensus between the participants. In addition, we performed computational experiments to study the behavior of our partial agreement formula and mathematically show that existing agreement metrics are a special case of our approach. Our methodology was evaluated through a gesture elicitation study conducted with a group of neurosurgeons. Nevertheless, our formulation can be applied to any other user-elicitation study. Results show that the level of agreement obtained by SAR metric is 2.64 times higher than the existing metrics. In addition to the most agreed gesture, SAR formulation also provides the mostly agreed descriptors which can potentially help the designers to come up with a final gesture set.
Naveen Madapana, Glebys T. Gonzalez, Juan P. Wachs
FG3
2020 Feature Selection for Zero-Shot Gesture Recognition
abstract
Existing classification techniques assign a predetermined categorical label to each sample and cannot recognize the new categories that might appear after the training stage. This limitation has led to the advent of new paradigms in machine learning such as zero-shot learning (ZSL). ZSL aims to recognize unseen categories by having a high-level description of them. While deep learning has pushed the limits of ZSL for object recognition, ZSL for temporal problems such as unfamiliar gesture recognition (ZSGL) remain unexplored. Previous attempts to address ZSGL were focused on the creation of gesture attributes, attribute-based datasets, and algorithmic improvements, and there is little or no research concerned with feature selection for ZSGL problems. It is indisputable that deep learning has obviated the need for feature engineering for the problems with large datasets. However, when the data is scarce, it is critical to leverage the domain information to create discriminative input features. The main goal of this work is to study the effect of three different feature extraction techniques (raw features, engineered features, and deep learning features) on the performance of ZSGL. Next, we propose a new approach for ZSGL that jointly minimizes the reconstruction loss, semantic and classification losses. Our methodology yields an unseen class accuracy of (38%) which parallels the accuracies obtained through state-of-the-art approaches.
Naveen Madapana, Juan P. Wachs
FG2
2020 Beyond MAGIC: Matching Collaborative Gestures using an optimization-based Approach
abstract
Gestures are a key aspect of communication during collaboration: through gestures we can express ideas, inquires and formalize instructions as we collaborate. Nevertheless, gesture analysis is not currently used to assess quality of task collaboration. One possible reason for this is that there is no consensus on how to represent and compare gestures from the semantic standpoint. To address this, this paper introduces three novel approaches to compare gestures performed by individuals as they collaborate to complete a physical task. Our approach relies on solving three variations of an integer optimization assignment problem, i.e. based on gesture similarity, based on temporal synchrony, and based on a combination of both. We collected the gestures of 40 participants (divided into 20 pairs) as they performed two collaborative tasks, and generated a human baseline that compared and matched their gestures. Afterwards, our gesture comparison approach was evaluated against other gestures comparison approaches based on how well they replicated the human baseline. Our approach outperformed the other approaches, agreeing with the human baseline over 85% of the times. Thus, the obtained results support the proposed technique for gesture comparison. This in turn can lead to the development of better methods to evaluate collaborative physical tasks.
Edgar Rojas-Muñoz, Juan P. Wachs
FG2
2020 The MAGIC of E-Health: A Gesture-Based Approach to Estimate Understanding and Performance in Remote Ultrasound Tasks
abstract
This work presents an approach to estimate task understanding and performance during a remote ultrasound training task via gestures. These task understanding insights are obtained through the PIA metric, a score that represents how well are gestures being used to complete a shared task. To evaluate our hypothesis, 20 participants performed a remote ultrasound training task consisting of three subtasks: vessel detection, blood extraction, and foreign body detection. Afterwards, their task understanding and performance was estimated using our PIA metric and three other metrics: error rate, idle time rate, and task completion percentage. After performing a correlation analysis, we found significant correlations between the PIA metric and all the other metrics for task understanding estimation. In addition, the insights generated from our PIA score explained inconsistencies in the participants' scores that were not expressed using the other metrics. Finally, we used two post-experiment questionnaires to subjectively evaluate the participants' perceived understanding and performance, and found that the PIA score was significantly correlated with the participants' overall task understanding. All these results indicate that a gesture-based metric can be used to estimate task understanding, which can have a positive impact in the way remote ultrasound tasks are performed and assessed.
Edgar Rojas-Muñoz, Juan P. Wachs
FG2
2020 Message from the General and Program Chairs FG 2020
abstract
Welcome to the 15th IEEE International Conference on Automatic Face and Gesture Recognition (FG 2020). FG is the premier international conference on vision-based automatic face and body behavior analysis and applications. Since its first meeting in Zurich in 1994, the conference has been held fourteen times throughout the world. This is the 15th conference.
Juan P. Wachs, Sergio Escalera, Jeffrey F. Cohn, Albert Ali Salah, Arun Ross
FG1
2020 The AI-Medic: A Multimodal Artificial Intelligent Mentor for Trauma Surgery
abstract
Telementoring generalist surgeons as they treat patients can be essential when in situ expertise is not readily available. However, adverse cyber-attacks, unreliable network conditions, and remote mentors' predisposition can significantly jeopardize the remote intervention. To provide medical practitioners with guidance when mentors are unavailable, we present the AI-Medic, the initial steps towards the development of a multimodal intelligent artificial system for autonomous medical mentoring. The system uses a tablet device to acquire the view of an operating field. This imagery is provided to an encoder-decoder neural network trained to predict medical instructions from the current view of a surgery. The network was training using DAISI, a dataset including images and instructions providing step-by-step demonstrations of surgical procedures. The predicted medical instructions are conveyed to the user via visual and auditory modalities.
Edgar Rojas-Muñoz, Kyle Couperus, Juan P. Wachs
ICMI3
2020 How About the Mentor? Effective Workspace Visualization in AR Telementoring
abstract
Augmented Reality (AR) benefits telementoring by enhancing the communication between the mentee and the remote mentor with mentor authored graphical annotations that are directly integrated into the mentee’s view of the workspace. An important problem is conveying the workspace to the mentor effectively, such that they can provide adequate guidance. AR headsets now incorporate a frontfacing video camera, which can be used to acquire the workspace. However, simply providing to the mentor this video acquired from the mentee’s first-person view is inadequate. As the mentee moves their head, the mentor’s visualization of the workspace changes frequently, unexpectedly, and substantially. This paper presents a method for robust high-level stabilization of a mentee first-person video to provide effective workspace visualization to a remote mentor. The visualization is stable, complete, up to date, continuous, distortion free, and rendered from the mentee’s typical viewpoint, as needed to best inform the mentor of the current state of the workspace. In one study, the stabilized visualization had significant advantages over unstabilized visualization, in the context of three number matching tasks. In a second study, stabilization showed good results, in the context of surgical telementoring, specifically for cricothyroidotomy training in austere settings.
Chengyuan Lin 0001, Edgar Rojas-Muñoz, Maria E. Cabrera, Natalia Sanchez-Tamayo, Daniel Andersen, Voicu Popescu, Juan Barragan Noguera, Ben Zarzaur, Kathryn Anderson, Thomas Douglas, Clare Griffis, Juan P. Wachs
VR13
2020 Agreement Study Using Gesture Description Analysis
abstract
Choosing adequate gestures for touchless interfaces is a challenging task that has a direct impact on human-computer interaction. Such gestures are commonly determined by the designer, ad-hoc, rule-based, or agreement-based methods. Previous approaches to assess agreement grouped the gestures into equivalence classes and ignored the integral properties that are shared between them. In this article, we propose a generalized framework that inherently incorporates the gesture descriptors into the agreement analysis. In contrast to previous approaches, we represent gestures using binary description vectors and allow them to be partially similar. In this context, we introduce a new metric referred to as soft agreement rate (SAR) to measure the level of agreement and provide a mathematical justification for this metric. Furthermore, we perform computational experiments to study the behavior of SAR and demonstrate that existing agreement metrics are a special case of our approach. Our method is evaluated and tested through a guessability study conducted with a group of neurosurgeons. Nevertheless, our formulation can be applied to any other user-elicitation study. Results show that the level of agreement obtained by SAR is 2.64 times higher than the previous metrics. Finally, we show that our approach complements the existing agreement techniques by generating an artificial lexicon based on the most agreed properties.
Naveen Madapana, Glebys T. Gonzalez, Lingsong Zhang, Richard Rodgers, Juan P. Wachs
IEEE Trans. Hum. Mach. Syst.5
2020 Multimodal Physiological Signals for Workload Prediction in Robot-assisted Surgery
abstract
Monitoring surgeon workload during robot-assisted surgery can guide allocation of task demands, adapt system interfaces, and assess the robotic system's usability. Current practices for measuring cognitive load primarily rely on questionnaires that are subjective and disrupt surgical workflow. To address this limitation, a computational framework is demonstrated to predict user workload during telerobotic surgery. This framework leverages wireless sensors to monitor surgeons’ cognitive load and predict their cognitive states. Continuous data across multiple physiological modalities (e.g., heart rate variability, electrodermal, and electroencephalogram activity) were simultaneously recorded for twelve surgeons performing surgical skills tasks on the validated da Vinci Skills Simulator. These surgical tasks varied in difficulty levels, e.g., requiring varying visual processing demand and degree of fine motor control. Collected multimodal physiological signals were fused using independent component analysis, and the predicted results were compared to the ground-truth workload level. Results compared performance of different classifiers, sensor fusion schemes, and physiological modality (i.e., prediction with single vs. multiple modalities). It was found that our multisensor approach outperformed individual signals and can correctly predict cognitive workload levels 83.2% of the time during basic and complex surgical skills tasks.
Tian Zhou 0005, Jackie S. Cha, Glebys T. Gonzalez, Juan P. Wachs, Chandru Sundaram, Denny Yu
ACM Trans. Hum. Robot Interact.4
2019 Database of Gesture Attributes: Zero Shot Learning for Gesture Recognition
abstract
Existing gesture classification techniques assign a categorical label to each gesture instance and learn to recognize only a predetermined set of gesture classes. These techniques lack adaptability to new or unseen gestures which is the premise of zero shot learning (ZSL). Hence we propose to identify the properties of gestures and thereby infer the categorical label instead of recognizing the class label directly. ZSL for gesture recognition has hardly been studied in the pattern recognition research. The reason is partly due to the lack of benchmarks and specialized datasets consisting of annotations for gesture attributes. In this regard, this paper presents the first annotated database of attributes for the gestures present in ChaLearn 2013 and MSRC - 12 datasets. This was achieved as follows; First, we identified a finite set of 64 discriminative and representative high level attributes of gestures from the literature. Further, we performed crowdsourced human studies using Amazon Mechanical Turk to obtain attribute annotations for 28 gesture classes. Next, we used our dataset to train existing ZSL classifiers to predict attribute labels. Finally, we provide benchmarks for unseen gesture class prediction on CGD2013 and MSRC-12. We have made this dataset publicly available to encourage researchers to further investigate this problem.
Naveen Madapana, Juan P. Wachs
FG2
2019 MAGIC: A Fundamental Framework for Gesture Representation, Comparison and Assessment
abstract
Gestures play a fundamental role in instructional processes between agents. However, effectively transferring this non-verbal information becomes complex when the agents are not physically co-located. Recently, remote collaboration systems that transfer gestural information have been developed. Nonetheless, these systems relegate gestures to an illustrative role: only a representation of the gesture is transmitted. We argue that further comparisons between the gestures can provide information of how well the tasks are being understood and performed. While gesture comparison frameworks exist, they only rely on gesture's appearance, leaving semantics and pragmatical aspects aside. This work introduces the Multi-Agent Gestural Instructions Comparer (MAGIC), an architecture that represents and compares gestures at the morphological, semantical and pragmatical levels. MAGIC abstracts gestures via a three-stage pipeline based on a taxonomy classification, a dynamic semantics framework and a constituency parsing; and utilizes a comparison scheme based on subtrees intersections to describe gesture similarity. This work shows the feasibility of the framework by assessing MAGIC's gesture matching accuracy against other gesture comparison frameworks during a mentor-mentee remote collaborative physical task scenario.
Edgar Rojas-Muñoz, Juan P. Wachs
FG2
2019 DESK: A Robotic Activity Dataset for Dexterous Surgical Skills Transfer to Medical Robots
abstract
Datasets are an essential component for training effective machine learning models. In particular, surgical robotic datasets have been key to many advances in semi-autonomous surgeries, skill assessment, and training. Simulated surgical environments can enhance the data collection process by making it faster, simpler and cheaper than real systems. In addition, combining data from multiple robotic domains can provide rich and diverse training data for transfer learning algorithms. In this paper, we present the DESK (DExterous Surgical SKills) dataset. It comprises a set of surgical robotic skills collected during a surgical training task using three robotic platforms: the Taurus II robot, Taurus II simulated robot, and the YuMi robot. This dataset was used to test the idea of transferring knowledge across different domains (e.g. from Taurus to YuMi robot) for a surgical gesture classification task with seven gestures/surgemes. We explored two different scenarios: 1) No transfer and 2) Domain transfer (simulated Taurus to real Taurus and YuMi robots). We conducted extensive experiments with three supervised learning models and provided baselines in each of these scenarios. Results show that using simulation data during training enhances the performance on the real robots, where limited real data is available. In particular, we obtained an accuracy of 55% on the real Taurus data using a model that is trained only on the simulator data, but that accuracy improved to 82% when the ratio of real to simulated data was increased to 0.18 in the training set.
Naveen Madapana, Thomas Low, Richard M. Voyles, Yexiang Xue, Juan P. Wachs, Md. Masudur Rahman 0001, Natalia Sanchez-Tamayo, Mythra V. Balakuntala, Glebys T. Gonzalez, Jyothsna Padmakumar Bindu, L. N. Vishnunandan Venkatesh, Xingguang Zhang, Juan Barragan Noguera
IROS5
2019 JISAP: Joint Inference for Surgeon Attributes Prediction during Robot-Assisted Surgery
abstract
In Robot-Assisted Surgery, predicting surgeon attributes such as task workload, operation performance, and expertise levels is important in providing tailored assistance. This paper proposes Joint Inference for Surgeon Attributes Prediction (JISAP), a computational framework to jointly infer surgeon attributes (i.e., task workload, operation performance, and expertise level) from multimodal physiological signals (heart rate variability, wrist motion, electrodermal, electromyography, and electroencephalogram activity). JISAP was evaluated with a dataset of twelve surgeons operating on the da Vinci Skills Simulator. It was found that JISAP can simultaneously predict surgeon attributes with a percentage error of 11.05%. Additionally, joint inference was found to outperform isolated inference with a boost of 10%.
Tian Zhou 0005, Jackie S. Cha, Glebys T. Gonzalez, Chandru Sundaram, Juan P. Wachs, Denny Yu
IROS5
2019 Extending Policy from One-Shot Learning through Coaching
abstract
Humans generally teach their fellow collaborators to perform tasks through a small number of demonstrations, often followed by episodes of coaching that tune and refine the execution during practice. Adopting a similar framework for teaching robots through demonstrations makes teaching tasks highly intuitive and imitating the refinement of complex tasks through coaching improves the efficacy. Unlike traditional Learning from Demonstration (LfD) approaches which rely on multiple demonstrations to train a task, we present a novel one-shot learning from demonstration approach, augmented by coaching, to transfer the task from task expert to robot. The demonstration is automatically segmented into a sequence of a priori skills (the task policy) parametrized to match task goals. During practice, the robotic skills self-evaluate their performances and refine the task policy to locally optimize cumulative performance. Then, human coaching further refines the task policy to explore and globally optimize the net performance. Both the self-evaluation and coaching are implemented using reinforcement learning (RL) methods. The proposed approach is evaluated using the task of scooping and unscooping granular media. The self-evaluator of the scooping skill uses the realtime force signature and resistive force theory to minimize scooping resistance similar to how humans scoop. Coaching feedback focuses modifications to sub-domains of the action space, using RL to converge to desired performance. Thus, the proposed method provides a framework for learning tasks from one demonstration and generalizing it using human feedback through coaching achieving a success rate of ≈90%.
Mythra V. Balakuntala, L. N. Vishnunandan Venkatesh, Jyothsna Padmakumar Bindu, Richard M. Voyles, Juan P. Wachs
RO-MAN5
2019 Transferring Dexterous Surgical Skill Knowledge between Robots for Semi-autonomous Teleoperation
abstract
In the future, deployable, teleoperated surgical robots can save the lives of critically injured patients in battlefield environments. These robotic systems will need to have autonomous capabilities to take over during communication delays and unexpected environmental conditions during critical phases of the procedure. Understanding and predicting the next surgical actions (referred as “surgemes”) is essential for autonomous surgery. Most approaches for surgeme recognition cannot cope with the high variability associated with austere environments and thereby cannot “transfer” well to field robotics. We propose a methodology that uses compact image representations with kinematic features for surgeme recognition in the DESK dataset. This dataset offers samples for surgical procedures over different robotic platforms with a high variability in the setup. We performed surgeme classification in two setups: 1) No transfer, 2) Transfer from a simulated scenario to two real deployable robots. Then, the results were compared with recognition accuracies using only kinematic data with the same experimental setup. The results show that our approach improves the recognition performance over kinematic data across different domains. The proposed approach produced a transfer accuracy gain up to 20% between the simulated and the real robot, and up to 31% between the simulated robot and a different robot. A transfer accuracy gain was observed for all cases, even those already above 90%.
Md. Masudur Rahman 0001, Natalia Sanchez-Tamayo, Glebys T. Gonzalez, Mridul Agarwal, Vaneet Aggarwal, Richard M. Voyles, Yexiang Xue, Juan P. Wachs
RO-MAN8
2019 Robust High-Level Video Stabilization for Effective AR Telementoring
abstract
This poster presents the design, implementation, and evaluation of a method for robust high-level stabilization of mentees first-person video in augmented reality (AR) telementoring. This video is captured by the front-facing built-in camera of an AR headset and stabilized by rendering from a stationary view a planar proxy of the workspace projectively texture mapped with the video feed. The result is stable, complete, up to date, continuous, distortion free, and rendered from the mentee's default viewpoint. The stabilization method was evaluated in two user studies, in the context of number matching and for cricothyroidotomy training, respectively. Both showed a significant advantage of our method compared with unstabilized visualization.
Chengyuan Lin 0001, Edgar Rojas-Muñoz, Maria E. Cabrera, Natalia Sanchez-Tamayo, Daniel Andersen, Voicu Popescu, Juan Barragan Noguera, Ben Zarzaur, Kathryn Anderson, Thomas Douglas, Clare Griffis, Juan P. Wachs
VR13
2018 Biomechanical-Based Approach to Data Augmentation for One-Shot Gesture Recognition
abstract
Most common approaches to one-shot gesture recognition have leveraged mainly conventional machine learning solutions and image based data augmentation techniques, ignoring the mechanisms that are used by humans to perceive and execute gestures, a key contextual component in this process. The novelty of this work consists on modeling the process that leads to the creation of gestures, rather than observing the gesture alone. In this approach, the context considered involves the way in which humans produce the gestures - the kinematic and biomechanical characteristics associated with gesture production and execution. By understanding the main "modes" of variation we can replicate the single observation many times. Consequently, the main strategy proposed in this paper includes generating a data set of human-like examples based on "naturalistic" features extracted from a single gesture sample while preserving fundamentally human characteristics like visual saliency, smooth transitions and economy of motion. The availability of a large data set of realistic samples allows the use state-of-the-art classifiers for further recognition. Several classifiers were trained and their recognition accuracies were assessed and compared to previous one-shot learning approaches. An average recognition accuracy of 95% among all classifiers highlights the relevance of keeping the human "in the loop" to effectively achieve one-shot gesture recognition.
Maria E. Cabrera, Juan P. Wachs
FG2
2018 Hard Zero Shot Learning for Gesture Recognition
abstract
Gesture based systems allow humans to interact with devices and robots in a natural way. Yet, current gesture recognition systems can not recognize the gestures outside a limited lexicon. This opposes the idea of lifelong learning which require systems to adapt to unseen object classes. These issues can be best addressed using Zero Shot Learning (ZSL), a paradigm in machine learning that leverages the semantic information to recognize new classes. ZSL systems developed in the past used hundreds of training examples to detect new classes and assumed that test examples come from unseen classes. This work introduces two complex and more realistic learning problems referred as Hard Zero Shot Learning (HZSL) and Generalized HZSL (G-HZSL) necessary to achieve Life Long Learning. The main objective of these problems is to recognize unseen classes with limited training information and relax the assumption that test instances come from unseen classes. We propose to leverage one shot learning (OSL) techniques coupled with ZSL approaches to address and solve the problem of HZSL for gesture recognition. Further, supervised clustering techniques are used to discriminate seen classes from unseen classes. We assessed and compared the performance of various existing algorithms on HZSL for gestures using two standard datasets: MSRC-12 and CGD2011. For four unseen classes, results show that the marginal accuracy of HZSL-15.2% and G-HZSL-14.39% are comparable to the performance of conventional ZSL. Given that we used only one instance and do not assume that test classes are unseen, the performance of HZSL and G-HZSL models were remarkable.
Naveen Madapana, Juan P. Wachs
ICPR2
2018 Image Exploration Procedure Classification with Spike-timing Neural Network for the Blind
abstract
Individuals who are blind use exploration procedures (EPs) to navigate and understand digital images. The ability to model and detect these EPs can help the assistive technologies' community build efficient and accessible interfaces for the blind and overall enhance human-machine interaction. In this paper, we propose a framework to classify various EPs using spike-timing neural networks (SNNs). While users interact with a digital image using a haptic device, rotation and translation-invariant features are computed directly from exploration trajectories acquired from the haptic control. These features are further encoded as model strings through trained SNNs. A classification scheme is then proposed to distinguish these model strings to identify the EPs. The framework adapted a modified Dynamic Time Wrapping (DTW) for spatial-temporal matching with Dempster-Shafer Theory (DST) for multimodal fusion. Experimental results (87.05% as EPs' detection accuracy) indicate the effectiveness of the proposed framework and its potential application in human-machine interfaces.
Ting Zhang 0013, Tian Zhou 0005, Bradley S. Duerstock, Juan P. Wachs
ICPR4
2018 Early Turn-Taking Prediction with Spiking Neural Networks for Human Robot Collaboration
abstract
Turn-taking is essential to the structure of human teamwork. Humans are typically aware of team members' intention to keep or relinquish their turn before a turn switch, where the responsibility of working on a shared task is shifted. Future co-robots are also expected to provide such competence. To that end, this paper proposes the Cognitive Turn-taking Model (CTTM), which leverages cognitive models (i.e., Spiking Neural Network) to achieve early turn-taking prediction. The CTTM framework can process multimodal human communication cues (both implicit and explicit) and predict human turn-taking intentions in an early stage. The proposed framework is tested on a simulated surgical procedure, where a robotic scrub nurse predicts the surgeon's turn-taking intention. It was found that the proposed CTTM framework outperforms the state-of-the-art turn-taking prediction algorithms by a large margin. It also outperforms humans when presented with partial observations of communication cues (i.e., less than 40 % of full actions). This early prediction capability enables robots to initiate turn-taking actions at an early stage, which facilitates collaboration and increases overall efficiency.
Tian Zhou 0005, Juan P. Wachs
ICRA2
2018 Variability Analysis on Gestures for People With Quadriplegia
abstract
Gesture-based interfaces have become an effective control modality within the human computer interaction realm to assist individuals with mobility impairments in accessing technologies for daily living to entertainment. Recent studies have shown that gesture-based interfaces in tandem with gaming consoles are being used to complement physical therapies at rehabilitation hospitals and in their homes. Because the motor movements of individuals with physical impairments are different from persons without disabilities, the gesture sets required to operate those interfaces must be customized. This limits significantly the number and quality of available software environments for users with motor impairments. Previous work presented an analytic approach to convert an existing gesture-based interface designed for individuals without disabilities to be usable by people with motor disabilities. The objective of this paper is to include gesture variability analysis into the existing framework using robotics as an additional validation framework. Based on this, a physical metric (referred as work) was empirically obtained to compare the physical effort of each gesture. An integration method was presented to determine the accessible gesture set based on stability and empirical robot execution. For all the gesture types, the accessible gestures were found to lie within 34% of the optimality of stability and work. Lastly, the gesture set determined by the proposed methodology was practically evaluated by target users in experiments while solving a spatial navigational problem.
Hairong Jiang, Bradley S. Duerstock, Juan P. Wachs
IEEE Trans. Cybern.3
2017 What Makes a Gesture a Gesture? Neural Signatures Involved in Gesture Recognition
abstract
Previous work in the area of gesture production, has made the assumption that machines can replicate humanlike gestures by connecting a bounded set of salient points in the motion trajectory. Those inflection points were hypothesized to also display cognitive saliency. The purpose of this paper is to validate that claim using electroencephalography (EEG). That is, this paper attempts to find neural signatures of gestures (also referred as placeholders) in human cognition, which facilitate the understanding, learning and repetition of gestures. Further, it is discussed whether there is a direct mapping between the placeholders and kinematic salient points in the gesture trajectories. These are expressed as relationships between inflection points in the gestures trajectories with oscillatory mu rhythms (8-12 Hz) in the EEG. This is achieved by correlating fluctuations in mu power during gesture observation with salient motion points found for each gesture. Peaks in the EEG signal at central electrodes (motor cortex; C3/Cz/C4) and occipital electrodes (visual cortex; O3/Oz/O4) were used to isolate the salient events within each gesture. We found that a linear model predicting mu peaks from motion inflections fits the data well. Increases in EEG power were detected 380 and 500ms after inflection points at occipital and central electrodes, respectively. These results suggest that coordinated activity in visual and motor cortices is sensitive to motion trajectories during gesture observation, and it is consistent with the proposal that inflection points operate as placeholders in gesture recognition.
Maria E. Cabrera, Keisha Novak, Daniel Foti, Richard M. Voyles, Juan P. Wachs
FG5
2017 One-Shot Gesture Recognition: One Step Towards Adaptive Learning
abstract
User's intentions may be expressed through spontaneous gesturing, which have been seen only a few times or never before. Recognizing such gestures involves one shot gesture learning. While most research has focused on the recognition of the gestures themselves, recently new approaches were proposed to deal with gesture perception and production as part of the recognition problem. The framework presented in this work focuses on learning the process that leads to gesture generation, rather than treating the gestures as the outcomes of a stochastic process only. This is achieved by leveraging kinematic and cognitive aspects of human interaction. These factors enable the artificial production of realistic gesture samples originated from a single observation, which in turn are used as training sets for state-of-the-art classifiers. Classification performance is evaluated in terms of recognition accuracy and coherency; the latter being a novel metric that determines the level of agreement between humans and machines. Specifically, the referred machines are robots which perform artificially generated examples. Coherency in recognition was determined at 93.8%, corresponding to a recognition accuracy of 89.2% for the classifiers and 92.5% for human participants. A proof of concept was performed towards the expansion of the proposed one shot learning approach to adaptive learning, and the results are presented and the implications discussed.
Maria E. Cabrera, Natalia Sanchez-Tamayo, Richard M. Voyles, Juan P. Wachs
FG4
2017 A Semantical & Analytical Approach for Zero Shot Gesture Learning
abstract
Zero shot learning (ZSL) is about being able to recognize gesture classes that were never seen before. This type of recognition involves the understanding that the presented gesture is a new form of expression from those observed so far, and yet carries embedded information universal to all the other gestures (also referred as context). As part of the same problem, it is required to determine what action/command this new gesture conveys, in order to react to the command autonomously. Research in this area may shed light to areas where ZSL occurs, such as spontaneous gestures. People perform gestures that may be new to the observer. This occurs when the gesturer is learning, solving a problem or acquiring a new language. The ability of having a machine recognizing spontaneous gesturing, in the same manner as humans do, would enable more fluent human-machine interaction. In this paper, we describe a new paradigm for ZSL based on adaptive learning, where it is possible to determine the amount of transfer learning carried out by the algorithm and how much knowledge is acquired from a new gesture observation. Another contribution is a procedure to determine what are the best semantic descriptors for a given command and how to use those as part of the ZSL approach proposed.
Naveen Madapana, Juan P. Wachs
FG2
2017 ZSGL: zero shot gestural learning
abstract
Gesture recognition systems enable humans to interact with machines in an intuitive and a natural way. Humans tend to create the gestures on the fly and conventional systems lack adaptability to learn new gestures beyond the training stage. This problem can be best addressed using Zero Shot Learning (ZSL), a paradigm in machine learning that aims to recognize unseen objects by just having a description of them. ZSL for gestures has hardly been addressed in computer vision research due to the inherent ambiguity and the contextual dependency associated with the gestures. This work proposes an approach for Zero Shot Gestural Learning (ZSGL) by leveraging the semantic information that is embedded in the gestures. First, a human factors based approach has been followed to generate semantic descriptors for gestures that can generalize to the existing gesture classes. Second, we assess the performance of various existing state-of-the-art algorithms on ZSL for gestures using two standard datasets: MSRC-12 and CGD2011 dataset. The obtained results (26.35% - unseen class accuracy) parallel the benchmark accuracies of attribute-based object recognition and justifies our claim that ZSL is a desirable paradigm for gesture based systems.
Naveen Madapana, Juan P. Wachs
ICMI2
2017 The Effect of Embodied Interaction in Visual-Spatial Navigation
abstract
This article aims to assess the effect of embodied interaction on attention during the process of solving spatio-visual navigation problems. It presents a method that links operator's physical interaction, feedback, and attention. Attention is inferred through networks called Bayesian Attentional Networks (BANs). BANs are structures that describe cause-effect relationship between attention and physical action. Then, a utility function is used to determine the best combination of interaction modalities and feedback. Experiments involving five physical interaction modalities (vision-based gesture interaction, glove-based gesture interaction, speech, feet, and body stance) and two feedback modalities (visual and sound) are described. The main findings are: (i) physical expressions have an effect in the quality of the solutions to spatial navigation problems; (ii) the combination of feet gestures with visual feedback provides the best task performance.
Ting Zhang 0013, Yu-Ting Li, Juan P. Wachs
ACM Trans. Interact. Intell. Syst.3
2016 Multi-target detection and tracking from a single camera in Unmanned Aerial Vehicles (UAVs)
abstract
Despite the recent flight control regulations, Unmanned Aerial Vehicles (UAVs) are still gaining popularity in civilian and military applications, as much as for personal use. Such emerging interest is pushing the development of effective collision avoidance systems. Such systems play a critical role UAVs operations especially in a crowded airspace setting. Because of cost and weight limitations associated with UAVs payload, camera based technologies are the de-facto choice for collision avoidance navigation systems. This requires multi-target detection and tracking algorithms from a video, which can be run on board efficiently. While there has been a great deal of research on object detection and tracking from a stationary camera, few have attempted to detect and track small UAVs from a moving camera. In this paper, we present a new approach to detect and track UAVs from a single camera mounted on a different UAV. Initially, we estimate background motions via a perspective transformation model and then identify distinctive points in the background subtracted image. We find spatio-temporal traits of each moving object through optical flow matching and then classify those candidate targets based on their motion patterns compared with the background. The performance is boosted through Kalman filter tracking. This results in temporal consistency among the candidate detections. The algorithm was validated on video datasets taken from a UAV. Results show that our algorithm can effectively detect and track small UAVs with limited computing resources.
Jing Li 0169, Dong Hye Ye, Timothy H. Chung, Mathias Kölsch, Juan P. Wachs, Charles A. Bouman
IROS5
2016 Embodied gesture learning from one-shot
abstract
This paper discusses the problem of one shot gesture recognition. This is relevant to the field of human-robot interaction, where the user's intentions are indicated through spontaneous gesturing (one shot) to the robot. The novelty of this work consists of learning the process that leads to the creation of a gesture, rather on the gesture itself. In our case, the context involves the way in which humans produce the gestures - the kinematic and anthropometric characteristics and the users' proxemics (the use of the space around them). In the method presented, the strategy is to generate a dataset of realistic samples based on biomechanical features extracted from a single gesture sample. These features, called “the gist of a gesture”, are considered to represent what humans remember when seeing a gesture and the cognitive process involved when trying to replicate it. By adding meaningful variability to these features, a large training data set is created while preserving the fundamental structure of the original gesture. Having a large dataset of realistic samples enables training classifiers for future recognition. Three classifiers were trained and tested using a subset of ChaLearn dataset, resulting in all three classifiers showing rather similar performance around 80% recognition rate Our classification results show the feasibility and adaptability of the presented technique regardless of the classifier.
Maria E. Cabrera, Juan P. Wachs
RO-MAN2
2016 Enhanced control of a wheelchair-mounted robotic manipulator using 3-D vision and multimodal interaction
Hairong Jiang, Ting Zhang 0013, Juan P. Wachs, Bradley S. Duerstock
Comput. Vis. Image Underst.3
2016 Introduction to Special Issue on Body Tracking and Healthcare
abstract
This Special Issue on Body Tracking and Healthcare highlights the exciting possibilities that sensor technologies are opening up in health and well-being. From the assessment and monitoring of medi...
Kenton O'Hara, Abigail Sellen, Juan P. Wachs
Hum. Comput. Interact.3
2016 Optimal Modality Selection for Cooperative Human-Robot Task Completion
abstract
Human-robot cooperation in complex environments must be fast, accurate, and resilient. This requires efficient communication channels where robots need to assimilate information using a plethora of verbal and nonverbal modalities such as hand gestures, speech, and gaze. However, even though hybrid human-robot communication frameworks and multimodal communication have been studied, a systematic methodology for designing multimodal interfaces does not exist. This paper addresses the gap by proposing a novel methodology to generate multimodal lexicons which maximizes multiple performance metrics over a wide range of communication modalities (i.e., lexicons). The metrics are obtained through a mixture of simulation and real-world experiments. The methodology is tested in a surgical setting where a robot cooperates with a surgeon to complete a mock abdominal incision and closure task by delivering surgical instruments. Experimental results show that predicted optimal lexicons significantly outperform predicted suboptimal lexicons (p <; 0.05) in all metrics validating the predictability of the methodology. The methodology is validated in two scenarios (with and without modeling the risk of a human-robot collision) and the differences in the lexicons are analyzed.
Mithun George Jacob, Juan P. Wachs
IEEE Trans. Cybern.2
2016 User-Centered and Analytic-Based Approaches to Generate Usable Gestures for Individuals With Quadriplegia
abstract
Hand gesture-based interfaces have become increasingly popular as a form to interact with computing devices. Unfortunately, standard gesture interfaces are not very usable by individuals with upper limb motor impairments, including quadriplegics due to spinal cord injury (SCI). The objective of this paper is to convert an existing interface to be usable by users with motor impairments. The key idea is to project existing patterns of gestural behavior to match those exhibited by users with quadriplegia due to common cervical SCIs. Two complementary approaches (a user-centered and an analytic approach) have been developed and validated to provide both subjective and quantitative solutions to interface design. The feasibility of the proposed methodology was validated through user-based experimental paradigms. Through this study, subjects with upper extremity motor impairments preferred (gave a significantly lower Borg scale) the use of alternative constrained gestures generated by the proposed approach rather than the standard gestures.
Hairong Jiang, Bradley S. Duerstock, Juan P. Wachs
IEEE Trans. Hum. Mach. Syst.3
2016 A comparative study for telerobotic surgery using free hand gestures
abstract
This research presents an exploratory study among touch-based and touchless interfaces selected to teleoperate a highly dexterous surgical robot. The possibility of incorporating touchless interfaces into the surgical arena may provide surgeons with the ability to engage in telerobotic surgery similarly as if they were operating with their bare hands. On the other hand, precision and sensibility may be lost. To explore the advantages and drawbacks of these modalities, five interfaces were selected to send navigational commands to the Taurus robot in the system: Omega, Hydra, and a keyboard. The first represented touch-based, while Leap Motion and Kinect were selected as touchless interfaces. Three experimental designs were selected to test the system, based on standardized surgically related tasks and clinically relevant performance metrics measured to evaluate the user's performance, learning rates, control stability, and interaction naturalness. The current work provides a benchmark and validation framework for the comparison of these two groups of interfaces and discusses their potential for current and future adoption in the surgical setting.
Tian Zhou 0005, Maria E. Cabrera, Juan P. Wachs, Thomas Low, Chandru Sundaram
J. Hum. Robot Interact.3
2016 Virtual annotations of the surgical field through an augmented reality transparent display
Daniel Andersen, Voicu Popescu, Maria E. Cabrera, Aditya Shanghavi, Gerardo Gómez, Sherri Marley, Brian Mullis, Juan P. Wachs
Vis. Comput.8
2015 Touchless Telerobotic Surgery - Is It Possible at All?
abstract
This paper presents a comprehensive evaluation among touchless, vision-based hand tracking interfaces (Kinect and Leap Motion) and the feasibility of their adoption into the surgical theater compared to traditional interfaces.
Tian Zhou 0005, Maria E. Cabrera, Juan P. Wachs
AAAI3
2015 Determining natural and accessible gestures using uncontrolled manifolds and cybernetics
abstract
Recent studies revealed that hand gesture-based interfaces can complement therapies for individuals with upper motor impairments and reduce the need of traditional rehabilitation sessions through hospital visits. Unfortunately, existing gesture-based interfaces have been developed without considering the physical limitations of users with motor impairments. An analytic approach was presented in our previous work to convert existing gesture-based interfaces designed for able-bodied individuals to be usable by individuals with quadriplegia using the Laban Theory of Movement. This paper extends the previous work by including gesture variability analysis (based on Uncontrolled Manifolds theory) and robotic execution. A WAM robotic arm was used to mimic gesture trajectories and a physical metric was empirically obtained to evaluate the physical effort of each gesture. At last, an integration method was presented to determine the accessible gesture set based on both the stability and empirical robot execution. For all the gesture classes, the accessible gestures were found to lie within 31% of the optimality of stability and work, respectively.
Hairong Jiang, Chun-Hao Hsu, Bradley S. Duerstock, Juan P. Wachs
IROS4
2015 Model-Based System Specification With Tesperanto: Readable Text From Formal Graphics
abstract
Technical reports and papers may be represented by a fundamental model, which can take the form of a block diagram, a state-machine, a flow diagram, or alternatively some ad hoc chart. This basic scheme can convey better the true value of otherwise verbose and potentially encumbered narrative-based specifications. We present a model-based methodology for authoring technical documents. The underlying idea is to first formalize the system to be specified using a conceptual model, and then automatically generate from the tested and verified model a humanly-readable text in a subset of English we call Tesperanto. This technical documents' authoring methodology is carried out in an integrated bimodal text-graphics document authoring environment. The methodology was evaluated with the International Organization for Standardization standards and a medical robotics case study. The evaluation resulted in tangible improvements in the quality and consistency of international standards. Further, it can serve to document complex dynamics among agents, such as interaction between an operation room technician robot and the surgeon, suggesting that it could be applied to represent and bring value to other types of technical documents.
Alex Blekhman, Juan P. Wachs, Dov Dori
IEEE Trans. Syst. Man Cybern. Syst.2
2014 Optimal modality selection for multimodal human-machine systems using RIMAG
abstract
Interpersonal communication in human teams is multimodal by nature and hybrid robot-human teams should be capable of utilizing diverse verbal and non-verbal communication channels (e.g. gestures, speech, and gaze). Additionally, this interaction must fulfill requirements such as speed, accuracy and resilience. While multimodal communication has been researched and human-robot mixed team communication frameworks have been developed, the computation of an effective combination of communication modalities (multimodal lexicon) to maximize effectiveness is an untapped area of research. The proposed framework objectively determines the set of optimal lexicons through multiobjective optimization of performance metrics over all feasible lexicons. The methodology is applied to the surgical setting, where a robotic nurse can collaborate with a surgical team by delivering surgical instruments as required. In this time-sensitive, high-risk context, performance metrics are obtained through a mixture of real-world experiments and simulation. Experimental results validate the predictability of the method since predicted optimal lexicons significantly (p <; 0.01) outperform predicted suboptimal lexicons in time, error rate and false positive rates.
Mithun George Jacob, Juan P. Wachs
SMC2
2014 An analytic approach to decipher usable gestures for quadriplegic users
abstract
With the advent of new gaming technologies, hand gestures are gaining popularity as an effective communication channel for human computer interaction (HCI). This is particularly relevant for patients recovering from mobility-related injuries or debilitating conditions who use gesture-based gaming for rehabilitation therapy. Unfortunately, most gesture-based gaming systems are designed for able-bodied users and are difficult and costly to adapt to people with upper extremity mobility impairments. While interface customization is an active area of work in assistive technologies (AT), there is no existing formal and analytical grounded methodology to adapt gesture-based control systems for quadriplegics. The goal of this work is to solve this hurdle by developing a mathematical framework to project the patterns of gestural behavior designed for existing gesture systems to those exhibited by quadriplegic subjects due to spinal cord injury (SCI). A key component of our framework relied on Laban movement analysis (LMA) theory, and consisted of four steps: acquiring and preprocessing gesture trajectories, extracting feature vectors, training transform functions, and generating constrained gestures. The feasibility of this framework was validated through user-based experimental paradigms and subject validation. It was found that 100% of gestures selected by subjects with high-level SCIs came from the constrained gesture set. Even for the low-level quadriplegic subject, the alternative gestures were preferred.
Hairong Jiang, Bradley S. Duerstock, Juan P. Wachs
SMC3
2014 Linking attention to physical action in complex decision making problems
abstract
Embodied interaction concerns the way that user senses the environment, acquires information, and exhibits intention by means of physical action. In a complex decision making scenario, which requires maintaining high level of attention continuously and deep understanding about the task and its context, the use of embodied interaction has the potential to promote thinking and learning. Creating a framework that allows decision makers to interact with information using the whole body in intuitive ways may offer cognitive advantages and greater efficiency. This paper proposes such a computational framework based on a Bayesian approach (coined BAN) to infer operators' focus of attention based on the operators' physical expressions. Then, utility theory is adopted in order to determine the best combinations of interaction modalities and feedback. Experiments involving five physical interaction modalities (touchless, glove-based, and step gestures, speech, and body balance) and two feedback modalities (visual and sound) were conducted to assess the proposed framework's performance. This also includes the likelihood of assessed attention from enhanced BANs and task performance as a function of the interaction and control modalities. Results show that physical expressions have a determining factor in the quality of the solutions in spatio-navigational type of problems.
Yu-Ting Li, Juan P. Wachs
SMC2
2014 An augmented reality approach to surgical telementoring
abstract
Optimal surgery and trauma treatment integrates different surgical skills frequently unavailable in rural/field hospitals. Telementoring can provide the missing expertise, but current systems require the trainee to focus on a nearby telestrator, fail to illustrate coming surgical steps, and give the mentor an incomplete picture of the ongoing surgery. A new telementoring system is presented that utilizes augmented reality to enhance the sense of co-presence. The system allows a mentor to add annotations to be displayed for a mentee during surgery. The annotations are displayed on a tablet held between the mentee and the surgical site as a heads-up display. As it moves, the system uses computer vision algorithms to track and align the annotations with the surgical region. Tracking is achieved through feature matching. To assess its performance, comparisons are made between SURF and SIFT detector, brute force and FLANN matchers, and hessian blob thresholds. The results show that the combination of a FLANN matcher and a SURF detector with a 1500 hessian threshold can optimize this system across scenarios of tablet movement and occlusion.
Timo Loescher, Shih Yu Lee, Juan P. Wachs
SMC3
2014 Multimodal approach to image perception of histology for the blind or visually impaired
abstract
Currently there is no suitable substitute technology to enable blind or visually impaired people (BVI) to interpret visual scientific data commonly generated during lab experimentation in real time, such as performing light microscopy, spectrometry, and observing chemical reactions. This reliance upon visual interpretation of scientific data certainly impedes BVIs from advancing in careers in medicine, biology and chemistry. To address this challenge, a real-time multimodal image perception system is developed to transform the standard lab blood smear image for persons with BVI to perceive, employing a combination of auditory, haptic, and vibrotactile feedbacks. These sensory feedbacks are used to convey visual information in appropriate perceptual channels, thus creating a palette of multimodal, sensorial information. A Bayesian network is developed to characterize images through two groups of features of interest: primary and peripheral features. Then, a method is conceived for optimal matching between primary features and sensory modalities. Experimental results confirmed this real-time approach of higher accuracy in recognizing and analyzing objects within images compared to tactile papers.
Ting Zhang 0013, Greg J. Williams, Bradley S. Duerstock, Juan P. Wachs
SMC4
2014 Operation room tool handling and miscommunication scenarios: An object-process methodology conceptual model
Juan P. Wachs, Boaz Frenkel, Dov Dori
Artif. Intell. Medicine1
2014 HEGM: A hierarchical elastic graph matching for hand gesture recognition
Yu-Ting Li, Juan P. Wachs
Pattern Recognit.2
2014 Guest Editorial - Special Issue on Robust Recognition Methods for Multimodal Interaction
Luís Gómez Déniz, Juan P. Wachs, Julio Jacobo-Berlles
Pattern Recognit. Lett.2
2014 Context-based hand gesture recognition for the operating room
Mithun George Jacob, Juan P. Wachs
Pattern Recognit. Lett.2
2014 A Machine Vision-Based Gestural Interface for People With Upper Extremity Physical Impairments
abstract
A machine vision-based gestural interface was developed to provide individuals with upper extremity physical impairments an alternative way to perform laboratory tasks that require physical manipulation of components. A color and depth based 3-D particle filter framework was constructed with unique descriptive features for face and hands representation. This framework was integrated into an interaction model utilizing spatial and motion information to deal efficiently with occlusions and its negative effects. More specifically, the suggested method proposed solves the false merging and false labeling problems characteristic in tracking through occlusion. The same feature encoding technique was subsequently used to detect, track and recognize users' hands. Experimental results demonstrated that the proposed approach was superior to other state-of-the-art tracking algorithms when interaction was present (97.52% accuracy). For gesture encoding, dynamic motion models were created employing the dynamic time warping method. The gestures were classified using a conditional density propagation-based trajectory recognition method. The hand trajectories were classified into different classes (commands) with a recognition accuracy of 95.9%. In addition, the new approach was validated with the “one shot learning” paradigm with comparable results to those reported in 2012. In a validation experiment, the gestures were used to control a mobile service robot and a robotic arm in a laboratory chemistry experiment. Effective control policies were selected to achieve optimal performance for the presented gestural control system through comparison of task completion time between different control modes.
Hairong Jiang, Bradley S. Duerstock, Juan P. Wachs
IEEE Trans. Syst. Man Cybern. Syst.3
2013 Surgical instrument handling and retrieval in the operating room with a multimodal robotic assistant
abstract
A robotic scrub nurse (RSN) designed for safe human-robot collaboration in the operating room (OR) is presented. The RSN assists the surgical staff in the OR by delivering instruments to the surgeon and operates through a multimodal interface allowing instruments to be requested through verbal commands or touchless gestures. A machine vision algorithm was designed to recognize the hand gestures performed by the user. To ensure safe human-robot collaboration, tool-tip trajectories are planned and executed to avoid collisions with the user. Experiments were conducted to test the system when speech and gesture modalities were used to interact with the robot, separately and together. The average system times were compared while performing a mock surgical task for each modality of interaction. The effects of modality training on task completion time were also studied. It was found that training results in a significant drop of 12.92% in task completion time. Experimental results show that 95.96% of the gestures used to interact with the robot were recognized correctly, and collisions with the user were completely avoided when using a new active obstacle avoidance algorithm.
Mithun George Jacob, Yu-Ting Li, Juan P. Wachs
ICRA3
2013 Recognizing hand gestures using the weighted elastic graph matching (WEGM) method
Yu-Ting Li, Juan P. Wachs
Image Vis. Comput.2
2012 Intention, Context and Gesture Recognition for Sterile MRI Navigation in the Operating Room
Mithun George Jacob, Christopher Cange, Rebecca Packer, Juan P. Wachs
CIARP4
2012 Facilitated Gesture Recognition Based Interfaces for People with Upper Extremity Physical Impairments
Hairong Jiang, Juan P. Wachs, Bradley S. Duerstock
CIARP2
2012 Hierarchical Elastic Graph Matching for Hand Gesture Recognition
Yu-Ting Li, Juan P. Wachs
CIARP2
2012 Robot, Pass Me the Scissors! How Robots Can Assist Us in the Operating Room
Juan P. Wachs
CIARP1
2012 Gestonurse: a multimodal robotic scrub nurse
abstract
A novel multimodal robotic scrub nurse (RSN) system for the operating room (OR) is presented. The RSN assists the main surgeon by passing surgical instruments. Experiments were conducted to test the system with speech and gesture modalities and average instrument acquisition times were compared. Experimental results showed that 97% of the gestures were recognized correctly under changes in scale and rotation and that the multimodal system responded faster than the unimodal systems. A relationship similar in form to Fitts's law for instrument picking accuracy is also presented.
Mithun George Jacob, Yu-Ting Li, Juan P. Wachs
HRI3
2011 A gesture driven robotic scrub nurse
abstract
A gesture driven robotic scrub nurse (GRSN) for the operating room (OR) is presented. The GRSN passes surgical instruments to the surgeon during surgery which reduces the workload of a human scrub nurse. This system offers several advantages such as freeing human nurses to perform concurrent tasks, and reducing errors in the OR due to miscommunication or absence of surgical staff. Hand gestures are recognized from a video stream, converted to instructions, and sent to a robotic arm which passes the required surgical instruments to the surgeon. Experimental results show that 95% of the gestures were recognized correctly. The gesture recognition algorithm presented is robust to changes in scale and rotation of the hand gestures. The system was compared to human task performance and was found to be only 0.83 seconds slower on average.
Mithun George Jacob, Yu-Ting Li, Juan P. Wachs
SMC3
2009 The Multi-level Learning and Classification of Multi-class Parts-Based Representations of U.S. Marine Postures
Deborah Goshorn, Juan P. Wachs, Mathias Kölsch
CIARP2
2008 Technical Brief: A Gesture-based Tool for Sterile Browsing of Radiology Images
abstract
The use of doctor-computer interaction devices in the operation room (OR) requires new modalities that support medical imaging manipulation while allowing doctors' hands to remain sterile, supporting their focus of attention, and providing fast response times. This paper presents "Gestix," a vision-based hand gesture capture and recognition system that interprets in real-time the user's gestures for navigation and manipulation of images in an electronic medical record (EMR) database. Navigation and other gestures are translated to commands based on their temporal trajectories, through video capture. "Gestix" was tested during a brain biopsy procedure. In the in vivo experiment, this interface prevented the surgeon's focus shift and change of location while achieving a rapid intuitive reaction and easy interaction. Data from two usability tests provide insights and implications regarding human-computer interaction based on nonverbal conversational modalities.
Juan P. Wachs, Helman Stern, Yael Edan, Michael Gillam, Jonathan A. Handler, Craig Feied, Mark S. Smith
J. Am. Medical Informatics Assoc.1
2006 A Real-Time Gesture Interface for Hands-Free Control of Electronic Medical Records
Craig Feied, Michael Gillam, Juan P. Wachs, Jonathan A. Handler, Helman Stern, Mark S. Smith
AMIA3
2006 Human Factors for Design of Hand Gesture Human - Machine Interaction
abstract
A global approach to hand gesture vocabulary design is proposed which includes human as well as technical design factors. The method of selecting gestures for preconceived command vocabularies has not been addressed in a systematic manner. Present methods are ad hoc. In an analytical approach technological factors of gesture recognition accuracy are easily obtained and well studied. Conversely, it is difficult to obtain measures of human centered desires (intuitiveness, comfort). These factors, being subjective, are costly and time consuming to obtain, and hence we have developed automated methods for acquisition of these data through specially designed applications. Results of the intuitiveness experiments showed when commands are presented as stimuli the gestural responses vary widely over a population of subjects. This result refutes the hypothesis that there exist universal common gestures to express user intentions or commands.
Helman Stern, Juan P. Wachs, Yael Edan
SMC2
2005 Cluster labeling and parameter estimation for the automated setup of a hand-gesture recognition system
abstract
In this work, we address the issue of reconfigurability of a hand-gesture recognition system. The calibration or setup of the operational parameters of such a system is a time-consuming effort, usually performed by trial and error, and often causing system performance to suffer because of designer impatience. In this work, we suggest a methodology using a neighborhood-search algorithm for tuning system parameters. Thus, the design of hand-gesture recognition systems is transformed into an optimization problem. To test the methodology, we address the difficult problem of simultaneous calibration of the parameters of the image processing/fuzzy C-means (FCM) components of a hand-gesture recognition system. In addition, we proffer a method for supervising the FCM algorithm using linear programming and heuristic labeling. Resulting solutions exhibited fast convergence (in the order of ten iterations) to reach recognition accuracies within several percent of the optimal. Comparative performance testing using three gesture databases (BGU, American Sign Language and Gripsee), and a real-time implementation (Tele-Gest) are reported on.
Juan P. Wachs, Helman Stern, Yael Edan
IEEE Trans. Syst. Man Cybern. Part A1
2003 Parameter search for an image processing fuzzy C-means hand gesture recognition system
abstract
This work describes a hand gesture recognition system using an optimized image processing-fuzzy C-means (FCM) algorithm. The parameters of the image processing and clustering algorithm were simultaneously found using a neighborhood parameter search routine, resulting in solutions within 1-2% of optimal. Comparison of user dependent and user independent systems, when tested with their own trainers, resulted in recognition accuracies of 98.9% and 98.2%, respectively. For experienced users, the opposite was true, testing recognition accuracies where better for user independent than user dependent systems (98.2% over 96.0%). These results are statistically significant at the .007 levels.
Juan P. Wachs, Helman Stern, Yael Edan
ICIP (3)1