Babak Taati

dblp:32/4116 · DBLP profile ↗
← Back
29ranked-venue papers
4as first author
17since 2021 · last 2026
0000-0001-9763-4293ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 1 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 4 since 2021Security and privacy · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Pain in 3D: Controllable Generation of Synthetic Faces for Automated Pain Assessment
Xin Lei Lin, Soroush Mehraban, Abhishek Moturu, Babak Taati
ICPR (15)4
2026 GAITGen: Disentangled Motion-Pathology Impaired Gait Generative Model - Bringing Motion Generation to the Clinical Domain
abstract
Gait analysis is crucial for the diagnosis and monitoring of movement disorders like Parkinson's Disease. While computer vision models have shown potential for objectively evaluating parkinsonian gait, their effectiveness is limited by scarce clinical datasets and the challenge of collecting large and well-labelled data, impacting model accuracy and risk of bias. To address these gaps, we propose GAITGen, a novel framework that generates realistic gait sequences conditioned on specified pathology severity levels. GAITGen employs a Conditional Residual Vector Quantized Variational Autoencoder to learn disentangled representations of motion dynamics and pathology-specific factors, coupled with Mask and Residual Transformers for conditioned sequence generation. GAITGen generates realistic, diverse gait sequences across severity levels, enriching datasets and enabling large-scale model training in parkinsonian gait analysis. Experiments on our new PD-GaM (real) dataset demonstrate that GAITGen outperforms adapted state-of-the-art models in both reconstruction fidelity and generation quality, accurately capturing critical pathology-specific gait features. A clinical user study confirms the realism and clinical relevance of our generated sequences. Moreover, incorporating GAITGen-generated data into downstream tasks improves parkinsonian gait severity estimation, highlighting its potential for advancing clinical gait analysis.
Vida Adeli, Soroush Mehraban, Majid Mirmehdi, Alan L. Whone, Benjamin Filtjens, Amirhossein Dadashzadeh, Alfonso Fasano, Andrea Iaboni, Babak Taati
WACV9
2026 FastHMR: Accelerating Human Mesh Recovery via Token and Layer Merging with Diffusion Decoding
abstract
Recent transformer-based models for 3D Human Mesh Recovery (HMR) have achieved strong performance but often suffer from high computational cost and complexity due to deep transformer architectures and redundant tokens. In this paper, we introduce two HMR-specific merging strategies: Error-Constrained Layer Merging (ECLM) and Mask-guided Token Merging (Mask-ToMe). ECLM selectively merges transformer layers that have minimal impact on the Mean Per Joint Position Error (MPJPE), while Mask-ToMe focuses on merging background tokens that contribute little to the final prediction. To further address the potential performance drop caused by merging, we propose a diffusion-based decoder that incorporates temporal context and leverages pose priors learned from large-scale motion capture datasets. Experiments across multiple benchmarks demonstrate that our method achieves up to 2.3× speed-up while slightly improving performance over the baseline.
Soroush Mehraban, Andrea Iaboni, Babak Taati
WACV3
2026 STARS: Self-supervised Tuning for 3D Action Recognition in Skeleton Sequences
abstract
Self-supervised pretraining methods with masked prediction demonstrate remarkable within-dataset performance in skeleton-based action recognition. However, we show that, unlike contrastive learning approaches, they do not produce well-separated clusters. Additionally, these methods struggle with generalization in few-shot settings. To address these issues, we propose Self-supervised Tuning for 3D Action Recognition in Skeleton sequences (STARS). Specifically, STARS first uses a masked prediction stage using an encoder-decoder architecture. It then employs nearest-neighbor contrastive learning to partially tune the weights of the encoder, enhancing the formation of semantic clusters for different actions. By tuning the encoder for a few epochs, and without using hand-crafted data augmentations, STARS achieves state-of-the-art self-supervised results in various benchmarks, including NTU-60, NTU-120, and PKU-MMD. In addition, STARS exhibits significantly better results than masked prediction models in few-shot settings, where the model has not seen the actions throughout pretraining.
Soroush Mehraban, Mohammad Javad Rajabi, Andrea Iaboni, Babak Taati
WACV4
2025 LIFT: Latent Implicit Functions for Task- and Data-Agnostic Encoding
abstract
Implicit Neural Representations (INRs) are proving to be a powerful paradigm in unifying task modeling across diverse data domains, offering key advantages such as memory efficiency and resolution independence. Conventional deep learning models are typically modality-dependent, often requiring custom architectures and objectives for different types of signals. However, existing INR frameworks frequently rely on global latent vectors or exhibit computational inefficiencies that limit their broader applicability. We introduce LIFT, a novel, high-performance framework that addresses these challenges by capturing multiscale information through meta-learning. LIFT leverages multiple parallel localized implicit functions alongside a hierarchical latent generator to produce unified latent representations that span local, intermediate, and global features. This architecture facilitates smooth transitions across local regions, enhancing expressivity while maintaining inference efficiency. Additionally, we introduce ReLIFT, an enhanced variant of LIFT that incorporates residual connections and expressive frequency encodings. With this straightforward approach, ReLIFT effectively addresses the convergence-capacity gap found in comparable methods, providing an efficient yet powerful solution to improve capacity and speed up convergence. Empirical results show that LIFT achieves state-of-the-art (SOTA) performance in generative modeling and classification tasks, with notable reductions in computational costs. Moreover, in single-task settings, the streamlined ReLIFT architecture proves effective in signal representations and inverse problem tasks.
Amirhossein Kazerouni, Soroush Mehraban, Michael Brudno, Babak Taati
ICCV4
2025 Care-PD: A Multi-Site Anonymized Clinical Dataset for Parkinson's Disease Gait Assessment
abstract
Objective gait assessment in Parkinson’s Disease (PD) is limited by the absence of large, diverse, and clinically annotated motion datasets. We introduce Care-PD, the largest publicly available archive of 3D mesh gait data for PD, and the first multi-site collection spanning 9 cohorts from 8 clinical centers. All recordings (RGB video or motion capture) are converted into anonymized SMPL meshes via a harmonized preprocessing pipeline. Care-PD supports two key benchmarks: supervised clinical score prediction (estimating Unified Parkinson’s Disease Rating Scale, UPDRS, gait scores) and unsupervised motion pretext tasks (2D-to-3D keypoint lifting and full-body 3D reconstruction). Clinical prediction is evaluated under four generalization protocols: within-dataset, cross-dataset, leave-one-dataset-out, and multi-dataset in-domain adaptation.To assess clinical relevance, we compare state-of-the-art motion encoders with a traditional gait-feature baseline, finding that encoders consistently outperform handcrafted features. Pretraining on Care-PD reduces MPJPE (from 60.8mm to 7.5mm) and boosts PD severity macro-F1 by 17\%, underscoring the value of clinically curated, diverse training data. Care-PD and all benchmark code are released for non-commercial research (Code, Data).
Vida Adeli, Ivan Klabucar, Javad Rajabi, Benjamin Filtjens, Soroush Mehraban, Diwei Wang, Trung-Hieu Hoang, Minh N. Do, Hyewon Seo, Candice Müller, Daniel Boari Coelho, Claudia de Oliveira, Pieter Ginis, Moran Gilat, Alice Nieuwboer, Joke Spildooren, J. Lucas McKay, Hyeokhyen Kwon, Gari D. Clifford, Christine D. Esper, Stewart A. Factor, Imari Genias, Amirhossein Dadashzadeh, Leia C. Shum, Alan L. Whone, Majid Mirmehdi, Andrea Iaboni, Babak Taati
NeurIPS28
2025 Token Perturbation Guidance for Diffusion Models
abstract
Classifier-free guidance (CFG) has become an essential component of modern diffusion models to enhance both generation quality and alignment with input conditions. However, CFG requires specific training procedures and is limited to conditional generation. To address these limitations, we propose Token Perturbation Guidance (TPG), a novel method that applies perturbation matrices directly to intermediate token representations within the diffusion network. TPG employs a norm-preserving shuffling operation to provide effective and stable guidance signals that improve generation quality without architectural changes. As a result, TPG is training-free and agnostic to input conditions, making it readily applicable to both conditional and unconditional generation. We also analyze the guidance term provided by TPG and show that its effect on sampling more closely resembles CFG compared to existing training-free guidance techniques. We extensively evaluate TPG on SDXL and Stable Diffusion 2.1, demonstrating nearly a 2x improvement in FID for unconditional generation over the SDXL baseline and showing that TPG closely matches CFG in prompt alignment. Thus, TPG represents a general, condition-agnostic guidance method that extends CFG-like benefits to a broader class of diffusion models.
Javad Rajabi, Soroush Mehraban, Seyedmorteza Sadat, Babak Taati
NeurIPS4
2025 SUM: Saliency Unification Through Mamba for Visual Attention Modeling
abstract
Visual attention modeling, important for interpreting and prioritizing visual stimuli, plays a significant role in applications such as marketing, multimedia, and robotics. Traditional saliency prediction models, especially those based on Convolutional Neural Networks (CNNs) or Transformers, achieve notable success by leveraging large-scale annotated datasets. However, the current state-of-the-art (SOTA) models that use Transformers are computationally expensive. Additionally, separate models are often required for each image type, lacking a unified approach. In this paper, we propose Saliency Unification through Mamba (SUM), a novel approach that integrates the efficient long-range dependency modeling of Mamba with U-Net to provide a unified model for diverse image types. Using a novel Conditional Visual State Space (C- VSS) block, SUM dynamically adapts to various image types, including natural scenes, web pages, and commercial imagery, ensuring universal applicability across different data types. Our comprehen-sive evaluations across five benchmarks demonstrate that SUM seamlessly adapts to different visual characteristics and consistently outperforms existing models. These results position SUM as a versatile and powerful tool for advancing visual attention modeling, offering a robust solution universally applicable across different types of visual content. Our codebase and pretrained models are publicly accessible on the https://arhosseini77.github.io/sum_page/.
Alireza Hosseini, Amirhossein Kazerouni, Saeed Akhavan, Michael Brudno, Babak Taati
WACV5
2024 Benchmarking Skeleton-based Motion Encoder Models for Clinical Applications: Estimating Parkinson's Disease Severity in Walking Sequences
abstract
Parkinson's Disease is a degenerative disorder for which precise motor assessment is critical. This study investigates the application of general human motion encoders trained on large-scale human motion datasets for analyzing gait patterns in PD patients. Although these models have learned a wealth of human biomechanical knowledge, their effectiveness in analyzing pathological movements, such as parkinsonian gait, has yet to be fully validated. We propose a comparative framework and evaluate six pre-trained state-of-the-art human motion encoder models on their ability to predict the Movement Disorder Society - Unified Parkinson's Disease Rating Scale (MDS-UPDRS-III) gait scores from motion capture data. We compare these against a traditional gait feature-based predictive model in a recently released large public PD dataset, including PD patients on and off medication. The feature-based model currently shows higher weighted average accuracy, precision, recall, and F1-score. Motion encoder models with closely comparable results demonstrate promise for scalability and efficiency in clinical settings. This potential is underscored by the enhanced performance of the encoder model upon fine-tuning on PD training set. Four of the six human motion models examined provided prediction scores that were significantly different between on- and off-medication states. This finding reveals the sensitivity of motion encoder models to nuanced clinical changes, emphasizing their potential utility in clinical settings. It also underscores the necessity for continued customization of these models to better capture disease-specific features, thereby reducing the reliance on labor-intensive feature engineering. Lastly, we establish a benchmark for the analysis of skeleton-based motion encoder models in clinical settings. To the best of our knowledge, this is the first study to provide a benchmark that enables state-of-the-art models to be tested and compete in a clinical context. Codes and benchmark leaderboard are available at code.
Vida Adeli, Soroush Mehraban, Irene Ballester, Yasamin Zarghami, Andrea Sabo, Andrea Iaboni, Babak Taati
FG7
2024 Evaluating Recent 2D Human Pose Estimators for 2D-3D Pose Lifting
abstract
Monocular 3D human pose estimation involves predicting the 3D pixel coordinates of key body joints from a 2D image or video. Typically, a 2D estimation model is employed to initially determine joint locations in an image, followed by training a separate model to lift these positions to 3D coordinates. In this paper, we evaluate the performance of recently proposed 2D human pose estimation models as different inputs for training and evaluation of 2D-3D lifting models. In addition, we propose four simple merging strategies to combine the outputs of these 2D human pose estimators and generate less noisy 2D inputs. To evaluate, four recent 2D pose estimators—ViTPose, PCT, MogaNet, and TransPose—are selected, and their corresponding 2D outputs are generated on the Human3.6M dataset. Subsequently, MotionAGFormer and PoseFormerV2 are trained and evaluated using each created 2D input and its corresponding 3D motion-capture ground truth. ViTPose stands out as the top-performing 2D estimator, and employing all merging strategies proves beneficial in generating a less noisy 2D input. Code and data are available at https://github.com/TaatiTeam/2DEstimatorEval.
Soroush Mehraban, Yiqian Qin, Babak Taati
FG3
2024 MotionAGFormer: Enhancing 3D Human Pose Estimation with a Transformer-GCNFormer Network
abstract
Recent transformer-based approaches have demonstrated excellent performance in 3D human pose estimation. However, they have a holistic view and by encoding global relationships between all the joints, they do not capture the local dependencies precisely. In this paper, we present a novel Attention-GCNFormer (AGFormer) block that divides the number of channels by using two parallel transformer and GCNFormer streams. Our proposed GCNFormer module exploits the local relationship between adjacent joints, outputting a new representation that is complementary to the transformer output. By fusing these two representation in an adaptive way, AGFormer exhibits the ability to better learn the underlying 3D structure. By stacking multiple AGFormer blocks, we propose MotionAGFormer in four different variants, which can be chosen based on the speed-accuracy trade-off. We evaluate our model on two popular benchmark datasets: Human3.6M and MPI-INF-3DHP. MotionAGFormer-B achieves state-of-the-art results, with P1 errors of 38.4 mm and 16.2 mm, respectively. Remarkably, it uses a quarter of the parameters and is three times more computationally efficient than the previous leading model on Human3.6M dataset. Code and models are available at https://github.com/TaatiTeam/MotionAGFormer.
Soroush Mehraban, Vida Adeli, Babak Taati
WACV3
2023 Pain Detection in Masked Faces during Procedural Sedation
abstract
Pain monitoring is essential to the quality of care for patients undergoing a medical procedure with sedation. An automated mechanism for detecting pain could improve sedation dose titration. Previous studies on facial pain detection have shown the viability of computer vision methods in detecting pain in unoccluded faces. However, the faces of patients undergoing procedures are often partially occluded by medical devices and face masks. A previous preliminary study on pain detection on artificially occluded faces has shown a feasible approach to detect pain from a narrow band around the eyes. This study has collected video data from masked faces of 14 patients undergoing procedures in an interventional radiology department and has trained a deep learning model using this dataset. The model was able to detect expressions of pain accurately and, after causal temporal smoothing, achieved an average precision (AP) of 0.72 and an area under the receiver operating characteristic curve (AVC) of 0.82. These results outperform baseline models and show viability of computer vision approaches for pain detection of masked faces during procedural sedation. Cross-dataset performance is also examined when a model is trained on a publicly available dataset and tested on the sedation videos. The ways in which pain expressions differ in the two datasets are qualitatively examined.
Yasamin Zarghami, S. Mafeld, Aaron Conway, Babak Taati
FG4
2023 Large Language Models are Fixated by Red Herrings: Exploring Creative Problem Solving and Einstellung Effect using the Only Connect Wall Dataset
abstract
The quest for human imitative AI has been an enduring topic in AI research since inception. The technical evolution and emerging capabilities of the latest cohort of large language models (LLMs) have reinvigorated the subject beyond academia to cultural zeitgeist. While recent NLP evaluation benchmark tasks test some aspects of human-imitative behaviour (e.g., BIG-bench's `human-like behavior' tasks), few, if not none, examine creative problem solving abilities. Creative problem solving in humans is a well-studied topic in cognitive neuroscience with standardized tests that predominantly use ability to associate (heterogeneous) connections among clue words as a metric for creativity. Exposure to misleading stimuli --- distractors dubbed red herrings --- impede human performance in such tasks via the fixation effect and Einstellung paradigm. In cognitive neuroscience studies, such fixations are experimentally induced by pre-exposing participants to orthographically similar incorrect words to subsequent word-fragments or clues. The popular British quiz show Only Connect's Connecting Wall segment essentially mimics Mednick's Remote Associates Test (RAT) formulation with built-in, deliberate red herrings, that makes it an ideal proxy dataset to explore and study fixation effect and Einstellung paradigm from cognitive neuroscience in LLMs. In addition to presenting the novel Only Connect Wall (OCW) dataset, we also report results from our evaluation of selected pre-trained language models and LLMs (including OpenAI's GPT series) on creative problem solving tasks like grouping clue words by heterogeneous connections, and identifying correct open knowledge domain connections in respective groups. We synthetically generate two additional datasets: OCW-Randomized, OCW-WordNet to further analyze our red-herrings hypothesis in language models.The code and link to the dataset is available at url.
Saeid Alavi Naeini, Raeid Saqur, Mozhgan Saeidi, John M. Giorgi, Babak Taati
NeurIPS5
2023 Ambient Monitoring of Gait and Machine Learning Models for Dynamic and Short-Term Falls Risk Assessment in People With Dementia
abstract
Falls are a leading cause of morbidity and mortality in older adults with dementia residing in long-term care. Having access to a frequently updated and accurate estimate of the likelihood of a fall over a short time frame for each resident will enable care staff to provide targeted interventions to prevent falls and resulting injuries. To this end, machine learning models to estimate and frequently update the risk of a fall within the next 4 weeks were trained on longitudinal data from 54 older adult participants with dementia. Data from each participant included baseline clinical assessments of gait, mobility, and fall risk at the time of admission, daily medication intake in three medication categories, and frequent assessments of gait performed via a computer vision-based ambient monitoring system. Systematic ablations investigated the effects of various hyperparameters and feature sets and experimentally identified differential contributions from baseline clinical assessments, ambient gait analysis, and daily medication intake. In leave-one-subject-out cross-validation, the best performing model predicts the likelihood of a fall over the next 4 weeks with a sensitivity and specificity of 72.8 and 73.2, respectively, and achieved an area under the receiver operating characteristic curve (AUROC) of 76.2. By contrast, the best model excluding ambient gait features achieved an AUROC of 56.2 with a sensitivity and specificity of 51.9 and 54.0, respectively. Future research will focus on externally validating these findings to prepare for the implementation of this technology to reduce fall and fall-related injuries in long-term care.
Vida Adeli, Navid Korhani, Andrea Sabo, Sina Mehdizadeh, Avril Mansfield, Alastair Flint, Andrea Iaboni, Babak Taati
IEEE J. Biomed. Health Informatics8
2022 Estimating Parkinsonism Severity in Natural Gait Videos of Older Adults With Dementia
abstract
Drug-induced parkinsonism affects many older adults with dementia, often causing gait disturbances. New advances in vision-based human pose- estimation have opened possibilities for frequent and unobtrusive analysis of gait in long-term care settings. This work leverages spatial-temporal graph convolutional network (ST-GCN) architectures and training procedures to predict clinical scores of parkinsonism in gait from video of individuals with dementia. We propose a two-stage training approach consisting of a self-supervised pretraining stage that encourages the ST-GCN model to learn about gait patterns before predicting clinical scores in the finetuning stage. The proposed ST-GCN models are evaluated on joint trajectories extracted from video and are compared against traditional (ordinal, linear, random forest) regression models and temporal convolutional network baselines. Three 2D human pose-estimation libraries (OpenPose, Detectron, AlphaPose) and the Microsoft Kinect (2D and 3D) are used to extract joint trajectories of 4787 natural walking bouts from 53 older adults with dementia. A subset of 399 walks from 14 participants is annotated with scores of parkinsonism severity on the gait criteria of the Unified Parkinson's Disease Rating Scale (UPDRS) and the Simpson-Angus Scale (SAS). Our results demonstrate that ST-GCN models operating on 3D joint trajectories extracted from the Kinect consistently outperform all other models and feature sets. Prediction of parkinsonism scores in natural walking bouts of unseen participants remains a challenging task, with the best models achieving macro-averaged F1-scores of 0.53 ± 0.03 and 0.40 ± 0.02 for UPDRS-gait and SAS-gait, respectively. Pre-trained model and demo code for this work is available.1
Andrea Sabo, Sina Mehdizadeh, Andrea Iaboni, Babak Taati
IEEE J. Biomed. Health Informatics4
2021 A New Dataset for Facial Motion Analysis in Individuals With Neurological Disorders
abstract
We present the first public dataset with videos of oro-facial gestures performed by individuals with oro-facial impairment due to neurological disorders, such as amyotrophic lateral sclerosis (ALS) and stroke. Perceptual clinical scores from trained clinicians are provided as metadata. Manual annotation of facial landmarks is also provided for a subset of over 3300 frames. Through extensive experiments with multiple facial landmark detection algorithms, including state-of-the-art convolutional neural network (CNN) models, we demonstrated the presence of bias in the landmark localization accuracy of pre-trained face alignment approaches in our participant groups. The pre-trained models produced higher errors in the two clinical groups compared to age-matched healthy control subjects. We also investigated how this bias changes when the existing models are fine-tuned using data from the target population. The release of this dataset aims to propel the development of face alignment algorithms robust to the presence of oro-facial impairment, support the automatic analysis and recognition of oro-facial gestures, enhance the automatic identification of neurological diseases, as well as the estimation of disease severity from videos and images.
Andrea Bandini, Sia Rezaei, Diego L. Guarin, Madhura Kulkarni, Derrick Lim, Mark I. Boulos, Lorne Zinman, Yana Yunusova, Babak Taati
IEEE J. Biomed. Health Informatics9
2021 Unobtrusive Pain Monitoring in Older Adults With Dementia Using Pairwise and Contrastive Training
abstract
Although pain is frequent in old age, older adults are often undertreated for pain. This is especially the case for long-term care residents with moderate to severe dementia who cannot report their pain because of cognitive impairments that accompany dementia. Nursing staff acknowledge the challenges of effectively recognizing and managing pain in long-term care facilities due to lack of human resources and, sometimes, expertise to use validated pain assessment approaches on a regular basis. Vision-based ambient monitoring will allow for frequent automated assessments so care staff could be automatically notified when signs of pain are displayed. However, existing computer vision techniques for pain detection are not validated on faces of older adults or people with dementia, and this population is not represented in existing facial expression datasets of pain. We present the first fully automated vision-based technique validated on a dementia cohort. Our contributions are threefold. First, we develop a deep learning-based computer vision system for detecting painful facial expressions on a video dataset that is collected unobtrusively from older adult participants with and without dementia. Second, we introduce a pairwise comparative inference method that calibrates to each person and is sensitive to changes in facial expression while using training data more efficiently than sequence models. Third, we introduce a fast contrastive training method that improves cross-dataset performance. Our pain estimation model outperforms baselines by a wide margin, especially when evaluated on faces of people with dementia. Pre-trained model and demo code available at https://github.com/TaatiTeam/pain_detection_demo.
Siavash Rezaei, Abhishek Moturu, Shun Zhao, Kenneth M. Prkachin, Thomas Hadjistavropoulos, Babak Taati
IEEE J. Biomed. Health Informatics6
2020 Estimation of Orofacial Kinematics in Parkinson's Disease: Comparison of 2D and 3D Markerless Systems for Motion Tracking
abstract
Orofacial deficits are common in people with Parkinson's disease (PD) and their evolution might represent an important biomarker of disease progression. We are developing an automated system for assessment of orofacial function in PD that can be used in-home or in-clinic and can provide useful and objective clinical information that informs disease management. Our current approach relies on color and depth cameras for the estimation of 3D facial movements. However, depth cameras are not commonly available, might be expensive, and require specialized software for control and data processing. The objective of this paper was to evaluate if depth cameras are needed to differentiate between healthy controls and PD patients based on features extracted from orofacial kinematics. Results indicate that 2D features, extracted from color cameras only, are as informative as 3D features, extracted from color and depth cameras, differentiating healthy controls from PD patients. These results pave the way for the development of a universal system for automatic and objective assessment of orofacial function in PD.
Diego L. Guarin, Aidan Dempster, Andrea Bandini, Yana Yunusova, Babak Taati
FG5
2019 Pain Expression Recognition Using Occluded Faces
abstract
Frequent pain monitoring in intensive care unit (ICU) settings has been shown to have achieved improvement in patient outcomes. Specifically, the nursing staff is required to assess the pain of the patients approximately every four hours, and make necessary pain medication adjustments. Shortage and overburdening of nursing staff is well acknowledged in many hospital settings worldwide. In this paper, we present a preliminary study on automatic recognition of facial expression of pain in ICU settings. There has been considerable work on computer based pain recognition for settings beyond the ICU, wherein, usually the entire face of the person is visible without occlusions. The ICU setting, however, brings along unique challenges for recognizing facial expression of pain – the patient’s face is often covered by a respirator mask, the face is partially occluded by accessories attached to the respirator, and the patient is occasionally in a transient state of consciousness. In this paper we investigate the use of computer vision techniques for recognizing pain from partially visible faces. To establish a proof-of-concept, we simulate the occlusions most likely to happen in an ICU setting by masking out the nose, cheeks, and mouth regions, and use only a narrow band around the eyes of the patient as an input to our method. We generate these simulated occlusions using data from the UNBC-McMaster pain expression archive −a previously used dataset for pain expression recognition based on fully visible faces. We investigate multiple feature representations by extracting features only on the basis of the eye-region and train classifiers to detect the presence of pain using leave-one-person-out cross validation. Our results suggest a potential viability of automatic pain monitoring in the ICU settings involving face occlusions.
Ahmed Ashraf 0001, Anqi Yang, Babak Taati
FG3
2018 Automatic Detection of Amyotrophic Lateral Sclerosis (ALS) from Video-Based Analysis of Facial Movements: Speech and Non-Speech Tasks
abstract
The analysis of facial movements in patients with amyotrophic lateral sclerosis (ALS) can provide important information about early diagnosis and tracking disease progression. However, the use of expensive motion tracking systems has limited the clinical utility of the assessment. In this study, we propose a marker-less video-based approach to discriminate patients with ALS from neurotypical subjects. Facial movements were recorded using a depth sensor (Intel® RealSense" SR300) during speech and nonspeech tasks. A small set of kinematic features of lips was extracted in order to mirror the perceptual evaluation performed by clinicians, considering the following aspects: (1) range of motion, (2) speed of motion, (3) symmetry, and (4) shape. Our results demonstrate that it is possible to distinguish patients with ALS from neurotypical subjects with high overall accuracy (up to 88.9%) during repetitions of sentences, syllables, and labial non-speech movements (e.g., lip spreading). This paper provides strong rationale for the development of automated systems to detect neurological diseases from facial movements. This work has a high social impact, as it opens new possibilities to develop intelligent systems to support clinicians in their diagnosis, introducing novel standards for assessing the oro-facial impairment in ALS, and tracking disease progression remotely from home.
Andrea Bandini, Jordan R. Green, Babak Taati, Silvia Orlandi, Lorne Zinman, Yana Yunusova
FG3
2017 Subspace selection to suppress confounding source domain information in AAM transfer learning
abstract
Active appearance models (AAMs) have seen tremendous success in face analysis. However, model learning depends on the availability of detailed annotation of canonical landmark points. As a result, when accurate AAM fitting is required on a different set of variations (expression, pose, identity), a new dataset is collected and annotated. To overcome the need for time consuming data collection and annotation, transfer learning approaches have received recent attention. The goal is to transfer knowledge from previously available datasets (source) to a new dataset (target). We propose a subspace transfer learning method, in which we select a subspace from the source that best describes the target space. We propose a metric to compute the directional similarity between the source eigenvectors and the target subspace. We show an equivalence between this metric and the variance of target data when projected onto source eigenvectors. Using this equivalence, we select a subset of source principal directions that capture the variance in target data. To define our model, we augment the selected source subspace with the target subspace learned from a handful of target examples. In experiments done on six public datasets, we show that our approach outperforms the state of the art in terms of the RMS fitting error as well as the percentage of test examples for which AAM fitting converges to the ground truth.
Azin Asgarian, Ahmed Ashraf 0001, David J. Fleet, Babak Taati
IJCB4
2017 Detecting unseen falls from wearable devices using channel-wise ensemble of autoencoders
Shehroz S. Khan, Babak Taati
Expert Syst. Appl.2
2017 Mixture-Model Clustering of Pathological Gait Patterns
abstract
This study applies mixture-model clustering to spatiotemporal gait parameters in order to characterize the pathological gait pattern and to generate a composite measure indicative of overall gait performance. Gait data from 68 adults with stroke (age: 61.5 ± 13.6 years) and 20 healthy adults (age: 28.8 ± 7.1 years) were used in this study. Participants performed three passes across a GAITRite mat at different time points following stroke (poststroke adults only). Mixture-model clustering grouped participants' gait patterns based on their spatiotemporal gait features including symmetry, speed, and variability. Mixture-models with different covariance matrix parameterizations and numbers of clusters were examined. The selected clustering model successfully categorized participants' spatiotemporal gait data into three clinically meaningful groups. Based on the clustering results, gait speed, and variability measures varied across the three groups. Individuals in Group 1 are all symmetric and had the fastest and lowest gait velocity and variability, respectively. As expected, healthy participants were assigned to Group 1. All gait parameters were at an intermediate level in Group 2 and worse condition in Group 3. Moreover, resulting cluster centers were in line with previously published clinical studies on gait. In addition to clustering, each individual was given an indexed membership (ranged 0-1) to each of three groups. These indexed memberships were proposed as a single measure to encompass information about multiple gait parameters (symmetry, speed, and variability) and as a measure that is sensitive and responsive to improvement or deterioration and rehabilitation over time.
Elham Dolatabadi, Avril Mansfield, Kara K. Patterson, Babak Taati, Alex Mihailidis
IEEE J. Biomed. Health Informatics4
2017 Noncontact Vision-Based Cardiopulmonary Monitoring in Different Sleeping Positions
abstract
Individuals with obstructive sleep apnea (OSA) can experience partial or complete collapse of the upper airway during sleep. This condition affects between 10-17% of adult men and 3-9% of adult women, requiring arousal to resume regular breathing. Frequent arousals disrupt proper sleeping patterns and cause daytime sleepiness. Untreated OSA has been linked to serious medical issues including cardiovascular disease and diabetes. Unfortunately, diagnosis rates are low (∼20%) and current sleep monitoring options are expensive, time consuming, and uncomfortable. Toward the development of a convenient, noncontact OSA monitoring system, this paper presents a simple, computer vision-based method to monitor cardiopulmonary signals (respiratory and heart rates) during sleep. System testing was performed with 17 healthy participants in five different simulated sleep positions. To monitor cardiopulmonary rates, distinctive points are automatically detected and tracked in infrared image sequences. Blind source separation is applied to extract candidate signals of interest. The optimal respiratory and heart rates are determined using periodicity measures based on spectral analysis. Estimates were validated by comparison to polysomnography recordings. The system achieved a mean percentage error of 3.4% and 5.0% for respiratory rate and heart rate, respectively. This study represents an important step in building an accessible, unobtrusive solution for sleep apnea diagnosis.
Michael H. Li, Azadeh Yadollahi, Babak Taati
IEEE J. Biomed. Health Informatics3
2016 Automated Video Analysis of Handwashing Behavior as a Potential Marker of Cognitive Health in Older Adults
abstract
The identification of different stages of cognitive impairment can allow older adults to receive timely care and plan for the level of caregiving. People with existing diagnosis of cognitive impairment go through episodic phases of dementia requiring different levels of care at different times. Monitoring the cognitive status of existing patients is, thus, critical to deciding the level of care required by older adults. In this paper, we present a system to assess the cognitive status of older adults by monitoring a common activity of daily living, namely handwashing. Specifically, we extract features from handwashing trials of participants diagnosed with different levels of dementia ranging from cognitively intact to severe cognitive impairment, as assessed by the mini-mental state exam (MMSE). Based on videos of handwashing trials, we extract two classes of features: one characterizing the occupancy of different sink regions by the participant, and the other capturing the path tortuosity of the motion trajectory of participant's hands. We perform correlation analysis to assess univariate capacity of individual features to predict MMSE scores. To assess multivariate performance, we use machine learning methods to train models that predict the cognitive status (aware, mild, moderate, severe), as well as the MMSE scores. We present results demonstrating that features derived from hand washing behavior can be potential surrogate markers of a person's dementia, which can be instrumental in developing automated tools for continuously monitoring the cognitive status of older adults.
Ahmed Ashraf 0001, Babak Taati
IEEE J. Biomed. Health Informatics2
2014 Data Mining in Bone Marrow Transplant Records to Identify Patients With High Odds of Survival
abstract
Patients undergoing a bone marrow stem cell transplant (BMT) face various risk factors. Analyzing data from past transplants could enhance the understanding of the factors influencing success. Records up to 120 measurements per transplant procedure from 1751 patients undergoing BMT were collected (Shariati Hospital). Collaborative filtering techniques allowed the processing of highly sparse records with 22.3% missing values. Ten-fold cross-validation was used to evaluate the performance of various classification algorithms trained on predicting the survival status. Modest accuracy levels were obtained in predicting the survival status (AUC = 0.69). More importantly, however, operations that had the highest chances of success were shown to be identifiable with high accuracy, e.g., 92% or 97% when identifying 74 or 31 recipients, respectively. Identifying the patients with the highest chances of survival has direct application in the prioritization of resources and in donor matching. For patients where high-confidence prediction is not achieved, assigning a probability to their survival odds has potential applications in probabilistic decision support systems and in combination with other sources of information.
Babak Taati, Jasper Snoek, Dionne M. Aleman, Ardeshir Ghavamzadeh
IEEE J. Biomed. Health Informatics1
2013 Video analysis for identifying human operation difficulties and faucet usability assessment
Babak Taati, Jasper Snoek, Alex Mihailidis
Neurocomputing1
2011 Local shape descriptor selection for object recognition in range data
Babak Taati, Michael A. Greenspan
Comput. Vis. Image Underst.1
2007 Variable Dimensional Local Shape Descriptors for Object Recognition in Range Data
abstract
We propose a new set of highly descriptive local shape descriptors (LSDs) for model-based object recognition and pose determination in input range data. Object recognition is performed in three phases: point matching, where point correspondences are established between range data and the complete model using local shape descriptors; pose recovery, where a computationally robust algorithm generates a rough alignment between the model and its instance in the scene, if such an instance is present; and pose refinement. While previously developed LSDs take a minimalist approach, in that they try to construct low dimensional and compact descriptors, we use high (up to 9) dimensional descriptors as the key to more accurate and robust point correspondence. Our strategy significantly simplifies the computational burden of the pose recovery phase by investing more time in the point matching phase. Experiments with Lidar and dense stereo range data illustrate the effectiveness of the approach by providing a higher percentage of correct matches in the candidate point matches list than a leading minimalist technique. Consequently, the number of RANSAC iterations required for recognition and pose determination is drastically smaller in our approach.
Babak Taati, Michel Bondy, Piotr Jasiobedzki, Michael A. Greenspan
ICCV1