Ifeoma Nwogu

dblp:98/3822 · DBLP profile ↗
← Back
45ranked-venue papers
10as first author
21since 2021 · last 2026
0000-0003-1414-6433ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 32 · 7 first-author · 18 since 2021Artificial intelligence and machine learning · 31 · 6 first-author · 17 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-authorSecurity and privacy · 2Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 A Large-Scale 3D Representation Dataset and Benchmark for Continuous Sign Language Understanding
Lipisha Chaudhary, Enjamamul Hoq, Lu Dong 0004, Henry Adler, Ifeoma Nwogu
FG5
2025 Exploring the Differences between Deaf and Hearing Infant Cries
abstract
In this study, we propose an interpretable AI approach for analyzing infant cry signals, focusing on distinguishing between deaf and hearing infants. Using the Baby Chillanto dataset, we first conduct a human test to determine how well humans can detect the differences between hearing and deaf infants. We then explore the use of typical features used in audio signal analysis - chromagram, Log-Mel spectrogram and Mel frequency cepstral coefficients (MFCC), to determine whether a deep learning model can perform as well or better than humans. To explain the model’s predictions, we apply SHapley Additive exPlanations(SHAP) which reveal the most significant portion of each feature set that contributes to the model’s decisions. This interpretable approach provides valuable insights into how acoustic characteristics differ between deaf and hearing infants, potentially aiding early detection and understanding of hearing impairments through non-invasive cry signal analysis. While the combined human test results were near random, our deep learning approach demonstrated both high classification performance and enhanced explainability.
Enjamamul Hoq, Ifeoma Nwogu
ICASSP2
2025 FUSE-MOS: Fusion of Speech Embeddings for MOS Prediction with Uncertainty Quantification
Enjamamul Hoq, Danielle Omondi, Ifeoma Nwogu
INTERSPEECH4
2025 AutoMisty: A Multi-Agent LLM Framework for Automated Code Generation in the Misty Social Robot
abstract
The social robot’s open API allows users to customize open-domain interactions. However, it remains inaccessible to those without programming experience. We introduce AutoMisty, the first LLM-powered multi-agent framework that converts natural-language commands into executable Misty robot code by decomposing high-level instructions, generating sub-task code, and integrating everything into a deployable program. Each agent employs a two-layer optimization mechanism: first, a self-reflective loop that instantly validates and automatically executes the generated code, regenerating whenever errors emerge; second, human review for refinement and final approval, ensuring alignment with user preferences and preventing error propagation. To evaluate AutoMisty’s effectiveness, we designed a benchmark task set spanning four levels of complexity and conducted experiments in a real Misty robot environment. Extensive evaluations demonstrate that AutoMisty not only consistently generates high-quality code but also enables precise code control, significantly outperforming direct reasoning with ChatGPT-4o and ChatGPT-o1. All code, optimized APIs, and experimental videos will be publicly released through the webpage: AutoMisty.
Lu Dong 0004, Sahana Rangasrinivasan, Ifeoma Nwogu, Srirangaraj Setlur, Venu Govindaraju
IROS4
2025 MCAD: Multimodal Context-Aware Audio Description Generation for Soccer
abstract
Audio Descriptions (AD) are essential for making visual content accessible to individuals with visual impairments. Recent works have shown a promising step towards automating AD, but they have been limited to describing high-quality movie content using human-annotated ground truth AD in the process. In this work, we present an end-to-end pipeline, MCAD, that extends AD generation beyond movies to the domain of sports, with a focus on soccer games, without relying on ground truth AD. To address the absence of domain-specific AD datasets, we fine-tune a Video Large Language Model on publicly available movie AD datasets so that it learns the narrative structure and conventions of AD. During inference, MCAD incorporates multimodal contextual cues such as player identities, soccer events/actions, and commentary from the game. These cues, combined with input prompts to the fine-tuned VideoLLM, allow the system to produce complete AD text for each video segment. We further introduce a new evaluation metric,$A R G E-A D$, designed to accurately assess the quality of generated AD. ARGE-AD evaluates the generated AD for the presence of five characteristics: (i) usage of people's names, (ii) mention of actions/events, (iii) appropriate length of AD, (iv) absence of pronouns, and ($v$) overlap from commentary/subtitles. We present an in-depth analysis of our approach on both movie and soccer datasets. We also validate the use of this metric to quantitatively comment on the quality of generated AD using our metric across domains. Additionally, we contribute audio descriptions for 100 soccer game clips annotated by two AD experts.
Lipisha Chaudhary, Trisha Mittal, Subhadra Gopalakrishnan, Ifeoma Nwogu, Jaclyn Pytlarz
ISM4
2024 Towards Open Domain Text-Driven Synthesis of Multi-person Motions
Mengyi Shan, Lu Dong 0004, Yutao Han, Yuan Yao 0001, Ifeoma Nwogu, Guo-Jun Qi, Mitch Hill
ECCV (65)6
2024 SignAvatar: Sign Language 3D Motion Reconstruction and Generation
abstract
Achieving expressive 3D motion reconstruction and automatic generation for isolated sign words can be challenging, due to the lack of real-world 3D sign-word data, the complex nuances of signing motions, and the cross-modal understanding of sign language semantics. To address these challenges, we introduce SignAvatar, a framework capable of both word-level sign language reconstruction and generation. SignAvatar employs a transformer-based conditional variational autoencoder architecture, effectively establishing relationships across different semantic modalities. Additionally, this approach incorporates a curriculum learning strategy to enhance the model's robustness and generalization, resulting in more realistic motions. Furthermore, we contribute the ASL3DWord dataset, composed of 3D joint rotation data for the body, hands, and face, for unique sign words. We demonstrate the effectiveness of SignAvatar through extensive experiments, showcasing its superior reconstruction and automatic generation capabilities. The code and dataset are available on the project page1.
Lu Dong 0004, Lipisha Chaudhary, Mason Lary, Ifeoma Nwogu
FG6
2024 Dataset Infant Anonymization with Pose and Emotion Retention
abstract
We demonstrate a procedure for the anonymization of infant subjects in videos such that salient behavioral information is retained. This method also creates a new identity that is consistent temporally across video frames. We present an overview of this anonymization process, which involves moving through the latent space of a generative model with an infant specific latent space traversal technique. We apply the technique on videos of infants, a historically difficult source of data, and make comparisons to other state-of-the-art anonymization systems. Metrics demonstrate an improved ability to retain emotional content of videos during the anonymization process, even during extreme emotions or poses, while maintaining a consistent identity throughout.
Mason Lary, Matthew Klawonn, Daniel S. Messinger, Ifeoma Nwogu
FG4
2024 A Comparative Study of Video-Based Human Representations for American Sign Language Alphabet Generation
abstract
Sign language is a complex visual language, and automatic interpretations of sign language can facilitate communication involving deaf individuals. As one of the essential components of sign language, fingerspelling connects the natural spoken languages to the sign language and expands the scale of sign language vocabulary. In practice, it is challenging to analyze fingerspelling alphabets due to their signing speed and small motion range. The usage of synthetic data has the potential of further improving fingerspelling alphabets analysis at scale. In this paper, we evaluate how different video-based human representations perform in a framework for Alphabet Generation for American Sign Language (ASL). We tested three mainstream video-based human representations: two-stream inflated 3D ConvNet, 3D landmarks of body joints, and rotation matrices of body joints. We also evaluated the effect of different skeleton graphs and selected body joints. The generation process of ASL fingerspelling used a transformer-based Conditional Variational Autoencoder. To train the model, we collected ASL alphabet signing videos from 17 signers with dynamic alphabet signing. The generated alphabets were evaluated using automatic metrics of quality such as FID, and we also considered supervised metrics by recognizing the generated entries using Spatio-Temporal Graph Convolutional Networks. Our experiments show that using the rotation matrices of the upper body joints and the signing hand give the best results for the generation of ASL alphabet signing. Going forward, our goal is to produce articulated fingerspelling words by combining individual alphabets learned in this work.
Lipisha Chaudhary, Lu Dong 0004, Srirangaraj Setlur, Venu Govindaraju, Ifeoma Nwogu
FG6
2024 Cross-Attention Based Influence Model for Manual and Nonmanual Sign Language Analysis
Lipisha Chaudhary, Ifeoma Nwogu
ICPR (21)3
2023 SignNet: Single Channel Sign Generation using Metric Embedded Learning
abstract
A true interpreting agent not only understands sign language and translates to text, but also understands text and translates to signs. Much of the AI work in sign language translation to date has focused mainly on translating from signs to text. Towards the latter goal, we propose a text-to-sign translation model, SignNet, which exploits the notion of similarity (and dissimilarity) of visual signs in translating. This module presented is only one part of a dual-learning two task process involving text-to-sign (T2S) as well as sign-to-text (S2T). We currently implement SignNet as a single channel architecture so that the output of the T2S task can be fed into S2T in a continuous dual learning framework. By single channel, we refer to a single modality, the body pose joints. In this work, we present SignNet, a T2S task using a novel metric embedding learning process, to preserve the distances between sign embeddings relative to their dissimilarity. We also describe how to choose positive and negative examples of signs for similarity testing. From our analysis, we observe that metric embedding learning-based model perform significantly better than the other models with traditional losses, when evaluated using BLEU scores. In the task of gloss to pose, SignNet performed as well as its state-of-the-art (SoTA) counterparts and outperformed them in the task of text to pose, by showing noteworthy enhancements in BLEU 1 - BLEU 4 scores (BLEU 1: 31 → 39; ≈26% improvement and BLEU 4: 10.43 →11.84; ≈14% improvement) when tested on the popular RWTH PHOENIX-Weather-2014T benchmark dataset
Tejaswini Ananthanarayana, Lipisha Chaudhary, Ifeoma Nwogu
FG3
2023 Analyzing Interactions in Paired Egocentric Videos
abstract
As wearable devices become more popular, ego-centric information recorded with these devices can be used to better understand the behaviors of the wearer and other people the wearer is interacting with. Data such as the voice, head movement, galvanic skin responses (GSR) to measure arousal levels, etc., obtained from such devices can provide a window into the underlying affect of both the wearer and his/her conversant. In this study, we examine the characteristics of two types of dyadic conversations. In one case, the interlocutors discuss a topic on which they agree, while the other situation involves interlocutors discussing a topic on which they disagree, even if they are friends. The range of topics is mostly politically motivated. The egocentric information is collected using a pair of wearable smart glasses for video data and a smart wristband for physiological data, including GSR. Using this data, various features are extracted including the facial expressions of the conversant and the 3D motion from the wearer's camera within the environment - this motion is termed as egomotion. The goal of this work is to investigate whether the nature of a discussion could be better determined either by evaluating the behavior of an individual in the conversation or by evaluating the pairing/coupling of the behaviors of the two people in the conversation. The pairing is accomplished using a modified formulation of the dynamic time warping (DTW) algorithm. A random forest classifier is implemented to evaluate the nature of the interaction (agreement versus disagreement) using individualistic and paired features separately. The study found that in the presence of the limited data used in this work, individual behaviors were slightly more indicative of the type of discussion (85.43% accuracy) than the paired behaviors (83.33% accuracy).
Ajeeta Khatri, Zachary Butler, Ifeoma Nwogu
FG3
2023 AI-Driven Sign Language Interpretation for Nigerian Children at Home
abstract
As many as three million school age children between the ages of 5 and 14 years, live with severe to profound hearing loss in Nigeria. Many of these Deaf or Hard of Hearing (DHH) children developed their hearing loss later in life, non-congenitally, hence their parents are hearing. While their teachers in the Deaf schools they attend can often communicate effectively with them in "dialects" of American Sign Language (ASL), the unofficial sign lingua franca in Nigeria, communication at home with other family members is challenging and sometimes non-existent. This results in adverse social consequences including stigmatization, for the students. With the recent successes of AI in natural language understanding, the goal of automated sign language understanding is becoming more realistic using neural deep learning technologies. To this effect, the proposed project aims at co-designing and developing an ongoing AI-driven two-way sign language interpretation tool that can be deployed in homes, to improve language accessibility and communication between the DHH students and other family members. This ensures inclusive and equitable social interactions and can promote lifelong learning opportunities for them outside of the school environment.
Ifeoma Nwogu, Roshan Lalintha Peiris, Karthik Dantu, Ruchi Gamta, Emma Asonye
IJCAI1
2023 Language-guided Human Motion Synthesis with Atomic Actions
abstract
Language-guided human motion synthesis has been a challenging task due to the inherent complexity and diversity of human behaviors. Previous methods face limitations in generalization to novel actions, often resulting in unrealistic or incoherent motion sequences. In this paper, we propose ATOM (ATomic mOtion Modeling) to mitigate this problem, by decomposing actions into atomic actions, and employing a curriculum learning strategy to learn atomic action composition. First, we disentangle complex human motions into a set of atomic actions during learning, and then assemble novel actions using the learned atomic actions, which offers better adaptability to new actions. Moreover, we introduce a curriculum learning training strategy that leverages masked motion modeling with a gradual increase in the mask ratio, and thus facilitates atomic action assembly. This approach mitigates the overfitting problem commonly encountered in previous methods while enforcing the model to learn better motion representations. We demonstrate the effectiveness of ATOM through extensive experiments, including text-to-motion and action-to-motion synthesis tasks. We further illustrate its superiority in synthesizing plausible and coherent text-guided human motion sequences.
Yuanhao Zhai 0001, Mingzhen Huang, Tianyu Luan, Lu Dong 0004, Ifeoma Nwogu, Siwei Lyu, David S. Doermann, Junsong Yuan 0001
ACM Multimedia5
2023 SignNet II: A Transformer-Based Two-Way Sign Language Translation Model
abstract
The role of a sign interpreting agent is to bridge the communication gap between the hearing-only and Deaf or Hard of Hearing communities by translating both from sign language to text and from text to sign language. Until now, much of the AI work in automated sign language processing has focused primarily on sign language to text translation, which puts the advantage mainly on the side of hearing individuals. In this article, we describe advances in sign language processing based on transformer networks. Specifically, we introduce SignNet II, a sign language processing architecture, a promising step towards facilitating two-way sign language communication. It is comprised of sign-to-text and text-to-sign networks jointly trained using a dual learning mechanism. Furthermore, by exploiting the notion of sign similarity, a metric embedding learning process is introduced to enhance the text-to-sign translation performance. Using a bank of multi-feature transformers, we analyzed several input feature representations and discovered that keypoint-based pose features consistently performed well, irrespective of the quality of the input videos. We demonstrated that the two jointly trained networks outperformed their singly-trained counterparts, showing noteworthy enhancements in BLEU-1 - BLEU-4 scores when tested on the largest available German Sign Language (GSL) benchmark dataset.
Lipisha Chaudhary, Tejaswini Ananthanarayana, Enjamamul Hoq, Ifeoma Nwogu
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Towards Understanding the Behaviors of Pretrained Compressed Convolutional Models
abstract
We investigate the behaviors that compressed convolutional models exhibit for two key areas within AI trust: (i) the ability for a model to be explained and (ii) its ability to be robust to adversarial attacks. While compression is known to shrink model size and decrease inference time, other properties of compression are not as well studied. We employ several compression methods on benchmark datasets, including ImageNet, to study how compression affects the convolutional aspects of an image model. We investigate explainability by studying how well compressed convolutional models can extract visual features with t-SNE, as well as visualizing localization ability of our models with class activation maps. We show that even with significantly compressed models, vital explainability is preserved and even enhanced. We find with applying the Carlini & Wagner attack algorithm on our compressed models, robustness is maintained and some forms of compression make attack more difficult or time-consuming.
Timothy Zee, Manohar Lakshmana, Ifeoma Nwogu
ICPR3
2022 Regression with Uncertainty Quantification in Large Scale Complex Data
abstract
While several methods for predicting uncertainty on deep networks have been recently proposed, they do not always readily translate to large and complex datasets without significant overhead. In this paper we utilize a special instance of the Mixture Density Networks (MDNs) to produce an elegant and compact approach to quantity uncertainty in regression problems. When applied to standard regression benchmark datasets, we show an improvement in predictive log-likelihood and root-mean-square-error when compared to existing state-of-the-art methods. We demonstrate the efficacy and practical usefulness of the method for (i) predicting future stock prices from stochastic, highly volatile time-series data; (ii) anomaly detection in real-life highly complex video segments; and (iii) the task of age estimation and data cleansing on the challenging IMDb-Wiki dataset of half a million face images.
Nicholas Wilkins, Ifeoma Nwogu
SMC3
2021 Dynamic Cross-Feature Fusion for American Sign Language Translation
abstract
While a significant amount of work has been done on the commonly used, tightly -constrained weather-based, German sign language (GSL) dataset, little has been done for continuous sign language translation (SLT) in more realistic settings, including American sign language (ASL) translation. Also, while CNN - based features have been consistently shown to work well on the GSL dataset, it is not clear whether such features will work as well in more realistic settings when there are more heterogeneous signers in non-uniform backgrounds. To this end, in this work, we introduce a new, realistic phrase-level ASL dataset (ASLing), and explore the role of different types of visual features (CNN embeddings, human body keypoints, and optical flow vectors) in translating it to spoken American English. We propose a novel Transformer-based, visual feature learning method for ASL translation. We demonstrate the explainability efficacy of our proposed learning methods by visualizing activation weights under various input conditions and discover that the body keypoints are consistently the most reliable set of input features. Using our model, we successfully transfer-learn from the larger GSL dataset to ASLing, resulting in significant BLEU score improvements. In summary, this work goes a long way in bringing together the AI resources required for automated ASL translation in unconstrained environments.
Tejaswini Ananthanarayana, Nikunj R. Kotecha, Priyanshu Srivastava, Lipisha Chaudhary, Nicholas Wilkins, Ifeoma Nwogu
FG6
2021 Coupled Systems for Modeling Rapport Between Interlocutors
abstract
This research work explores different machine learning techniques for recognizing the existence of rapport between two people engaged in a conversation, based on their facial expressions. First using artificially generated pairs of correlated data signals, a coupled gated recurrent unit (cGRU) neural network is developed to measure the extent of similarity between the temporal evolution of pairs of time-series signals. By pre-selecting their covariance values (between 0.1 and 1.0), pairs of coupled sequences are generated. Using the developed cGRU architecture, this covariance between the signals is successfully recovered. Using this and various other coupled architectures, tests for rapport (measured by the extent of mirroring and mimicking of behaviors) are conducted on real-life datasets. On fifty-nine (N = 59) pairs of interactants in an interview setting, a transformer based coupled architecture performs the best in determining the existence of rapport. To test for generalization, the models were applied on never-been-seen data collected 14 years prior, also to predict the existence of rapport. The coupled transformer model again performed the best for this transfer learning task, determining which pairs of interactants had rapport and which did not. The experiments and results demonstrate the advantages of coupled architectures for predicting an interactional process such as rapport, even in the presence of limited data.
Srijan Sharma, Kantha Girish Gangadhara, Anne Solbu Slowe, Mark G. Frank, Ifeoma Nwogu
FG6
2021 Towards the Synthesis of Parent-Infant Facial Interactions
abstract
This work is motivated by the need to automate the analysis of parent-infant interactions to better understand the existence of any potential behavioral patterns useful for the early diagnosis of autism spectrum disorder (ASD). It presents an approach for synthesizing the facial expression exchanges that occur during parent-infant interactions. This is accomplished by developing a novel approach that uses landmarks when synthesizing changing facial expressions. The proposed model consists of two components: (i) The first is a landmark converter that receives a set of facial landmarks and the target emotion as input and outputs a set of new landmarks transformed to match the emotion. (ii) The second component involves an image converter that takes in an input image, a target landmark and a target emotion and outputs a face transformed to match the input emotion. The inclusion of landmarks in the generation process proves useful in the generation of baby facial expressions; babies have somewhat different facial musculature and facial dynamics than adults. This paper presents a realistic-looking matrix of changing facial expressions sampled from a 2-D emotion continuum (valence and arousal) and displays successfully transferred facial expressions from real-life mother-infant dyads to novel ones.
Renke Wang, Amy Yeo-jin Ahn, Daniel S. Messinger, Ifeoma Nwogu
FG4
2021 Effects of Feature Scaling and Fusion on Sign Language Translation
Tejaswini Ananthanarayana, Lipisha Chaudhary, Ifeoma Nwogu
Interspeech3
2020 A Computational View of the Emotional Regulation of Disgust using Multimodal Sensors
abstract
Emotion regulation can be characterized by different activities that attempt to alter an emotional response, whether behavioral, physiological or neurological. The two most widely adopted strategies, cognitive reappraisal and expressive suppression are explored in this study, specifically in the context of disgust. Study participants (N = 21) experienced disgust via video exposure, and were instructed to either regulate their emotions or express them freely. If regulating, they were required to either cognitively reappraise or suppress their emotional experiences while viewing the videos. Video recordings of the participants' faces were taken during the experiment and electrocardiogram (ECG), electromyography (EMG), and galvanic skin response (GSR) readings were also collected for further analysis. We compared the participants behavioral (facial musculature movements) and physiological (GSR and heart rate) responses as they aimed to alter their emotional responses and computationally determined that when responding to disgust stimuli, the signals recorded during suppression and free expression were very similar, whereas those recorded during cognitive reappraisal were significantly different,. Thus, in the context of this study, from a signal analysis perspective, we conclude that emotion regulation via cognitive reappraisal significantly alters participants' physiological responses to disgust, unlike regulation via suppression.
Vidhathri Kota, Geeta Gali, Ifeoma Nwogu
FG3
2020 Interpretable Emotion Classification Using Temporal Convolutional Models
abstract
As with many problems solved by deep neural networks, existing solutions rarely explain, precisely, the important factors responsible for the predictions made by the model. This work looks to investigate how different spatial regions and landmark points change in position over time, to better explain the underlying factors responsible for various facial emotion expressions. By pinpointing the specific regions or points responsible for the classification of a particular facial expression, we gain better insight into the dynamics of the face when displaying that emotion. To accomplish this, we examine two spatiotemporal representations of moving faces, while expressing different emotions. The representations are then presented to a convolutional neural network for emotion classification. Class activation maps are used in highlighting the regions of interest and the results are qualitatively compared with the well known facial action units, using the facial action coding system. The model was originally trained and tested on the CK+ dataset for emotion classification, and then generalized to the SAMM dataset. In so doing, we successfully present an interpretable technique for understanding the dynamics that occur during convolutional-based prediction tasks on sequences of face data.
Manasi Gund, Abhiram Ravi Bharadwaj, Ifeoma Nwogu
ICPR3
2020 Modeling Global Body Configurations in American Sign Language
abstract
American Sign Language (ASL) is the fourth most commonly used language in the United States and is the language most commonly used by Deaf people in the United States and the English-speaking regions of Canada. Unfortunately, until recently, ASL received little research. This is due, in part, to its delayed recognition as a language until William C. Stokoe's publication in 1960. Limited data has been a long-standing obstacle to ASL research and computational modeling. The lack of large-scale datasets has prohibited many modern machine-learning techniques, such as Neural Machine Translation, from being applied to ASL. In addition, the modality required to capture sign language (i.e. video) is complex in natural settings (as one must deal with background noise, motion blur, and the curse of dimensionality). Finally, when compared with spoken languages, such as English, there has been limited research conducted into the linguistics of ASL. We realize a simplified version of Liddell and Johnson's Movement-Hold (MH) Model using a Probabilistic Graphical Model (PGM). We trained our model on ASLing, a dataset collected from three fluent ASL signers. We evaluate our PGM against other models to determine its ability to model ASL. Finally, we interpret various aspects of the PGM and draw conclusions about ASL phonetics. The main contributions of this paper are
Nicholas Wilkins, Max Cordes Galbraith, Ifeoma Nwogu
INTERSPEECH3
2020 Analyzing the Extent of Rapport in Groups of Triads Via Interactional Synchrony
abstract
Research in social psychology has extensively shown that in cohesive groups, individuals often mirror each other's prosody, facial expressions, and body movements. This mirroring effect can help determine the level of comfort or the extent of engagement and genuine interest between two or more interlocutors. In this work, using an annotated dataset consisting of videos of three-person conversations, we aim to analyze the extent of rapport in each of the triadic groups. We generate behavioral curves from features extracted from the participants' face and body movements. These are the sampled time series signals resulting from their multimodal features. Next, the extents of synchrony are analyzed by aligning the behavioral curves of pairs of participants. The alignment tests show that basic correlation coefficient measures outperform more advanced curve matching techniques when used to estimate the similarities between multidimensional behavior curves. They also show that in this dataset, synchrony is better observed from facial expressions than body movements. For this reason, using facial action units, we show that an end-to-end recursive neural network (RNN) trained using a regression loss yields good results in predicting the extent of synchrony in small groups.
Nicholas Wilkins, Ifeoma Nwogu
SMC2
2018 A Study on the Suppression of Amusement
abstract
In this work, we aim to gain better insights into the underlying behaviors of people when they attempt to suppress amusement, the positive emotion experienced, specifically from finding something funny. We aim to better understand the different physiological manifestations that occur when this suppression happens. We investigate this phenomenon by observing the presence/absence of action units (AUs) during amusement expression and suppression. We also record galvanic skin responses (GSR) to more closely observe if there are major differences in physiological manifestations during amusement expression versus suppression. This study was performed as a part of a larger one on deceit detection since amusement suppression can also be viewed as a form deceit, based on the context. We showed that the overall facial expression signatures were unsurprisingly very different for amusement expression and suppression; the features associated with positive emotions were clearly dominant during free expression and sadness was especially dominant during suppression. In observing the GSR readings, free expression and suppression manifested quite differently across individuals but similarly within the same individual, suggesting that the internal state or emotional arousal level of the participants were not altered significantly when trying to suppress amusement. Finally, in more than 75% of the cases, we found that arousal induced by amusement could not readily be eliminated, even when the individuals succeeded in masking it on their faces. This further validates the claim in the social psychology literature that suppression decreases expressive behavior, but does not decrease the emotional experience (measured via GSR).
Ifeoma Nwogu, Bryan Passino, Reynold J. Bailey
FG1
2018 Analyzing Skin Lesions in Dermoscopy Images Using Convolutional Neural Networks
abstract
In this paper, we discuss the problem of automatic skin lesion analysis, specifically melanoma detection and semantic segmentation. We accomplish this by using deep learning techniques to perform classification on publicly available dermoscopic images. Skin cancer, of which melanoma is a type, is the most prevalent form of cancer in the US and more than four million cases are diagnosed in the US every year. In this work, we present our efforts towards an accessible, deep learning-based system that can be used for skin lesion classification, thus leading to an improved melanoma screening system. For classification, a deep convolutional neural network architecture is first implemented over the raw images. In addition, hand-coded features such as 166-D color histogram distribution, edge histogram and Multiscale Color local binary patterns are extracted from the images and presented to a random forest classifier. The average of the outputs from the two mentioned classifiers is taken as the final classification result. The classification task achieves an accuracy of 80.3%, AUC score of 0.69 and a precision score of 0.81. For segmentation, we implement a convolutional-deconvolutional architecture and the segmentation model achieves a Dice coefficient of 73.5%.
Vatsala Singh, Ifeoma Nwogu
SMC2
2017 Cognitive-Biometric Recognition From Language Usage: A Feasibility Study
abstract
We propose a novel cognitive biometrics modality based on written language-usage of an individual. This is a feasibility study using the Internet-scale blogs, with tens of thousands of authors to create a cognitive fingerprint for an individual. Existing cognitive biometric modalities involve learning from obtrusive sensors placed on human body. Our modality is based on the characteristic pattern of how individuals express their thoughts through written language. The problems of cognitive authentication (1:1 comparison of genuine versus impostor) and identification (1:n search) are formulated. We detail the algorithms to learn a classifier to distinguish between genuine and impostor classes (for authentication) and multiple classes (for identification). We conclude that a cognitive fingerprint can be successfully learnt, using stylistic (writing style), semantic (themes), and syntactic (grammatical) features extracted from blogs. Our methodology shows promising results (with 79% as the area under the ROC (AUC) in case of authentication). For identification, the individual class accuracies are up to 90%. We performed stricter tests to see how our system performs for unseen user, and report the accuracies of 72% (genuine) and 71% (impostor). Such a study lays the groundwork for building alternative cognitive systems. The modality, presented here, is easy to obtain, unobtrusive and needs no additional hardware.
Neeti Pokhriyal, Kshitij Tayal, Ifeoma Nwogu, Venu Govindaraju
IEEE Trans. Inf. Forensics Secur.3
2016 Understanding Line Plots Using Bayesian Network
abstract
Information graphics, such as bar charts, graphs, plots etc. in scientific documents primarily facilitate better understanding of information. Graphics are a key component in technical documents as they are simplified representations of complex ideas. When the traditional optical character recognition (OCR) systems is used on digitized documents, we lose the ideas conveyed in these information graphics since OCRs typically work only on text. And although in more recent times, tools have been developed to extract information graphics from pdf files, they still do not intelligently interpret the contents of the extracted graphics. We therefore propose a method for identifying the intended messages of line plots using a Bayesian network. We accomplish this by first extracting a dense set of points in from a line plot and then represent the entire line plot as a sequence of trends. We then implement a Bayesian network for reasoning about the messages conveyed by the line plots and their trends. We validate our approach by performing experiments on a dataset obtained from computer science conference publications and evaluate the performance of the network against the messages generated by human end users. The resulting intended message gives holistic information about the line plot(s) as well as lower level information about the trends that make up the plot.
Rathin Radhakrishnan Nair, Nishant Sankaran, Ifeoma Nwogu, Venu Govindaraju
DAS3
2016 Segmentation of highly unstructured handwritten documents using a neural network technique
abstract
In recent years there has been a growing interest in digitizing the extensive amounts of books and documents that existed preceding the widespread adoption of digital technologies. Many of these digitizing initiatives deal with huge collections of handwritten documents, for which document image analysis techniques (page segmentation, keyword-spotting, optical character recognition (OCR), etc) are not yet as mature as for printed text. Thus, there is an imminent need to develop techniques to understand, archive, index and search the manuscripts. The antiquated approach of manually transcribing handwritten collections and then using standard text retrieval techniques can be very expensive for large collections. But many of the manuscripts in these collections, unlike machine-printed texts, contain unstructured information, cluttered group of texts and graphics that do not necessarily follow a pre-specified format, thus making it quite challenging to automatically process. Thus, in this paper we present a convolutional neural network (CNN) based implementation that is used to segment pages of handwritten documents into their constituent sections. We showcase a multiscale sliding window based network that is trained to predict the sections of the pages in handwritten manuscripts. The results of the network are post-processed with a novel region growing technique to further improve the segmentation results. The implementation is applied on the Marianne Moore archival collection, a body of handwritten notes and memos by the renowned author Marianne Moore (1887-1972), one of the foremost modernist poets of the early twentieth-century. We present our segmentation results both quantitatively and qualitatively.
Rathin Radhakrishnan Nair, Bhargava Urala Kota, Ifeoma Nwogu, Venu Govindaraju
ICPR3
2015 Automated analysis of line plots in documents
abstract
Information graphics, such as graphs and plots, are used in technical documents to convey information to humans and to facilitate greater understanding. Usually, graphics are a key component in a technical document, as they enable the author to convey complex ideas in a simplified visual format. However, in an automatic text recognition system, which are typically used to digitize documents, the ideas conveyed in a graphical format are lost. We contend that the message or extracted information can be used to help better understand the ideas conveyed in the document. In scientific papers, line plots are the most commonly used graphic to represent experimental results in the form of correlation present between values represented on the axes. The contribution of our work is in the series of image processing algorithms that are used to automatically extract relevant information, including text and plot from graphics found in technical documents. We validate the approach by performing the experiments on a dataset of line plots obtained from scientific documents from computer science conference papers and evaluate the variation of a reconstructed curve from the original curve. Our algorithm achieves a classification accuracy of 91% across the dataset and successfully extracts the axes from 92% of line plots. Axes label extraction and line curve tracing are performed successfully in about half the line plots as well.
Rathin Radhakrishnan Nair, Nishant Sankaran, Ifeoma Nwogu, Venu Govindaraju
ICDAR3
2014 Use of language as a cognitive biometric trait
abstract
This paper investigates whether the cognitive state of a person can be learnt and used as a novel biometric trait. We explore the idea of using language written by an author, as his/her cognitive fingerprint. The dataset consists of millions of blogs written by thousands of authors on the Internet. Our proposed method learns a classifier that can distinguish between genuine and impostor authors. Our results are encouraging (we report 72% Area under the ROC curve) and show that users do have a distinctive linguistic style, which is evident even when analyzing a corpora as large and diverse as the Internet. When we tested on new authors that the system had never encountered before, our methodology correctly identified genuine authors with 78% accuracy and impostors with 76% accuracy.
Neeti Pokhriyal, Ifeoma Nwogu, Venu Govindaraju
IJCB2
2014 Dimensionality Reduction with Subspace Structure Preservation
Devansh Arpit, Ifeoma Nwogu, Venu Govindaraju
NIPS2
2014 Shared features for multiple face-based biometrics
abstract
People often make instant judgments about the age, health, mood, personality and character of others based on their facial features. It is not clear from a cognitive aspect whether these different traits require different sets of features or a shared feature set. Till date, much of the computational face image analysis work such as face recognition, face-based deceit detection, age estimation, gender estimation, etc, have been developed on datasets and features specific only to the problem-at-hand. In this paper, we explore an approach for performing face image analysis using a shared set of features for different tasks. By performing unsupervised learning on a large collection of face images, we learn the parameters of a probabilistic generative face model, and by projecting a new face image into this probabilistic space, we obtain a set of face features not created for any specific face analysis tasks. We investigate the use of such shared features and successfully predict the level of attractiveness, whether or not a face is made-up, the facial expression, and the gender of a person, given any arbitrary, near-frontal face image.
Ifeoma Nwogu, Yingbo Zhou 0002
SMC1
2013 Language-motivated approaches to action recognition
Manavender R. Malgireddy, Ifeoma Nwogu, Venu Govindaraju
J. Mach. Learn. Res.2
2013 Labeling Spain With Stanford
abstract
We present an end-to-end framework for outdoor scene region decomposition, learned on a small set of randomly selected images that generalizes well to multiple data sets containing images from around the world. We discuss the different aspects of the framework especially a generalized variational inference method with better approximations to the true marginals of a graphical model. Experimentally, we explain why the framework is robust and performs competitively on many diverse scene data sets, including several unseen scene types. We have obtained high pixel-level accuracies (≈ 80%) in three of the four data sets, which include a benchmark data set known as the Stanford background data set. Our model obtained over 70% accuracy on the fourth data set, which contained a number of indoor and close-up images that are significantly different from our training examples.
Yingbo Zhou 0002, Ifeoma Nwogu, Venu Govindaraju
IEEE Trans. Image Process.2
2011 DISCO: Describing Images Using Scene Contexts and Objects
abstract
In this paper, we propose a bottom-up approach to generating short descriptive sentences from images, to enhance scene understanding. We demonstrate automatic methods for mapping the visual content in an image to natural spoken or written language. We also introduce a human-in-the-loop evaluation strategy that quantitatively captures the meaningfulness of the generated sentences. We recorded a correctness rate of 60.34% when human users were asked to judge the meaningfulness of the sentences generated from relatively challenging images. Also, our automatic methods compared well with the state-of-the-art techniques for the related computer vision tasks.
Ifeoma Nwogu, Yingbo Zhou 0002, Christopher Brown 0001
AAAI1
2011 Lie to Me: Deceit detection via online behavioral learning
abstract
Inspired by the the behavioral scientific discoveries of Dr. Paul Ekman in relation to deceit detection, along with the television drama series Lie to Me, also based on Dr. Ekman's work, we use machine learning techniques to study the underlying phenomena expressed when a person tells a lie. We build an automated framework which detects deceit by measuring the deviation from normal behavior, at a critical point in the course of an investigative interrogation. Behavioral psychologists have shown that the eyes (via either gaze aversion or gaze extension) can be good “reflectors” of the inner emotions, when a person tells a high-stake lie. Hence we develop our deceit detection framework around eye movement changes. A dynamic bayesian model of eye movements is trained during a normal course of conversation for each subject, to represent normal behavior. The remaining conversation is broken into sequences and each sequence is tested against the parameters of the model of normal behavior. At the critical points in the interrogations, the deviations from normalcy are observed and used to deduce verity/deceit. An analysis on 40 subjects gave an accuracy of 82.5% which strongly suggests that the latent parameters of eye movements successfully capture behavioral changes and could be viable for use in automated deceit detection.
Nisha Bhaskaran, Ifeoma Nwogu, Mark G. Frank, Venu Govindaraju
FG2
2011 A Shared Parameter Model for Gesture and Sub-gesture Analysis
Manavender R. Malgireddy, Ifeoma Nwogu, Subarna Ghosh, Venu Govindaraju
IWCIA2
2008 (BP)2: Beyond pairwise Belief Propagation labeling by approximating Kikuchi free energies
abstract
Belief propagation (BP) can be very useful and efficient for performing approximate inference on graphs. But when the graph is very highly connected with strong conflicting interactions, BP tends to fail to converge. Generalized Belief Propagation (GBP) provides more accurate solutions on such graphs, by approximating Kikuchi free energies, but the clusters required for the Kikuchi approximations are hard to generate. We propose a new algorithmic way of generating such clusters from a graph without exponentially increasing the size of the graph during triangulation. In order to perform the statistical region labeling, we introduce the use of superpixels for the nodes of the graph, as it is a more natural representation of an image than the pixel grid. This results in a smaller but much more highly interconnected graph where BP consistently fails. We demonstrate how our version of the GBP algorithm outperforms BP on synthetic and natural images and in both cases, GBP converges after only a few iterations.
Ifeoma Nwogu, Jason J. Corso
CVPR1
2008 Labeling Irregular Graphs with Belief Propagation
Ifeoma Nwogu, Jason J. Corso
IWCIA1
2008 Exploratory Identification of Image-Based Biomarkers for Solid Mass Pulmonary Tumors
Ifeoma Nwogu, Jason J. Corso
MICCAI (1)1
2007 PDE-Based Enhancement of Low Quality Documents
abstract
Partial Differential Equations are becoming one of the core tools for low-level image processing. They are especially functional in diffusion processes and variational models. In this paper, we exploit the regional smoothing that occurs in a nonlinear diffusion process and use this to enhance text in a degraded document image. The proposed smoothing method is robust when applied to either a highly corrupted text document or one with little degradation. The technique was tested on historical documents, carbon copies with highly varying grayscale backgrounds and on synthetic noisy documents. The PDE-based method far outperformed other industry-standard binarization techniques when compared quantitatively and qualitatively.
Ifeoma Nwogu, Zhixin Shi, Venu Govindaraju
ICDAR1
2007 Fast Temporal Tracking and 3D Reconstruction of a Single Coronary Vessel
abstract
Vessel extraction and vessel motion estimation from X-ray angiograms has been a challenging computer vision problem for several years. We have developed a fast and accurate method for extracting and tracking intravascular imaging data from X-ray angiograms. We accomplish this by reconstructing a moving 3D vessel, which contains more information than the static 2D snapshot image. Our approach involves identifying the vessel-of-interest in two biplane images, abstracting them into centerlines, and tracking them in ensuing images using deformable templates and graph techniques for optimization. When tested on fifteen patient datasets, the computational time was approximately 5 seconds per vessel per frame for vessels of length 80-100 mm.
Ifeoma Nwogu, Liana M. Lorigo
ICIP (5)1
2005 Word Separation of Unconstrained Handwritten Text Lines in PCR Forms
abstract
An approach for segmenting handwritten text in a pre-hospital care report (PCR) is presented. Segmentation of lines and words in a PCR is extremely challenging due to the nature of the environment in which the reports are created, giving rise to low quality, poorly written, loosely constrained data. Stroke analyses are performed and image primitives are extracted for word detection. A heuristics-based approach, involving gap spacing, height transitions, and the average stroke width of the writer is used in detecting word boundaries. Carbon copies of live PCRs are used for testing. Experiments show perfect segmentation of 69%, outperforming the more tested and proven algorithms by as much as 15%.
Ifeoma Nwogu, Gyeonghwan Kim
ICDAR1