VLDB 2026 Research / reviewers in the wild / expert
Nicholas B. Allen
dblp:146/4856
· DBLP profile ↗
11ranked-venue papers
0as first author
5since 2021 · last 2024
0000-0002-1086-6639ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5Human-computer interaction and ubiquitous computing · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | SMURF: Statistical Modality Uniqueness and Redundancy FactorizationabstractMultimodal late fusion is a well-performing fusion method that sums the outputs of separately processed modalities, so-called modality contributions, to create a prediction; for example, summing contributions from vision, acoustic, and language to predict affective states. In this paper, our primary goal is to improve the interpretability of what modalities contribute to the prediction in late fusion models. More specifically, we want to factorize modality contributions into what is consistently shared by at least two modalities (pairwise redundant contributions) and what the remaining modality-specific contributions are (unique contributions). Our secondary goal is to improve robustness to missing modalities by encouraging the model to learn redundant contributions. To achieve our two goals, we propose SMURF (Statistical Modality Uniqueness and Redundancy Factorization), a late fusion method that factorizes its outputs into a) unique contributions that are uncorrelated with all other modalities and b) pairwise redundant contributions that are maximally correlated between two modalities. For our primary goal, we 1) verify SMURF's factorization on a synthetic dataset, 2) ensure that its factorization does not degrade predictive performance on eight affective datasets, and 3) observe significant relationships between its factorization and human judgments on three datasets. For our secondary goal, we demonstrate that SMURF leads to more robustness to missing modalities at test time compared to three late fusion baselines. Torsten Wörtwein, Nicholas B. Allen, Jeffrey F. Cohn, Louis-Philippe Morency |
ICMI | 2 |
| 2023 | Quantifying & Modeling Multimodal Interactions: An Information Decomposition FrameworkabstractThe recent explosion of interest in multimodal applications has resulted in a wide selection of datasets and methods for representing and integrating information from different modalities. Despite these empirical advances, there remain fundamental research questions: How can we quantify the interactions that are necessary to solve a multimodal task? Subsequently, what are the most suitable multimodal models to capture these interactions? To answer these questions, we propose an information-theoretic approach to quantify the degree of redundancy, uniqueness, and synergy relating input modalities with an output task. We term these three measures as the PID statistics of a multimodal distribution (or PID for short), and introduce two new estimators for these PID statistics that scale to high-dimensional distributions. To validate PID estimation, we conduct extensive experiments on both synthetic datasets where the PID is known and on large-scale multimodal benchmarks where PID estimations are compared with human annotations. Finally, we demonstrate their usefulness in (1) quantifying interactions within multimodal datasets, (2) quantifying interactions captured by multimodal models, (3) principled approaches for model selection, and (4) three real-world case studies engaging with domain experts in pathology, mood prediction, and robotic perception where our framework helps to recommend strong multimodal models for each application. Paul Pu Liang, Chun Kai Ling, Suzanne Nie, Richard J. Chen, Nicholas B. Allen, Randy Auerbach, Faisal Mahmood 0001, Ruslan Salakhutdinov, Louis-Philippe Morency |
NeurIPS | 8 |
| 2022 | Language Use in Mother-Adolescent Dyadic Interaction: Preliminary ResultsabstractThis preliminary study applied a computer-assisted quantitative linguistic analysis to examine the effectiveness of language-based classification models to discriminate between mothers (n = 140) with and without history of treatment for depression (51% and 49%, respectively). Mothers were recorded during a problem-solving interaction with their adolescent child. Transcripts were manually annotated and analyzed using a dictionary-based, natural-language program approach (Linguistic Inquiry and Word Count). To assess the importance of linguistic features to correctly classify history of depression, we used Support Vector Machines (SVM) with interpretable features. Using linguistic features identified in the empirical literature, an initial SVM achieved nearly 63% accuracy. A second SVM using only the top 5 highest ranked SHAP features improved accuracy to 67.15%. The findings extend the existing literature base on understanding language behavior of depressed mood states, with a focus on the linguistic style of mothers with and without a history of treatment for depression and its potential impact on child development and trans-generational transmission of depression. Laura A. Cariola, Saurabh Hinduja, Maneesh Bilalpur, Lisa Sheeber, Nicholas B. Allen, Louis-Philippe Morency, Jeffrey F. Cohn |
ACII | 5 |
| 2021 | Learning Language and Multimodal Privacy-Preserving Markers of Mood from Mobile DataabstractPaul Pu Liang, Terrance Liu, Anna Cai, Michal Muszynski, Ryo Ishii, Nick Allen, Randy Auerbach, David Brent, Ruslan Salakhutdinov, Louis-Philippe Morency. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Paul Pu Liang, Terrance Liu, Anna Cai, Michal Muszynski, Ryo Ishii, Nicholas B. Allen, Randy Auerbach, David Brent, Ruslan Salakhutdinov, Louis-Philippe Morency |
ACL/IJCNLP (1) | 6 |
| 2021 | Human-Guided Modality Informativeness for Affective StatesabstractThis paper studies the hypothesis that not all modalities are always needed to predict affective states. We explore this hypothesis in the context of recognizing three affective states that have shown a relation to a future onset of depression: positive, aggressive, and dysphoric. In particular, we investigate three important modalities for face-to-face conversations: vision, language, and acoustic modality. We first perform a human study to better understand which subset of modalities people find informative, when recognizing three affective states. As a second contribution, we explore how these human annotations can guide automatic affect recognition systems to be more interpretable while not degrading their predictive performance. Our studies show that humans can reliably annotate modality informativeness. Further, we observe that guided models significantly improve interpretability, i.e., they attend to modalities similarly to how humans rate the modality informativeness, while at the same time showing a slight increase in predictive performance. Torsten Wörtwein, Lisa Sheeber, Nicholas B. Allen, Jeffrey F. Cohn, Louis-Philippe Morency |
ICMI | 3 |
| 2015 | Detection of depression in adolescents based on statistical modeling of emotional influences in parent-adolescent conversationsabstractThe current benchmark speech-based depression detection techniques rely on acoustic speech parameters collected from large sets of representative speech recordings. This study for the first time investigates depression detection based on the higher order influence model (HOIM) coefficients and emotional transition parameters derived from a relatively small set of conversational speech recordings representing 63 different parent-adolescent conversations of time duration 20 minutes each. The adolescents included 29 (24 female and 5 male) individuals diagnosed with major depressive disorder and 34 (24 female and 8 male) healthy individuals. The mental state of parents was not assessed. The model-based depression diagnosis was compared with benchmark techniques based on acoustic speech parameters (mel frequency cepstral coefficients (MFCC) and Teager energy operator (TEO)). The classification into depressed on non-depressed categories was performed using the Gaussian Mixture Model (GMM) for the acoustic parameters and the support vector machine (SVM) for the HOIM features. The model based technique led to the highest average classification accuracy of 94% of for the HOIM of order 4, whereas the best benchmark techniques scored 70% for the optimized MFCCs and 71% for the optimized TEO features. Melissa N. Stolar, Margaret Lech, Nicholas B. Allen |
ICASSP | 3 |
| 2013 | Introducing Emotions to the Modelingof Intra- and Inter-Personal Influencesin Parent-Adolescent ConversationsabstractAn understanding of the dynamics underlying emotional interactions between speakers is essential to the design of effective conversational strategies for interviews, mental health therapies, teaching and counseling, as well as the design of naturalistic human-machine communication systems. The present study introduces a new approach to the modeling of emotional influences during parent-adolescent conversations. The proposed dynamic influence model (DIM) estimates the joint conditional probabilities of speaker's states as a linear combination of simpler inter- and intra-speaker conditional probabilities. Contrary to the previously existing influence models (IMs), the DIM's coefficients are given not as static, constant values but as dynamically changing functions of the time delay between the current and the previous state. The speaker's states were annotated using four labels (speech with positive emotion, speech with negative emotion, emotionally neutral speech and silence with undefined emotion). Experimental results based on the audio recordings of 63 different naturalistic (not acted) parent-adolescent conversations showed that the proposed method leads to psychologically plausible observations. It was also demonstrated that the proposed DIM can achieve up to 20 percent higher accuracy of discriminating between emotional influence patterns of parents and adolescents when compared to the previously used static IM. Melissa N. Stolar, Margaret Lech, Lisa Sheeber, Ian S. Burnett, Nicholas B. Allen |
IEEE Trans. Affect. Comput. | 5 |
| 2012 | Early prediction of major depression in adolescents using glottal wave characteristics and Teager Energy parametersabstractPrevious studies of an automated detection of Major Depression in adolescents based on acoustic speech analysis identified the glottal and the Teager Energy features as the strongest correlates of depression. This study investigates the effectiveness of these features in an early prediction of Major Depression in adolescents using a fully automated speech analysis and classification system. The prediction was achieved through a binary classification of speech recordings from 15 adolescents who developed Major Depression within two years after these recordings were made and 15 adolescents who did not developed Major Depression within the same time period. The results provided a proof of concept that an acoustic speech analysis can be used in early prediction of depression. The glottal features made the strongest predictors of depression with 69% accuracy, 62% specificity and 76% sensitivity. The TEO feature derived from glottal wave also provided good results, specifically when calculated at the frequency range of 1.3 kHz to 5.5 kHz. Kuan Ee Brian Ooi, Lu-Shih Alex Low, Margaret Lech, Nicholas B. Allen |
ICASSP | 4 |
| 2010 | Influence of acoustic low-level descriptors in the detection of clinical depression in adolescentsabstractIn this paper, we report the influence that classification accuracies have in speech analysis from a clinical dataset by adding acoustic low-level descriptors (LLD) belonging to prosodic (i.e. pitch, formants, energy, jitter, shimmer) and spectral features (i.e. spectral flux, centroid, entropy and roll-off) along with their delta (Δ) and delta-delta (Δ-Δ) coefficients to two baseline features of Mel frequency cepstral coefficients and Teager energy critical-band based autocorrelation envelope. Extracted acoustic low-level descriptors (LLD) that display an increase in accuracy after being added to these baseline features were finally modeled together using Gaussian mixture models and tested. A clinical data set of speech from 139 adolescents, including 68 (49 girls and 19 boys) diagnosed as clinically depressed, was used in the classification experiments. For male subjects, the combination of (TEO-CB-Auto-Env + Δ + Δ-Δ) + F0 + (LogE + Δ + Δ-Δ) + (Shimmer + Δ) + Spectral Flux + Spectral Roll-off gave the highest classification rate of 77.82% while for the female subjects, using TEO-CB-Auto-Env gave an accuracy of 74.74%. Lu-Shih Alex Low, Namunu Chinthaka Maddage, Margaret Lech, Lisa Sheeber, Nicholas B. Allen |
ICASSP | 5 |
| 2010 | On the importance of glottal flow spectral energy for the recognition of emotions in speechabstractTwo new approaches to feature extraction for automatic emotion classification in speech are described and tested. The methods are based on recent laryngological experiments testing the glottal air flow during phonation. The proposed approach calculates the area under the spectral energy envelope of the speech signal (AUSEES) and the glottal waveform (AUSEEG). The new methods provided very high recognition rates for seven emotions (contempt, angry, anxious, dysphoric, pleasant, neutral and happy). The speech data included 170 adult speakers (95 female and 75 male). The classification results showed that the new features provided significantly higher classification results (89.95% for AUSEEG, 76.07% for AUSEES) compared to the baseline MFCC approach (37.81%). The glottal waveform based AUSEEG features provided better results than the speech based AUSEES features, indicating that the majority of the emotion information is likely to be added to speech during the glottal wave formation Margaret Lech, Nicholas B. Allen |
INTERSPEECH | 3 |
| 2008 | Recognition of stress in speech using wavelet analysis and Teager energy operatorabstractThe automatic recognition and classification of speech under stress has applications in behavioural and mental health sciences, human to machine communication and robotics. The majority of recent studies are based on a linear model of the speech signal. In this study, the nonlinear Teager Energy Operator (TEO) analysis was used to derive the classification features. Moreover, the TEO analysis was combined with the Discrete Wavelet Transform, Wavelet Packet and Perceptual Wavelet Packet transforms to produce the Normalised TEO Autocorrelation Envelope Area coefficients for the classification process. The classification was performed using a Gaussian Mixture Model under speaker-independent conditions. The speech was classified into two classes: neutral and stressed. The best overall performance was observed for the features extracted using TEO analysis in combination with the Perceptual Wavelet Packet method. The accuracy in this case ranges from 94% to 96% depending on the type of mother wavelet Margaret Lech, Sheeraz Memon, Nicholas B. Allen |
INTERSPEECH | 4 |