VLDB 2026 Research / reviewers in the wild / expert
Robert T. Schultz
dblp:59/3345
· DBLP profile ↗
30ranked-venue papers
0as first author
7since 2021 · last 2026
0000-0001-9817-3425ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 17 · 1 since 2021Artificial intelligence and machine learning · 11 · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Micro-DualNet: Dual-Path Spatio-Temporal Network for Micro-Action Recognition
Naga VS Raviteja Chappa, Evangelos Sariyanidi, Lisa Yankowitz, Gokul Nair, Casey Zampella, Robert T. Schultz, Birkan Tunç |
FG | 6 |
| 2025 | Beyond FACS: Data-driven Facial Expression Dictionaries, with Application to Predicting AutismabstractThe Facial Action Coding System (FACS) has been used by numerous studies to investigate the links between facial behavior and mental health. The laborious and costly process of FACS coding has motivated the development of machine learning frameworks for Action Unit (AU) detection. Despite intense efforts spanning three decades, the detection accuracy for many AUs is considered to be below the threshold needed for behavioral research. Also, many AUs are excluded altogether, making it impossible to fulfill the ultimate goal of FACSthe representation of any facial expression in its entirety. This paper considers an alternative approach. Instead of creating automated tools that mimic FACS experts, we propose to use a new coding system that mimics the key properties of FACS. Specifically, we construct a data-driven coding system called the Facial Basis, which contains units that correspond to localized and interpretable 3D facial movements, and overcomes three structural limitations of automated FACS coding. First, the proposed method is completely unsupervised, bypassing costly, laborious and variable manual annotation. Second, Facial Basis reconstructs all observable movement, rather than relying on a limited repertoire of recognizable movements (as in automated FACS). Finally, the Facial Basis units are additive, whereas AUs may fail detection when they appear in a non-additive combination. The proposed method outperforms the most frequently used AU detector in predicting autism diagnosis from in-person and remote conversations, highlighting the importance of encoding facial behavior comprehensively. To our knowledge, Facial Basis is the first alternative to FACS for deconstructing facial expressions in videos into localized movements. We provide an open source implementation of the method at github.com/sariyanidi/FacialBasis. Evangelos Sariyanidi, Lisa Yankowitz, Robert T. Schultz, John D. Herrington, Birkan Tunç, Jeffrey Cohn |
FG | 3 |
| 2025 | Phenotype Representation and Analysis via Discriminative Atypicality (PRADA) to Capture the Structural Heterogeneity of Autism Spectrum Disorder
Emre Onemli, Ahsan Mahmood, Omar Azrak, Dea Garic, Meghan R. Swanson, Rebecca Grzadzinski, Kattia Mata, Mark D. Shen, Jessica B. Girault, Tanya St. John, Juhi Pandey, Lonnie Zwaigenbaum, Annette M. Estes, Audrey M. Shen, Stephen Dager, Robert T. Schultz, Kelly N. Botteron, Alan C. Evans, Jed T. Elison, Essa Yacoub, Sun Hyung Kim, Robert C. McKinstry, Guido Gerig, Heather Cody Hazlett, Natasha Marrus, Joseph Piven, John R. Pruett Jr., Martin Styner |
MICCAI (2) | 16 |
| 2024 | Detecting Autism from Head Movements using KinesicsabstractHead movements play a crucial role in social interactions. The quantification of communicative movements such as nodding, shaking, orienting, and backchanneling is significant in behavioral and mental health research. However, automated localization of such head movements within videos remains challenging in computer vision due to their arbitrary start and end times, durations, and frequencies. In this work, we introduce a novel and efficient coding system for head movements, grounded in Birdwhistell’s kinesics theory, to automatically identify basic head motion units such as nodding and shaking. Our approach first defines the smallest unit of head movement, termed kine, based on the anatomical constraints of the neck and head. We then quantify the location, magnitude, and duration of kines within each angular component of head movement. Through defining possible combinations of identified kines, we define a higher-level construct, kineme, which corresponds to basic head motion units such as nodding and shaking. We validate the proposed framework by predicting autism spectrum disorder (ASD) diagnosis from video recordings of interacting partners. We show that the multi-scale property of the proposed framework provides a significant advantage, as collapsing behavior across temporal scales reduces performance consistently. Finally, we incorporate another fundamental behavioral modality, namely speech, and show that distinguishing between speaking- and listening-time head movements significantly improves ASD classification performance. Muhittin Gökmen, Evangelos Sariyanidi, Lisa Yankowitz, Casey Zampella, Robert T. Schultz, Birkan Tunç |
ICMI | 5 |
| 2024 | Inequality-Constrained 3D Morphable Face Model Fittingabstract3D morphable model (3DMM) fitting on 2D data is traditionally done via unconstrained optimization with regularization terms to ensure that the result is a plausible face shape and is consistent with a set of 2D landmarks. This paper presents inequality-constrained 3DMM fitting as the first alternative to regularization in optimization-based 3DMM fitting. Inequality constraints on the 3DMM's shape coefficients ensure face-like shapes without modifying the objective function for smoothness, thus allowing for more flexibility to capture person-specific shape details. Moreover, inequality constraints on landmarks increase robustness in a way that does not require per-image tuning. We show that the proposed method stands out with its ability to estimate person-specific face shapes by jointly fitting a 3DMM to multiple frames of a person. Further, when used with a robust objective function, namely gradient correlation, the method can work "in-the-wild" even with a 3DMM constructed from controlled data. Lastly, we show how to use the log-barrier method to efficiently implement the method. To our knowledge, we present the first 3DMM fitting framework that requires no learning yet is accurate, robust, and efficient. The absence of learning enables a generic solution that allows flexibility in the input image size, interchangeable morphable models, and incorporation of camera matrix. Evangelos Sariyanidi, Casey Zampella, Robert T. Schultz, Birkan Tunç |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Meta-evaluation for 3D Face Reconstruction Via Synthetic DataabstractThe standard benchmark metric for 3D face reconstruction is the geometric error between reconstructed meshes and the ground truth. Nearly all recent reconstruction methods are validated on real ground truth scans, in which case one needs to establish point correspondence prior to error computation, which is typically done with the Chamfer (i.e., nearest neighbor) criterion. However, a simple yet fundamental question have not been asked: Is the Chamfer error an appropriate and fair benchmark metric for 3D face reconstruction? More generally, how can we determine which error estimator is a better benchmark metric? We present a meta-evaluation framework that uses synthetic data to evaluate the quality of a geometric error estimator as a benchmark metric for face reconstruction. Further, we use this framework to experimentally compare four geometric error estimators. Results show that the standard approach not only severely underestimates the error, but also does so inconsistently across reconstruction methods, to the point of even altering the ranking of the compared methods. Moreover, although non-rigid ICP leads to a metric with smaller estimation bias, it could still not correctly rank all compared reconstruction methods, and is significantly more time consuming than Chamfer. In sum, we show several issues present in the current benchmarking and propose a procedure using synthetic data to address these issues. Evangelos Sariyanidi, Claudio Ferrari, Stefano Berretti, Robert T. Schultz, Birkan Tunç |
IJCB | 4 |
| 2023 | Automatically Predicting Perceived Conversation Quality in a Pediatric Sample Enriched for AutismabstractSocial interaction quality ratings derived from short natural conversations can differentiate children with and without autism at the group level. In this work, we explored conversations between children and an unfamiliar adult who rated their social interaction success on six dimensions. Using hand-crafted acoustic and lexical features, we built different classifiers to predict children's dimensional conversation quality. The best classifier achieved 61% accuracy, which outperformed human raters (49%). Follow-up analyses revealed that a subset of features determined communication quality scores. Additionally, we extracted acoustic features using a pretrained audio transformer and improved our prediction to 68%. This study suggests that automatically predicting conversation quality could be an inexpensive and objective way to monitor intervention progress in children with communication challenges, and could be used to identify intervention targets for improving conversational success. Yahan Yang, Sunghye Cho, Maxine Covello, Azia Knox, Osbert Bastani, James Weimer, Edgar Dobriban, Robert T. Schultz, Insup Lee 0001, Julia Parish-Morris |
INTERSPEECH | 8 |
| 2020 | Discovering Synchronized Subsets of Sequences: A Large Scale SolutionabstractFinding the largest subset of sequences (i.e., time series) that are correlated above a certain threshold, within large datasets, is of significant interest for computer vision and pattern recognition problems across domains, including behavior analysis, computational biology, neuroscience, and finance. Maximal clique algorithms can be used to solve this problem, but they are not scalable. We present an approximate, but highly efficient and scalable, method that represents the search space as a union of sets called ϵ-expanded clusters, one of which is theoretically guaranteed to contain the largest subset of synchronized sequences. The method finds synchronized sets by fitting a Euclidean ball on ϵ-expanded clusters, using Jung's theorem. We validate the method on data from the three distinct domains of facial behavior analysis, finance, and neuroscience, where we respectively discover the synchrony among pixels of face videos, stock market item prices, and dynamic brain connectivity data. Experiments show that our method produces results comparable to, but up to 300 times faster than, maximal clique algorithms, with speed gains increasing exponentially with the number of input sequences. Evangelos Sariyanidi, Casey Zampella, G. Keith Bartley, John D. Herrington, Theodore D. Satterthwaite, Robert T. Schultz, Birkan Tunç |
CVPR | 6 |
| 2020 | Can Facial Pose and Expression Be Separated With Weak Perspective Camera?abstractSeparating facial pose and expression within images requires a camera model for 3D-to-2D mapping. The weak perspective (WP) camera has been the most popular choice; it is the default, if not the only option, in state-of-the-art facial analysis methods and software. WP camera is justified by the supposition that its errors are negligible when the subjects are relatively far from the camera, yet this claim has never been tested despite nearly 20 years of research. This paper critically examines the suitability of WP camera for separating facial pose and expression. First, we theoretically show that WP causes pose-expression ambiguity, as it leads to estimation of spurious expressions. Next, we experimentally quantify the magnitude of spurious expressions. Finally, we test whether spurious expressions have detrimental effects on a common facial analysis application, namely Action Unit (AU) detection. Contrary to conventional wisdom, we find that severe pose-expression ambiguity exists even when subjects are not close to the camera, leading to large false positive rates in AU detection. We also demonstrate that the magnitude and characteristics of spurious expressions depend on the point distribution model used to model the expressions. Our results suggest that common assumptions about WP need to be revisited in facial expression modeling, and that facial analysis software should encourage and facilitate the use of the true camera model whenever possible. Evangelos Sariyanidi, Casey Zampella, Robert T. Schultz, Birkan Tunç |
CVPR | 3 |
| 2020 | Inequality-Constrained and Robust 3D Face Model Fitting
Evangelos Sariyanidi, Casey Zampella, Robert T. Schultz, Birkan Tunç |
ECCV (9) | 3 |
| 2019 | Automatic Detection of Autism Spectrum Disorder in Children Using Acoustic and Text Features from Brief Natural Conversations
Sunghye Cho, Mark Y. Liberman, Neville Ryant, Meredith Cola, Robert T. Schultz, Julia Parish-Morris |
INTERSPEECH | 5 |
| 2019 | Structural Connectivity Analysis Using Finsler GeometryabstractIn this work we demonstrate how Finsler geometry---and specifically the related geodesic tracto-graphy---can be levied to analyze structural connections between different brain regions. We present new theoretical developments which support the definition of a novel Finsler metric and associated connectivity measures, based on closely related works on the Riemannian framework for diffusion MRI. Using data from the Human Connectome Project, as well as population data from an autism spectrum disorder study, we demonstrate that this new Finsler metric, together with the new connectivity measures, results in connectivity maps that are much closer to known tract anatomy compared to previous geodesic connectivity methods. Our implementation can be used to compute geodesic distance and connectivity maps for segmented areas and is publicly available. Tom C. J. Dela Haije, Peter Savadjiev, Andrea Fuster, Robert T. Schultz, Ragini Verma, Luc Florack, Carl-Fredrik Westin |
SIAM J. Imaging Sci. | 4 |
| 2017 | On characterizing population commonalities and subject variations in brain networks
Yasser Ghanbari, Luke Bloy, Birkan Tunç, Varsha Shankar, Timothy P. L. Roberts, J. Christopher Edgar, Robert T. Schultz, Ragini Verma |
Medical Image Anal. | 7 |
| 2016 | Building Language Resources for Exploring Autism Spectrum Disorders
Julia Parish-Morris, Christopher Cieri, Mark Y. Liberman, Leila Bateman, Emily Ferguson, Robert T. Schultz |
LREC | 6 |
| 2014 | Functionally Driven Brain Networks Using Multi-layer Graph Clustering
Yasser Ghanbari, Luke Bloy, Varsha Shankar, J. Christopher Edgar, Timothy P. L. Roberts, Robert T. Schultz, Ragini Verma |
MICCAI (3) | 6 |
| 2014 | Identifying group discriminative and age regressive sub-networks from DTI-based connectivity via a unified framework of non-negative matrix factorization and graph embedding
Yasser Ghanbari, Alex R. Smith, Robert T. Schultz, Ragini Verma |
Medical Image Anal. | 3 |
| 2014 | Fusion of white and gray matter geometry: A framework for investigating brain development
Peter Savadjiev, Yogesh Rathi, Sylvain Bouix, Alex R. Smith, Robert T. Schultz, Ragini Verma, Carl-Fredrik Westin |
Medical Image Anal. | 5 |
| 2013 | Connectivity Subnetwork Learning for Pathology and Developmental Variations
Yasser Ghanbari, Alex R. Smith, Robert T. Schultz, Ragini Verma |
MICCAI (1) | 3 |
| 2013 | Combining Surface and Fiber Geometry: An Integrated Approach to Brain Morphology
Peter Savadjiev, Yogesh Rathi, Sylvain Bouix, Alex R. Smith, Robert T. Schultz, Ragini Verma, Carl-Fredrik Westin |
MICCAI (1) | 5 |
| 2011 | HARDI Based Pattern Classifiers for the Identification of White Matter Pathologies
Luke Bloy, Madhura Ingalhalikar, Harini Eavani, Timothy P. L. Roberts, Robert T. Schultz, Ragini Verma |
MICCAI (2) | 5 |
| 2008 | 3D Cerebral Cortical Morphometry in Autism: Increased Folding in Children and Adolescents in Frontal, Parietal, and Temporal Lobes
Suyash P. Awate, Lawrence Win, Paul A. Yushkevich, Robert T. Schultz, James C. Gee |
MICCAI (1) | 4 |
| 2008 | Bayesian Analysis of fMRI Data with ICA Based Spatial Prior
Deepti R. Bathula, Hemant D. Tagare, Lawrence H. Staib, Xenophon Papademetris, Robert T. Schultz, James S. Duncan |
MICCAI (2) | 5 |
| 2004 | Integrated Intensity and Point-Feature Nonrigid Registration
Xenophon Papademetris, Andrea Jackowski, Robert T. Schultz, Lawrence H. Staib, James S. Duncan |
MICCAI (1) | 3 |
| 2004 | Functional Brain Image Analysis Using Joint Function-Structure Priors
Jing Yang 0005, Xenophon Papademetris, Lawrence H. Staib, Robert T. Schultz, James S. Duncan |
MICCAI (2) | 4 |
| 2003 | Computing 3D Non-rigid Brain Registration Using Extended Robust Point Matching for Composite Multisubject fMRI Analysis
Xenophon Papademetris, Andrea Jackowski, Robert T. Schultz, Lawrence H. Staib, James S. Duncan |
MICCAI (2) | 3 |
| 2003 | A unified non-rigid feature registration method for brain mapping
Haili Chui, Lawrence Win, Robert T. Schultz, James S. Duncan, Anand Rangarajan 0001 |
Medical Image Anal. | 3 |
| 1999 | A New Approach to 3D Sulcal Ribbon Finding from MR Images
Xiaolan Zeng, Lawrence H. Staib, Robert T. Schultz, Hemant D. Tagare, Lawrence Win, James S. Duncan |
MICCAI | 3 |
| 1999 | Segmentation and Measurement of the Cortex from 3D MR Images Using Coupled Surfaces PropagationabstractThe cortex is the outermost thin layer of gray matter in the brain; geometric measurement of the cortex helps in understanding brain anatomy and function. In the quantitative analysis of the cortex from MR images, extracting the structure and obtaining a representation for various measurements are key steps. While manual segmentation is tedious and labor intensive, automatic reliable efficient segmentation and measurement of the cortex remain challenging problems, due to its convoluted nature. Here we present a new approach of coupled-surfaces propagation, using level set methods to address such problems. Our method is motivated by the nearly constant thickness of the cortical mantle and takes this tight coupling as an important constraint. By evolving two embedded surfaces simultaneously, each driven by its own image-derived information while maintaining the coupling, a final representation of the cortical bounding surfaces and an automatic segmentation of the cortex are achieved. Characteristics of the cortex, such as cortical surface area, surface curvature, and cortical thickness, are then evaluated. The level set implementation of surface propagation offers the advantage of easy initialization, computational efficiency, and the ability to capture deep sulcal folds. Results and validation from various experiments on both simulated and real three-dimensional (3-D) MR images are provided. Xiaolan Zeng, Lawrence H. Staib, Robert T. Schultz, James S. Duncan |
IEEE Trans. Medical Imaging | 3 |
| 1998 | Volumetric Layer Segmentation Using Coupled Surfaces PropagationabstractThe problem of segmenting a volumetric layer of finite thickness is encountered in several important areas within medical image analysis. Key examples include the extraction of the cortical gray matter of the brain and the left ventricle myocardium of the heart. The coupling between the two bounding surfaces of such a layer provides important information that helps to solve the segmentation problem. Here we propose a new approach of coupled surfaces propagation via level set methods, which takes into account coupling as an important constraint. By evolving two embedded surfaces simultaneously, each driven by its own image-derived information while maintaining the coupling, we capture a representation of the two bounding surfaces and achieve automatic segmentation on the layer. Characteristic gray level values, instead of image gradient information alone, are incorporated in deriving the useful image information to drive the surface propagation, which enables our approach to capture the homogeneity inside the layer. The level set implementation offers the advantage of easy initialization, computational efficiency and the ability to capture deep folds of the sulci. As a test example, we apply our approach to unedited 3D Magnetic Resonance (MR) brain images. Our algorithm automatically isolates the brain from non-brain structures and recovers the cortical gray matter. Xiaolan Zeng, Lawrence H. Staib, Robert T. Schultz, James S. Duncan |
CVPR | 3 |
| 1998 | Segmentation and Measurement of the Cortex from 3D MR Images
Xiaolan Zeng, Lawrence H. Staib, Robert T. Schultz, James S. Duncan |
MICCAI | 3 |