VLDB 2026 Research / reviewers in the wild / expert
Paul Sajda
dblp:03/2610
· DBLP profile ↗
39ranked-venue papers
9as first author
6since 2021 · last 2026
0000-0002-9738-1342ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 4 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 6 · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SwEYEpinch and Beyond: Exploring Intuitive, Efficient Text Entry for Extended Reality via Eye and Hand TrackingabstractDespite steady progress, text entry in Extended Reality (XR) often remains slower and more effortful than typing on a physical keyboard or touchscreen. We explore a simple idea: use gaze to swipe through a virtual keyboard for the fast, low-effort where and a manual pinch held throughout the swipe for the when, extending and validating it through a series of user studies. We first show that a basic version including a low-latency decoder with spatiotemporal Dynamic Time Warping and fixation filtering outperforms selecting individual keys sequentially, either by finger tapping each or gazing at each while pinching. We then add mid-swipe prediction and in-gesture cancellation, improving words per minute (WPM) without hurting accuracy. We show that this approach is faster and more preferred than previous gaze-swipe approaches, finger tapping with prediction, or hand swiping with the same additions. Furthermore, a seven-day, 30-session study demonstrates sustained learning, with peak performance reaching 64.7 WPM. Ziheng 'Leo' Li, Xichen He, Mengyuan Wu, Zeyi Tong, Haowen Wei, Benjamin Yang, Steven K. Feiner, Paul Sajda |
CHI | 8 |
| 2026 | Gaze patterns predict preference and confidence in pairwise AI image evaluationabstractPreference learning methods, such as Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO), rely on pairwise human judgments, yet little is known about the cognitive processes underlying these judgments. We investigate whether eye-tracking can reveal preference formation during pairwise AI-generated image evaluation. Thirty participants completed 1,800 trials while their gaze was recorded. We replicated the gaze cascade effect, with gaze shifting toward chosen images approximately one second before the decision. Cascade dynamics were consistent across confidence levels. Gaze features predicted binary choice (68% accuracy), with chosen images receiving more dwell time, fixations, and revisits. Gaze transitions distinguished high-confidence from uncertain decisions (66% accuracy), with low-confidence trials showing more image switches per second. These results show that gaze patterns predict both choice and confidence in pairwise image evaluations, suggesting that eye-tracking provides implicit signals relevant to the quality of preference annotations. Nikolas Papadopoulos, Shreenithi Navaneethan, Sheng Bai, Ankur Samanta, Paul Sajda |
ETRA | 5 |
| 2025 | Using Eye Tracking and AI-Powered Experimental Design to Create Patient-Centric Clinical Studies
Linbi Hong, Corbin Ping, Diego E. Arias, Nikolas Papadopoulos, Victoria Liu, Paul Sajda, Christopher Sege, Lisa M. McTeague |
ETRA | 6 |
| 2025 | Enabling Multi-Robot Collaboration from Single-Human GuidanceabstractLearning collaborative behaviors is essential for multi-agent systems. Traditionally, multi-agent reinforcement learning solves this implicitly through a joint reward and centralized observations, assuming collaborative behavior will emerge. Other studies propose to learn from demonstrations of a group of collaborative experts. Instead, we propose an efficient and explicit way of learning collaborative behaviors in multi-agent systems by leveraging expertise from only a single human. Our insight is that humans can naturally take on various roles in a team. We show that agents can effectively learn to collaborate by allowing a human operator to dynamically switch between controlling agents for a short period and incorporating a human-like theory-of-mind model of teammates. Our experiments showed that our method improves the success rate of a challenging collaborative hide-and-seek task by up to 58% with only 40 minutes of single-human guidance. We further demonstrate our findings transfer to the real world by conducting multi-robot experiments. Zhengran Ji, Paul Sajda, Boyuan Chen 0001 |
ICRA | 3 |
| 2024 | Circular Clustering With Polar Coordinate ReconstructionabstractThere is a growing interest in characterizing circular data found in biological systems. Such data are wide-ranging and varied, from the signal phase in neural recordings to nucleotide sequences in round genomes. Traditional clustering algorithms are often inadequate due to their limited ability to distinguish differences in the periodic component θ. Current clustering schemes for polar coordinate systems have limitations, such as being only angle-focused or lacking generality. To overcome these limitations, we propose a new analysis framework that utilizes projections onto a cylindrical coordinate system to represent objects in a polar coordinate system optimally. Using the mathematical properties of circular data, we show that our approach always finds the correct clustering result within the reconstructed dataset, given sufficient periodic repetitions of the data. This framework is generally applicable and adaptable to most state-of-the-art clustering algorithms. We demonstrate on synthetic and real data that our method generates more appropriate and consistent clustering results than standard methods. In summary, our proposed analysis framework overcomes the limitations of existing polar coordinate-based clustering methods and provides an accurate and efficient way to cluster circular data. Xiaoxiao Sun 0003, Paul Sajda |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2023 | Pupillary response is associated with the reset and switching of functional brain networks during salience processingabstractThe interface between processing internal goals and salient events in the environment involves various top-down processes. Previous studies have identified multiple brain areas for salience processing, including the salience network (SN), dorsal attention network, and the locus coeruleus-norepinephrine (LC-NE) system. However, interactions among these systems in salience processing remain unclear. Here, we simultaneously recorded pupillometry, EEG, and fMRI during an auditory oddball paradigm. The analyses of EEG and fMRI data uncovered spatiotemporally organized target-associated neural correlates. By modeling the target-modulated effective connectivity, we found that the target-evoked pupillary response is associated with the network directional couplings from late to early subsystems in the trial, as well as the network switching initiated by the SN. These findings indicate that the SN might cooperate with the pupil-indexed LC-NE system in the reset and switching of cortical networks, and shed light on their implications in various cognitive processes and neurological diseases. Hengda He, Linbi Hong, Paul Sajda |
PLoS Comput. Biol. | 3 |
| 2020 | Accelerated Robot Learning via Human Brain SignalsabstractIn reinforcement learning (RL), sparse rewards are a natural way to specify the task to be learned. However, most RL algorithms struggle to learn in this setting since the learning signal is mostly zeros. In contrast, humans are good at assessing and predicting the future consequences of actions and can serve as good reward/policy shapers to accelerate the robot learning process. Previous works have shown that the human brain generates an error-related signal, measurable using electroencephelography (EEG), when the human perceives the task being done erroneously. In this work, we propose a method that uses evaluative feedback obtained from human brain signals measured via scalp EEG to accelerate RL for robotic agents in sparse reward settings. As the robot learns the task, the EEG of a human observer watching the robot attempts is recorded and decoded into noisy error feedback signal. From this feedback, we use supervised learning to obtain a policy that subsequently augments the behavior policy and guides exploration in the early stages of RL. This bootstraps the RL learning process to enable learning from sparse reward. Using a simple robotic navigation task as a test bed, we show that our method achieves a stable obstacle-avoidance policy with high success rate, outperforming learning from sparse rewards only that struggles to achieve obstacle avoidance behavior or fails to advance to the goal. Iretiayo Akinola, Junyao Shi, Xiaomin He, Pawan Lapborisuth, Jingxi Xu 0002, David Watkins-Valls, Paul Sajda, Peter K. Allen |
ICRA | 8 |
| 2019 | A state-space model for inferring effective connectivity of latent neural dynamics from simultaneous EEG/fMRIabstractInferring effective connectivity between spatially segregated brain regions is important for understanding human brain dynamics in health and disease. Non-invasive neuroimaging modalities, such as electroencephalography (EEG) and functional magnetic resonance imaging (fMRI), are often used to make measurements and infer connectivity. However most studies do not consider integrating the two modalities even though each is an indirect measure of the latent neural dynamics and each has its own spatial and/or temporal limitations. In this study, we develop a linear state-space model to infer the effective connectivity in a distributed brain network based on simultaneously recorded EEG and fMRI data. Our method first identifies task-dependent and subject-dependent regions of interest (ROI) based on the analysis of fMRI data. Directed influences between the latent neural states at these ROIs are then modeled as a multivariate autogressive (MVAR) process driven by various exogenous inputs. The latent neural dynamics give rise to the observed scalp EEG measurements via a biophysically informed linear EEG forward model. We use a mean-field variational Bayesian approach to infer the posterior distribution of latent states and model parameters. The performance of the model was evaluated on two sets of simulations. Our results emphasize the importance of obtaining accurate spatial localization of ROIs from fMRI. Finally, we applied the model to simultaneously recorded EEG-fMRI data from 10 subjects during a Face-Car-House visual categorization task and compared the change in connectivity induced by different stimulus categories. Tao Tu 0001, John W. Paisley, Stefan Haufe, Paul Sajda |
NeurIPS | 4 |
| 2017 | Advanced Technologies for Brain Research [Scanning the Issue]abstractWe believe that this special issue will serve to increase the public awareness and foster discussions on the multiple worldwide BRAIN initiatives, both within and outside the IEEE, providing an impetus for development of long-term cost-effective healthcare solutions. We also believe that the topics presented in this special issue will serve as scientific evidence for health and policy advocates of the value of neurotechnologies for improving the neurological and mental health and wellbeing of the general population. Below we briefly highlight the papers and technologies in this special issue. Metin Akay, Paul Sajda, Silvestro Micera, Jose M. Carmena |
Proc. IEEE | 2 |
| 2017 | Fusing Multiple Neuroimaging Modalities to Assess Group Differences in Perception-Action CouplingabstractIn the last few decades, non-invasive neuroimaging has revealed macro-scale brain dynamics that underlie perception, cognition and action. Advances in non-invasive neuroimaging target two capabilities; 1) increased spatial and temporal resolution of measured neural activity, and 2) innovative methodologies to extract brain-behavior relationships from evolving neuroimaging technology. We target the second. Our novel methodology integrated three neuroimaging methodologies and elucidated expertise-dependent differences in functional (fused EEG-fMRI) and structural (dMRI) brain networks for a perception-action coupling task. A set of baseball players and controls performed a Go/No-Go task designed to mimic the situation of hitting a baseball. In the functional analysis, our novel fusion methodology identifies 50ms windows with predictive EEG neural correlates of expertise and fuses these temporal windows with fMRI activity in a whole-brain 2mm voxel analysis, revealing time-localized correlations of expertise at a spatial scale of millimeters. The spatiotemporal cascade of brain activity reflecting expertise differences begins as early as 200ms after the pitch starts and lasting up to 700ms afterwards. Network differences are spatially localized to include motor and visual processing areas, providing evidence for differences in perception-action coupling between the groups. Furthermore, an analysis of structural connectivity revealed that the players have significantly more connections between cerebellar and left frontal/motor regions, and many of the functional activation differences between the groups are located within structurally defined network modules that differentiate expertise. In short, our novel method illustrates how multimodal neuroimaging can provide specific macro-scale insights into the functional and structural correlates of expertise development. Jordan Muraskin, Jason Sherwin, Gregory Lieberman, Javier O. Garcia, Timothy D. Verstynen, Jean M. Vettel, Paul Sajda |
Proc. IEEE | 7 |
| 2016 | Closed-loop regulation of user state during a boundary avoidance taskabstractPilot induced oscillations (PIOs) are potentially catastrophic events that occur during flight when pilots attempt to control an aircraft close to a performance or physical boundary. PIO-like behavior is typically observed in boundary avoidance tasks (BAT), which simulate tight performance or physical boundaries and induce high cognitive workload. Our previous research linked the occurrence of PIO-like behavior to network level activity in the brain, where higher states of arousal reduce the flexibility of decision making networks such that less environmental information was incorporated to dynamically adjust action. This led us to hypothesize that down regulating arousal via closed-loop audio feedback of a user state could improve piloting performance by enabling increased decision flexibility. Here we show our initial results testing this hypothesis, where we use a hybrid brain computer interface (hBCI) to dynamically provide feedback to a “pilot” that facilitates their ability to reduce their state of arousal. We conduct a systematic comparison relative to control and sham conditions and test to see if this feedback increases the time a “pilot” can fly before a catastrophic PIO. We find that hBCI feedback, which includes central nervous system components consistent with theta activity in the anterior cingulate cortex (ACC), enables prolonged flight relative to closed-loop control and sham feedback. We also find that this feedback induces changes in pupil diameter which are absent in openloop conditions and closed-loop conditions when feedback is not veridical. Pupil diameter has been reported as a surrogate measure of activity in the locus coeruleus-norepinephrine (LC-NE) system which is also linked to a circuit that includes the ACC. We conclude that the feedback we induce with our hBCI provides preliminary evidence that self-regulation of LC-NE/ACC is possible and can be used to dynamically increase decision flexibility when under high cognitive workload. Josef Faller, Sameer Saproo, Victor Shih, Paul Sajda |
SMC | 4 |
| 2016 | Predicting decision accuracy and certainty in complex brain-machine interactionsabstractA promising application of brain machine interfaces (BMIs) is predicting user cognitive state, particularly in complex and demanding scenarios, so that automation can dynamically and adaptively adjust task parameters to optimize joint human-machine performance. In this paper we analyze neural, physiological and behavioral data recorded during a complex two-person “crew station” task and investigate whether these measures provide information for inferring user decision state. Specifically, we investigate how measures of EEG, pupil dilation, heart rate and response time, can be fused to infer decision confidence and accuracy in two side-tasks occurring throughout a three hour experimental session. One side-task is an auditory task, the other a visual task, both occurring within the context of the crew station scenario (auditory alert and a visual satellite map N-back task). We find that the best prediction performance always fuses EEG and pupil dilation measures, with results yielding between 70%–75% accuracy with respect to whether the subject(s) will skip making the decision (i.e. have high uncertainty) or whether he/she makes an error. Interestingly, the results suggest a possible mechanistic explanation for the utility of the fused measures, specifically the interaction between the locus coeruleus (LC), whose activity is linked to arousal state and can be inferred from pupil dilation, and the anterior cingulate (ACC), which has been linked to decision formation and monitoring and whose activity is typically measured via EEG. In general, our results demonstrate the potential in using fused neuro/physio measures to infer and track human operator decision uncertainty during demanding complex tasks, possibly enabling BMIs to eventually be employed as “cognitive orthotics” for improving man-machine interaction and performance. Victor Shih, Ludan Zhang, Christian Kothe, Scott Makeig, Paul Sajda |
SMC | 5 |
| 2016 | Unsupervised adaptive transfer learning for Steady-State Visual Evoked Potential brain-computer interfacesabstractRecent advances in signal processing for the detection of Steady-State Visual Evoked Potentials (SSVEPs) have moved away from traditionally calibrationless methods, such as canonical correlation analysis, and towards algorithms that require substantial training data. In general, this has improved detection rates, but SSVEP-based brain-computer interfaces (BCIs) now suffer from the requirement of costly calibration sessions. Here, we address this issue by applying transfer learning techniques to SSVEP detection. Our novel Adaptive-C3A method incorporates an unsupervised adaptation algorithm that requires no calibration data. Our approach learns SSVEP templates for the target user and provides robust class separation in feature space leading to increased classification accuracy. Our method achieves significant improvements in performance over a standard CCA method as well as a transfer variant of the state-of-the art Combined-CCA method for calibrationless SSVEP detection. Nicholas R. Waytowich, Josef Faller, Javier O. Garcia, Jean M. Vettel, Paul Sajda |
SMC | 5 |
| 2014 | Correlating Speaker Gestures in Political Debates with Audience Engagement Measured via EEGabstractWe hypothesize that certain speaker gestures can convey significant information that are correlated to audience engagement. We propose gesture attributes, derived from speakers' tracked hand motions to automatically quantify these gestures from video. Then, we demonstrate a correlation between gesture attributes and an objective method of measuring audience engagement: electroencephalography (EEG) in the domain of political debates. We collect 47 minutes of EEG recordings from each of 20 subjects watching clips of the 2012 U.S. Presidential debates. The subjects are examined in aggregate and in subgroups according to gender and political affiliation. We find statistically significant correlations between gesture attributes (particularly extremal pose) and our feature of engagement derived from EEG both with and without audio. For some stratifications, the Spearman rank correlation reaches as high as rho = 0.283 with p < 0.05, Bonferroni corrected. From these results, we identify those gestures that can be used to measure engagement, principally those that break habitual gestural patterns. John R. Zhang, Jason Sherwin, Jacek Dmochowski, Paul Sajda, John R. Kender |
ACM Multimedia | 4 |
| 2010 | Second-Order Bilinear Discriminant Analysis
Christoforos Christoforou, Robert M. Haralick, Paul Sajda, Lucas C. Parra |
J. Mach. Learn. Res. | 3 |
| 2010 | Maximum Likelihood in Cost-Sensitive Learning: Model Specification, Approximations, and Upper Bounds
Jacek Dmochowski, Paul Sajda, Lucas C. Parra |
J. Mach. Learn. Res. | 2 |
| 2010 | A Fast Hybrid Algorithm for Large-Scale l1-Regularized Logistic Regression
Jianing Shi, Wotao Yin, Stanley J. Osher, Paul Sajda |
J. Mach. Learn. Res. | 4 |
| 2010 | In a Blink of an Eye and a Switch of a Transistor: Cortically Coupled Computer VisionabstractOur society's information technology advancements have resulted in the increasingly problematic issue of information overload—i.e., we have more access to information than we can possibly process. This is nowhere more apparent than in the volume of imagery and video that we can access on a daily basis—for the general public, availability of YouTube video and Google Images, or for the image analysis professional tasked with searching security video or satellite reconnaissance. Which images to look at and how to ensure we see the images that are of most interest to us, begs the question of whether there are smart ways to triage this volume of imagery. Over the past decade, computer vision research has focused on the issue of ranking and indexing imagery. However, computer vision is limited in its ability to identify interesting imagery, particularly as “interesting” might be defined by an individual. In this paper we describe our efforts in developing brain–computer interfaces (BCIs) which synergistically integrate computer vision and human vision so as to construct a system for image triage. Our approach exploits machine learning for real-time decoding of brain signals which are recorded noninvasively via electroencephalography (EEG). The signals we decode are specific for events related to imagery attracting a user's attention. We describe two architectures we have developed for this type of cortically coupled computer vision and discuss potential applications and challenges for the future. Paul Sajda, Eric Pohlmeyer, Jun Wang 0006, Lucas C. Parra, Christoforos Christoforou, Jacek Dmochowski, Barbara Hanna, Claus Bahlmann, Maneesh Kumar Singh 0001, Shih-Fu Chang |
Proc. IEEE | 1 |
| 2009 | Brain state decoding for rapid image retrievalabstractHuman visual perception is able to recognize a wide range of targets under challenging conditions, but has limited throughput. Machine vision and automatic content analytics can process images at a high speed, but suffers from inadequate recognition accuracy for general target classes. In this paper, we propose a new paradigm to explore and combine the strengths of both systems. A single trial EEG-based brain machine interface (BCI) subsystem is used to detect objects of interest of arbitrary classes from an initial subset of images. The EEG detection outcomes are used as input to a graph-based pattern mining subsystem to identify, refine, and propagate the labels to retrieve relevant images from a much larger pool. The combined strategy is unique in its generality, robustness, and high throughput. It has great potential for advancing the state of the art in media retrieval applications. We have evaluated and demonstrated significant performance gains of the proposed system with multiple and diverse image classes over several data sets, including those from Internet (Caltech 101) and remote sensing images. In this paper, we will also present insights learned from the experiments and discuss future research directions. Jun Wang 0006, Eric Pohlmeyer, Barbara Hanna, Yu-Gang Jiang 0001, Paul Sajda, Shih-Fu Chang |
ACM Multimedia | 5 |
| 2007 | Second Order Bilinear Discriminant Analysis for single trial EEG analysisabstractTraditional analysis methods for single-trial classification of electro-encephalography (EEG) focus on two types of paradigms: phase locked methods, in which the amplitude of the signal is used as the feature for classification, i.e. event related potentials; and second order methods, in which the feature of interest is the power of the signal, i.e event related (de)synchronization. The process of deciding which paradigm to use is ad hoc and is driven by knowledge of neurological findings. Here we propose a unified method in which the algorithm learns the best first and second order spatial and temporal features for classification of EEG based on a bilinear model. The efficiency of the method is demonstrated in simulated and real EEG from a benchmark data set for Brain Computer Interface. Christoforos Christoforou, Paul Sajda, Lucas C. Parra |
NIPS | 2 |
| 2006 | Classifying Single-Trial ERPs from Visual and Frontal Cortex during Free ViewingabstractEvent-related potentials (ERPs) recorded at the scalp are indicators of brain activity associated with event-related information processing; hence they may be suitable for the assessment of changes in cognitive processing load. While the measurement of ERPs in a laboratory setting and classifying those ERPs is trivial, such a task presents major challenges in a "real world" setting where the EEG signals are recorded when subjects freely move their eyes and the sensory inputs are continuously, as opposed to discretely presented. Here we demonstrate that with the aid of second-order blind identification (SOBI), a blind source separation (BSS) algorithm: (1) we can extract ERPs from such challenging data sets; (2) we were able to obtain meaningful single-trial ERPs in addition to averaged ERPs; and (3) we were able to estimate the spatial origins of these ERPs. Finally, using back-propagation neural networks as classifiers, we show that these single-trial ERPs from specific brain regions can be used to determine moment-to-moment changes in cognitive processing load during a complex "real world" task. Akaysha C. Tang, Matthew T. Sutherland, Christopher J. McKinney, Jingyu Liu 0001, Lucas C. Parra, Adam D. Gerson, Paul Sajda |
IJCNN | 8 |
| 2006 | HiRes - a tool for comprehensive assessment and interpretation of metabolomic dataabstractUNLABELLED: The increasing role of metabolomics in system biology is driving the development of tools for comprehensive analysis of high-resolution NMR spectral datasets. This task is quite challenging since unlike the datasets resulting from other 'omics', a substantial preprocessing of the data is needed to allow successful identification of spectral patterns associated with relevant biological variability. HiRes is a unique stand-alone software tool that combines standard NMR spectral processing functionalities with techniques for multi-spectral dataset analysis, such as principal component analysis and non-negative matrix factorization. In addition, HiRes contains extensive abilities for data cleansing, such as baseline correction, solvent peak suppression, removal of frequency shifts owing to experimental conditions as well as auxiliary information management. Integration of these components together with multivariate analytical procedures makes HiRes very capable of addressing the challenges for assessment and interpretation of large metabolomic datasets, greatly simplifying this otherwise lengthy and difficult process and assuring optimal information retrieval. AVAILABILITY: HiRes is freely available for research purposes at http://hatch.cpmc.columbia.edu/highresmrs.html Radka Stoyanova, Shuyan Du, Paul Sajda, Truman R. Brown |
Bioinform. | 4 |
| 2006 | Varying complexity in tree-structured image distribution modelsabstractProbabilistic models of image statistics underlie many approaches in image analysis and processing. An important class of such models have variables whose dependency graph is a tree. If the hidden variables take values on a finite set, most computations with the model can be performed exactly, including the likelihood calculation, training with the EM algorithm, etc. Crouse et al. developed one such model, the hidden Markov tree (HMT). They took particular care to limit the complexity of their model. We argue that it is beneficial to allow more complex tree-structured models, describe the use of information theoretic penalties to choose the model complexity, and present experimental results to support these proposals. For these experiments, we use what we call the hierarchical image probability (HIP) model. The differences between the HIP and the HMT models include the use of multivariate Gaussians to model the distributions of local vectors of wavelet coefficients and the use of different numbers of hidden states at each resolution. We demonstrate the broad utility of image distributions by applying the HIP model to classification, synthesis, and compression, across a variety of image types, namely, electrooptical, synthetic aperture radar, and mammograms (digitized X-rays). In all cases, we compare with the HMT. Clay Spence, Lucas C. Parra, Paul Sajda |
IEEE Trans. Image Process. | 3 |
| 2005 | Neural mechanisms of contrast dependent receptive field size in V1abstractBased on a large scale spiking neuron model of the input layers 4C and of macaque, we identify neural mechanisms for the observed contrast dependent receptive field size of V1 cells. We observe a rich variety of mechanisms for the phenomenon and analyze them based on the relative gain of excitatory and inhibitory synaptic inputs. We observe an average growth in the spatial extent of excitation and inhibition for low contrast, as predicted from phenomenological models. However, contrary to phenomenological models, our simulation results suggest this is neither sufficient nor necessary to explain the phenomenon. Jim Wielaard, Paul Sajda |
NIPS | 2 |
| 2004 | Integration of form and motion within a generative model of visual cortex
Paul Sajda, Kyungim Baek |
Neural Networks | 1 |
| 2004 | Nonnegative matrix factorization for rapid recovery of constituent spectra in magnetic resonance chemical shift imaging of the brainabstractWe present an algorithm for blindly recovering constituent source spectra from magnetic resonance (MR) chemical shift imaging (CSI) of the human brain. The algorithm, which we call constrained nonnegative matrix factorization (cNMF), does not enforce independence or sparsity, instead only requiring the source and mixing matrices to be nonnegative. It is based on the nonnegative matrix factorization (NMF) algorithm, extending it to include a constraint on the positivity of the amplitudes of the recovered spectra. This constraint enables recovery of physically meaningful spectra even in the presence of noise that causes a significant number of the observation amplitudes to be negative. We demonstrate and characterize the algorithm's performance using 31P volumetric brain data, comparing the results with two different blind source separation methods: Bayesian spectral decomposition (BSD) and nonnegative sparse coding (NNSC). We then incorporate the cNMF algorithm into a hierarchical decomposition framework, showing that it can be used to recover tissue-specific spectra given a processing hierarchy that proceeds coarse-to-fine. We demonstrate the hierarchical procedure on 1H brain data and conclude that the computational efficiency of the algorithm makes it well-suited for use in diagnostic work-up. Paul Sajda, Shuyan Du, Truman R. Brown, Radka Stoyanova, Dikoma C. Shungu, Xiangling Mao, Lucas C. Parra |
IEEE Trans. Medical Imaging | 1 |
| 2003 | Single-trial detection in EEG and MEG: Keeping it linear
Lucas C. Parra, Christopher V. Alvino, Akaysha C. Tang, Barak A. Pearlmutter, Nick Yeung, Allen Osman, Paul Sajda |
Neurocomputing | 7 |
| 2003 | Blind Source Separation via Generalized Eigenvalue Decomposition
Lucas C. Parra, Paul Sajda |
J. Mach. Learn. Res. | 2 |
| 2003 | A multi-scale probabilistic network model for detection, synthesis and compression in mammographic image analysis
Paul Sajda, Clay Spence, Lucas C. Parra |
Medical Image Anal. | 1 |
| 2002 | Learning contextual relationships in mammograms using a pyramid neural networkabstractThis paper describes a pattern recognition architecture, which we term hierarchical pyramid/neural network (HPNN), that learns to exploit image structure at multiple resolutions for detecting clinically significant features in digital/digitized mammograms. The HPNN architecture consists of a hierarchy of neural networks, each network receiving feature inputs at a given scale as well as features constructed by networks lower in the hierarchy. Networks are trained using a novel error function for the supervised learning of image search/detection tasks when the position of the objects to be found is uncertain or ill defined. We have evaluated the HPNN's ability to eliminate false positive (FP) regions of interest generated by the University of Chicago's (UofC) Computer-aided diagnosis (CAD) systems for microcalcification and mass detection. Results show that the HPNN architecture, trained using the uncertain object position (UOP) error function, reduces the FP rate of a mammographic CAD system by approximately 50% without significant loss in sensitivity. Investigation into the types of FPs that the HPNN eliminates suggests that the pattern recognizer is automatically learning and exploiting contextual information. Clinical utility is demonstrated through the evaluation of an integrated system in a clinical reader study. We conclude that the HPNN architecture learns contextual relationships between features at multiple scales and integrates these features for detecting microcalcifications and breast masses. Paul Sajda, Clay Spence, John C. Pearson |
IEEE Trans. Medical Imaging | 1 |
| 2000 | Hierarchical Image Probability (HIP) ModelsabstractWe formulate a model for probability distributions on image spaces. We show that any distribution of images can be factored exactly into conditional distributions of feature vectors at one resolution (pyramid level) conditioned on the image information at lower resolutions. We would like to factor this over positions in the pyramid levels to make it tractable, but such factoring may miss long-range dependencies. To capture long-range dependencies, we introduce hidden class labels at each pixel in the pyramid. The result is a hierarchical mixture of conditional probabilities, similar to a hidden Markov model on a tree. The model parameters can be found with maximum likelihood estimation using the EM algorithm. We have obtained encouraging preliminary results on the problems of detecting various objects in SAR images and target recognition in optical aerial images. Clay Spence, Lucas C. Parra, Paul Sajda |
ICIP | 3 |
| 2000 | Higher-Order Statistical Properties Arising from the Non-Stationarity of Natural SignalsabstractWe present evidence that several higher-order statistical proper(cid:173) ties of natural images and signals can be explained by a stochastic model which simply varies scale of an otherwise stationary Gaus(cid:173) sian process. We discuss two interesting consequences. The first is that a variety of natural signals can be related through a com(cid:173) mon model of spherically invariant random processes, which have the attractive property that the joint densities can be constructed from the one dimensional marginal. The second is that in some cas(cid:173) es the non-stationarity assumption and only second order methods can be explicitly exploited to find a linear basis that is equivalent to independent components obtained with higher-order methods. This is demonstrated on spectro-temporal components of speech. Lucas C. Parra, Clay Spence, Paul Sajda |
NIPS | 3 |
| 1999 | Unmixing Hyperspectral Data
Lucas C. Parra, Clay Spence, Paul Sajda, Andreas Ziehe, Klaus-Robert Müller |
NIPS | 3 |
| 1998 | Applications of Multi-Resolution Neural Networks to Mammography
Clay Spence, Paul Sajda |
NIPS | 2 |
| 1995 | A hierarchical neural network architecture that learns target context: applications to digital mammographyabstractAn important problem in image analysis is finding small objects in large images. The problem is challenging because: 1) searching a large image is computationally expensive; and 2) small targets (on the order of a few pixels in size) have relatively few distinctive features which enable them to be distinguished from non-targets. To overcome these challenges the authors have developed a hierarchical neural network architecture which combines multiresolution pyramid processing with neural networks. Here the authors discuss the application of their hierarchical neural network architecture to the problem of detecting microcalcifications in digital mammograms. Microcalcifications are cues for breast tumors. 30% to 50% of breast carcinomas have microcalcifications visible in mammograms while 60% to 80% of all breast tumors eventually show microcalcifications via histology. Similar to the building/ATR problem, microcalcifications are generally very small point-like objects (<10 pixels in mammograms) which are hard to detect. Radiologists must often exploit other information in the imagery (e.g. location of blood vessels, ducts, etc.) in order to detect these microcalcifications. Here the authors examine how well their hierarchical neural network architecture learns and exploits contextual information in mammograms. Paul Sajda, Clay Spence, John C. Pearson |
ICIP (3) | 1 |
| 1995 | Integrating neural networks with image pyramids to learn target context
Paul Sajda, Clay Spence, Steven C. Hsu, John C. Pearson |
Neural Networks | 1 |
| 1993 | Dual Mechanisms for Neural Binding and Segmentation
Paul Sajda, Leif H. Finkel |
NIPS | 1 |
| 1992 | Object segmentation and binding within a biologically-based neural network model of depth-from-occlusionabstractThe problems of object segmentation and binding are addressed within a biologically based network model capable of determining depth from occlusion. In particular, the authors discuss two subprocesses most relevant to segmentation and binding: contour binding and figure direction. They propose that these two subprocesses have intrinsic constraints that allow several underdetermined problems in occlusion processing and object segmentation to be uniquely solved. Simulations that demonstrate the role these subprocesses play in discriminating objects and stratifying them in depth are reported. The network is tested on illusory stimuli, with the network's response indicating the existence of robust psychological properties in the system.> Paul Sajda, Leif H. Finkel |
CVPR | 1 |
| 1992 | Object Discrimination Based on Depth-from-OcclusionabstractWe present a model of how objects can be visually discriminated based on the extraction of depth-from-occlusion. Object discrimination requires consideration of both the binding problem and the problem of segmentation. We propose that the visual system binds contours and surfaces by identifying “proto-objects”—compact regions bounded by contours. Proto-objects can then be linked into larger structures. The model is simulated by a system of interconnected neural networks. The networks have biologically motivated architectures and utilize a distributed representation of depth. We present simulations that demonstrate three robust psychophysical properties of the system. The networks are able to stratify multiple occluding objects in a complex scene into separate depth planes. They bind the contours and surfaces of occluded objects (for example, if a tree branch partially occludes the moon, the two "half-moons" are bound into a single object). Finally, the model accounts for human perceptions of illusory contour stimuli. Leif H. Finkel, Paul Sajda |
Neural Comput. | 2 |