Miguel P. Eckstein

dblp:56/975 · DBLP profile ↗
← Back
25ranked-venue papers
0as first author
7since 2021 · last 2023
0000-0002-5528-355XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 since 2021Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Planning, search and constraint satisfaction · 27% Image recognition and object detection · 22% Vision and language · 21%
Computer graphics and multimedia
6 papers
Visual content generation and editing · 61% Image and video processing · 16% Computational photography and imaging · 14%
Human-computer interaction and pervasive computing
5 papers
Human-AI interaction · 32% Wearable and physiological sensing · 28% User interface design and tools · 16%

Topics — the 24 heaviest of 29, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning › hybrid planning
neuro-symbolic planning
0.712023
Neuro-Symbolic Procedural Planning with Commonsense Prompting · ICLR 2023
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › task planning
procedure planning
0.712023
Neuro-Symbolic Procedural Planning with Commonsense Prompting · ICLR 2023
Visual content generation and editing › image generation
text-to-image generation
0.712023
Collaborative Generative AI: Integrating GPT-k for Efficient Editing in Text-to-Image Generation · EMNLP 2023
Computer vision › Image recognition and object detection
object detection
0.632019
Assessment of Faster R-CNN in Man-Machine Collaborative Search · CVPR 2019
From Where and How to What We See · ICCV 2013
Can Peripheral Representations Improve Clutter Metrics on Complex Scenes? · NIPS 2016
Computer vision › Vision and language › vision-language generation
language-guided video editing
0.612022
M3L: Language-based Video Editing via Multi-Modal Multi-Level Transformers · CVPR 2022
Machine learning › Deep learning architectures and training › transformer
multimodal transformer
0.612022
M3L: Language-based Video Editing via Multi-Modal Multi-Level Transformers · CVPR 2022
Wearable and physiological sensing
eye tracking
0.532017
Attention Allocation Aid for Visual Search · CHI 2017
From Where and How to What We See · ICCV 2013
Eye tracking assisted extraction of attentionally important objects from videos · CVPR 2015
Knowledge, reasoning and agents › Knowledge representation and reasoning › causal reasoning
counterfactual reasoning
0.412020
SSCR: Iterative Language-Based Image Editing via Self-Supervised Counterfactual Reasoning · EMNLP (1) 2020
Computer vision › Vision and language
vision-and-language navigation
0.412020
Counterfactual Vision-and-Language Navigation via Adversarial Path Sampler · ECCV (6) 2020
Visual content generation and editing › image editing › interactive image editing
iterative image editing
0.412020
SSCR: Iterative Language-Based Image Editing via Self-Supervised Counterfactual Reasoning · EMNLP (1) 2020
Computer vision › Image recognition and object detection › object detection › region-based detection
Faster R-CNN
0.412019
Assessment of Faster R-CNN in Man-Machine Collaborative Search · CVPR 2019
Computational photography and imaging › color science
metamerism
0.412019
Towards Metamerism via Foveated Style Transfer · ICLR (Poster) 2019
Visual content generation and editing
style transfer
0.412019
Towards Metamerism via Foveated Style Transfer · ICLR (Poster) 2019
Human-AI interaction › AI-assisted decision-making
human-AI collaborative decision making
0.412019
Assessment of Faster R-CNN in Man-Machine Collaborative Search · CVPR 2019
User interface design and tools
attention allocation
0.312017
Attention Allocation Aid for Visual Search · CHI 2017
Image and video processing › video segmentation
video object extraction
0.212015
Eye tracking assisted extraction of attentionally important objects from videos · CVPR 2015
Image and video processing
video segmentation
0.212015
Eye tracking assisted extraction of attentionally important objects from videos · CVPR 2015
Knowledge, reasoning and agents › Knowledge representation and reasoning
commonsense reasoning
0.212023
Neuro-Symbolic Procedural Planning with Commonsense Prompting · ICLR 2023
Visual content generation and editing
video generation
0.212022
M3L: Language-based Video Editing via Multi-Modal Multi-Level Transformers · CVPR 2022
Learning and educational technologies
attention prediction
0.212013
From Where and How to What We See · ICCV 2013
Usability and user experience research
human performance evaluation
0.112019
Assessment of Faster R-CNN in Man-Machine Collaborative Search · CVPR 2019
Human-robot interaction › teleoperation
supervisory control
0.112017
Attention Allocation Aid for Visual Search · CHI 2017
Interaction techniques and input
visual search
0.112017
Attention Allocation Aid for Visual Search · CHI 2017
Computer vision › Image recognition and object detection
visual search
0.112016
Can Peripheral Representations Improve Clutter Metrics on Complex Scenes? · NIPS 2016

Methods — techniques the papers use, named apart from their topics

large language model prompting · 1.3human study · 1.3counterfactual reasoning · 1.3multimodal transformer · 1.1cross-modal alignment · 1.1self-supervised learning · 0.9cross-task consistency · 0.9Faster R-CNN · 0.8eye tracking · 0.7large language model · 0.7commonsense prompting · 0.7adversarial sampling · 0.4foveated style transfer · 0.4eye-tracking · 0.4VGG16 · 0.4probabilistic modeling · 0.3feature congestion · 0.2eccentricity modulation · 0.2
YearPublicationVenuePosition
2023 Collaborative Generative AI: Integrating GPT-k for Efficient Editing in Text-to-Image Generation
abstract
The field of text-to-image (T2I) generation has garnered significant attention both within the research community and among everyday users.Despite the advancements of T2I models, a common issue encountered by users is the need for repetitive editing of input prompts in order to receive a satisfactory image, which is time-consuming and labor-intensive.Given the demonstrated text generation power of largescale language models, such as GPT-k, we investigate the potential of utilizing such models to improve the prompt editing process for T2I generation.We conduct a series of experiments to compare the common edits made by humans and GPT-k, evaluate the performance of GPT-k in prompting T2I, and examine factors that may influence this process.We found that GPT-k models focus more on inserting modifiers while humans tend to replace words and phrases, which includes changes to the subject matter.Experimental results show that GPT-k are more effective in adjusting modifiers rather than predicting spontaneous changes in the primary subject matters.Adopting the edit suggested by GPT-k models may reduce the percentage of remaining edits by 20-30%. 1 Our experiments are conducted upon StableDiffusion since it is a wide-adopted open-source large text-to-image generative model with SoTA performance.
Wanrong Zhu, Xinyi Wang 0003, Tsu-Jui Fu, Xin Wang 0061, Miguel P. Eckstein, William Yang Wang
EMNLP6
2023 Neuro-Symbolic Procedural Planning with Commonsense Prompting
Weixi Feng, Wanrong Zhu, Wenda Xu, Xin Wang 0061, Miguel P. Eckstein, William Yang Wang
ICLR6
2023 A 2D Synthesized Image Improves the 3D Search for Foveated Visual Systems
abstract
Current medical imaging increasingly relies on 3D volumetric data making it difficult for radiologists to thoroughly search all regions of the volume. In some applications (e.g., Digital Breast Tomosynthesis), the volumetric data is typically paired with a synthesized 2D image (2D-S) generated from the corresponding 3D volume. We investigate how this image pairing affects the search for spatially large and small signals. Observers searched for these signals in 3D volumes, 2D-S images, and while viewing both. We hypothesize that lower spatial acuity in the observers' visual periphery hinders the search for the small signals in the 3D images. However, the inclusion of the 2D-S guides eye movements to suspicious locations, improving the observer's ability to find the signals in 3D. Behavioral results show that the 2D-S, used as an adjunct to the volumetric data, improves the localization and detection of the small (but not large) signal compared to 3D alone. There is a concomitant reduction in search errors as well. To understand this process at a computational level, we implement a Foveated Search Model (FSM) that executes human eye movements and then processes points in the image with varying spatial detail based on their eccentricity from fixations. The FSM predicts human performance for both signals and captures the reduction in search errors when the 2D-S supplements the 3D search. Our experimental and modeling results delineate the utility of 2D-S in 3D search-reduce the detrimental impact of low-resolution peripheral processing by guiding attention to regions of interest, effectively reducing errors.
Devi Klein, Miguel A. Lago, Craig K. Abbey, Miguel P. Eckstein
IEEE Trans. Medical Imaging4
2022 M3L: Language-based Video Editing via Multi-Modal Multi-Level Transformers
abstract
Video editing tools are widely used nowadays for digital design. Although the demand for these tools is high, the prior knowledge required makes it difficult for novices to get started. Systems that could follow natural language instructions to perform automatic editing would significantly improve accessibility. This paper introduces the language-based video editing (LBVE) task, which allows the model to edit, guided by text instruction, a source video into a target video. LBVE contains two features: 1) the scenario of the source video is preserved instead of generating a completely different video; 2) the semantic is presented differently in the target video, and all changes are controlled by the given instruction. We propose a Multi-Modal Multi-Level Transformer (M3L) to carry out LBVE. M3L dynamically learns the correspondence between video perception and language semantic at different levels, which benefits both the video understanding and video frame synthesis. We build three new datasets for evaluation, including two diagnostic and one from natural videos with human-labeled text. Extensive experimental results show that M3L is effective for video editing and that LBVE can lead to a new field toward vision-and-language research.
Tsu-Jui Fu, Xin Wang 0061, Scott T. Grafton, Miguel P. Eckstein, William Yang Wang
CVPR4
2022 Imagination-Augmented Natural Language Understanding
abstract
Yujie Lu, Wanrong Zhu, Xin Wang, Miguel Eckstein, William Yang Wang. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Wanrong Zhu, Xin Wang 0061, Miguel P. Eckstein, William Yang Wang
NAACL-HLT4
2022 Diagnosing Vision-and-Language Navigation: What Really Matters
abstract
Wanrong Zhu, Yuankai Qi, Pradyumna Narayana, Kazoo Sone, Sugato Basu, Xin Wang, Qi Wu, Miguel Eckstein, William Yang Wang. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Wanrong Zhu, Yuankai Qi, Pradyumna Narayana, Kazoo Sone, Sugato Basu, Xin Wang 0061, Qi Wu 0001, Miguel P. Eckstein, William Yang Wang
NAACL-HLT8
2021 Foveated Model Observers for Visual Search in 3D Medical Images
abstract
Model observers have a long history of success in predicting human observer performance in clinically-relevant detection tasks. New 3D image modalities provide more signal information but vastly increase the search space to be scrutinized. Here, we compared standard linear model observers (ideal observers, non-pre-whitening matched filter with eye filter, and various versions of Channelized Hotelling models) to human performance searching in 3D 1/f2.8filtered noise images and assessed its relationship to the more traditional location known exactly detection tasks and 2D search. We investigated two different signal types that vary in their detectability away from the point of fixation (visual periphery). We show that the influence of 3D search on human performance interacts with the signal's detectability in the visual periphery. Detection performance for signals difficult to detect in the visual periphery deteriorates greatly in 3D search but not in 3D location known exactly and 2D search. Standard model observers do not predict the interaction between 3D search and signal type. A proposed extension of the Channelized Hotelling model (foveated search model) that processes the image with reduced spatial detail away from the point of fixation, explores the image through eye movements, and scrolls across slices can successfully predict the interaction observed in humans and also the types of errors in 3D search. Together, the findings highlight the need for foveated model observers for image quality evaluation with 3D search.
Miguel A. Lago, Craig K. Abbey, Miguel P. Eckstein
IEEE Trans. Medical Imaging3
2020 Counterfactual Vision-and-Language Navigation via Adversarial Path Sampler
Tsu-Jui Fu, Xin Wang 0061, Matthew F. Peterson, Scott T. Grafton, Miguel P. Eckstein, William Yang Wang
ECCV (6)5
2020 SSCR: Iterative Language-Based Image Editing via Self-Supervised Counterfactual Reasoning
abstract
Iterative Language-Based Image Editing (IL-BIE) tasks follow iterative instructions to edit images step by step.Data scarcity is a significant issue for ILBIE as it is challenging to collect large-scale examples of images before and after instruction-based changes.However, humans still accomplish these editing tasks even when presented with an unfamiliar image-instruction pair.Such ability results from counterfactual thinking and the ability to think about alternatives to events that have happened already.In this paper, we introduce a Self-Supervised Counterfactual Reasoning (SSCR) framework that incorporates counterfactual thinking to overcome data scarcity.SSCR allows the model to consider out-ofdistribution instructions paired with previous images.With the help of cross-task consistency (CTC), we train these counterfactual instructions in a self-supervised scenario.Extensive results show that SSCR improves the correctness of ILBIE in terms of both object identity and position, establishing a new state of the art (SOTA) on two IBLIE datasets (i-CLEVR and CoDraw).Even with only 50% of the training data, SSCR achieves a comparable result to using complete data.
Tsu-Jui Fu, Xin Wang 0061, Scott T. Grafton, Miguel P. Eckstein, William Yang Wang
EMNLP (1)4
2019 Assessment of Faster R-CNN in Man-Machine Collaborative Search
abstract
With the advent of modern expert systems driven by deep learning that supplement human experts (e.g. radiologists, dermatologists, surveillance scanners), we analyze how and when do such expert systems enhance human performance in a fine-grained small target visual search task. We set up a 2 session factorial experimental design in which humans visually search for a target with and without a Deep Learning (DL) expert system. We evaluate human changes of target detection performance and eye-movements in the presence of the DL system. We find that performance improvements with the DL system (computed via a Faster R-CNN with a VGG16) interacts with observer's perceptual abilities (e.g., sensitivity). The main results include: 1) The DL system reduces the False Alarm rate per Image on average across observer groups of both high/low sensitivity; 2) Only human observers with high sensitivity perform better than the DL system, while the low sensitivity group does not surpass individual DL system performance, even when aided with the DL system itself; 3) Increases in number of trials and decrease in viewing time were mainly driven by the DL system only for the low sensitivity group. 4) The DL system aids the human observer to fixate at a target by the 3rd fixation. These results provide insights of the benefits and limitations of deep learning systems that are collaborative or competitive with humans.
Arturo Deza, Amit Surana, Miguel P. Eckstein
CVPR3
2019 Towards Metamerism via Foveated Style Transfer
Arturo Deza, Aditya Jonnalagadda, Miguel P. Eckstein
ICLR (Poster)3
2017 Attention Allocation Aid for Visual Search
abstract
This paper outlines the development and testing of a novel, feedback-enabled attention allocation aid (AAAD), which uses real-time physiological data to improve human performance in a realistic sequential visual search task. Indeed, by optimizing over search duration, the aid improves efficiency, while preserving decision accuracy, as the operator identifies and classifies targets within simulated aerial imagery. Specifically, using experimental eye-tracking data and measurements about target detectability across the human visual field, we develop functional models of detection accuracy as a function of search time, number of eye movements, scan path, and image clutter. These models are then used by the AAAD in conjunction with real time eye position data to make probabilistic estimations of attained search accuracy and to recommend that the observer either move on to the next image or continue exploring the present image. An experimental evaluation in a scenario motivated from human supervisory control in surveillance missions confirms the benefits of the AAAD.
Arturo Deza, Jeffrey Russel Peters, Grant S. Taylor, Amit Surana, Miguel P. Eckstein
CHI5
2017 Object detection through search with a foveated visual system
abstract
Humans and many other species sense visual information with varying spatial resolution across the visual field (foveated vision) and deploy eye movements to actively sample regions of interests in scenes. The advantage of such varying resolution architecture is a reduced computational, hence metabolic cost. But what are the performance costs of such processing strategy relative to a scheme that processes the visual field at high spatial resolution? Here we first focus on visual search and combine object detectors from computer vision with a recent model of peripheral pooling regions found at the V1 layer of the human visual system. We develop a foveated object detector that processes the entire scene with varying resolution, uses retino-specific object detection classifiers to guide eye movements, aligns its fovea with regions of interest in the input image and integrates observations across multiple fixations. We compared the foveated object detector against a non-foveated version of the same object detector which processes the entire image at homogeneous high spatial resolution. We evaluated the accuracy of the foveated and non-foveated object detectors identifying 20 different objects classes in scenes from a standard computer vision data set (the PASCAL VOC 2007 dataset). We show that the foveated object detector can approximate the performance of the object detector with homogeneous high spatial resolution processing while bringing significant computational cost savings. Additionally, we assessed the impact of foveation on the computation of bottom-up saliency. An implementation of a simple foveated bottom-up saliency model with eye movements showed agreement in the selection of top salient regions of scenes with those selected by a non-foveated high resolution saliency model. Together, our results might help explain the evolution of foveated visual systems with eye movements as a solution that preserves perceptual performance in visual search while resulting in computational and metabolic savings to the brain.
Emre Akbas, Miguel P. Eckstein
PLoS Comput. Biol.2
2016 Can Peripheral Representations Improve Clutter Metrics on Complex Scenes?
abstract
Previous studies have proposed image-based clutter measures that correlate with human search times and/or eye movements. However, most models do not take into account the fact that the effects of clutter interact with the foveated nature of the human visual system: visual clutter further from the fovea has an increasing detrimental influence on perception. Here, we introduce a new foveated clutter model to predict the detrimental effects in target search utilizing a forced fixation search task. We use Feature Congestion (Rosenholtz et al.) as our non foveated clutter model, and we stack a peripheral architecture on top of Feature Congestion for our foveated model. We introduce the Peripheral Integration Feature Congestion (PIFC) coefficient, as a fundamental ingredient of our model that modulates clutter as a non-linear gain contingent on eccentricity. We finally show that Foveated Feature Congestion (FFC) clutter scores (r(44) = −0.82 ± 0.04, p < 0.0001) correlate better with target detection (hit rate) than regular Feature Congestion (r(44) = −0.19 ± 0.13, p = 0.0774) in forced fixation search; and we extend foveation to other clutter models showing stronger correlations in all cases. Thus, our model allows us to enrich clutter perception research by computing fixation specific clutter maps. Code for building peripheral representations is available.
Arturo Deza, Miguel P. Eckstein
NIPS2
2015 Scene Inversion Slows the Rejection of False Positives through Saccade Exploration During Search
Kathryn Koehler, Miguel P. Eckstein
CogSci2
2015 Eye tracking assisted extraction of attentionally important objects from videos
abstract
Visual attention is a crucial indicator of the relative importance of objects in visual scenes to human viewers. In this paper, we propose an algorithm to extract objects which attract visual attention from videos. As human attention is naturally biased towards high level semantic objects in visual scenes, this information can be valuable to extract salient objects. The proposed algorithm extracts dominant visual tracks using eye tracking data from multiple subjects on a video sequence by a combination of mean-shift clustering and Hungarian algorithm. These visual tracks guide a generic object search algorithm to get candidate object locations and extents in every frame. Further, we propose a novel multiple object extraction algorithm by constructing a spatio-temporal mixed graph over object candidates. Bounding box based object extraction inference is performed using binary linear integer programming on a cost function defined over the graph. Finally, the object boundaries are refined using grabcut segmentation. The proposed technique outperforms state-of-the-art video segmentation using eye tracking prior and obtains favorable object extraction over algorithms which do not utilize eye tracking data.
S. Karthikeyan 0001, Thuyen Ngo, Miguel P. Eckstein, B. S. Manjunath
CVPR3
2015 Derivation of an Observer Model Adapted to Irregular Signals Based on Convolution Channels
abstract
Anthropomorphic model observers are mathe- matical algorithms which are applied to images with the ultimate goal of predicting human signal detection and classification accuracy across varieties of backgrounds, image acquisitions and display conditions. A limitation of current channelized model observers is their inability to handle irregularly-shaped signals, which are common in clinical images, without a high number of directional channels. Here, we derive a new linear model observer based on convolution channels which we refer to as the "Filtered Channel observer" (FCO), as an extension of the channelized Hotelling observer (CHO) and the nonprewhitening with an eye filter (NPWE) observer. In analogy to the CHO, this linear model observer can take the form of a single template with an external noise term. To compare with human observers, we tested signals with irregular and asymmetrical shapes spanning the size of lesions down to those of microcalfications in 4-AFC breast tomosynthesis detection tasks, with three different contrasts for each case. Whereas humans uniformly outperformed conventional CHOs, the FCO observer outperformed humans for every signal with only one exception. Additive internal noise in the models allowed us to degrade model performance and match human performance. We could not match all the human performances with a model with a single internal noise component for all signal shape, size and contrast conditions. This suggests that either the internal noise might vary across signals or that the model cannot entirely capture the human detection strategy. However, the FCO model offers an efficient way to apprehend human observer performance for a non-symmetric signal.
Iván Díaz, Craig K. Abbey, Pontus Timberg, Miguel P. Eckstein, Francis R. Verdun, Cyril Castella, François O. Bochud
IEEE Trans. Medical Imaging4
2014 Single-Trial Classification of Event-Related Potentials in Rapid Serial Visual Presentation Tasks Using Supervised Spatial Filtering
abstract
Accurate detection of single-trial event-related potentials (ERPs) in the electroencephalogram (EEG) is a difficult problem that requires efficient signal processing and machine learning techniques. Supervised spatial filtering methods that enhance the discriminative information in EEG data are commonly used to improve single-trial ERP detection. We propose a convolutional neural network (CNN) with a layer dedicated to spatial filtering for the detection of ERPs and with training based on the maximization of the area under the receiver operating characteristic curve (AUC). The CNN is compared with three common classifiers: 1) Bayesian linear discriminant analysis; 2) multilayer perceptron (MLP); and 3) support vector machines. Prior to classification, the data were spatially filtered with xDAWN (for the maximization of the signal-to-signal-plus-noise ratio), common spatial pattern, or not spatially filtered. The 12 analytical techniques were tested on EEG data recorded in three rapid serial visual presentation experiments that required the observer to discriminate rare target stimuli from frequent nontarget stimuli. Classification performance discriminating targets from nontargets depended on both the spatial filtering method and the classifier. In addition, the nonlinear classifier MLP outperformed the linear methods. Finally, training based AUC maximization provided better performance than training based on the minimization of the mean square error. The results support the conclusion that the choice of the systems architecture is critical and both spatial filtering and classification must be considered together.
Hubert Cecotti, Miguel P. Eckstein, Barry Giesbrecht
IEEE Trans. Neural Networks Learn. Syst.2
2013 From Where and How to What We See
abstract
Eye movement studies have confirmed that overt attention is highly biased towards faces and text regions in images. In this paper we explore a novel problem of predicting face and text regions in images using eye tracking data from multiple subjects. The problem is challenging as we aim to predict the semantics (face/text/background) only from eye tracking data without utilizing any image information. The proposed algorithm spatially clusters eye tracking data obtained in an image into different coherent groups and subsequently models the likelihood of the clusters containing faces and text using a fully connected Markov Random Field (MRF). Given the eye tracking data from a test image, it predicts potential face/head (humans, dogs and cats) and text locations reliably. Furthermore, the approach can be used to select regions of interest for further analysis by object detectors for faces and text. The hybrid eye position/object detector approach achieves better detection performance and reduced computation time compared to using only the object detection algorithm. We also present a new eye tracking dataset on 300 images selected from ICDAR, Street-view, Flickr and Oxford-IIIT Pet Dataset from 15 subjects.
S. Karthikeyan 0001, Vignesh Jagadeesh, Renuka Shenoy, Miguel P. Eckstein, B. S. Manjunath
ICCV4
2012 A signal detection analysis of the effects of repeated context on visual search
Ryan W. Kasper, Miguel P. Eckstein, Barry Giesbrecht
CogSci2
2010 Evolution and Optimality of Similar Neural Mechanisms for Perception and Action during Search
abstract
A prevailing theory proposes that the brain's two visual pathways, the ventral and dorsal, lead to differing visual processing and world representations for conscious perception than those for action. Others have claimed that perception and action share much of their visual processing. But which of these two neural architectures is favored by evolution? Successful visual search is life-critical and here we investigate the evolution and optimality of neural mechanisms mediating perception and eye movement actions for visual search in natural images. We implement an approximation to the ideal Bayesian searcher with two separate processing streams, one controlling the eye movements and the other stream determining the perceptual search decisions. We virtually evolved the neural mechanisms of the searchers' two separate pathways built from linear combinations of primary visual cortex receptive fields (V1) by making the simulated individuals' probability of survival depend on the perceptual accuracy finding targets in cluttered backgrounds. We find that for a variety of targets, backgrounds, and dependence of target detectability on retinal eccentricity, the mechanisms of the searchers' two processing streams converge to similar representations showing that mismatches in the mechanisms for perception and eye movements lead to suboptimal search. Three exceptions which resulted in partial or no convergence were a case of an organism for which the targets are equally detectable across the retina, an organism with sufficient time to foveate all possible target locations, and a strict two-pathway model with no interconnections and differential pre-filtering based on parvocellular and magnocellular lateral geniculate cell properties. Thus, similar neural mechanisms for perception and eye movement actions during search are optimal and should be expected from the effects of natural selection on an organism with limited time to search for food that is not equi-detectable across its retina and interconnected perception and action neural pathways.
Sheng Zhang 0007, Miguel P. Eckstein
PLoS Comput. Biol.2
2006 The effect of nonlinear human visual system components on performance of a channelized Hotelling observer in structured backgrounds
abstract
Linear model observers based on statistical decision theory have been used successfully to predict human visual detection of aperiodic signals in a variety of noisy backgrounds. However, some models have included nonlinearities such as a transducer or nonlinear decision rules to handle intrinsic uncertainty. In addition, masking models used to predict human visual detection of signals superimposed on one of two identical backgrounds (masks) usually include a number of nonlinear components in the channels that reflect properties of the firing of cells in the primary visual cortex (V1). The effect of these nonlinearities on the ability of linear model observers to predict human signal detection in real patient structured backgrounds is unknown. We evaluate the effect of including different nonlinear human visual system components into a linear channelized Hotelling observer (CHO) using a signal known exactly but variable (SKEV) task. In particular, we evaluate whether the rank order of two compression algorithms (JPEG versus JPEG 2000) and two compression encoder settings (JPEG 2000 default versus JPEG 2000 optimized) based on model observer signal detection performance in X-ray coronary angiograms is altered by inclusion of nonlinear components. The results show: 1) the simpler linear CHO model observer outperforms CHO model with the nonlinear components; 2) the rank order of model observer performance for the compression algorithms/parameters does not change when the nonlinear components are included. For the present task and images, the results suggest that the addition of the nonlinearities to a channelized Hotelling model may add complexity to the model observers without great impact on rank order evaluation of image processing and/or acquisition algorithms.
Binh Pham 0002, Miguel P. Eckstein
IEEE Trans. Medical Imaging3
2004 Automated optimization of JPEG 2000 encoder options based on model observer performance for detecting variable signals in X-ray coronary angiograms
abstract
Image compression is indispensable in medical applications where inherently large volumes of digitized images are presented. JPEG 2000 has recently been proposed as a new image compression standard. The present recommendations on the choice of JPEG 2000 encoder options were based on nontask-based metrics of image quality applied to nonmedical images. We used the performance of a model observer [non-prewhitening matched filter with an eye filter (NPWE)] in a visual detection task of varying signals [signal known exactly but variable (SKEV)] in X-ray coronary angiograms to optimize JPEG 2000 encoder options through a genetic algorithm procedure. We also obtained the performance of other model observers (Hotelling, Laguerre-Gauss Hotelling, channelized-Hotelling) and human observers to evaluate the validity of the NPWE optimized JPEG 2000 encoder settings. Compared to the default JPEG 2000 encoder settings, the NPWE-optimized encoder settings improved the detection performance of humans and the other three model observers for an SKEV task. In addition, the performance also was improved for a more clinically realistic task where the signal varied from image to image but was not known a priori to observers [signal known statistically (SKS)]. The highest performance improvement for humans was at a high compression ratio (e.g., 30:1) which resulted in approximately a 75% improvement for both the SKEV and SKS tasks.
Binh Pham 0002, Miguel P. Eckstein
IEEE Trans. Medical Imaging3
2004 Evaluation of JPEG 2000 encoder options: human and model observer detection of variable signals in X-ray coronary angiograms
abstract
Previous studies have evaluated the effect of the new still image compression standard JPEG 2000 using nontask based image quality metrics, i.e., peak-signal-to-noise-ratio (PSNR) for nonmedical images. In this paper, the effect of JPEG 2000 encoder options was investigated using the performance of human and model observers (nonprewhitening matched filter with an eye filter, square-window Hotelling, Laguerre-Gauss Hotelling and channelized Hotelling model observer) for clinically relevant visual tasks. Two tasks were investigated: the signal known exactly but variable task (SKEV) and the signal known statistically task (SKS). Test images consisted of real X-ray coronary angiograms with simulated filling defects (signals) inserted in one of the four simulated arteries. The signals varied in size and shape. Experimental results indicated that the dependence of task performance on the JPEG 2000 encoder options was similar for all model and human observers. Model observer performance in the more tractable and computationally economic SKEV task can be used to reliably estimate performance in the complex but clinically more realistic SKS task. JPEG 2000 encoder settings different from the default ones resulted in greatly improved model and human observer performance in the studied clinically relevant visual tasks using real angiography backgrounds.
Binh Pham 0002, Miguel P. Eckstein
IEEE Trans. Medical Imaging3
2002 Optimal shifted estimates of human-observer templates in two-alternative forced-choice experiments
abstract
For performing simple detection and discrimination tasks in image noise, human observers are often modeled by a cross-correlation between the image and an observer template followed by the injection of the observer's internal noise. This paper is concerned with estimating this template using the two-alternative forced-choice (2AFC) experimental paradigm. The basic idea behind the estimation procedure is to average the noise fields of the images used in a 2AFC experiment with a weight that depends on whether the observer got the trial correct or incorrect. We describe a method that produces unbiased estimates of the observer template up to a constant of proportionality under the linear cross-correlation model. The method proposed here is different from some previous methods in the way it assigns weights to the noise fields and we show that the resulting errors in the estimated template are minimized. We also propose and validate a formula for approximating the error covariance associated with the template estimates.
Craig K. Abbey, Miguel P. Eckstein
IEEE Trans. Medical Imaging2