Zygmunt Pizlo

dblp:10/6112 · DBLP profile ↗
← Back
24ranked-venue papers
5as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 10 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
3D vision · 85% Trustworthy machine learning · 15%
Computer graphics and multimedia
3 papers
Multimedia analysis and retrieval · 46% Visualization and visual analytics · 39% Audio and music processing · 10%

Topics — the 13 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
robustness
0.312025
The 3D-PC: a benchmark for visual perspective taking in humans and machines · ICLR 2025
Visualization and visual analytics › perception
perceptual studies
0.212016
Using virtual environments to evaluate assumptions of the human visual system · VR 2016
Computer vision › 3D vision
3d reconstruction
0.212014
Detecting 3-D Mirror Symmetry in a 2-D Camera Image for 3-D Shape Recovery · Proc. IEEE 2014
Computer vision › 3D vision
3d shape reconstruction
0.212014
Detecting 3-D Mirror Symmetry in a 2-D Camera Image for 3-D Shape Recovery · Proc. IEEE 2014
Computer vision › 3D vision › 3d reconstruction › geometric reconstruction
symmetry-based reconstruction
0.212014
Detecting 3-D Mirror Symmetry in a 2-D Camera Image for 3-D Shape Recovery · Proc. IEEE 2014
Multimedia analysis and retrieval
video summarization
0.222010
Camera Motion-Based Analysis of User Generated Video · IEEE Trans. Multim. 2010
Automated Video Program Summarization Using Speech Transcripts · IEEE Trans. Multim. 2006
Multimedia analysis and retrieval › video content analysis
user-generated video analysis
0.112010
Camera Motion-Based Analysis of User Generated Video · IEEE Trans. Multim. 2010
Computer vision › 3D vision
3d scene reconstruction
0.112016
Using virtual environments to evaluate assumptions of the human visual system · VR 2016
Audio and music processing
speech recognition
0.112006
Automated Video Program Summarization Using Speech Transcripts · IEEE Trans. Multim. 2006
Image and video processing › saliency detection
saliency map
0.012010
Camera Motion-Based Analysis of User Generated Video · IEEE Trans. Multim. 2010
Multimedia analysis and retrieval › multimedia browsing
video browsing
0.012006
Automated Video Program Summarization Using Speech Transcripts · IEEE Trans. Multim. 2006
Electronic design automation
design for manufacturability
0.011985
Tolerance Assignment for IC Selection Tests · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1985
Electronic design automation › hardware verification and test › VLSI testing
manufacturing test
0.011985
Tolerance Assignment for IC Selection Tests · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 1985

Methods — techniques the papers use, named apart from their topics

text prompting · 0.9linear probing · 0.9fine-tuning · 0.9deep neural network · 0.9object counting experiment · 0.5CAVE simulation · 0.5perspective projection · 0.4orthographic projection · 0.4a priori constraints · 0.4saliency map · 0.1eye tracking · 0.1camera motion analysis · 0.1word and bigram frequency scoring · 0.1pause detection · 0.1statistical optimization · 0.0
YearPublicationVenuePosition
2025 The 3D-PC: a benchmark for visual perspective taking in humans and machines
abstract
Visual perspective taking (VPT) is the ability to perceive and reason about the perspectives of others. It is an essential feature of human intelligence, which develops over the first decade of life and requires an ability to process the 3D structure of visual scenes. A growing number of reports have indicated that deep neural networks (DNNs) become capable of analyzing 3D scenes after training on large image datasets. We investigated if this emergent ability for 3D analysis in DNNs is sufficient for VPT with the 3D perception challenge (3D-PC): a novel benchmark for 3D perception in humans and DNNs. The 3D-PC is comprised of three 3D-analysis tasks posed within natural scene images: (i.) a simple test of object depth order, (ii.) a basic VPT task (VPT-basic), and (iii.) a more challenging version of VPT (VPT-perturb) designed to limit the effectiveness of "shortcut" visual strategies. We tested human participants (N=33) and linearly probed or text-prompted over 300 DNNs on the challenge and found that nearly all of the DNNs approached or exceeded human accuracy in analyzing object depth order. Surprisingly, DNN accuracy on this task correlated with their object recognition performance. In contrast, there was an extraordinary gap between DNNs and humans on VPT-basic. Humans were nearly perfect, whereas most DNNs were near chance. Fine-tuning DNNs on VPT-basic brought them close to human performance, but they, unlike humans, dropped back to chance when tested on VPT-perturb. Our challenge demonstrates that the training routines and architectures of today's DNNs are well-suited for learning basic 3D properties of scenes and objects but are ill-suited for reasoning about these properties like humans do. We release our 3D-PC datasets and code to help bridge this gap in 3D perception between humans and machines.
Drew Linsley, Peisen Zhou, Alekh Karkada Ashok, Akash Nagaraj, Gaurav Gaonkar, Francis E. Lewis, Zygmunt Pizlo, Thomas Serre
ICLR7
2024 Automatic segmentation and implicit surface representation of dynamic cardiac data
abstract
Abstract Segmentation of anatomical structures on 2D images of cardiac exams is necessary for performing 3D volumetric analysis, enabling the computation of parameters for diagnosing cardiovascular disease. In this work, we present robust algorithms to automatically segment cardiac imaging data and generate a volumetric anatomical reconstruction of a patient-specific heart model by propagating active contour output within a patient stack through a self-supervised learning model. Contour initializations are automatically generated, then output segmentations on sparse image slices are transferred and merged across a stack of images within the same heart data set during the segmentation process. We demonstrate whole-heart segmentation and compare the results with ground truth manual annotations. Additionally, we provide a framework to represent segmented heart data in the form of implicit surfaces, allowing interpolation operations to generate intermediary models of heart sections and volumes throughout the cardiac cycle and to estimate ejection fraction.
Andy Thai, Irmina Gradus-Pizlo, Zygmunt Pizlo, Hakan Sahin, Meenakshisundaram Gopi
Vis. Comput.3
2019 Investigating the role of the visual system in solving the traveling salesperson problem
Zahra Sajedinia, Zygmunt Pizlo, Sébastien Hélie
CogSci2
2016 Using virtual environments to evaluate assumptions of the human visual system
abstract
Virtual reality applications provide an opportunity to test human vision in well-controlled scenarios that would be difficult or impossible to generate in real physical spaces. This paper presents a study intended to evaluate the importance of possible assumptions made by the human visual system. Using a CAVE simulation, participants viewed and counted virtual furniture objects in a variety of experimental manipulations. The assumption of uprightness against inversion, or the `gravity constraint,' was identified as a significant assumption of the visual system (p <; 0.001). Monocular vs. binocular vision was also demonstrated as an important factor in this study (p = 0.01), while color vs. grayscale did not have a significant impact on task performance (p = 0.16). By including the binocular cue, and the assumption about the direction of gravity, the scene reconstruction produced by our computer vision model is reliable. The model can detect and count symmetrical objects in a 3D real scene and then recover their 3D shapes.
Eric Palmer, Aaron Michaux, Zygmunt Pizlo
VR3
2014 Detecting 3-D Mirror Symmetry in a 2-D Camera Image for 3-D Shape Recovery
abstract
In this paper, we take up the long-standing problem of how to recover 3-D shapes represented by a 2-D image, such as the image on the retina of the eye, or in a video camera. Our approach is biologically grounded in a theory of how the human visual system solves this problem, focusing on shapes that are mirror symmetrical in 3-D. A 3-D mirror-symmetrical shape can be recovered from a single 2-D orthographic or perspective image by applying several a priori constraints: 3-D mirror symmetry, 3-D compactness, and planarity of contours. From the computational point of view, the application of a 3-D symmetry constraint is challenging because it requires establishing 3-D symmetry correspondence among features of a 2-D image, which itself is asymmetrical for almost all viewing directions relative to the 3-D symmetrical shape. We describe new invariants of a 3-D to 2-D projection for the case of a pair of mirror-symmetrical planar contours, and we formally state and prove the necessary and sufficient conditions for detection of this type of symmetry in a single orthographic and perspective image.
Tadamasa Sawada, Zygmunt Pizlo
Proc. IEEE3
2012 Navigation toward Non-static Target Object Using Footprint Detection Based Tracking
Meng Yi, Yinfei Yang, Wenjing Qi, Yu Zhou 0016, Zygmunt Pizlo, Longin Jan Latecki
ACCV (3)6
2012 A simplified subjective video quality assessment method based on signal detection theory
abstract
A simplified protocol and associated metrics based on Signal Detection Theory (SDT) for subjective Video Quality Assessment (VQA) is proposed with the aim of filling the gap existing between the lack of discrimination abilities of objective Quality Estimates (specially when perceptually motivated processing methods are involved) and the costly normative subjective quality tests. The proposed protocol employs a reduced number of assessors and provides a quality ranking of the methods being evaluated. It is intended for providing the rapid experimental turn around necessary for developing algorithms. We have validated our proposal by corroborating with our test a well-known result for the video coding community: the quality benefits of including an in-loop deblocking filter. A software interface to design and administrate the test is also made publicly available.
Manuel de-Frutos-López, Ana Belén Mejía-Ocaña, Sergio Sanz Rodríguez, Carmen Peláez-Moreno, Fernando Díaz-de-María, Zygmunt Pizlo
PCS6
2011 Perceptually Based Appearance Modification for Compliant Appearance Editing
abstract
Abstract Projection‐based appearances are used in a variety of computer graphics applications to impart different appearances onto physical surfaces using digitally controlled projector light. To achieve a compliant appearance, all points on the physical surface must be altered to the colours of the desired target appearance; otherwise, an incompliant appearance results in a misleading visualization. Previous systems typically assume to operate with compliant appearances or restrict themselves to the simpler case of white surfaces. To achieve compliancy, one may change the physical surface's albedo, increase the amount of projector light radiance available or modify the target appearance's colours. This paper presents an approach to modify a target appearance to achieve compliant appearance editing without altering the physical surface or the projector setup. Our system minimally alters the target appearance's colours while maintaining cues important for perceptual similarity (e.g. colour constancy). First, we discuss how to measure colour compliancy. Next, we describe our approach to partition the physical surface into patches based on the surface's colours and the target appearance's colours. Finally, we describe our appearance optimization process, which computes a compliant appearance that is as perceptually similar as possible to the target appearance's colours. We perform several real‐world projection‐based appearances and compare our results to naïve approaches, which either ignore compliancy or simply reduce the appearance's overall brightness.
Alvin J. Law, Daniel G. Aliaga, Behzad Sajadi, Aditi Majumder, Zygmunt Pizlo
Comput. Graph. Forum5
2010 Camera Motion-Based Analysis of User Generated Video
abstract
In this paper we propose a system for the analysis of user generated video (UGV). UGV often has a rich camera motion structure that is generated at the time the video is recorded by the person taking the video, i.e., the ¿camera person.¿ We exploit this structure by defining a new concept known ascamera viewfor temporal segmentation of UGV. The segmentation provides a video summary with unique properties that is useful in applications such as video annotation. Camera motion is also a powerful feature for identification of keyframes and regions of interest (ROIs) since it is an indicator of the camera person's interests in the scene and can also attract the viewers' attention. We propose a new location-based saliency map which is generated based on camera motion parameters. This map is combined with other saliency maps generated using features such as color contrast, object motion and face detection to determine the ROIs. In order to evaluate our methods we conducted several user studies. A subjective evaluation indicated that our system produces results that is consistent with viewers' preferences. We also examined the effect of camera motion on human visual attention through an eye tracking experiment. The results showed a high dependency between the distribution of fixation points of the viewers and the direction of camera movement which is consistent with our location-based saliency map.
Golnaz Abdollahian, Cüneyt M. Taskiran, Zygmunt Pizlo, Edward J. Delp
IEEE Trans. Multim.3
2009 Approximative graph pyramid solution of the E-TSP
Yll Haxhimusa, Walter G. Kropatsch, Zygmunt Pizlo, Adrian Ion
Image Vis. Comput.3
2008 A study on the effect of camera motion on human visual attention
abstract
The aim of this paper is to examine the effect of camera motion in user generated video with respect to human visual attention. Having a more accurate human attention model is particularly useful in applications such as video summarization where the identification of visual importance is crucial to the quality of the results. Most models proposed thus far, have not considered camera motion as an independent factor in identifying visual saliency. In this study eye movement was recorded while subjects watched videos with different types of camera motion. The distribution of the fixation points indicated high correlation between the visual saliency and the type and direction of camera movement. The statistical tests confirmed that the results are statistically significant.
Golnaz Abdollahian, Zygmunt Pizlo, Edward J. Delp
ICIP2
2007 Human Perception of 3D Shapes
Zygmunt Pizlo
CAIP1
2006 Automated Video Program Summarization Using Speech Transcripts
abstract
Compact representations of video data greatly enhances efficient video browsing. Such representations provide the user with information about the content of the particular sequence being examined while preserving the essential message. We propose a method to automatically generate video summaries using transcripts obtained by automatic speech recognition. We divide the full program into segments based on pause detection and derive a score for each segment, based on the frequencies of the words and bigrams it contains. Then, a summary is generated by selecting the segments with the highest score to duration ratios while at the same time maximizing the coverage of the summary over the full program. We developed an experimental design and a user study to judge the quality of the generated video summaries. We compared the informativeness of the proposed algorithm with two other algorithms for three different programs. The results of the user study demonstrate that the proposed algorithm produces more informative summaries than the other two algorithms
Cüneyt M. Taskiran, Zygmunt Pizlo, Arnon Amir, Dulce B. Ponceleon, Edward J. Delp
IEEE Trans. Multim.2
2005 Vision pyramids that do not grow too high
Walter G. Kropatsch, Yll Haxhimusa, Zygmunt Pizlo, Georg Langs
Pattern Recognit. Lett.3
2004 Integral Trees: Subtree Depth and Diameter
Walter G. Kropatsch, Yll Haxhimusa, Zygmunt Pizlo
IWCIA3
2000 Recognition of a solid shape from its single perspective image obtained by a calibrated camera
Zygmunt Pizlo, Kirk Loubier
Pattern Recognit.1
1999 Binocular Shape Reconstruction: Psychological Plausibility of the 8-Point Algorithm
Moses W. Chan, Zygmunt Pizlo, David M. Chelberg
Comput. Vis. Image Underst.2
1997 The Geometry of Visual Space: About the Incompatibility between Science and Mathematics
Zygmunt Pizlo, Azriel Rosenfeld, Isaac Weiss
Comput. Vis. Image Underst.1
1997 Visual Space: Mathematics, Engineering, and Science
Zygmunt Pizlo, Azriel Rosenfeld, Isaac Weiss
Comput. Vis. Image Underst.1
1996 Video and image systems engineering education for the 21st century
abstract
We are developing a new graduate program at Purdue in Video and Image Systems Engineering (VISE). The project is comprised of three parts: a new curriculum centered around a degree option in VISE to be earned as part of the Masters or Ph.D. degrees; a state-of-the-art lecture/laboratory facility for instruction, laboratory experiments, and project and homework activities in VISE courses; and enhancement of existing courses and development of new courses in the VISE area.
Jan P. Allebach, Charles A. Bouman, Edward J. Coyle, Edward J. Delp, David A. Landgrebe, Anthony A. Maciejewski, Zygmunt Pizlo, Ness Shroff, Michael D. Zoltowski
ICIP (1)7
1996 Issues in the design of studies to test the effectiveness of stereo imaging
abstract
Recently, there has been a great increase in interest in using three dimensional stereoscopic displays to provide viewers with realistic 3D views of objects of interest. Some applications where stereoscopic displays are becoming popular include medical visualization, visualization of meteorological data, and various virtual reality applications. To quantify the effectiveness of stereoscopic systems over conventional monoscopic systems, well-designed experiments and data analysis methods are necessary. This task requires the combined effort of application scientists and experts in experimental design. Lack of interdisciplinary collaboration is a primary weakness of many stereoscopic display studies, resulting in the neglect of many important but subtle experimental issues. In this paper, we discuss specific issues that arise in the design of studies to determine the effectiveness of digital stereo imagery. Issues concerning statistical analysis of the experimental data are also discussed. References to related literature from engineering, computer graphics, and psychophysics are given. The issues developed herein provide a guideline for the design of studies to compare observer performance when using different imaging modalities.
Jean Hsu, Zygmunt Pizlo, David M. Chelberg, Charles F. Babbs, Edward J. Delp
IEEE Trans. Syst. Man Cybern. Part A2
1995 Preclinical ROC studies of digital stereomammography
abstract
Reports the diagnostic performance of observers in detecting abnormalities in computer-generated mammogram-like images. A mathematical model of the human breast is defined in which breast tissues are simulated by spheres of different sizes and densities. Images are generated by casting rays from a specified source, through the model, and onto an image plane. Observer performance when using two viewing modalities (stereo versus mono) is compared. In the stereo viewing mode, images are presented to the observer (wearing liquid-crystal display glasses), such that the left eye sees the left image only and the right eye sees the right image only. In this way, the images can be fused by the observer to obtain a sense of depth. In the mono viewing mode, identical images are presented to the left and right eyes so that no binocular disparities will be produced by the images. Observer response data are evaluated using receiver operating characteristic (ROC) analysis to characterize any difference in detectability of abnormalities (in either the density or the arrangement of simulated tissue densities) using the two viewing modes. The authors' experimental results indicate the clear superiority of stereo viewing for detection of arrangement abnormalities. For detection of density abnormalities, the performance of the two viewing modes is similar. These preliminary results suggest that stereomammography may permit easier detection of certain tissue abnormalities, perhaps providing a route to earlier tumor detection in cases of breast cancer.
Jean Hsu, David M. Chelberg, Charles F. Babbs, Zygmunt Pizlo, Edward J. Delp
IEEE Trans. Medical Imaging4
1992 Recognition of planar shapes from perspective images using contour-based invariants
Zygmunt Pizlo, Azriel Rosenfeld
CVGIP Image Underst.1
1985 Tolerance Assignment for IC Selection Tests
abstract
IC chips or manufactured wafers run through selection processes at various stages of the fabrication process. Typically, the most important is the selection performed on IC dice which are tested directly on manufacturing wafers. This paper deals with the problem of the optimal assignment of the upper and lower selection thresholds applied for selecting dice during the wafer measurements. The tolerance assignment is defined as a statistical optimization problem, where the optimization objective function is a measure of the manufacturing profit. In the paper a method for computing a solution of this optimization problem is proposed, and an example of the industrial application of this method in the Computer-Aided Manufacturing (CAM) area is given.
Wojciech Maly, Zygmunt Pizlo
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2