VLDB 2026 Research / reviewers in the wild / expert
Shmuel Peleg
dblp:p/ShmuelPeleg
· DBLP profile ↗
112ranked-venue papers
18as first author
4since 2021 · last 2024
0000-0002-4468-2619ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 78 · 7 first-author · 4 since 2021Artificial intelligence and machine learning · 75 · 13 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 1 first-authorComputer networks · 1Theory of computation · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
24 papers |
3D vision · 42% Video understanding and tracking · 24% Segmentation and scene understanding · 17% | |
| Computer graphics and multimedia
39 papers |
Image and video processing · 33% Computational photography and imaging · 31% Multimedia analysis and retrieval · 17% | |
| Network and information security
3 papers |
Security and privacy of machine learning · 66% Privacy and data protection · 33% Digital forensics and information hiding · 0% |
Topics — the 30 heaviest of 103, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Security and privacy of machine learning
membership inference |
0.5 | 1 | 2021 | Membership Inference Attacks are Easier on Difficult Problems · ICCV 2021 |
Computational photography and imaging
image stitching |
0.3 | 8 | 2007 | Dynamosaicing: Mosaicing of Dynamic Scenes · IEEE Trans. Pattern Anal. Mach. Intell. 2007 Seamless image stitching by minimizing false edges · IEEE Trans. Image Process. 2006 Seamless Image Stitching in the Gradient Domain · ECCV (4) 2004 |
Computer vision › 3D vision
camera calibration |
0.2 | 1 | 2016 | Camera Calibration from Dynamic Silhouettes Using Motion Barcodes · CVPR 2016 |
Computer vision › Image recognition and object detection
character recognition |
0.2 | 1 | 2016 | Visual Learning of Arithmetic Operation · AAAI 2016 |
Computer vision › Video understanding and tracking
egocentric video understanding |
0.2 | 1 | 2016 | An Egocentric Look at Video Photographer Identity · CVPR 2016 |
Computer vision › 3D vision › multi-view geometry
epipolar geometry estimation |
0.2 | 1 | 2016 | Camera Calibration from Dynamic Silhouettes Using Motion Barcodes · CVPR 2016 |
Computer vision › 3D vision › multi-view geometry › epipolar geometry estimation
fundamental matrix estimation |
0.2 | 1 | 2016 | Fundamental Matrices from Moving Objects Using Line Motion Barcodes · ECCV (2) 2016 |
Privacy and data protection
anonymity |
0.2 | 1 | 2016 | An Egocentric Look at Video Photographer Identity · CVPR 2016 |
Image and video processing
video stabilization |
0.2 | 1 | 2015 | EgoSampling: Fast-forward and stereo for egocentric videos · CVPR 2015 |
Computer vision › Video understanding and tracking › temporal understanding
temporal segmentation |
0.2 | 1 | 2014 | Temporal Segmentation of Egocentric Videos · CVPR 2014 |
Computer vision › Segmentation and scene understanding
video segmentation |
0.2 | 1 | 2014 | Temporal Segmentation of Egocentric Videos · CVPR 2014 |
Wearable and physiological sensing › wearable camera
egocentric video |
0.2 | 1 | 2014 | Temporal Segmentation of Egocentric Videos · CVPR 2014 |
Multimedia analysis and retrieval
video indexing |
0.2 | 2 | 2008 | Nonchronological Video Synopsis and Indexing · IEEE Trans. Pattern Anal. Mach. Intell. 2008 Webcam Synopsis: Peeking Around the World · ICCV 2007 |
Computer vision › Segmentation and scene understanding
medical image segmentation |
0.1 | 1 | 2021 | Membership Inference Attacks are Easier on Difficult Problems · ICCV 2021 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.1 | 1 | 2021 | Membership Inference Attacks are Easier on Difficult Problems · ICCV 2021 |
Multimedia analysis and retrieval › video summarization
video synopsis |
0.1 | 2 | 2007 | Webcam Synopsis: Peeking Around the World · ICCV 2007 Making a Long Video Short: Dynamic Video Synopsis · CVPR (1) 2006 |
Computational photography and imaging
panoramic imaging |
0.1 | 5 | 2001 | Omnistereo: Panoramic Stereo Imaging · IEEE Trans. Pattern Anal. Mach. Intell. 2001 Rectified Mosaicing: Mosaics without the Curl · CVPR 2000 Cameras for Stereo Panoramic Imaging · CVPR 2000 |
Image and video processing › image sequence processing
time manipulation |
0.1 | 2 | 2007 | Dynamosaicing: Mosaicing of Dynamic Scenes · IEEE Trans. Pattern Anal. Mach. Intell. 2007 Dynamosaics: Video Mosaics with Non-Chronological Time · CVPR (1) 2005 |
Graph algorithms and graph theory › graph theory
graph labeling |
0.1 | 2 | 2009 | Shift-map image editing · ICCV 2009 A New Probabilistic Relaxation Scheme · IEEE Trans. Pattern Anal. Mach. Intell. 1980 |
Visual content generation and editing
image editing |
0.1 | 1 | 2009 | Shift-map image editing · ICCV 2009 |
Image and video processing › image restoration
image inpainting |
0.1 | 1 | 2009 | Shift-map image editing · ICCV 2009 |
Visual content generation and editing
image retargeting |
0.1 | 1 | 2009 | Shift-map image editing · ICCV 2009 |
Computer vision › 3D vision
image mosaicing |
0.1 | 1 | 2008 | Minimal Aspect Distortion (MAD) Mosaicing of Long Scenes · Int. J. Comput. Vis. 2008 |
Computer vision › Video understanding and tracking › video summarization
video synopsis |
0.1 | 1 | 2008 | Nonchronological Video Synopsis and Indexing · IEEE Trans. Pattern Anal. Mach. Intell. 2008 |
Image and video processing
image registration |
0.1 | 2 | 2007 | Dynamosaicing: Mosaicing of Dynamic Scenes · IEEE Trans. Pattern Anal. Mach. Intell. 2007 Image sequence enhancement using sub-pixel displacements · CVPR 1988 |
Computer vision › Video understanding and tracking
motion analysis |
0.1 | 1 | 2016 | Fundamental Matrices from Moving Objects Using Line Motion Barcodes · ECCV (2) 2016 |
Computational photography and imaging › image stitching
video mosaicing |
0.1 | 2 | 2005 | Dynamosaics: Video Mosaics with Non-Chronological Time · CVPR (1) 2005 Universal Mosaicing using Pipe Projection · ICCV 1998 |
Virtual and augmented reality › immersive video
360-degree video |
0.1 | 1 | 2007 | Dynamosaicing: Mosaicing of Dynamic Scenes · IEEE Trans. Pattern Anal. Mach. Intell. 2007 |
Visual content generation and editing
video editing |
0.1 | 1 | 2007 | Dynamosaicing: Mosaicing of Dynamic Scenes · IEEE Trans. Pattern Anal. Mach. Intell. 2007 |
Image and video processing › gradient-domain image processing
gradient-domain blending |
0.1 | 1 | 2006 | Seamless image stitching by minimizing false edges · IEEE Trans. Image Process. 2006 |
Methods — techniques the papers use, named apart from their topics
reconstruction error · 1.0predictability error · 1.0gait recognition · 0.5camera motion analysis · 0.5cumulative displacement curves · 0.4graph cuts · 0.3temporal signature · 0.2neural network · 0.2motion barcode · 0.2line motion barcodes · 0.2energy minimization · 0.2smoothness term · 0.2hierarchical optimization · 0.2activity condensation · 0.2object-based video representation · 0.1video stream processing · 0.1object detection · 0.1max-flow · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Audio Enhancement from Multiple Crowdsourced Recordings: A Simple and Effective Baseline
Shiran Aziz, Yossi Adi, Shmuel Peleg |
INTERSPEECH | 3 |
| 2022 | Deep Audio Waveform PriorabstractConvolutional neural networks contain strong priors for generating natural looking images [1]. These priors enable image denoising, super resolution, and inpainting in an unsupervised manner. Previous attempts to demonstrate similar ideas in audio, namely deep audio priors, (i) use hand picked architectures such as harmonic convolutions, (ii) only work with spectrogram input, and (iii) have been used mostly for eliminating Gaussian noise [2]. In this work we show that existing SOTA architectures for audio source separation contain deep priors even when working with the raw waveform. Deep priors can be discovered by training a neural network to generate a single corrupted signal when given white noise as input. A network with relevant deep priors is likely to generate a cleaner version of the signal before converging on the corrupted signal. We demonstrate this restoration effect with several corruptions: background noise, reverberations, and a gap in the signal (audio inpainting). Arnon Turetzky, Tzvi Michelson, Yossi Adi, Shmuel Peleg |
INTERSPEECH | 4 |
| 2021 | Crypto-Oriented Neural Architecture DesignabstractSending private data to Neural Network applications raises many privacy concerns. The cryptography community developed a variety of secure computation methods to address such privacy issues. As generic techniques for secure computation are typically prohibitively expensive, efforts focus on optimizing these cryptographic tools. Differently, we propose to optimize the design of crypto-oriented neural architectures, introducing a novel Partial Activation layer. The proposed layer is much faster for secure computation as it contains fewer non linear computations. Evaluating our method on three state-of-the-art architectures (SqueezeNet, ShuffleNetV2, and MobileNetV2) demonstrates significant improvement to the efficiency of secure inference on common evaluation metrics. Avital Shafran, Gil Segev 0001, Shmuel Peleg, Yedid Hoshen |
ICASSP | 3 |
| 2021 | Membership Inference Attacks are Easier on Difficult ProblemsabstractMembership inference attacks (MIA) try to detect if data samples were used to train a neural network model, e.g. to detect copyright abuses. We show that models with higher dimensional input and output are more vulnerable to MIA, and address in more detail models for image translation and semantic segmentation, including medical image segmentation. We show that reconstruction-errors can lead to very effective MIA attacks as they are indicative of memorization. Unfortunately, reconstruction error alone is less effective at discriminating between non-predictable images used in training and easy to predict images that were never seen before. To overcome this, we propose using a novel predictability error that can be computed for each sample, and its computation does not require a training set. Our membership error, obtained by subtracting the predictability error from the reconstruction error, is shown to achieve high MIA accuracy on an extensive number of benchmarks.1 Avital Shafran, Shmuel Peleg, Yedid Hoshen |
ICCV | 2 |
| 2019 | Dynamic Temporal Alignment of Speech to LipsabstractMany speech segments in movies are re-recorded in a studio during post-production, to compensate for poor sound quality as recorded on location. We present an audio-to-video method for automating speech to lips alignment, stretching and compressing the audio signal to match the lip movements. This alignment is based on deep audio-visual features, mapping the lips video and the speech signal to a shared representation. Using this representation we compute the lip-sync error between every short speech period and every video frame, followed by the determination of the optimal corresponding frame for each short sound period over the entire video clip. We demonstrate successful alignment both quantitatively, using a human perception-inspired metric, as well as qualitatively. The strongest advantage of our audio-to-video approach is in cases where the original voice in unclear. In these cases state-of-the-art audio only methods will fail. Tavi Halperin, Ariel Ephrat, Shmuel Peleg |
ICASSP | 3 |
| 2018 | Seeing Through Noise: Visually Driven Speaker Separation And EnhancementabstractIsolating the voice of a specific person while filtering out other voices or background noises is challenging when video is shot in noisy environments. We propose audio-visual methods to isolate the voice of a single speaker and eliminate unrelated sounds. First, face motions captured in the video are used to estimate the speaker's voice, by passing the silent video frames through a video-to-speech neural network-based model. Then the speech predictions are applied as a filter on the noisy input audio. This approach avoids using mixtures of sounds in the learning process, as the number of such possible mixtures is huge, and would inevitably bias the trained model. We evaluate our method on two audio-visual datasets, GRID and TCD-TIMIT, and show that our method attains significant SDR and PESQ improvements over the raw video-to-speech predictions, and a well-known audio-only method. Aviv Gabbay, Ariel Ephrat, Tavi Halperin, Shmuel Peleg |
ICASSP | 4 |
| 2018 | Visual Speech EnhancementabstractWhen video is shot in noisy environment, the voice of a speaker seen in the video can be enhanced using the visible mouth movements, reducing background noise. While most existing methods use audio-only inputs, improved performance is obtained with our visual speech enhancement, based on an audio-visual neural network. We include in the training data videos to which we added the voice of the target speaker as background noise. Since the audio input is not sufficient to separate the voice of a speaker from his own voice, the trained model better exploits the visual input and generalizes well to different noise types. The proposed model outperforms prior audio visual methods on two public lipreading datasets. It is also the first to be demonstrated on a dataset not designed for lipreading, such as the weekly addresses of Barack Obama. Aviv Gabbay, Asaph Shamir, Shmuel Peleg |
INTERSPEECH | 3 |
| 2018 | EgoSampling: Wide View Hyperlapse From Egocentric VideosabstractThe possibility of sharing one's point of view makes the use of wearable cameras compelling. These videos are often long, boring, and coupled with extreme shaking, as the camera is worn on a moving person. Fast-forwarding (i.e., frame sampling) is a natural choice for quick video browsing. However, this accentuates the shake caused by natural head motion in an egocentric video, making the fast-forwarded video useless. We propose EgoSampling, an adaptive frame sampling that gives stable, fast-forwarded, hyperlapse videos. Adaptive frame sampling is formulated as an energy minimization problem, whose optimal solution can be found in polynomial time. We further turn the camera shake from a drawback into a feature, enabling the increase in field of view of the output video. This is obtained when each output frame is mosaiced from several input frames. The proposed technique also enables the generation of a single hyperlapse video from multiple egocentric videos, allowing even faster video consumption. Tavi Halperin, Yair Poleg, Chetan Arora 0001, Shmuel Peleg |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | Vid2speech: Speech reconstruction from silent videoabstractSpeechreading is a notoriously difficult task for humans to perform. In this paper we present an end-to-end model based on a convolutional neural network (CNN) for generating an intelligible acoustic speech signal from silent video frames of a speaking person. The proposed CNN generates sound features for each frame based on its neighboring frames. Waveforms are then synthesized from the learned speech features to produce intelligible speech. We show that by leveraging the automatic feature learning capabilities of a CNN, we can obtain state-of-the-art word intelligibility on the GRID dataset, and show promising results for learning out-of-vocabulary (OOV) words. Ariel Ephrat, Shmuel Peleg |
ICASSP | 2 |
| 2016 | Visual Learning of Arithmetic OperationabstractA simple Neural Network model is presented for end-to-end visual learning of arithmetic operations from pictures of numbers. The input consists of two pictures, each showing a 7-digit number. The output, also a picture, displays the number showing the result of an arithmetic operation (e.g., addition or subtraction) on the two input numbers. The concepts of a number, or of an operator, are not explicitly introduced. This indicates that addition is a simple cognitive task, which can be learned visually using a very small number of neurons. Other operations, e.g., multiplication, were not learnable using this architecture. Some tasks were not learnable end-to-end (e.g., addition with Roman numerals), but were easily learnable once broken into two separate sub-tasks: a perceptual Character Recognition and cognitive Arithmetic sub-tasks. This indicates that while some tasks may be easily learnable end-to-end, other may need to be broken into sub-tasks. Yedid Hoshen, Shmuel Peleg |
AAAI | 2 |
| 2016 | Camera Calibration from Dynamic Silhouettes Using Motion BarcodesabstractComputing the epipolar geometry between cameras with very different viewpoints is often problematic as matching points are hard to find. In these cases, it has been proposed to use information from dynamic objects in the scene for suggesting point and line correspondences. We propose a speed up of about two orders of magnitude, as well as an increase in robustness and accuracy, to methods computing epipolar geometry from dynamic silhouettes. This improvement is based on a new temporal signature: motion barcode for lines. Motion barcode is a binary temporal sequence for lines, indicating for each frame the existence of at least one foreground pixel on that line. The motion barcodes of two corresponding epipolar lines are very similar, so the search for corresponding epipolar lines can be limited only to lines having similar barcodes. The use of motion barcodes leads to increased speed, accuracy, and robustness in computing the epipolar geometry. Gil Ben-Artzi, Yoni Kasten, Shmuel Peleg, Michael Werman |
CVPR | 3 |
| 2016 | An Egocentric Look at Video Photographer IdentityabstractEgocentric cameras are being worn by an increasing number of users, among them many security forces worldwide. GoPro cameras already penetrated the mass market, reporting substantial increase in sales every year. As headworn cameras do not capture the photographer, it may seem that the anonymity of the photographer is preserved even when the video is publicly distributed. We show that camera motion, as can be computed from the egocentric video, provides unique identity information. The photographer can be reliably recognized from a few seconds of video captured when walking. The proposed method achieves more than 90% recognition accuracy in cases where the random success rate is only 3%. Applications can include theft prevention by locking the camera when not worn by its rightful owner. Searching video sharing services (e.g. YouTube) for egocentric videos shot by a specific photographer may also become possible. An important message in this paper is that photographers should be aware that sharing egocentric video will compromise their anonymity, even when their face is not visible. Yedid Hoshen, Shmuel Peleg |
CVPR | 2 |
| 2016 | Fundamental Matrices from Moving Objects Using Line Motion Barcodes
Yoni Kasten, Gil Ben-Artzi, Shmuel Peleg, Michael Werman |
ECCV (2) | 3 |
| 2016 | Epipolar geometry based on line similarityabstractIt is known that epipolar geometry can be computed from three epipolar line correspondences but this computation is rarely used in practice since there are no simple methods to find corresponding lines. Instead, methods for finding corresponding points are widely used. This paper proposes a similarity measure between lines that indicates whether two lines are corresponding epipolar lines and enables finding epipolar line correspondences as needed for the computation of epipolar geometry. A similarity measure between two lines, suitable for video sequences of a dynamic scene, has been previously described. This paper suggests a stereo matching similarity measure suitable for images. It is based on the quality of stereo matching between the two lines, as corresponding epipolar lines yield a good stereo correspondence. Instead of an exhaustive search over all possible pairs of lines, the search space is substantially reduced when two corresponding point pairs are given. We validate the proposed method using real-world images and compare it to state-of-the-art methods. We found this method to be more accurate by a factor of five compared to the standard method using seven corresponding points and comparable to the 8-point algorithm. Gil Ben-Artzi, Tavi Halperin, Michael Werman, Shmuel Peleg |
ICPR | 4 |
| 2016 | Compact CNN for indexing egocentric videosabstractWhile egocentric video is becoming increasingly popular, browsing it is very difficult. In this paper we present a compact 3D Convolutional Neural Network (CNN) architecture for long-term activity recognition in egocentric videos. Recognizing long-term activities enables us to temporally segment (index) long and unstructured egocentric videos. Existing methods for this task are based on hand tuned features derived from visible objects, location of hands, as well as optical flow. Given a sparse optical flow volume as input, our CNN classifies the camera wearer's activity. We obtain classification accuracy of 89%, which outperforms the current state-of-the-art by 19%. Additional evaluation is performed on an extended egocentric video dataset, classifying twice the amount of categories than current state-of-the-art. Furthermore, our CNN is able to recognize whether a video is egocentric or not with 99.2% accuracy, up by 24% from current state-of-the-art. To better understand what the network actually learns, we propose a novel visualization of CNN kernels as flow fields. Yair Poleg, Ariel Ephrat, Shmuel Peleg, Chetan Arora 0001 |
WACV | 3 |
| 2015 | EgoSampling: Fast-forward and stereo for egocentric videosabstractWhile egocentric cameras like GoPro are gaining popularity, the videos they capture are long, boring, and difficult to watch from start to end. Fast forwarding (i.e. frame sampling) is a natural choice for faster video browsing. However, this accentuates the shake caused by natural head motion, making the fast forwarded video useless. We propose EgoSampling, an adaptive frame sampling that gives more stable fast forwarded videos. Adaptive frame sampling is formulated as energy minimization, whose optimal solution can be found in polynomial time. In addition, egocentric video taken while walking suffers from the left-right movement of the head as the body weight shifts from one leg to another. We turn this drawback into a feature: Stereo video can be created by sampling the frames from the left most and right most head positions of each step, forming approximate stereo-pairs. Yair Poleg, Tavi Halperin, Chetan Arora 0001, Shmuel Peleg |
CVPR | 4 |
| 2015 | Event retrieval using motion barcodesabstractWe introduce a simple and effective method for retrieval of videos showing a specific event, even when the videos of that event were captured from significantly different viewpoints. Appearance-based methods fail in such cases, as appearances change with large changes of viewpoints. Our method is based on a pixel-based feature, “motion barcode”, which records the existence/non-existence of motion as a function of time. While appearance, motion magnitude, and motion direction can vary greatly between disparate viewpoints, the existence of motion is viewpoint invariant. Based on the motion barcode, a similarity measure is developed for videos of the same event taken from very different viewpoints. This measure is robust to occlusions common under different viewpoints, and can be computed efficiently. Event retrieval is demonstrated using challenging videos from stationary and hand held cameras. Gil Ben-Artzi, Michael Werman, Shmuel Peleg |
ICIP | 3 |
| 2015 | Live video synopsis for multiple camerasabstractVideo surveillance cameras generate most of recorded video, and there is far more recorded video than operators can watch. Much progress has recently been made using summarization of recorded video, but such techniques do not have much impact on live video surveillance. We assume a camera hierarchy where a Master camera observes the decision-critical region, and one or more Slave cameras observe regions where past activity is important for making the current decision. We propose that when people appear in the live Master camera, the Slave cameras will display their past activities, and the operator could use past information for real-time decision making. The basic units of our method are action tubes, representing objects and their trajectories over time. Our object-based method has advantages over frame based methods, as it can handle multiple people, multiple activities for each person, and can address re-identification uncertainty. Yedid Hoshen, Shmuel Peleg |
ICIP | 2 |
| 2015 | The Information in Temporal HistogramsabstractIn many circumstances the limitation for use of video cameras is energy. The energy needed for compression and transmission of video is substantial, and is linear with the number of transmitted frames. Time-lapse photography, a drastic reduction of transmitted frame rate, is an obvious solution, say by transmitting one frame every several minutes. The temporal resolution of the video is lost. Can we reduce the number of transmitted frames but still keep some information in the original frame rate? In this work we examine a new paradigm for static cameras, the Histogram Camera. Frames are examined (but not coded or transmitted) at video frame rate, and for each pixel a temporal histogram of the intensity values is maintained. These temporal histograms, one per pixel, are transmitted at the reduced frame rates. It is shown that objects that change status from moving to stationary or vice versa can be extracted from the pixel-wise temporal histograms at high temporal accuracy. A storyboard summary of the video between frame transmissions can be generated. In addition, objects extracted from temporal histograms enable both background reconstruction, and matching across cameras with very different viewpoints. These benefits suggests that the Histogram Camera may be an important part of future very low frame rate cameras. Yedid Hoshen, Shmuel Peleg |
WACV | 2 |
| 2014 | Head Motion Signatures from Egocentric Videos
Yair Poleg, Chetan Arora 0001, Shmuel Peleg |
ACCV (3) | 3 |
| 2014 | Scene geometry from moving objectsabstractIt has been observed that in most videos recorded by surveillance cameras the image size of an object is a linear function of the y coordinate of its image location. This simple linear relationship holds in the most common surveillance camera configurations, where objects move on a planar surface and the camera's X axis is parallel to that plane. This linear relationship enables us to easily perform and enhance several geometric tasks based on tracking an object over a few frames: (i) computing the horizon; (ii) computing the relative real world sizes of objects in the scene based on their image appearance; (iii) improving tracking by constraining an object's location and size. When the the camera's X axis is not parallel to the ground plane, after tracking a couple of objects it is possible to find the rotation which rectifies the video so that its new X axis is parallel to the ground plane. Eitan Richardson, Shmuel Peleg, Michael Werman |
AVSS | 2 |
| 2014 | Temporal Segmentation of Egocentric VideosabstractThe use of wearable cameras makes it possible to record life logging egocentric videos. Browsing such long unstructured videos is time consuming and tedious. Segmentation into meaningful chapters is an important first step towards adding structure to egocentric videos, enabling efficient browsing, indexing and summarization of the long videos. Two sources of information for video segmentation are (i) the motion of the camera wearer, and (ii) the objects and activities recorded in the video. In this paper we address the motion cues for video segmentation. Motion based segmentation is especially difficult in egocentric videos when the camera is constantly moving due to natural head movement of the wearer. We propose a robust temporal segmentation of egocentric videos into a hierarchy of motion classes using a new Cumulative Displacement Curves. Unlike instantaneous motion vectors, segmentation using integrated motion vectors performs well even in dynamic and crowded scenes. No assumptions are made on the underlying scene structure and the method works in indoor as well as outdoor situations. We demonstrate the effectiveness of our approach using publicly available videos as well as choreographed videos. We also suggest an approach to detect the fixation of wearer's gaze in the walking portion of the egocentric videos. Yair Poleg, Chetan Arora 0001, Shmuel Peleg |
CVPR | 3 |
| 2013 | Efficient representation of distributions for background subtractionabstractMulti dimensional probability distributions are used in many surveillance tasks such as modeling color distribution of background pixels for Background Subtraction. Accurate representation of such distributions, e.g. in a histogram, requires much memory that may not be available when a histogram is computed for each pixel. Parametric representations such as Gaussian Mixture Models (GMM) are very efficient in memory but may not be accurate enough when the distribution is not from the assumed model. We propose a memory efficient representation for distributions. Histograms cells usually have equal width, and count the hits in each cell (Equi-width histograms). In most cases a 1D distribution can be represented more efficiently when cell sizes change so that each cell will have same number of hits (Equi-depth histograms). We propose to describe compactly multi-dimensional distributions (e.g. color) using an equi-depth histograms. Online computation of such histograms is described, and examples are given for background subtraction. Yedid Hoshen, Chetan Arora 0001, Yair Poleg, Shmuel Peleg |
AVSS | 4 |
| 2013 | Keynote lecture 2: "Video synopsis"abstractSummary form only given. Surveillance video is practically never used. There are claims that 0.5% of the video is watched, but the true number is probably much smaller. The reason is clear: There are too many hours of surveillance video for people to watch. Most attempts to deal with the overflow of surveillance video involve the development of automatic video understanding: object recognition and activity understanding. Video Synopsis is complementary to video understanding. After objects are detected by background subtraction, video synopsis changes the time of display of each object such that more objects are "packed" into a shorter time. The resulting video is a shorter summary of the original video, where the objects are shown more densely than in the original video. While video synopsis can reduce, on the average, an hour of video into a minute, the synopsis loses causality: Objects that appear together in the original video may appear at different time in the synopsis, and vice versa. The combination of video synopsis and video understanding is expected to give the maximum benefit. As video understanding is still not fool proof, people need to examine its results. Since the video showing all objects of interest will be too long, video synopsis is an excellent tool to display efficiently the results of video understanding, for video examination and even for training classifiers. Shmuel Peleg |
AVSS | 1 |
| 2012 | Alignment and mosaicing of non-overlapping imagesabstractImage alignment and mosaicing are usually performed on a set of overlapping images, using features in the area of overlap for alignment and for seamless stitching. Without image overlap current methods are helpless, and this is the case we address in this paper. So if a traveler wants to create a panoramic mosaic of a scene from pictures he has taken, but realizes back home that his pictures do not overlap, there is still hope. The proposed process has three stages: (i) Images are extrapolated beyond their original boundaries, hoping that the extrapolated areas will cover the gaps between them. This extrapolation becomes more blurred as we move away from the original image. (ii) The extrapolated images are aligned and their relative positions recovered. (iii) The gaps between the images are inpainted to create a seamless mosaic image. Yair Poleg, Shmuel Peleg |
ICCP | 2 |
| 2011 | Real-Time Stereo Mosaicing Using Feature TrackingabstractReal-time creation of video mosaics needs fast and accurate motion computation. While most mosaicing methods can use 2D image motion, the creation of multi view stereo mosaics needs more accurate 3D motion computation. Fast and accurate computation of 3D motion is challenging in the case of unstabilized cameras moving in 3D scenes, which is always the case when stereo mosaics are used. Efficient blending of the mosaic strip is also essential. Most cases of stereo mosaicing satisfy the assumption of limited camera motion, with no forward motion and no change in internal parameters. Under these assumptions uniform sideways motion creates straight epipolar lines. When the 3D motion is computed correctly, images can be aligned in space-time volume to give straight epipolar lines, a method which is depth invariant. We propose to align the video sequence in a space-time volume based on efficient feature tracking, and in this paper we used Kernel Tracking. Computation is fast as the motion in computed only for a few regions of the image, yet giving accurate 3D motion. This computation is faster and more accurate than the previously used direct approach. We also present "Barcode Blending", a new approach for using pyramid blending in video mosaics, which is very efficient. Barcode Blending overcomes the complexity of building pyramids for multiple narrow strips, combining all strips in a single blending step. The entire stereo mosaicing process is highly efficient in computation and in memory, and can be performed on mobile devices. Marc Vivet, Shmuel Peleg, Xavier Binefa |
ISM | 2 |
| 2010 | Identifying Surprising Events in Videos Using Bayesian Topic Models
Avishai Hendel, Daphna Weinshall, Shmuel Peleg |
ACCV (3) | 3 |
| 2010 | Bacteria-Filters: Persistent particle filters for background subtractionabstractMoving objects are usually detected by measuring the appearance change from a background model. The background model should adapt to slow changes such as illumination, but detect faster changes caused by moving objects. Particle filters do an excellent task in modeling non parametric distributions as needed for a background model, but may adapt too quickly to the foreground objects. A persistent particle filter is proposed, following bacterial persistence. Bacterial persistence is linked to the random switch of bacteria between two states: A normal growing cell and a dormant but persistent cell. The dormant cells can survive stress such as antibiotics. When a dormant cell switches to a normal status after the stress is over, bacterial growth continues. Similar to bacteria, particles will switch between dormant and active states, where dormant particles will not adapt to the changing environment. A further modification of particle filters allows discontinuous jumps into new parameters enabling foreground objects to join the background when they stop moving. This can also quickly build multi-modal distributions. Yair Movshovitz-Attias, Shmuel Peleg |
ICIP | 2 |
| 2009 | Clustered Synopsis of Surveillance VideoabstractMillions of surveillance cameras record video around the clock, producing huge video archives. Even when a video archive is known to include critical activities, finding them is like finding a needle in a haystack, making the archive almost worthless. Two main approaches were proposed to address this problem: action recognition and video summarization. Methods for automatic detection of activities still face problems in many scenarios. The video synopsis approach to video summarization is very effective, but may produce confusing summaries by the simultaneous display of multiple activities.A new methodology for the generation of short and coherent video summaries is presented, based on clustering of similar activities. Objects with similar activities are easy to watch simultaneously, and outliers can be spotted instantly. Clustered synopsis is also suitable for efficient creation of ground truth data. Yael Pritch, Sarit Ratovitch, Avishai Hendel, Shmuel Peleg |
AVSS | 4 |
| 2009 | Shift-map image editingabstractGeometric rearrangement of images includes operations such as image retargeting, inpainting, or object rearrangement. Each such operation can be characterized by a shiftmap: the relative shift of every pixel in the output image from its source in an input image. We describe a new representation of these operations as an optimal graph labeling, where the shift-map represents the selected label for each output pixel. Two terms are used in computing the optimal shift-map: (i) A data term which indicates constraints such as the change in image size, object rearrangement, a possible saliency map, etc. (ii) A smoothness term, minimizing the new discontinuities in the output image caused by discontinuities in the shift-map. This graph labeling problem can be solved using graph cuts. Since the optimization is global and discrete, it outperforms state of the art methods in most cases. Efficient hierarchical solutions for graph-cuts are presented, and operations on 1M images can take only a few seconds. Yael Pritch, Eitam Kav-Venaki, Shmuel Peleg |
ICCV | 3 |
| 2008 | Minimal Aspect Distortion (MAD) Mosaicing of Long Scenes
Alex Rav-Acha, Giora Engel, Shmuel Peleg |
Int. J. Comput. Vis. | 3 |
| 2008 | Nonchronological Video Synopsis and IndexingabstractThe amount of captured video is growing with the increased numbers of video cameras, especially the increase of millions of surveillance cameras that operate 24 hours a day. Since video browsing and retrieval is time consuming, most captured video is never watched or examined. Video synopsis is an effective tool for browsing and indexing of such a video. It provides a short video representation, while preserving the essential activities of the original video. The activity in the video is condensed into a shorter period by simultaneously showing multiple activities, even when they originally occurred at different times. The synopsis video is also an index into the original video by pointing to the original time of each activity. Video Synopsis can be applied to create a synopsis of an endless video streams, as generated by webcams and by surveillance cameras. It can address queries like "Show in one minute the synopsis of this camera broadcast during the past day''. This process includes two major phases: (i) An online conversion of the endless video stream into a database of objects and activities (rather than frames). (ii) A response phase, generating the video synopsis as a response to the user's query. Yael Pritch, Alex Rav-Acha, Shmuel Peleg |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2007 | Webcam Synopsis: Peeking Around the WorldabstractThe world is covered with millions of Webcams, many transmit everything in their field of view over the Internet 24 hours a day. A Web search finds public webcams in airports, intersections, classrooms, parks, shops, ski resorts, and more. Even more private surveillance cameras cover many private and public facilities. Webcams are an endless resource, but most of the video broadcast will be of little interest due to lack of activity. We propose to generate a short video that will be a synopsis of an endless video streams, generated by webcams or surveillance cameras. We would like to address queries like "I would like to watch in one minute the highlights of this camera broadcast during the past day". The process includes two major phases: (i) An online conversion of the video stream into a database of objects and activities (rather than frames), (ii) A response phase, generating the video synopsis as a response to the user's query. To include maximum information in a short synopsis we simultaneously show activities that may have happened at different times. The synopsis video can also be used as an index into the original video stream. Yael Pritch, Alex Rav-Acha, Avital Gutman, Shmuel Peleg |
ICCV | 4 |
| 2007 | Dynamosaicing: Mosaicing of Dynamic ScenesabstractThis paper explores the manipulation of time in video editing, enabling to control the chronological time of events. These time manipulations include slowing down (or postponing) some dynamic events while speeding up (or advancing) others. When a video camera scans a scene, aligning all the events to a single time interval will result in a panoramic movie. Time manipulations are obtained by first constructing an aligned space-time volume from the input video, and then sweeping a continuous 2D slice (time front) through that volume, generating a new sequence of images. For dynamic scenes, aligning the input video frames poses an important challenge. We propose to align dynamic scenes using a new notion of "dynamics constancy", which is more appropriate for this task than the traditional assumption of "brightness constancy". Another challenge is to avoid visual seams inside moving objects and other visual artifacts resulting from sweeping the space-time volumes with time fronts of arbitrary geometry. To avoid such artifacts, we formulate the problem of finding optimal time front geometry as one of finding a minimal cut in a 4D graph, and solve it using max-flow methods. Alex Rav-Acha, Yael Pritch, Dani Lischinski, Shmuel Peleg |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2006 | Making a Long Video Short: Dynamic Video SynopsisabstractThe power of video over still images is the ability to represent dynamic activities. But video browsing and retrieval are inconvenient due to inherent spatio-temporal redundancies, where some time intervals may have no activity, or have activities that occur in a small image region. Video synopsis aims to provide a compact video representation, while preserving the essential activities of the original video. We present dynamic video synopsis, where most of the activity in the video is condensed by simultaneously showing several actions, even when they originally occurred at different times. For example, we can create a "stroboscopic movie", where multiple dynamic instances of a moving object are played simultaneously. This is an extension of the still stroboscopic picture. Previous approaches for video abstraction addressed mostly the temporal redundancy by selecting representative key-frames or time intervals. In dynamic video synopsis the activity is shifted into a significantly shorter period, in which the activity is much denser. Video examples can be found online in http://www.vision.huji.ac.il/synopsis Alex Rav-Acha, Yael Pritch, Shmuel Peleg |
CVPR (1) | 3 |
| 2006 | Lucas-Kanade without Iterative WarpingabstractMany methods for motion computation and object tracking are based on the Lucas-Kanade (LK) framework. We present a method which substantially speeds up the LK approach while preserving its accuracy. This acceleration is obtained by avoiding the iterative image warping, inherent to the LK framework. A three-fold speedup is observed on standard image alignment tasks. Our second contribution focuses on adopting a multi-frame approach in order to increase alignment accuracy and robustness. By utilizing the acceleration procedure, the complexity of this multi-frame alignment becomes comparable to that of the two-frame approach. Alex Rav-Acha, Shmuel Peleg |
ICIP | 2 |
| 2006 | Seamless image stitching by minimizing false edgesabstractVarious applications such as mosaicing and object insertion require stitching of image parts. The stitching quality is measured visually by the similarity of the stitched image to each of the input images, and by the visibility of the seam between the stitched images. In order to define and get the best possible stitching, we introduce several formal cost functions for the evaluation of the stitching quality. In these cost functions the similarity to the input images and the visibility of the seam are defined in the gradient domain, minimizing the disturbing edges along the seam. A good image stitching will optimize these cost functions, overcoming both photometric inconsistencies and geometric misalignments between the stitched images. We study the cost functions and compare their performance for different scenarios both theoretically and practically. Our approach is demonstrated in various applications including generation of panoramic images, object blending and removal of compression artifacts. Comparisons with existing methods show the benefits of optimizing the measures in the gradient domain. Assaf Zomet, Anat Levin, Shmuel Peleg, Yair Weiss |
IEEE Trans. Image Process. | 3 |
| 2005 | Dynamosaics: Video Mosaics with Non-Chronological TimeabstractWith the limited field of view of human vision, our perception of most scenes is built over time while our eyes are scanning the scene. In the case of static scenes, this process can be modeled by panoramic mosaicing: stitching together images into a panoramic view. Can a dynamic scene, scanned by a video camera, be represented with a dynamic panoramic video even though different regions were visible at different times? In this paper, we explore time flow manipulation in video, such as the creation of new videos in which events that occurred at different times are displayed simultaneously. More general changes in the time flow are also possible, which enable re-scheduling the order of dynamic events in the video, for example. We generate dynamic mosaics by sweeping the aligned space-time volume of the input video by a time front surface and generating a sequence of time slices in the process. Various sweeping strategies and different time front evolutions manipulate the time flow in the video, enabling many unexplored and powerful effects, such as panoramic movies. Alex Rav-Acha, Yael Pritch, Dani Lischinski, Shmuel Peleg |
CVPR (1) | 4 |
| 2005 | Two motion-blurred images are better than one
Alex Rav-Acha, Shmuel Peleg |
Pattern Recognit. Lett. | 2 |
| 2004 | Seamless Image Stitching in the Gradient Domain
Anat Levin, Assaf Zomet, Shmuel Peleg, Yair Weiss |
ECCV (4) | 3 |
| 2004 | Fast panoramic stereo matching using cylindrical maximum surfacesabstractThis paper presents a fast panoramic stereo matching algorithm using a cylindrical maximum surface technique. The disparity for a pair of panoramic images is found in a cylindrical shaped correlation coefficient volume by obtaining the maximum surface rather than simply choosing a position that gives the maximum correlation coefficient value. The use of our cylindrical maximum surface technique ensures that the disparities obtained at the left and the right columns of the panoramic stereo images are properly constrained. Typical running time for a pair of 1324 x 120 images is about 0.33 s on a 1.7-GHz PC. A variety of real images have been tested, and good results have been obtained. Changming Sun, Shmuel Peleg |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2003 | Mosaicing New Views: The Crossed-Slits ProjectionabstractWe introduce anew kind of mosaicing, where the position of the sampling strip varies as a function of the input camera location. The new images that are generated this way correspond to a new projection model defined by two slits, termed here the Crossed-Slits (X-Slits) projection. In this projection model, every 3D point is projected by a ray defined as the line that passes through that point and intersects the two slits. The intersection of the projection rays with the imaging surface defines the image. X-Slits mosaicing provides two benefits. First, the generated mosaics are closer to perspective images than traditional pushbroom mosaics. Second, by simple manipulations of the strip sampling function, we can change the location of one of the virtual slits, providing a virtual walkthrough of a X-Slits camera; all this can be done without recovering any 3D geometry and without calibration. A number of examples where we translate the virtual camera and change its orientation are given; the examples demonstrate realistic changes in parallax, reflections, and occlusions. Assaf Zomet, Doron Feldman, Shmuel Peleg, Daphna Weinshall |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2002 | Multi-sensor Super-ResolutionabstractImage sensing is usually done with multiple sensors, like the RGB sensors in color imaging, the IR and EO sensors in surveillance and satellite imaging, etc. The resolution of each sensor can be increased by considering the images of the other sensors, and using the statistical redundancy among the sensors. Particularly, we use the fact that most discontinuities in the image of one sensor correspond to discontinuities in the other sensors. Two applications are presented: Increasing the resolution of a single color image by using the correlation among the three color channels, and enhancing noisy IR images. Assaf Zomet, Shmuel Peleg |
WACV | 2 |
| 2001 | Robust Super-ResolutionabstractA robust approach for super-resolution is, presented, which is especially valuable in the presence of outliers. Such outliers may be due to motion errors, inaccurate blur models, noise, moving objects, motion blur etc. This robustness is needed since super-resolution methods are very sensitive to such errors. A robust median estimator is combined in an iterative process to achieve a super resolution algorithm. This process can increase resolution even in regions with outliers, where other super resolution methods actually degrade the image. Assaf Zomet, Alex Rav-Acha, Shmuel Peleg |
CVPR (1) | 3 |
| 2001 | Omnistereo: Panoramic Stereo ImagingabstractAn omnistereo panorama consists of a pair of panoramic images, where one panorama is for the left eye and another panorama is for the right eye. The panoramic stereo pair provides a stereo sensation up to a full 360 degrees. Omnistereo panoramas can be constructed by mosaicing images from a single rotating camera. This approach also enables the control of stereo disparity, giving larger baselines for faraway scenes, and a smaller baseline for closer scenes. Capturing panoramic omnistereo images with a rotating camera makes it impossible to capture dynamic scenes at video rates and limits omnistereo imaging to stationary scenes. We present two possibilities for capturing omnistereo panoramas using optics without any moving parts. A special mirror is introduced such that viewing the scene through this mirror creates the same rays as those used with the rotating cameras. The lens used for omnistereo panorama is also introduced, together with the design of the mirror. Omnistereo panoramas can also be rendered by computer graphics methods to represent virtual environments. Shmuel Peleg, Moshe Ben-Ezra, Yael Pritch |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2000 | Cameras for Stereo Panoramic ImagingabstractA panorama for visual stereo consists of a pair of panoramic images, where one panorama is for the left eye, and another panorama is for the right eye. A panoramic stereo pair provides a stereo sensation lip to a full 360 degrees. A stereo panorama cannot be photographed by two omnidirectional cameras from two viewpoints. It is normally constructed by mosaicing together images from a rotating stereo pair, or from a single moving camera. Capturing stereo panoramic images by a rotating camera makes it impossible to capture dynamic scenes at video rates, and limits stereo panoramic imaging to stationary scenes. This paper presents two possibilities for capturing stereo panoramic images using optics, without any moving parts. A special mirror is introduced such that viewing the scene through this mirror creates the same rays as those used with the rotating cameras. Such a mirror enables the capture of stereo panoramic movies with a regular video camera. A lens for stereo panorama is also introduced. The designs of the mirror and of the lens are based on curves whose caustic is a circle. Shmuel Peleg, Yael Pritch, Moshe Ben-Ezra |
CVPR | 1 |
| 2000 | Rectified Mosaicing: Mosaics without the CurlabstractWhen images captured by a tilted camera are mosaiced into a panorama, the resulting mosaic is curled. This happens, for example, with a panning camera that is not perfectly horizontal, and with a translating camera facing a tilted planar surface. The tilt of the camera causes differences in image velocity between the top and bottom parts of the image, causing the curled mosaic. In rectified mosaicing these distortions are overcome by warping the strips into rectangles, while keeping some image feature invariant. This warping equalizes the image motion at the different image parts, and the resulting mosaic is straight. Mosaicing is done without camera calibration or knowledge of the scene, and the process adapts automatically to smooth changes in the scene and the imaging conditions. Shmuel Peleg, Assaf Zomet, Chetan Arora 0001 |
CVPR | 1 |
| 2000 | Model Based Pose Estimator Using Linear-Programming
Moshe Ben-Ezra, Shmuel Peleg, Michael Werman |
ECCV (1) | 2 |
| 2000 | Efficient Super-Resolution and Applications to MosaicsabstractMosaicing and super resolution are two ways to combine information from multiple frames in video sequences. Mosaicing displays the information of multiple frames in a single panoramic image. Super-resolution uses regions which appear in multiple frames to improve resolution and reduce noise. The aim of this work is constructing a high resolution mosaic from a video sequence in an efficient way. Simple combination of the two methods is problematic, since the alignment used in mosaicing may not be accurate enough for super resolution. Another issue is the efficiency of the super resolution algorithm, which requires heavy computations, especially when applied to large images such as panoramic mosaics. This paper introduces two novelties. First, a framework for super resolution algorithms is presented, which enables the development of very efficient algorithms. Second, a method for applying super resolution to panoramic mosaics is presented. This method preserves the geometry of the original mosaic image, while improving its resolution. Assaf Zomet, Shmuel Peleg |
ICPR | 2 |
| 2000 | Restoration of multiple images with motion blur in different directionsabstractImages degraded by motion blur can be restored when several blurred images are given, and the direction of motion blur in each image is different. Given two motion blurred images, best restoration is obtained when the directions of motion blur in the two images are orthogonal. Motion blur at different directions is common, for example, in the case of small hand-held digital cameras due to fast hand trembling and the light weight of the camera. Restoration examples are given on simulated data as well as on images with real motion blur. Alex Rav-Acha, Shmuel Peleg |
WACV | 2 |
| 2000 | Real-Time Motion Analysis with Linear Programming
Moshe Ben-Ezra, Shmuel Peleg, Michael Werman |
Comput. Vis. Image Underst. | 2 |
| 2000 | Mosaicing on Adaptive ManifoldsabstractImage mosaicing is commonly used to increase the visual field of view by pasting together many images or video frames. Existing mosaicing methods are based on projecting all images onto a predetermined single manifold: A plane is commonly used for a camera translating sideways, a cylinder is used for a panning camera, and a sphere is used for a camera which is both panning and tilting. While different mosaicing methods should therefore be used for different types of camera motion, more general types of camera motion, such as forward motion, are practically impossible for traditional mosaicing. A new methodology to allow image mosaicing in more general cases of camera motion is presented. Mosaicing is performed by projecting thin strips from the images onto manifolds which are adapted to the camera motion. While the limitations of existing mosaicing techniques are a result of using predetermined manifolds, the use of more general manifolds overcomes these limitations. Shmuel Peleg, Benny Rousso, Alex Rav-Acha, Assaf Zomet |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1999 | Stereo Panorama with a Single CameraabstractFull panoramic images, covering 360 degrees, can be created either by using panoramic cameras or by mosaicing together many regular images. Creating panoramic views in stereo, where one panorama is generated for the left eye, and another panorama is generated for the right eye is more problematic. Earlier attempts to mosaic images from a rotating pair of stereo cameras faced severe problems of parallax and of scale changes. A new family of multiple viewpoint image projections, the Circular Projections, is developed. Two panoramic images taken using such projections can serve as a panoramic stereo pair. A system is described to generates a stereo panoramic image using circular projections from images or video taken by a single rotating camera. The system works in real-time on a PC. It should be noted that the stereo images are created without computation of 3D structure, and the depth effects are created only in the viewer's brain. Shmuel Peleg, Moshe Ben-Ezra |
CVPR | 1 |
| 1999 | Real-Time Motion Analysis with Linear-ProgrammingabstractA method to compute motion models in real time from point-to-line correspondences using linear programming is presented. Point-to-line correspondences are the most reliable motion measurements given the aperture effect, and it is shown how they can approximate other motion measurements as well. Using an L/sub 1/ error measure for image alignment based on point-to-line correspondences and minimizing this measure using linear programming, achieves results which are more robust than the commonly used L/sub 2/ metric. While estimators based on L/sub 1/ are not theoretically robust, experiments show that the proposed method is robust enough to allow accurate motion recovery in hundreds of consecutive frames. The entire computation is performed in real-time on a PC with no special hardware. Moshe Ben-Ezra, Shmuel Peleg, Michael Werman |
ICCV | 2 |
| 1998 | Universal Mosaicing using Pipe ProjectionabstractVideo mosaicing is commonly used to increase the visual field by pasting together many video frames. Existing mosaicing methods are effective only in very limited cases where the image motion is almost a uniform translation or the camera performs a pure pan. Forward camera motion or camera zoom are very problematic for traditional mosaicing. A mosaicing methodology to allow image mosaicing in the most general cases is presented, where frames in the video sequence are transformed such that the optical flow becomes parallel. This transformation is an oblique projection of the image into a "viewing pipe" whose central axis is the trajectory of the camera. The "pipe projection" enables to define high quality mosaicing even for the most challenging cases of forward motion and of zoom. In addition view interpolation, generating dense intermediate views is used to overcome parallax effects. Benny Rousso, Shmuel Peleg, Ilan Finci, Alex Rav-Acha |
ICCV | 2 |
| 1998 | Efficient computation of the most probable motion from fuzzy correspondencesabstractAn algorithm is presented for finding the most probable image motion between two images from fuzzy point correspondences. In fuzzy correspondence a point in one image is assigned to a region in the other image. Such a region can be line (aperture effect) or a convex polygon. Noise and outliers are always present, and points may belong to different motions. The presented algorithm, which uses linear programming, recovers the motion parameters and performs outlier rejection and motion-segmentation at the same time. The linear program computes the global optimum without a need for initial guess. Moshe Ben-Ezra, Shmuel Peleg, Michael Werman |
WACV | 2 |
| 1998 | Applying super-resolution to panoramic mosaicsabstractMosaicing and super resolution are two ways to combine information from multiple frames in video sequences. Mosaicing displays the information of multiple frames in a single panoramic image. Super-resolution uses regions which appear in multiple frames to improve resolution and reduce noise. Simple combination of the two methods is problematic since the alignment used for mosaicing may not be accurate enough over the entire region for super resolution. In this paper we introduce a 2-step process: First we create a panoramic mosaic from the images. We then align small image regions to the panorama, and apply super resolution resulting in a panoramic image with higher resolution. Assaf Zomet, Shmuel Peleg |
WACV | 2 |
| 1997 | Video Mosaicing using Manifold Projection
Benny Rousso, Shmuel Peleg, Ilan Finci |
BMVC | 2 |
| 1997 | Panoramic mosaics by manifold projectionabstractAs the field of view of a picture is much smaller than our own visual field of view, it is common to paste together several pictures to create a panoramic mosaic having a larger field of view. Images with a wider field of view can be generated by using fish-eye lens, or panoramic mosaics can be created by special devices which rotate around the camera's optical center (Quicktime VR, Surround Video), or by aligning, and pasting, frames in a video sequence to a single reference frame. Existing mosaicing methods have strong limitations on imaging conditions, and distortions are common. Manifold projection enables the creation of panoramic mosaics from video sequences under more general conditions, and in particular the unrestricted motion of a hand-held camera. The panoramic mosaic is a projection of the scene into a virtual manifold whose structure depends on the camera's motion. This manifold is more general than the customary projections onto a single image plane or onto a cylinder. In addition to being more general than traditional mosaics, manifold projection is also computationally efficient, as the only image deformations used are image plane translations and rotations. Real-time, software only, implementation on a Pentium-PC, proves the superior quality and speed of this approach. Shmuel Peleg, Joshua Herman |
CVPR | 1 |
| 1997 | Recovery of Ego-Motion Using Region AlignmentabstractA method for computing the 3D camera motion (the ego-motion) in a static scene is described, where initially a detected 2D motion between two frames is used to align corresponding image regions. We prove that such a 2D registration removes all effects of camera rotation, even for those image regions that remain misaligned. The resulting residual parallax displacement field between the two region-aligned images is an epipolar field centered at the FOE (Focus-of-Expansion). The 3D camera translation is recovered from the epipolar field. The 3D camera rotation is recovered from the computed 3D translation and the detected 2D motion. The decomposition of image motion into a 2D parametric motion and residual epipolar parallax displacements avoids many of the inherent ambiguities and instabilities associated with decomposing the image motion into its rotational and translational components, and hence makes the computation of ego-motion or 3D structure estimation more robust. Michal Irani, Benny Rousso, Shmuel Peleg |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 1996 | Robust Recovery of Camera Rotation from Three FramesabstractComputing camera rotation from image sequences can be used for image stabilization, and when the camera rotation is known the computation of translation and scene structure are much simplified as well. A robust approach for recovering camera rotation is presented, which does not assume any specific scene structure (e.g. no planar surface is required), and which avoids prior computation of the epipole. Given two images taken from two different viewing positions, the rotation matrix between the images can be computed from any three homography matrices. The homographies are computed using the trilinear tensor which describes the relations between the projections of a 3D point into three images. The entire computation is linear for small angles, and is therefore fast and stable. Iterating the linear computation can then be used to recover larger rotations as well. Benny Rousso, Shai Avidan, Amnon Shashua, Shmuel Peleg |
CVPR | 4 |
| 1996 | Local Quantitative Measurements for Cardiac Motion Analysis
Serge Benayoun, Dany Kharitonsky, Avraham Zilberman, Shmuel Peleg |
ECCV (2) | 4 |
| 1995 | Symmetry as a Continuous FeatureabstractSymmetry is treated as a continuous feature and a continuous measure of distance from symmetry in shapes is defined. The symmetry distance (SD) of a shape is defined to be the minimum mean squared distance required to move points of the original shape in order to obtain a symmetrical shape. This general definition of a symmetry measure enables a comparison of the "amount" of symmetry of different shapes and the "amount" of different symmetries of a single shape. This measure is applicable to any type of symmetry in any dimension. The symmetry distance gives rise to a method of reconstructing symmetry of occluded shapes. The authors extend the method to deal with symmetries of noisy and fuzzy data. Finally, the authors consider grayscale images as 3D shapes, and use the symmetry distance to find the orientation of symmetric objects from their images, and to find locally symmetric regions in images. Hagit Hel-Or, Shmuel Peleg, David Avnir |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1994 | Recovery of ego-motion using image stabilizationabstractA method for computing the 3D camera motion (the ego-motion) in a static scene is introduced, which is based on computing the 2D image motion of a single image region directly from image intensities. The computed image motion of this image region is used to register the images so that the detected image region appears stationary. The resulting displacement field for the entire scene between the registered frames is affected only by the 3D translation of the camera. After canceling the effects of the camera rotation by using such 2D image registration, the 3D camera translation is computed by finding the focus-of-expansion in the translation-only set of registered frames. This step is followed by computing the camera rotation to complete the computation of the ego-motion. The presented method avoids the inherent problems in the computation of optical flow and of feature matching, and does not assume any prior feature detection or feature correspondence.> Michal Irani, Benny Rousso, Shmuel Peleg |
CVPR | 3 |
| 1994 | Accurate computation of optical flow by using layered motion representationsabstractThis paper presents a framework combining two prevailing approaches to motion analysis: optical flows which describes motion at each point, and methods that define global motions for larger regions. Image motion is represented by layers-image regions whose coherent motion can be approximated by some parametric motion model. The motion at every point is obtained by the parametric motion estimate of the entire layer, corrected by a residual flow which captures the difference between the real image motion and the layer's motion model. The new approach is able to construct accurate flow fields in the presence of multiple motions, motion boundaries, and transparent motions. Steven C. Hsu, P. Anandan 0001, Shmuel Peleg |
ICPR (1) | 3 |
| 1994 | Symmetry of fuzzy dataabstractSymmetry is usually viewed as a discrete feature: an object is either symmetric or non-symmetric. Following the view that symmetry is a continuous feature, a continuous symmetry measure (CSM) has been developed to evaluate symmetries of shapes and objects. In this paper the authors extend the symmetry measure to evaluate the imperfect symmetry of fuzzy shapes, i.e. shapes with uncertain point localization. The authors find the probability distribution of symmetry values for a given fuzzy shape. Additionally, for every such fuzzy shape, the authors find the most probable symmetric shape. Hagit Hel-Or, Shmuel Peleg, David Avnir |
ICPR (1) | 2 |
| 1994 | Computing occluding and transparent motions
Michal Irani, Benny Rousso, Shmuel Peleg |
Int. J. Comput. Vis. | 3 |
| 1993 | Robust Recovery of Ego-Motion
Michal Irani, Benny Rousso, Shmuel Peleg |
CAIP | 3 |
| 1993 | Completion of occluded shapes using symmetryabstractFollowing the view that symmetry is a continuous feature, a continuous symmetry measure (CSM) is developed to evaluate symmetrics of shapes and objects. The symmetry measure is extended to evaluate the symmetry of occluded shapes. Using the symmetry measure, occluded shapes are reconstructed by locating the center of symmetry of the shape.> Hagit Hel-Or, Shmuel Peleg, David Avnir |
CVPR | 2 |
| 1993 | Motion Analysis for Image Enhancement: Resolution, Occlusion, and Transparency
Michal Irani, Shmuel Peleg |
J. Vis. Commun. Image Represent. | 2 |
| 1992 | Image sequence enhancement using multiple motions analysisabstractA method for detecting and tracking multiple moving objects, using both a large spatial region and a large temporal region, without assuming temporal motion constancy is described. When the large spatial region of analysis has multiple moving objects, the motion parameters and the locations of the objects are computed for one object after another. A method for segmenting the image plane into differently moving objects and computing their motions using two frames is presented. The tracking of detected objects using temporal integration and the algorithms for enhancement of tracked objects by filling-in occluded regions and by improving the spatial resolution of the imaged objects are described.> Michal Irani, Shmuel Peleg |
CVPR | 2 |
| 1992 | A measure of symmetry based on shape similarityabstractThe authors view symmetry as a continuous feature, and define a continuous symmetry measure (CSM) of shapes. The general definition of symmetry measure allows a comparison of the amount of symmetry of different shapes and the amount of different symmetries of a single shape. Furthermore, the CSM is associated with the symmetric shape that is closest to the given one, enabling visual evaluation of the CSM.> Hagit Hel-Or, Shmuel Peleg, David Avnir |
CVPR | 2 |
| 1992 | Detecting and Tracking Multiple Moving Objects Using Temporal Integration
Michal Irani, Benny Rousso, Shmuel Peleg |
ECCV | 3 |
| 1992 | Surface reconstruction from derivativesabstractMost methods to reconstruct surfaces from their derivatives assume two orthogonal derivatives in perfect registration. The authors propose an approach to use derivatives in arbitrary directions and to register orthogonal derivatives if they are not registered. They also develop a method to integrate second derivatives in the reconstruction.> Ran Bronstein, Michael Werman, Shmuel Peleg |
ICPR (1) | 3 |
| 1992 | Hierarchical symmetryabstractThe authors view symmetry as a continuous feature and dependent on resolution. Combining a continuous symmetry measure (CSM) with a multiresolution scheme, the authors present a method that hierarchically detects symmetric and almost symmetric patterns. Evaluation of symmetry at low frequencies guides the process to find the symmetry at higher frequencies.> Hagit Hel-Or, Shmuel Peleg, David Avnir |
ICPR (3) | 2 |
| 1992 | A Three-Frame Algorithm for Estimating Two-Component Image MotionabstractA fundamental assumption made in formulating optical-flow algorithms, that motion at any point in an image can be represented as a single pattern component undergoing a simple translation, fails for a number of situations that commonly occur in real-world images. An alternative formulation of the local motion assumption in which there may be two distinct patterns undergoing coherent (e.g. affine) motion within a given local analysis region is proposed. An algorithm for the analysis of two-component motion in which tracking and nulling mechanisms applied to three consecutive image frames separate and estimate the individual components is given. Precise results are obtained, even for components that differ only slightly in velocity as well as for a faint component in the presence of a dominant, masking component. The algorithm provides precise motion estimates for a set of elementary two-motion configurations and is robust in the presence of noise.> James R. Bergen, Peter J. Burt, Rajesh Hingorani, Shmuel Peleg |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 1991 | Improving resolution by image registration
Michal Irani, Shmuel Peleg |
CVGIP Graph. Model. Image Process. | 2 |
| 1991 | Characterization of right-handed and left-handed shapes
Yacov Hel-Or, Shmuel Peleg, David Avnir |
CVGIP Image Underst. | 2 |
| 1990 | Transparent-Motion Analysis
James R. Bergen, Peter J. Burt, Rajesh Hingorani, Shmuel Peleg |
ECCV | 4 |
| 1990 | Computing two motions from three framesabstractA fundamental assumption made in formulating optical-flow algorithms is that motion at any point in any image can be represented as a single pattern undergoing a simple translation: even complex motion will appear as a uniform displacement when viewed through a sufficiently small window. This assumption fails in a number of common situations. The authors propose an alternative formulation in which there may be two distinct patterns undergoing coherent motion within a given local analysis region. They then present an algorithm for the analysis of two-component motion. They also demonstrate that the algorithm provides precise motion estimates for a set of elementary two-motion configurations, and show that it is robust in the presence of noise.> James R. Bergen, Peter J. Burt, Rajesh Hingorani, Shmuel Peleg |
ICCV | 4 |
| 1990 | Super resolution from image sequencesabstractAn iterative algorithm to increase image resolution is described. Examples are shown for low-resolution gray-level pictures, with an increase of resolution clearly observed after only a few iterations. The same method can also be used for deblurring a single blurred image. The approach is based on the resemblance of the presented problem to the reconstruction of a 2-D object from its 1-D projections in computer-aided tomography. The algorithm performed well for both computer-simulated and real images and is shown, theoretically and practically, to converge quickly. The algorithm can be executed in parallel for faster hardware implementation.> Michal Irani, Shmuel Peleg |
ICPR (2) | 2 |
| 1990 | Segmentation by minimum length encodingabstractA digitized waveform is approximated by segments whose total description length is minimal for a given error bound. This approximation can be computed efficiently and can be used for segmentation. Some applications involving the use of one-dimensional methods to segment two-dimensional gray-scale and range images are shown.> Daniel Keren, Ruth Marcus, Michael Werman, Shmuel Peleg |
ICPR (1) | 4 |
| 1990 | Motion based segmentationabstractAn iterative method is described for segmenting image sequences into independently moving regions while computing the motion parameters of each region. In each iteration, image points are classified into regions based on their consistency with the different motion estimates, and motion estimates are then updated using the obtained regions. The motion estimates and the segmentation improve with every iteration, and the iteration stops when a stable segmentation is obtained. Accurate motion parameters are recovered for each segment. The process is performed directly on gray-level images and does not require detection of special feature points and the computation of point correspondence. It is also faster and more robust than optical-flow-based segmentation methods.> Shmuel Peleg, Hillel Rom |
ICPR (1) | 1 |
| 1990 | Attentive transmission
Hagit Hel-Or, Shmuel Peleg |
J. Vis. Commun. Image Represent. | 2 |
| 1990 | Nonlinear Multiresolution: A Shape-From-Shading ExampleabstractA method for image resolution reduction that simulates the reduction of shape resolution and is to be used with shape-from-shading algorithms is presented. Resolution reduction is used for accelerating shape-from-shading algorithms using multiresolution pyramids. The goal of this correspondence is to show that resolution reduction based only on image intensities can be inferior to resolution reduction using knowledge of the surface reflectance. This is true for all cases of parameter estimation from intensity images, when the intensity is not a linear function of the parameters. Since optical reduction of images, as when moving a camera away from the object, is a simple gray-level averaging a question is cast on the general scheme of shape-from-shading. It is concluded that for accurate shape-from-shading one needs to know the distance from the object, and the shape should also be known in advance.> Shmuel Peleg, Gad Ron |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1990 | Stereo by Incremental Matching of ContoursabstractContours made of sequences of adjacent edge points are used as primitives in stereo pair matching. Matching contour segments, rather than the traditional epipolar edge points, can greatly reduce possible ambiguity. This is done by reformulating point-matching constraints to apply to contour matching, and by introducing a unique incremental matching scheme. Best-matched contours are paired first, constraining through neighborhood support their neighboring contours. Examples of the proposed stereo matching scheme are shown, with very few errors, for aerial images of natural terrain.> Doron Sherman, Shmuel Peleg |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1989 | Multiresolution shape from shadingabstractA method to reduce image resolution that simulates the reduction of shape resolution is presented. Such resolution reduction is critical for shape-from-shading algorithms. This method of resolution reduction is used in accelerating the shape from shading by using small multiresolution pyramids. A Lambertian reflectance model is used, but a similar approach can be developed for other reflectance models. The goal is to show that the resolution reduction based only on image intensities is inferior to resolution reduction using knowledge on the surface reflectance.> Gad Ron, Shmuel Peleg |
CVPR | 2 |
| 1989 | A Unified Approach to the Change of Resolution: Space and Gray-LevelabstractIt is shown that by defining a suitable measure for the comparison of images, changes in resolution can be treated with the same tool as changes in color resolution. A gray-tone image, for example, can be compared to a half-tone image having only two colors (black and white), but of higher spatial resolution. A graph-theoretical definition of the basic measure used is introduced. This is followed by application to spatial resampling and gray-level requantization. This results in a hybrid treatment of resolution, and the possibility of trading spatial for gray-level resolution and vice versa.> Shmuel Peleg, Michael Werman, Hillel Rom |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1988 | How to tell right from left [chirality for 2-D binary shapes]abstractThe authors study the notion of chirality for two-dimensional binary shapes, and introduce measures to test whether a shape is symmetric, and if not whether it is left-handed or right-handed. The measures are based on boundary analysis, and perform well even when digital images of left-handed shapes differ from the mirror images of right-handed shapes. Such situations may occur due to natural variations and digitization errors. The measures can also successfully treat partially occluded shapes, and provide indications on the change of chirality as resolution changes.> Yacov Hel-Or, Shmuel Peleg, Hagit Hel-Or |
CVPR | 2 |
| 1988 | Image sequence enhancement using sub-pixel displacementsabstractGiven a sequence of images taken from a moving camera, they are registered with subpixel accuracy in respect to translation and rotation. The subpixel registration allows image enhancement with respect to improved resolution and noise cleaning. Both the registration and the enhancement procedures are described. The methods are particularly useful for image sequences taken from an aircraft or satellite where images in a sequence differ mostly by translation and rotation. In these cases, the process results in images that are stable, clean, and sharp.> Danny Keren, Shmuel Peleg, Rafi Brada |
CVPR | 2 |
| 1988 | Image representation using Voronoi tessellation: adaptive and secureabstractAn image is represented by the Voronoi tessellation generated from selected sampling points. Using a multiresolution approach, the density of the sampling points can be adaptive to image properties: smoother regions will have fewer sampling points than more detailed regions. The adaptation property results in better image quality than nonadaptive Voronoi representations, while preserving the property that only the holder of the seed of the pseudorandom number generator can reconstruct the original image.> Hillel Rom, Shmuel Peleg |
CVPR | 2 |
| 1988 | Gray level requantization
Michael Werman, Shmuel Peleg |
Comput. Vis. Graph. Image Process. | 2 |
| 1987 | Improving image resolution using subpixel motion
Shmuel Peleg, Danny Keren, Limor Schweitzer |
Pattern Recognit. Lett. | 1 |
| 1987 | Inversion of picture operators
Haim Schweitzer, Shmuel Peleg |
Pattern Recognit. Lett. | 2 |
| 1987 | Representation of patterns of symbols by equations with applications to puzzle solving
Haim Schweitzer, Shmuel Peleg |
Pattern Recognit. Lett. | 2 |
| 1985 | A distance metric for multidimensional histograms
Michael Werman, Shmuel Peleg, Azriel Rosenfeld |
Comput. Vis. Graph. Image Process. | 2 |
| 1985 | Fuzzy and probability vectors as elements of a vector space
Haim Schweitzer, Shmuel Peleg |
Inf. Sci. | 2 |
| 1985 | Min-Max Operators in Texture AnalysisabstractA signature is generated for a given picture by operating on it with different masks. The operations are gray level generalizations of ``shrink'' and ``expand'' for binary pictures using Serra's morphological methods [10]. The signature is a set of numbers, each corresponding to an application of an operator at a certain scale and direction, and can be used to analyze the discriminate textures. It is shown that this family of signatures includes as special cases several currently used texture descriptors. Michael Werman, Shmuel Peleg |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1985 | Noisy image restoration by cost function minimization
Eli Harouche, Shmuel Peleg, Haim Schweitzer, Larry Davis 0001 |
Pattern Recognit. Lett. | 2 |
| 1984 | Classification by discrete optimization
Shmuel Peleg |
Comput. Vis. Graph. Image Process. | 1 |
| 1984 | Multiple Resolution Texture Analysis and ClassificationabstractTextures are classified based on the change in their properties with changing resolution. The area of the gray level surface is measured at serveral resolutions. This area decreases at coarser resolutions since fine details that contribute to the area disappear. Fractal properties of the picture are computed from the rate of this decrease in area, and are used for texture comparison and classification. The relation of a texture picture to its negative, and directional properties, are also discussed. Shmuel Peleg, Joseph Naor, Ralph Hartley, David Avnir |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1983 | Image Compression and Filtering Using Pyramid Data Structures
Joseph Naor, Shmuel Peleg |
IJCAI | 2 |
| 1983 | Hierarchical image representation for compression, filtering and normalization
Joseph Naor, Shmuel Peleg |
Pattern Recognit. Lett. | 2 |
| 1983 | Digital Image Compression by Outer Product ExpansionabstractWe approximate a digital image as a sum of outer products dxyTwheredis a real number but the vectorsxandyhave elements +1, -1, or 0 only. The expansion gives a least squares approximation. Work is proportional to the number of pixels; reconstruction involves only additions. Dianne P. O'Leary, Shmuel Peleg |
IEEE Trans. Commun. | 2 |
| 1983 | Analysis of relaxation processes: The two-node two-label caseabstractSeveral relaxation processes are analyzed in the simple case of two nodes, each having two possible labels. It is shown that the choice of coefficients is very important. For certain values of the coefficients, some processes will have a single nontrivial convergence point regardless of the initial labeling. For other choices of the coefficients, there can be more than one possible convergence point, and different solutions can be obtained for different initial labelings. In the probabilistic approach where the coefficients are predefined in terms of joint probabilities, there are always two nontrivial convergence points for all possible coefficients. The results are also compared to the Bayesian analysis that can be obtained in this simple case of two nodes. Since certain selections of coefficients can give unacceptable results even in this simple case, it can be expected that the proper selection of coefficients will be much more important in the general case involving larger numbers of nodes and labels. Dianne P. O'Leary, Shmuel Peleg |
IEEE Trans. Syst. Man Cybern. | 2 |
| 1983 | Classification and tracking using local optimizationabstractWhen a heuristic function is available to evaluate classification, a special search procedure is applied to find a classification optimizing this function. A specific application to image segmentation is presented, including several examples. The major difference between this approach and previous optimization attempts is the use of deterministic rather than probabilistic classifications. The approach is also applied to object tracking in image sequences. Shmuel Peleg, Allon Nathan |
IEEE Trans. Syst. Man Cybern. | 1 |
| 1982 | Image Smoothing and Segmentation by Multiresolution Pixel Linking: Further Experiments and ExtensionsabstractA recently developed method of image smoothing and segmentation makes use of a "pyramid" of images at successively lower resolutions. It establishes links between pixels at successive levels of the pyramid; the subtrees of the pyramid defined by these links yield a segmentation of the image into regions over which the smoothing takes place. This paper investigates several variations on the basic linking process with regard to such factors as initialization, criteria for linking, and iteration scheme used. It also studies generalizations in which the links are weighted rather than forced, and in which interactions among the pixels at a given level are also allowed. Finally, it extends the approach to links based on more than one feature of a pixel, e.g., on color components or local property values. Tsai Hong, K. A. Narayanan, Shmuel Peleg, Azriel Rosenfeld, Teresa M. Silberberg |
IEEE Trans. Syst. Man Cybern. | 3 |
| 1981 | A Min-Max Medial Axis TransformationabstractBlum's medial axis transformation (MAT) of the set S of 1's in a binary picture can be defined by an iterative shrinking and reexpanding process which detects ``corners'' on the contours of constant distance from S¿, and thereby yields a ``skeleton'' of S. For unsegmented (gray level) pictures, one can use an analogous definition, in which local MIN and MAX operations play the roles of shrinking and expanding, to compute a ``MMMAT value'' at each point of the picture. The set of points having high values defines a good ``skeleton'' for the set of high-gray level points in the given picture. Shmuel Peleg, Azriel Rosenfeld |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1981 | Shape Segmentation Using RelaxationabstractRelaxation is applied to the segmentation of closed boundary curves of shapes. The ambiguous segmentation of the boundary is represented by a directed graph structure whose nodes represent segments, where two nodes are joined by an arc if the segments are consecutive along the boundary. A probability vector is associated with each node; each component of this vector provides an estimate of the probability that the corresponding segment is a particular part of the object. Relaxation is used to eliminate impossible sequences of parts, or reduce the probabilities of unlikely ones. In experiments involving airplane shapes, this almost always results in a drastic simplification of the graph with only good interpretations surviving. The approach is also extended to include curve linking and gap filling. A chain coded input image is broken into segments based on a measure of local curvature. Gap completions linking pairs of segments are then proposed and represented in a graph structure. A second graph, whose nodes consist of paths in the above graph, is constructed, and the nodes of the second graph are probabilistically classified as various object parts. Relaxation is then applied to increase the probability of mutually supporting classifications, and decrease the probability of unsupported decisions. A modified relaxation process using information about the size, spatial position, and orientation of the object parts yielded a high degree of disambiguation. Wallace S. Rutkowski, Shmuel Peleg, Azriel Rosenfeld |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1980 | Labeling evaluation in probabilistic networks
Shmuel Peleg |
Inf. Sci. | 1 |
| 1980 | A New Probabilistic Relaxation SchemeabstractLet a vector of probabilities be associated with every node of a graph. These probabilities define a random variable representing the possible labels of the node. Probabilities at neighboring nodes are used iteratively to update the probabilities at a given node based on statistical relations among node labels. The results are compared with previous work on probabilistic relaxation labeling, and examples are given from the image segmentation domain. References are also given to applications of the new scheme in text processing. Shmuel Peleg |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1979 | Maximal Derivations for Probabilistic Strings in Stochastic Languages
Shmuel Peleg |
Inf. Control. | 1 |