Daphna Weinshall

dblp:93/1568 · DBLP profile ↗
← Back
95ranked-venue papers
24as first author
8since 2021 · last 2026
0000-0001-8893-8586ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 90 · 22 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 43 · 11 first-author · 2 since 2021Security and privacy · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
68 papers
Learning theory · 21% Efficient and distributed learning · 17% Kernel, tree and ensemble methods · 10%
Computer graphics and multimedia
18 papers
Image and video processing · 37% Multimedia analysis and retrieval · 26% Geometric modeling and processing · 14%
Databases, data mining, and information retrieval
9 papers
Data mining · 80% Information retrieval · 10% Machine learning and data management · 7%

Topics — the 30 heaviest of 141, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Learning theory
overfitting
2.632026
Forget Me Not: Fighting Local Overfitting With Knowledge Fusion and Distillation · IEEE Trans. Pattern Anal. Mach. Intell. 2026
On Local Overfitting and Forgetting in Deep Neural Networks · AAAI 2025
United We Stand: Using Epoch-Wise Agreement of Ensembles to Combat Overfit · AAAI 2024
Machine learning › Kernel, tree and ensemble methods
ensemble learning
1.932026
On Local Overfitting and Forgetting in Deep Neural Networks · AAAI 2025
United We Stand: Using Epoch-Wise Agreement of Ensembles to Combat Overfit · AAAI 2024
Forget Me Not: Fighting Local Overfitting With Knowledge Fusion and Distillation · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Machine learning › Efficient and distributed learning
active learning
1.222023
How to Select Which Active Learning Strategy is Best Suited for Your Specific Problem and Budget · NeurIPS 2023
Active Learning Through a Covering Lens · NeurIPS 2022
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
1.012026
Forget Me Not: Fighting Local Overfitting With Knowledge Fusion and Distillation · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Machine learning › Learning paradigms › continual learning
catastrophic forgetting
0.912025
On Local Overfitting and Forgetting in Deep Neural Networks · AAAI 2025
Machine learning › Deep learning architectures and training › regularization
early stopping
0.812024
United We Stand: Using Epoch-Wise Agreement of Ensembles to Combat Overfit · AAAI 2024
Machine learning › Learning paradigms
curriculum learning
0.722019
On The Power of Curriculum Learning in Training Deep Networks · ICML 2019
Curriculum Learning by Transfer Learning: Theory and Experiments with Deep Networks · ICML 2018
Machine learning › Reinforcement learning
strategy selection
0.712023
How to Select Which Active Learning Strategy is Best Suited for Your Specific Problem and Budget · NeurIPS 2023
Machine learning › Representation and self-supervised learning
word representation
0.622018
Coming to Your Senses: on Controls and Evaluation Sets in Polysemy Research · EMNLP 2018
Outta Control: Laws of Semantic Change and Inherent Biases in Word Representation Models · EMNLP 2017
Machine learning › Efficient and distributed learning › active learning
low-budget active learning
0.612022
Active Learning Through a Covering Lens · NeurIPS 2022
Natural language and speech › Language models and text generation
neural language model
0.612022
The Grammar-Learning Trajectories of Neural Language Models · ACL (1) 2022
Machine learning › Learning theory
neural network theory
0.612022
Principal Components Bias in Over-parameterized Linear Models, and its Manifestation in Deep Neural Networks · J. Mach. Learn. Res. 2022
Machine learning › Learning theory › inductive bias
simplicity bias
0.612022
Principal Components Bias in Over-parameterized Linear Models, and its Manifestation in Deep Neural Networks · J. Mach. Learn. Res. 2022
Machine learning › Efficient and distributed learning
subset selection
0.612022
Active Learning Through a Covering Lens · NeurIPS 2022
Natural language and speech › Language models and text generation
text representation
0.612022
The Grammar-Learning Trajectories of Neural Language Models · ACL (1) 2022
Empirical software engineering › benchmarking
benchmark dataset
0.412020
Let's Agree to Agree: Neural Networks Share Classification Order on Real Datasets · ICML 2020
Machine learning › Learning theory
online learning
0.432013
Hierarchical Regularization Cascade for Joint Learning · ICML (3) 2013
Online Learning in the Embedded Manifold of Low-rank Matrices · J. Mach. Learn. Res. 2012
Online Learning in The Manifold of Low-Rank Matrices · NIPS 2010
Machine learning › Optimization for machine learning
optimization landscape
0.412019
On The Power of Curriculum Learning in Training Deep Networks · ICML 2019
Machine learning › Optimization for machine learning
stochastic gradient descent
0.312018
Curriculum Learning by Transfer Learning: Theory and Experiments with Deep Networks · ICML 2018
Machine learning › Deep learning architectures and training
convolutional neural network
0.312017
Hidden Layers in Perceptual Learning · CVPR 2017
Natural language and speech › Information extraction and text analysis › lexical semantics
semantic change detection
0.312017
Outta Control: Laws of Semantic Change and Inherent Biases in Word Representation Models · EMNLP 2017
Machine learning › Representation and self-supervised learning › representation learning
metric learning
0.352007
Learning distance function by coding similarity · ICML 2007
Analyzing Auditory Neurons by Learning Distance Functions · NIPS 2005
Boosting margin based distance functions for clustering · ICML 2004
Knowledge, reasoning and agents › Knowledge representation and reasoning › ontology › concept hierarchy
label hierarchy
0.222012
Beyond Novelty Detection: Incongruent Events, When General and Specific Classifiers Disagree · IEEE Trans. Pattern Anal. Mach. Intell. 2012
Beyond Novelty Detection: Incongruent Events, when General and Specific Classifiers Disagree · NIPS 2008
Machine learning › Trustworthy machine learning
novelty detection
0.222012
Beyond Novelty Detection: Incongruent Events, When General and Specific Classifiers Disagree · IEEE Trans. Pattern Anal. Mach. Intell. 2012
Beyond Novelty Detection: Incongruent Events, when General and Specific Classifiers Disagree · NIPS 2008
Machine learning › Optimization for machine learning › constrained optimization
budget-constrained learning
0.212023
How to Select Which Active Learning Strategy is Best Suited for Your Specific Problem and Budget · NeurIPS 2023
Computer vision › Image recognition and object detection
image classification
0.212022
Active Learning on a Budget: Opposite Strategies Suit High and Low Budgets · ICML 2022
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning
self-supervised representation learning
0.212022
Active Learning Through a Covering Lens · NeurIPS 2022
Machine learning › Learning paradigms
multi-task learning
0.212013
Hierarchical Regularization Cascade for Joint Learning · ICML (3) 2013
Natural language and speech › Information extraction and text analysis
topic model
0.212013
Modeling Musical Influence with Topic Models · ICML (2) 2013
Data mining › text mining › topic modeling
latent dirichlet allocation
0.212013
LDA Topic Model with Soft Assignment of Descriptors to Words · ICML (3) 2013

Methods — techniques the papers use, named apart from their topics

knowledge distillation · 2.6forgetting rate score · 1.0checkpoint ensembling · 1.0theoretical analysis · 1.0ensemble learning · 0.8transfer learning · 0.7derivative-based selection · 0.7clustering · 0.7semi-supervised learning · 0.6word embeddings · 0.6synthetic dataset · 0.4empirical observation · 0.4variational inference · 0.3dynamic texture model · 0.3topic modeling · 0.2latent dirichlet allocation · 0.2mahalanobis distance · 0.1fisher linear discriminant · 0.1
YearPublicationVenuePosition
2026 Forget Me Not: Fighting Local Overfitting With Knowledge Fusion and Distillation
abstract
Overfitting in deep neural networks occurs less frequently than expected. This is a puzzling observation, as theory predicts that greater model capacity should eventually lead to overfitting - yet this is rarely seen in practice. But what if overfitting does occur, not globally, but in specific sub-regions of the data space? In this work, we introduce a novel score that measures the forgetting rate of deep models on validation data, capturing what we term local overfitting: a performance degradation confined to certain regions of the input space. We demonstrate that local overfitting can arise even without conventional overfitting, and is closely linked to the double descent phenomenon. Building on these insights, we introduce a two-stage approach that leverages the training history of a single model to recover and retain forgotten knowledge: first, by aggregating checkpoints into an ensemble, and then by distilling it into a single model of the original size, thus enhancing performance without added inference cost. Extensive experiments across multiple datasets, modern architectures, and training regimes validate the effectiveness of our approach. Notably, in the presence of label noise, our method - Knowledge Fusion followed by Knowledge Distillation - outperforms both the original model and independently trained ensembles, achieving a rare win-win scenario: reduced training and inference complexity.
Uri Stern, Eli Corn, Daphna Weinshall
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 On Local Overfitting and Forgetting in Deep Neural Networks
abstract
The infrequent occurrence of overfitting in deep neural networks is perplexing: contrary to theoretical expectations, increasing model size often enhances performance in practice. But what if overfitting does occur, though restricted to specific sub-regions of the data space? In this work, we propose a novel score that captures the forgetting rate of deep models on validation data. We posit that this score quantifies local overfitting: a decline in performance confined to certain regions of the data space. We then show empirically that local overfitting occurs regardless of the presence of traditional overfitting. Using the framework of deep over-parametrized linear models, we offer a certain theoretical characterization of forgotten knowledge, and show that it correlates with knowledge forgotten by real deep models. Finally, we devise a new ensemble method that aims to recover forgotten knowledge, relying solely on the training history of a single network. When combined with knowledge distillation, this method will enhance the performance of a trained model without adding inference costs. Extensive empirical evaluations demonstrate the efficacy of our method across multiple datasets, contemporary neural network architectures, and training protocols.
Uri Stern, Tomer Yaacoby, Daphna Weinshall
AAAI3
2024 United We Stand: Using Epoch-Wise Agreement of Ensembles to Combat Overfit
abstract
Deep neural networks have become the method of choice for solving many classification tasks, largely because they can fit very complex functions defined over raw data. The downside of such powerful learners is the danger of overfit. In this paper, we introduce a novel ensemble classifier for deep networks that effectively overcomes overfitting by combining models generated at specific intermediate epochs during training. Our method allows for the incorporation of useful knowledge obtained by the models during the overfitting phase without deterioration of the general performance, which is usually missed when early stopping is used. To motivate this approach, we begin with the theoretical analysis of a regression model, whose prediction - that the variance among classifiers increases when overfit occurs - is demonstrated empirically in deep networks in common use. Guided by these results, we construct a new ensemble-based prediction method, where the prediction is determined by the class that attains the most consensual prediction throughout the training epochs. Using multiple image and text classification datasets, we show that when regular ensembles suffer from overfit, our method eliminates the harmful reduction in generalization due to overfit, and often even surpasses the performance obtained by early stopping. Our method is easy to implement and can be integrated with any training scheme and architecture, without additional prior knowledge beyond the training set. It is thus a practical and useful tool to overcome overfit.
Uri Stern, Daniel Shwartz, Daphna Weinshall
AAAI3
2023 How to Select Which Active Learning Strategy is Best Suited for Your Specific Problem and Budget
abstract
In the domain of Active Learning (AL), a learner actively selects which unlabeled examples to seek labels from an oracle, while operating within predefined budget constraints. Importantly, it has been recently shown that distinct query strategies are better suited for different conditions and budgetary constraints. In practice, the determination of the most appropriate AL strategy for a given situation remains an open problem. To tackle this challenge, we propose a practical derivative-based method that dynamically identifies the best strategy for a given budget. Intuitive motivation for our approach is provided by the theoretical analysis of a simplified scenario. We then introduce a method to dynamically select an AL strategy, which takes into account the unique characteristics of the problem and the available budget. Empirical results showcase the effectiveness of our approach across diverse budgets and computer vision tasks.
Guy Hacohen, Daphna Weinshall
NeurIPS2
2022 The Grammar-Learning Trajectories of Neural Language Models
abstract
The learning trajectories of linguistic phenomena in humans provide insight into linguistic representation, beyond what can be gleaned from inspecting the behavior of an adult speaker.To apply a similar approach to analyze neural language models (NLM), it is first necessary to establish that different models are similar enough in the generalizations they make.In this paper, we show that NLMs with different initialization, architecture, and training data acquire linguistic phenomena in a similar order, despite their different end performance.These findings suggest that there is some mutual inductive bias that underlies these models' learning of linguistic phenomena.Taking inspiration from psycholinguistics, we argue that studying this inductive bias is an opportunity to study the linguistic representation implicit in NLMs.Leveraging these findings, we compare the relative performance on different phenomena at varying learning stages with simpler reference models.Results suggest that NLMs exhibit consistent "developmental" stages.Moreover, we find the learning trajectory to be approximately one-dimensional: given an NLM with a certain overall performance, it is possible to predict what linguistic generalizations it has already acquired.Initial analysis of these stages presents phenomena clusters (notably morphological ones), whose performance progresses in unison, suggesting a potential link between the generalizations behind them.
Leshem Choshen, Guy Hacohen, Daphna Weinshall, Omri Abend
ACL (1)3
2022 Active Learning on a Budget: Opposite Strategies Suit High and Low Budgets
abstract
Investigating active learning, we focus on the relation between the number of labeled examples (budget size), and suitable querying strategies. Our theoretical analysis shows a behavior reminiscent of phase transition: typical examples are best queried when the budget is low, while unrepresentative examples are best queried when the budget is large. Combined evidence shows that a similar phenomenon occurs in common classification models. Accordingly, we propose TypiClust – a deep active learning strategy suited for low budgets. In a comparative empirical investigation of supervised learning, using a variety of architectures and image datasets, TypiClust outperforms all other active learning strategies in the low-budget regime. Using TypiClust in the semi-supervised framework, performance gets an even more significant boost. In particular, state-of-the-art semi-supervised methods trained on CIFAR-10 with 10 labeled examples selected by TypiClust, reach 93.2% accuracy – an improvement of 39.4% over random selection. Code is available at https://github.com/avihu111/TypiClust.
Guy Hacohen, Avihu Dekel, Daphna Weinshall
ICML3
2022 Active Learning Through a Covering Lens
abstract
Deep active learning aims to reduce the annotation cost for the training of deep models, which is notoriously data-hungry. Until recently, deep active learning methods were ineffectual in the low-budget regime, where only a small number of examples are annotated. The situation has been alleviated by recent advances in representation and self-supervised learning, which impart the geometry of the data representation with rich information about the points. Taking advantage of this progress, we study the problem of subset selection for annotation through a “covering” lens, proposing ProbCover – a new active learning algorithm for the low budget regime, which seeks to maximize Probability Coverage. We then describe a dual way to view the proposed formulation, from which one can derive strategies suitable for the high budget regime of active learning, related to existing methods like Coreset. We conclude with extensive experiments, evaluating ProbCover in the low-budget regime. We show that our principled active learning strategy improves the state-of-the-art in the low-budget regime in several image recognition benchmarks. This method is especially beneficial in the semi-supervised setting, allowing state-of-the-art semi-supervised methods to match the performance of fully supervised methods, while using much fewer labels nonetheless. Code is available at https://github.com/avihu111/TypiClust.
Ofer Yehuda, Avihu Dekel, Guy Hacohen, Daphna Weinshall
NeurIPS4
2022 Principal Components Bias in Over-parameterized Linear Models, and its Manifestation in Deep Neural Networks
abstract
Recent work suggests that convolutional neural networks of different architectures learn to classify images in the same order. To understand this phenomenon, we revisit the over-parametrized deep linear network model. Our analysis reveals that, when the hidden layers are wide enough, the convergence rate of this model's parameters is exponentially faster along the directions of the larger principal components of the data, at a rate governed by the corresponding singular values. We term this convergence pattern the Principal Components bias (PC-bias). Empirically, we show how the PC-bias streamlines the order of learning of both linear and non-linear networks, more prominently at earlier stages of learning. We then compare our results to the simplicity bias, showing that both biases can be seen independently, and affect the order of learning in different ways. Finally, we discuss how the PC-bias may explain some benefits of early stopping and its connection to PCA, and why deep networks converge more slowly with random labels.
Guy Hacohen, Daphna Weinshall
J. Mach. Learn. Res.2
2020 Let's Agree to Agree: Neural Networks Share Classification Order on Real Datasets
abstract
We report a series of robust empirical observations, demonstrating that deep Neural Networks learn the examples in both the training and test sets in a similar order. This phenomenon is observed in all the commonly used benchmarks we evaluated, including many image classification benchmarks, and one text classification benchmark. While this phenomenon is strongest for models of the same architecture, it also crosses architectural boundaries – models of different architectures start by learning the same examples, after which the more powerful model may continue to learn additional examples. We further show that this pattern of results reflects the interplay between the way neural networks learn benchmark datasets. Specifically, when fixing the architecture, we describe synthetic datasets for which this pattern is no longer observed. When fixing the dataset, we show that other learning paradigms may learn the data in a different order. We hypothesize that our results reflect how neural networks discover structure in natural datasets.
Guy Hacohen, Leshem Choshen, Daphna Weinshall
ICML3
2020 Generative Latent Implicit Conditional Optimization when Learning from Small Sample
abstract
We revisit the long-standing problem of learning from small sample, to which end we propose a novel method called GLICO (Generative Latent Implicit Conditional Optimization). GLICO learns a mapping from the training examples to a latent space, and a generator that generates images from vectors in the latent space. Unlike most recent works, which rely on access to large amounts of unlabeled data, GLICO does not require access to any additional data other than the small set of labeled points. In fact, GLICO learns to synthesize completely new samples for every class using as little as 5 or 10 examples per class, with as few as 10 such classes without imposing any prior. GLICO is then used to augment the small training set while training a classifier on the small sample. To this end our proposed method samples the learned latent space using spherical interpolation, and generates new examples using the trained generator. Empirical results show that the new sampled set is diverse enough, leading to improvement in image classification in comparison with the state of the art, when trained on small samples obtained from CIFAR-10, CIFAR-100, and CUB-200.
Idan Azuri, Daphna Weinshall
ICPR2
2020 Multi-Modal Deep Clustering: Unsupervised Partitioning of Images
abstract
The clustering of unlabeled raw images is a daunting task, which has recently been approached with some success by deep learning methods. Here we propose an unsupervised clustering framework, which learns a deep neural network in an end-to-end fashion, providing direct cluster assignments of images without additional processing. Multi-Modal Deep Clustering (MMDC), trains a deep network to align its image embeddings with target points sampled from a Gaussian Mixture Model distribution. The cluster assignments are then determined by mixture component association of image embeddings. Simultaneously, the same deep network is trained to solve an additional self-supervised task of predicting image rotations. This pushes the network to learn more meaningful image representations that facilitate a better clustering. Experimental results show that MMDC achieves or exceeds state-of-the-art performance on six challenging benchmarks. On natural image datasets we improve on previous results with significant margins of up to 20% absolute accuracy points, yielding an accuracy of 82% on CIFAR-10, 45% on CIFAR-100 and 69% on STL-10.
Guy Shiran, Daphna Weinshall
ICPR2
2019 On The Power of Curriculum Learning in Training Deep Networks
abstract
Training neural networks is traditionally done by providing a sequence of random mini-batches sampled uniformly from the entire training data. In this work, we analyze the effect of curriculum learning, which involves the non-uniform sampling of mini-batches, on the training of deep networks, and specifically CNNs trained for image recognition. To employ curriculum learning, the training algorithm must resolve 2 problems: (i) sort the training examples by difficulty; (ii) compute a series of mini-batches that exhibit an increasing level of difficulty. We address challenge (i) using two methods: transfer learning from some competitive “teacher" network, and bootstrapping. In our empirical evaluation, both methods show similar benefits in terms of increased learning speed and improved final performance on test data. We address challenge (ii) by investigating different pacing functions to guide the sampling. The empirical investigation includes a variety of network architectures, using images from CIFAR-10, CIFAR-100 and subsets of ImageNet. We conclude with a novel theoretical analysis of curriculum learning, where we show how it effectively modifies the optimization landscape. We then define the concept of an ideal curriculum, and show that under mild conditions it does not change the corresponding global minimum of the optimization function.
Guy Hacohen, Daphna Weinshall
ICML2
2018 Topic models for automated motor analysis in schizophrenia patients
abstract
Wearable devices fitted with various sensors are increasingly being used for the automatic and continuous tracking and monitoring of patients. Only first steps have been taken in the field of psychiatric care, where long term tracking of patient behavior holds the promise to help practitioners to better understand both individual patients, and the disorders in general. In this paper we use topic models for unsupervised analysis of movement activity of schizophrenia patients in a closed ward setting. Results demonstrate that features computed on the basis of this analysis differentially characterize interesting sub-populations of schizophrenia patients. Positive-signs schizophrenia sub-population was found to have high motor richness and low typicallity, while negative-signs patients had low motor richness and lower typicality. In addition we design a classifier which correctly classified up to 80% of the clinical sub-population (f-score=0.774) based on motor features.
Talia Tron, Yehezkel S. Resheff, Mikhail Bazhmin, Abraham Peled, Alexander Grinsphoon, Daphna Weinshall
BSN6
2018 Coming to Your Senses: on Controls and Evaluation Sets in Polysemy Research
abstract
The point of departure of this article is the claim that sense-specific vectors provide an advantage over normal vectors due to the polysemy that they presumably represent.This claim is based on performance gains observed in gold standard evaluation tests such as word similarity tasks.We demonstrate that this claim, at least as it is instantiated in prior art, is unfounded in two ways.Furthermore, we provide empirical data and an analytic discussion that may account for the previously reported improved performance.First, we show that ground-truth polysemy degrades performance in word similarity tasks.Therefore word similarity tasks are not suitable as an evaluation test for polysemy representation.Second, random assignment of words to senses is shown to improve performance in the same task.This and additional results point to the conclusion that performance gains as reported in previous work may be an artifact of random sense assignment, which is equivalent to sub-sampling and multiple estimation of word vector representations.Theoretical analysis shows that this may on its own be beneficial for the estimation of word similarity, by reducing the bias in the estimation of the cosine distance.
Haim Dubossarsky, Eitan Grossman, Daphna Weinshall
EMNLP3
2018 Curriculum Learning by Transfer Learning: Theory and Experiments with Deep Networks
abstract
We provide theoretical investigation of curriculum learning in the context of stochastic gradient descent when optimizing the convex linear regression loss. We prove that the rate of convergence of an ideal curriculum learning method is monotonically increasing with the difficulty of the examples. Moreover, among all equally difficult points, convergence is faster when using points which incur higher loss with respect to the current hypothesis. We then analyze curriculum learning in the context of training a CNN. We describe a method which infers the curriculum by way of transfer learning from another network, pre-trained on a different task. While this approach can only approximate the ideal curriculum, we observe empirically similar behavior to the one predicted by the theory, namely, a significant boost in convergence speed at the beginning of training. When the task is made more difficult, improvement in generalization performance is also observed. Finally, curriculum learning exhibits robustness against unfavorable conditions such as excessive regularization.
Daphna Weinshall, Gad Cohen, Dan Amir
ICML1
2017 Hidden Layers in Perceptual Learning
abstract
Studies in visual perceptual learning investigate the way human performance improves with practice, in the context of relatively simple (and therefore more manageable) visual tasks. Building on the powerful tools currently available for the training of Convolution Neural Networks (CNN), networks whose original architecture was inspired by the visual system, we revisited some of the open computational questions in perceptual learning. We first replicated two representative sets of perceptual learning experiments by training a shallow CNN to perform the relevant tasks. These networks qualitatively showed most of the characteristic behavior observed in perceptual learning, including the hallmark phenomena of specificity and its various manifestations in the forms of transfer or partial transfer, and learning enabling. We next analyzed the dynamics of weight modifications in the networks, identifying patterns which appeared to be instrumental for the transfer (or generalization) of learned skills from one task to another in the simulated networks. These patterns may identify ways by which the domain of search in the parameter space during network re-training can be significantly reduced, thereby accomplishing knowledge transfer.
Gad Cohen, Daphna Weinshall
CVPR2
2017 Outta Control: Laws of Semantic Change and Inherent Biases in Word Representation Models
abstract
This article evaluates three proposed laws of semantic change.Our claim is that in order to validate a putative law of semantic change, the effect should be observed in the genuine condition but absent or reduced in a suitably matched control condition, in which no change can possibly have taken place.Our analysis shows that the effects reported in recent literature must be substantially revised: (i) the proposed negative correlation between meaning change and word frequency is shown to be largely an artefact of the models of word representation used; (ii) the proposed negative correlation between meaning change and prototypicality is shown to be much weaker than what has been claimed in prior art; and (iii) the proposed positive correlation between meaning change and polysemy is largely an artefact of word frequency.These empirical observations are corroborated by analytical proofs that show that count representations introduce an inherent dependence on word frequency, and thus word frequency cannot be evaluated as an independent factor with these representations.
Haim Dubossarsky, Daphna Weinshall, Eitan Grossman
EMNLP2
2017 Implicit Media Tagging and Affect Prediction from RGB-D Video of Spontaneous Facial Expressions
abstract
We present a method that automatically evaluates emotional response from spontaneous facial activity. The automatic evaluation of emotional response, or affect, is a fascinating challenge with many applications. Our approach is based on the inferred activity of facial muscles over time, as automatically obtained from an RGB-D video recording of spontaneous facial activity. Our contribution is two-fold: First, we constructed a database of publicly available short video clips, which elicit a strong emotional response in a consistent manner across different individuals. Each video was tagged by its characteristic emotional response along 4 scales: Valence, Arousal, Likability and Rewatch (the desire to watch again). The second contribution is a two-step prediction method, based on learning, which was trained and tested using this database of tagged video clips. Our method was able to successfully predict the aforementioned 4 dimensional representation of affect, achieving high correlation (0.87-0.95) between the predicted scores and the affect tags. As part of the prediction algorithm we identified the period of strongest emotional response in the viewing recordings, in a method that was blind to the video clip being watched, showing high agreement between independent viewers. Finally, inspection of the relative contribution of different feature types to the prediction process revealed that temporal facets contributed more to the prediction of individual affect than to media tags.
Daniel Hadar, Talia Tron, Daphna Weinshall
FG3
2017 Optimized Linear Imputation
abstract
Often in real-world datasets, especially in high dimensional data, some feature values are missing. Since most data analysis and statistical methods do not handle gracefully missing values, the first step in the analysis requires the imputation of missing values. Indeed, there has been a long standing interest in methods for the imputation of missing values as a pre-processing step. One recent and effective approach, the IRMI stepwise regression imputation method, uses a linear regression model for each real-valued feature on the basis of all other features in the dataset. However, the proposed iterative formulation lacks convergence guarantee. Here we propose a closely related method, stated as a single optimization problem and a block coordinate-descent solution which is guaranteed to converge to a local minimum. Experiments show results on both synthetic and benchmark datasets, which are comparable to the results of the IRMI method whenever it converges. However, while in the set of experiments described here IRMI often does not converge, the performance of our methods is shown to be markedly superior in comparison with other methods.
Yehezkel S. Resheff, Daphna Weinshall
ICPRAM2
2015 Matrix factorization approach to behavioral mode analysis from acceleration data
abstract
The field of Movement Ecology is experiencing a period of rapid growth in availability of data, and like many other fields is turning to data science for tools and methods to cope with the new challenges and opportunities that this presents. One rich and interesting source of data is the bio-logger. These small electronic devices are attached to animals free to roam in their natural habitats, and report back readings from multiple sensors, including GPS and accelerometer bursts. A common use of this accelerometer data is for supervised learning of behavioral modes. However, there is a need for unsupervised analysis tools as well, due to the inherent difficulties of obtaining a labeled dataset, which in some cases is either infeasible or does not successfully encompass the full repertoire of behavioral modes of interest. Here we present a matrix factorization based clustering method that allows either a soft or a hard partitioning of acceleration measurements, as well as a straight-forward way of drawing insight into the complex movements themselves. The method is validated by comparing the partitions with a labeled dataset, and is further compared to standard methods highlighting the advantages of the new method.
Yehezkel S. Resheff, Shay Rotics, Ran Nathan, Daphna Weinshall
DSAA4
2013 Modeling Musical Influence with Topic Models
abstract
The role of musical influence has long been debated by scholars and critics in the humanities, but never in a data-driven way. In this work we approach the question of influence by applying topic-modeling tools (Blei & Lafferty, 2006; Gerrish & Blei, 2010) to a dataset of 24941 songs by 9222 artists, from the years 1922 to 2010. We find the models to be significantly correlated with a human-curated influence measure, and to clearly outperform a baseline method. Further using the learned model to study properties of influence, we find that musical influence and musical innovation are not monotonically correlated. However, we do find that the most influential songs were more innovative during two time periods: the early 1970’s and the mid 1990’s.
Uri Shalit, Daphna Weinshall, Gal Chechik
ICML (2)2
2013 LDA Topic Model with Soft Assignment of Descriptors to Words
abstract
The LDA topic model is being used to model corpora of documents that can be represented by bags of words. Here we extend the LDA model to deal with documents that are represented more naturally by bags of continuous descriptors. Given a finite dictionary of words which are generative models of descriptors, our extended LDA model allows for the soft assignment of descriptors to (many) dictionary words. We derive variational inference and parameter estimation procedures for the extended model, which closely resemble those obtained for the original model, with two important differences: First, the histogram of word counts is replaced by a histogram of pseudo word counts, or sums of responsibilities over all descriptors. Second, parameter estimation now depends on the average covariance matrix between these pseudo-counts, reflecting the fact that with soft assignment words are not independent. We use this approach to address novelty detection, where we seek to identify video events with low posterior probability. Video events are described by a generative dynamic texture model, from which we naturally derive a dictionary of generative words. Using a benchmark dataset for novelty detection, we show a very significant improvement in the detection of novel events when using our extended LDA model with soft assignment to words as against hard assignment (the original model), achieving state of the art novelty detection results.
Daphna Weinshall, Gal Levi, Dmitri Hanukaev
ICML (3)1
2013 Hierarchical Regularization Cascade for Joint Learning
abstract
As the sheer volume of available benchmark datasets increases, the problem of joint learning of classifiers and knowledge-transfer between classifiers, becomes more and more relevant. We present a hierarchical approach which exploits information sharing among different classification tasks, in multi-task and multi-class settings. It engages a top-down iterative method, which begins by posing an optimization problem with an incentive for large scale sharing among all classes. This incentive to share is gradually decreased,until there is no sharing and all tasks are considered separately. The method therefore exploits different levels of sharing within a given group of related tasks, without having to make hard decisions about the grouping of tasks. In order to deal with large scale problems, with many tasks and many classes, we extend our batch approach to an online setting and provide regret analysis of the algorithm. We tested our approach extensively on synthetic and real datasets, showing significant improvement over baseline and state-of-the-art methods.
Alon Zweig, Daphna Weinshall
ICML (3)2
2012 Online Learning in the Embedded Manifold of Low-rank Matrices
Uri Shalit, Daphna Weinshall, Gal Chechik
J. Mach. Learn. Res.2
2012 Beyond Novelty Detection: Incongruent Events, When General and Specific Classifiers Disagree
abstract
Unexpected stimuli are a challenge to any machine learning algorithm. Here, we identify distinct types of unexpected events when general-level and specific-level classifiers give conflicting predictions. We define a formal framework for the representation and processing of incongruent events: Starting from the notion of label hierarchy, we show how partial order on labels can be deduced from such hierarchies. For each event, we compute its probability in different ways, based on adjacent levels in the label hierarchy. An incongruent event is an event where the probability computed based on some more specific level is much smaller than the probability computed based on some more general level, leading to conflicting predictions. Algorithms are derived to detect incongruent events from different types of hierarchies, different applications, and a variety of data types. We present promising results for the detection of novel visual and audio objects, and new patterns of motion in video. We also discuss the detection of Out-Of- Vocabulary words in speech recognition, and the detection of incongruent events in a multimodal audiovisual scenario.
Daphna Weinshall, Alon Zweig, Hynek Hermansky, Stefan Kombrink, Frank W. Ohl, Jörn Anemüller, Jörg-Hendrik Bach, Luc Van Gool, Fabian Nater, Tomás Pajdla, Michal Havlena, Misha Pavel
IEEE Trans. Pattern Anal. Mach. Intell.1
2011 Extracting foreground masks towards object recognition
abstract
Effective segmentation prior to recognition has been shown to improve recognition performance. However, most segmentation algorithms adopt methods which are not explicitly linked to the goal of object recognition. Here we solve a related but slightly different problem in order to assist object recognition more directly - the extraction of a foreground mask, which identifies the locations of objects in the image. We propose a novel foreground/background segmentation algorithm that attempts to segment the interesting objects from the rest of the image, while maximizing an objective function which is tightly related to object recognition. We do this in a manner which requires no class-specific knowledge of object categories, using a probabilistic formulation which is derived from manually segmented images. The model includes a geometric prior and an appearance prior, whose parameters are learnt on the fly from images that are similar to the query image. We use graph-cut based energy minimization to enforce spatial coherence on the model's output. The method is tested on the challenging VOC09 and VOC10 segmentation datasets, achieving excellent results in providing a foreground mask. We also provide comparisons to the recent segmentation method of [7].
Amir Rosenfeld, Daphna Weinshall
ICCV2
2011 Monaural Azimuth Localization Using Spectral Dynamics of Speech
abstract
We tackle the task of localizing speech signals on the horizontal plane using monaural cues. We show that monaural cues as incorporated in speech are efficiently captured by amplitude modulation spectra patterns. We demonstrate that by using these patterns, a linear Support Vector Machine can use directionality related information to learn to discriminate and classify sound location at high resolution. We propose a straightforward and robust way of integrating information from two ears: treating each ear as an independent processor and integrate the information at the decision level by doing that ambiguity is to a large extent resolved.
Roi Kliper, Hendrik Kayser, Daphna Weinshall, Israel Nelken, Jörn Anemüller
INTERSPEECH3
2010 Identifying Surprising Events in Videos Using Bayesian Topic Models
Avishai Hendel, Daphna Weinshall, Shmuel Peleg
ACCV (3)2
2010 Online Learning in The Manifold of Low-Rank Matrices
abstract
When learning models that are represented in matrix forms, enforcing a low-rank constraint can dramatically improve the memory and run time complexity, while providing a natural regularization of the model. However, naive approaches for minimizing functions over the set of low-rank matrices are either prohibitively time consuming (repeated singular value decomposition of the matrix) or numerically unstable (optimizing a factored representation of the low rank matrix). We build on recent advances in optimization over manifolds, and describe an iterative online learning procedure, consisting of a gradient step, followed by a second-order retraction back to the manifold. While the ideal retraction is hard to compute, and so is the projection operator that approximates it, we describe another second-order retraction that can be computed efficiently, with run time and memory complexity of O((n+m)k) for a rank-k matrix of dimension m x n, given rank one gradients. We use this algorithm, LORETA, to learn a matrix-form similarity measure over pairs of documents represented as high dimensional vectors. LORETA improves the mean average precision over a passive- aggressive approach in a factorized model, and also improves over a full model trained over pre-selected features using the same memory requirements. LORETA also showed consistent improvement over standard methods in a large (1600 classes) multi-label image classification task.
Uri Shalit, Daphna Weinshall, Gal Chechik
NIPS2
2008 Beyond Novelty Detection: Incongruent Events, when General and Specific Classifiers Disagree
abstract
Unexpected stimuli are a challenge to any machine learning algorithm. Here we identify distinct types of unexpected events, focusing on 'incongruent events' - when 'general level' and 'specific level' classifiers give conflicting predictions. We define a formal framework for the representation and processing of incongruent events: starting from the notion of label hierarchy, we show how partial order on labels can be deduced from such hierarchies. For each event, we compute its probability in different ways, based on adjacent levels (according to the partial order) in the label hierarchy . An incongruent event is an event where the probability computed based on some more specific level (in accordance with the partial order) is much smaller than the probability computed based on some more general level, leading to conflicting predictions. We derive algorithms to detect incongruent events from different types of hierarchies, corresponding to class membership or part membership. Respectively, we show promising results with real data on two specific problems: Out Of Vocabulary words in speech recognition, and the identification of a new sub-class (e.g., the face of a new individual) in audio-visual facial object recognition.
Daphna Weinshall, Hynek Hermansky, Alon Zweig, Holly Brügge Jimison, Frank W. Ohl, Misha Pavel
NIPS1
2008 Efficient Learning of Relational Object Class Models
Aharon Bar-Hillel, Daphna Weinshall
Int. J. Comput. Vis.2
2008 Motion Segmentation and Depth Ordering Using an Occlusion Detector
abstract
We present a novel method for motion segmentation and depth ordering from a video sequence in general motion. We first compute motion segmentation based on differential properties of the spatio-temporal domain, and scale-space integration. Given a motion boundary, we describe two algorithms to determine depth ordering from two- and three- frame sequences. An remarkable characteristic of our method is its ability compute depth ordering from only two frames. The segmentation and depth ordering algorithms are shown to give good results on 6 real sequences taken in general motion. We use synthetic data to show robustness to high levels of noise and illumination changes; we also include cases where no intensity edge exists at the location of the motion boundary, or when no parametric motion model can describe the data. Finally, we describe human experiments showing that people, like our algorithm, can compute depth ordering from only two frames, even when the boundary between the layers is not visible in a single frame.
Doron Feldman, Daphna Weinshall
IEEE Trans. Pattern Anal. Mach. Intell.2
2007 Exploiting Object Hierarchy: Combining Models from Different Category Levels
abstract
We investigated the computational properties of natural object hierarchy in the context of constellation object class models, and its utility for object class recognition. We first observed an interesting computational property of the object hierarchy: comparing the recognition rate when using models of objects at different levels, the higher more inclusive levels (e.g., closed-frame vehicles or vehicles) exhibit higher recall but lower precision when compared with the class specific level (e.g., bus). These inherent differences suggest that combining object classifiers from different hierarchical levels into a single classifier may improve classification, as it appears like these models capture different aspects of the object. We describe a method to combine these classifiers, and analyze the conditions under which improvement can be guaranteed. When given a small sample of a new object class, we describe a method to transfer knowledge across the tree hierarchy, between related objects. Finally, we describe extensive experiments using object hierarchies obtained from publicly available datasets, and show that the combined classifiers significantly improve recognition results.
Alon Zweig, Daphna Weinshall
ICCV2
2007 Learning distance function by coding similarity
abstract
We consider the problem of learning a similarity function from a set of positive equivalence constraints, i.e. 'similar' point pairs. We define the similarity in information theoretic terms, as the gain in coding length when shifting from independent encoding of the pair to joint encoding. Under simple Gaussian assumptions, this formulation leads to a non-Mahalanobis similarity function which is efficient and simple to learn. This function can be viewed as a likelihood ratio test, and we show that the optimal similarity-preserving projection of the data is a variant of Fisher Linear Discriminant. We also show that under some naturally occurring sampling conditions of equivalence constraints, this function converges to a known Mahalanobis distance (RCA). The suggested similarity function exhibits superior performance over alternative Mahalanobis distances learnt from the same data. Its superiority is demonstrated in the context of image retrieval and graph based clustering, using a large number of data sets.
Aharon Bar-Hillel, Daphna Weinshall
ICML2
2006 Learning a kernel function for classification with small training samples
abstract
When given a small sample, we show that classification with SVM can be considerably enhanced by using a kernel function learned from the training data prior to discrimination. This kernel is also shown to enhance retrieval based on data similarity. Specifically, we describe KernelBoost - a boosting algorithm which computes a kernel function as a combination of 'weak' space partitions. The kernel learning method naturally incorporates domain knowledge in the form of unlabeled data (i.e. in a semi-supervised or transductive settings), and also in the form of labeled samples from relevant related problems (i.e. in a learning-to-learn scenario). The latter goal is accomplished by learning a single kernel function for all classes. We show comparative evaluations of our method on datasets from the UCI repository. We demonstrate performance enhancement on two challenging tasks: digit classification with kernel SVM, and facial image retrieval based on image similarity as measured by the learnt kernel.
Tomer Hertz, Aharon Bar-Hillel, Daphna Weinshall
ICML3
2006 Subordinate class recognition using relational object models
abstract
We address the problem of sub-ordinate class recognition, like the distinction between different types of motorcycles. Our approach is motivated by observations from cognitive psychology, which identify parts as the defining component of basic level categories (like motorcycles), while sub-ordinate categories are more often defined by part properties (like 'jagged wheels'). Accordingly, we suggest a two-stage algorithm: First, a relational part based object model is learnt using unsegmented object images from the inclusive class (e.g., motorcycles in general). The model is then used to build a class-specific vector representation for images, where each entry corresponds to a model's part. In the second stage we train a standard discriminative classifier to classify subclass instances (e.g., cross motorcycles) based on the class-specific vector representation. We describe extensive experimental results with several subclasses. The proposed algorithm typically gives better results than a competing one-step algorithm, or a two stage algorithm where classification is based on a model of the sub-ordinate class.
Aharon Bar-Hillel, Daphna Weinshall
NIPS2
2006 Cognitive Authentication Schemes Safe Against Spyware (Short Paper)
abstract
Can we secure user authentication against eavesdropping adversaries, relying on human cognitive functions alone, unassisted by any external computational device? To accomplish this goal, we propose challenge response protocols that rely on a shared secret set of pictures. Under the considered brute-force attack the protocols are safe against eavesdropping, in that a modestly powered adversary who fully records a series of successful interactions cannot compute the user's secret. Moreover, the protocols can be tuned to any desired level of security against random guessing, where security can be traded-off with authentication time. The proposed protocols have two drawbacks: First, training is required to familiarize the user with the secret set of pictures. Second, depending on the level of security required, entry time can be significantly longer than with alternative methods. We describe user studies showing that people can use these protocols successfully, and quantify the time it takes for training and for successful authentication. We show evidence that the secret can be maintained for a long time (up to a year) with relatively low loss
Daphna Weinshall
S&P1
2005 Object Class Recognition by Boosting a Part-Based Model
abstract
We propose a new technique for object class recognition, which learns a generative appearance model in a discriminative manner. The technique is based on the intermediate representation of an image as a set of patches, which are extracted using an interest point detector. The learning problem becomes an instance of supervised learning from sets of unordered features. In order to solve this problem, we designed a classifier based on a simple, part based, generative object model. Only the appearance of each part is modeled. When learning the model parameters, we use a discriminative boosting algorithm which minimizes the loss of the training error directly. The models thus learnt have clear probabilistic semantics, and also maintain good classification performance. The performance of the algorithm has been tested using publicly available benchmark data, and shown to be comparable to other state of the art algorithms for this task; our main advantage in these comparisons is speed (order of magnitudes faster) and scalability.
Aharon Bar-Hillel, Tomer Hertz, Daphna Weinshall
CVPR (1)3
2005 Efficient Learning of Relational Object Class Models
abstract
We present an efficient method for learning part-based object class models. The models include location and scale relations between parts, as well as part appearance. Models are learnt from raw object and background images, represented as an unordered set of features extracted using an interest point detector. The object class is generatively modeled using a simple Bayesian network with a central hidden node containing location and scale information, and nodes describing object parts. The model's parameters, however are optimized to reduce a loss function which reflects training error as in discriminative methods. Specifically, the optimization is done using a boosting-like technique with complexity linear in the number of parts and the number of features per image. This efficiency allows our method to learn relational models with many parts and features, and leads to improved results when compared with other methods. Extensive experimental results are described, using some common bench-mark datasets and three sets of newly collected data, showing the relative advantage of our method.
Aharon Bar-Hillel, Tomer Hertz, Daphna Weinshall
ICCV3
2005 Realtime IBR with Omnidirectional Crossed-Slits Projection
abstract
The crossed-slits (X-Slits) projection can be used to generate new views of a scene from a sequence of perspective images. Compared with other image-based rendering (IBR) techniques, X-Slits image generation is simple and requires a relatively small number of input images, which makes it suitable for realtime IBR. In this paper we extend this model to omnidirectional cameras and a circular slit. We show how it can be used for realtime image-based rendering of omnidirectional images, and how to optimize it for speed and quality. We analyze the inherent geometric distortions of the circular X-Slits projection, and describe a normalization mechanism to reduce distortions, creating a realistic virtual environment. Essentially the same mechanism is used to augment the X-Slits images with artificial objects, when using standard graphics tools which assume perspective projection.
Doron Feldman, Daphna Weinshall
ICCV2
2005 Analyzing Auditory Neurons by Learning Distance Functions
abstract
We present a novel approach to the characterization of complex sensory neurons. One of the main goals of characterizing sensory neurons is to characterize dimensions in stimulus space to which the neurons are highly sensitive (causing large gradients in the neural responses) or al- ternatively dimensions in stimulus space to which the neuronal response are invariant (defining iso-response manifolds). We formulate this prob- lem as that of learning a geometry on stimulus space that is compatible with the neural responses: the distance between stimuli should be large when the responses they evoke are very different, and small when the re- sponses they evoke are similar. Here we show how to successfully train such distance functions using rather limited amount of information. The data consisted of the responses of neurons in primary auditory cortex (A1) of anesthetized cats to 32 stimuli derived from natural sounds. For each neuron, a subset of all pairs of stimuli was selected such that the responses of the two stimuli in a pair were either very similar or very dissimilar. The distance function was trained to fit these constraints. The resulting distance functions generalized to predict the distances between the responses of a test stimulus and the trained stimuli.
Inna Weiner, Tomer Hertz, Israel Nelken, Daphna Weinshall
NIPS4
2005 Learning a Mahalanobis Metric from Equivalence Constraints
abstract
Many learning algorithms use a metric defined over the input space as a principal tool, and their performance critically depends on the quality of this metric. We address the problem of learning metrics using side-information in the form of equivalence constraints. Unlike labels, we demonstrate that this type of side-information can sometimes be automatically obtained without the need of human intervention. We show how such side-information can be used to modify the representation of the data, leading to improved clustering and classification. Specifically, we present the Relevant Component Analysis (RCA) algorithm, which is a simple and efficient algorithm for learning a Mahalanobis metric. We show that RCA is the solution of an interesting optimization problem, founded on an information theoretic basis. If dimensionality reduction is allowed within RCA, we show that it is optimally accomplished by a version of Fisher's linear discriminant that uses constraints. Moreover, under certain Gaussian assumptions, RCA can be viewed as a Maximum Likelihood estimation of the within class covariance matrix. We conclude with extensive empirical evaluations of RCA, showing its advantage over alternative methods.
Aharon Bar-Hillel, Tomer Hertz, Noam Shental, Daphna Weinshall
J. Mach. Learn. Res.4
2004 Learning Distance Functions for Image Retrieval
Tomer Hertz, Aharon Bar-Hillel, Daphna Weinshall
CVPR (2)3
2004 Boosting margin based distance functions for clustering
abstract
The performance of graph based clustering methods critically depends on the quality of the distance function used to compute similarities between pairs of neighboring nodes. In this paper we learn distance functions by training binary classifiers with margins. The classifiers are defined over the product space of pairs of points and are trained to distinguish whether two points come from the same class or not. The signed margin is used as the distance value. Our main contribution is a distance learning method (DistBoost), which combines boosting hypotheses over the product space with a weak learner based on partitioning the original feature space. Each weak hypothesis is a Gaussian mixture model computed using a semi-supervised constrained EM algorithm, which is trained using both unlabeled and labeled data. We also consider SVM and decision trees boosting as margin based classifiers in the product space. We experimentally compare the margin based distance functions with other existing metric learning methods, and with existing techniques for the direct incorporation of constraints into various clustering algorithms. Clustering performance is measured on some benchmark databases from the UCI repository, a sample from the MNIST database, and a data set of color images of animals. In most cases the DistBoost algorithm significantly and robustly outperformed its competitors.
Tomer Hertz, Aharon Bar-Hillel, Daphna Weinshall
ICML3
2003 Enhancing Image and Video Retrieval: Learning via Equivalence Constraint
abstract
The paper is about learning using partial information in the form of equivalence constraints. Equivalence constraints provide relational information about the labels of data points, rather than the labels themselves. Our work is motivated by the observation that in many real life applications partial information about the data can be obtained with very little cost. For example, in video indexing we may want to use the fact that a sequence of faces obtained from successive frames in roughly the same location is likely to contain the same unknown individual. Learning using equivalence constraints is different from learning using labels and poses new technical challenges. In this paper we present three novel methods for clustering and classification, which use equivalence constraints. We provide results of our methods on a distributed image querying system that works on a large facial image database, and on the clustering and retrieval of surveillance data. Our results show that we can significantly improve the performance of image retrieval by taking advantage of such assumptions as temporal continuity in the data. Significant improvement is also obtained by making the users of the system take the role of distributed teachers, which reduces the need for expensive labeling by paid human labor.
Tomer Hertz, Noam Shental, Aharon Bar-Hillel, Daphna Weinshall
CVPR (2)4
2003 On the Epipolar Geometry of the Crossed-Slits Projection
abstract
The Crossed-Slits (X-Slits) camera is defined by two nonintersecting slits, which replace the pinhole in the common perspective camera. Each point in space is projected to the image plane by a ray which passes through the point and the two slits. The X-Slits projection model includes the pushbroom camera as a special case. In addition, it describes a certain class of panoramic images, which are generated from sequences obtained by translating pinhole cameras. In this paper we develop the epipolar geometry of the X-Slits projection model. We show an object which is similar to the fundamental matrix; our matrix, however, describes a quadratic relation between corresponding image points (using the Veronese mapping). Similarly the equivalent of epipolar lines are conics in the image plane. Unlike the pin-hole case, epipolar surfaces do not usually exist in the sense that matching epipolar lines lie on a single surface; we analyze the cases when epipolar surfaces exist, and characterize their properties. Finally, we demonstrate the matching of points in pairs of X-Slits panoramic images.
Doron Feldman, Tomás Pajdla, Daphna Weinshall
ICCV3
2003 Learning Distance Functions using Equivalence Relations
Aharon Bar-Hillel, Tomer Hertz, Noam Shental, Daphna Weinshall
ICML4
2003 Computing Gaussian Mixture Models with EM Using Equivalence Constraints
abstract
Density estimation with Gaussian Mixture Models is a popular gener- ative technique used also for clustering. We develop a framework to incorporate side information in the form of equivalence constraints into the model estimation procedure. Equivalence constraints are defined on pairs of data points, indicating whether the points arise from the same source (positive constraints) or from different sources (negative con- straints). Such constraints can be gathered automatically in some learn- ing problems, and are a natural form of supervision in others. For the estimation of model parameters we present a closed form EM procedure which handles positive constraints, and a Generalized EM procedure us- ing a Markov net which handles negative constraints. Using publicly available data sets we demonstrate that such side information can lead to considerable improvement in clustering tasks, and that our algorithm is preferable to two other suggested methods using the same type of side information.
Noam Shental, Aharon Bar-Hillel, Tomer Hertz, Daphna Weinshall
NIPS4
2003 Mosaicing New Views: The Crossed-Slits Projection
abstract
We introduce anew kind of mosaicing, where the position of the sampling strip varies as a function of the input camera location. The new images that are generated this way correspond to a new projection model defined by two slits, termed here the Crossed-Slits (X-Slits) projection. In this projection model, every 3D point is projected by a ray defined as the line that passes through that point and intersects the two slits. The intersection of the projection rays with the imaging surface defines the image. X-Slits mosaicing provides two benefits. First, the generated mosaics are closer to perspective images than traditional pushbroom mosaics. Second, by simple manipulations of the strip sampling function, we can change the location of one of the virtual slits, providing a virtual walkthrough of a X-Slits camera; all this can be done without recovering any 3D geometry and without calibration. A number of examples where we translate the virtual camera and change its orientation are given; the examples demonstrate realistic changes in parallax, reflections, and occlusions.
Assaf Zomet, Doron Feldman, Shmuel Peleg, Daphna Weinshall
IEEE Trans. Pattern Anal. Mach. Intell.4
2002 Adjustment Learning and Relevant Component Analysis
Noam Shental, Tomer Hertz, Daphna Weinshall, Misha Pavel
ECCV (4)3
2002 New View Generation with a Bi-centric Camera
Daphna Weinshall, Mi-Suen Lee, Tomás Brodský, M. Trajkovic, Doron Feldman
ECCV (1)1
2001 A Computer Vision System for On-Screen Item Selection by Finger Pointing
abstract
Pointing at planar surfaces such as TV and computer monitors or projection screens can be a useful mode of interaction between humans and machines. To a large extent what seems to hinder the use of vision in such practical applications is the difficulty of the computational task, which is typically defined as 3-D reconstruction from uncalibrated 2-D images of a non-static scene. We describe below two designs where, using one or two cameras, the target of pointing on a flat monitor or screen is identified without 3-D inference, using only image morphing and line intersection. This is accomplished by registering the images with the target plane. When used to identify a pointing target on a surface hidden from the camera (e.g., a computer monitor which supports the camera itself as in most PC configurations), we add aperture(s) coplanar with the target surface in front of the camera(s). We describe experimental results showing a fully automated procedure for pointing target detection with high accuracy. The simplicity of our method and its robustness, as well as the relative accuracy of our results, can make pointing a practical means of human-machine interaction.
Mi-Suen Lee, Daphna Weinshall, Eric Cohen-Solal
CVPR (1)2
2001 Self-Organization in Vision: Stochastic Clustering for Image Segmentation, Perceptual Grouping, and Image Database Organization
abstract
We present a stochastic clustering algorithm which uses pairwise similarity of elements and show how it can be used to address various problems in computer vision, including the low-level image segmentation, mid-level perceptual grouping, and high-level image database organization. The clustering problem is viewed as a graph partitioning problem, where nodes represent data elements and the weights of the edges represent pairwise similarities. We generate samples of cuts in this graph, by using Karger's contraction algorithm (1996), and compute an "average" cut which provides the basis for our solution to the clustering problem. The stochastic nature of our method makes it robust against noise, including accidental edges and small spurious clusters. The complexity of our algorithm is very low: O(|E| log/sup 2/ N) for N objects, |E| similarity relations, and a fixed accuracy level. In addition, and without additional computational cost, our algorithm provides a hierarchy of nested partitions. We demonstrate the superiority of our method for image segmentation on a few synthetic and real images, both B&W and color. Our other examples include the concatenation of edges in a cluttered scene (perceptual grouping) and the organization of an image database for the purpose of multiview 3D object recognition.
Yoram Gdalyahu, Daphna Weinshall, Michael Werman
IEEE Trans. Pattern Anal. Mach. Intell.2
2000 Perceptual Grouping and Segmentation by Stochastic Clustering
abstract
We use cluster analysis as a unifying principle for problems from low, middle and high level vision. The clustering problem is viewed as graph partitioning, where nodes represent data elements and the weights of the edges represent pairwise similarities. Our algorithm generates samples of cuts in this graph, by using David Karger's contraction algorithm, and computes an "average" cut which provides the basis for our solution to the clustering problem. The stochastic nature of our method makes it robust against noise, including accidental edges and small spurious clusters. The complexity of our algorithm is very low: O(N log/sup 2/ N)for N objects and a fired accuracy level. Without additional computational cost, our algorithm provides a hierarchy of nested partitions. We demonstrate the superiority of our method for image segmentation on a few real color images. Our second application includes the concatenation of edges in a cluttered scene (perceptual grouping), where we show that the same clustering algorithm achieves as good a grouping, if not better as more specialized methods.
Yoram Gdalyahu, Noam Shental, Daphna Weinshall
CVPR3
2000 Classification with Nonmetric Distances: Image Retrieval and Class Representation
abstract
A key problem in appearance-based vision is understanding how to use a set of labeled images to classify new images. Systems that model human performance, or that use robust image matching methods, often use nonmetric similarity judgments; but when the triangle inequality is not obeyed, most pattern recognition techniques are not applicable. Exemplar-based (nearest-neighbor) methods can be applied to a wide class of nonmetric similarity functions. The key issue, however, is to find methods for choosing good representatives of a class that accurately characterize it. We show that existing condensing techniques are ill-suited to deal with nonmetric dataspaces. We develop techniques for solving this problem, emphasizing two points: First, we show that the distance between images is not a good measure of how well one image can represent another in nonmetric spaces. Instead, we use the vector correlation between the distances from each image to other previously seen images. Second, we show that in nonmetric spaces, boundary points are less significant for capturing the structure of a class than in Euclidean spaces. We suggest that atypical points may be more important in describing classes. We demonstrate the importance of these ideas to learning that generalizes from experience by improving performance. We also suggest ways of applying parametric techniques to supervised learning problems that involve a specific nonmetric distance functions, showing how to generalize the idea of linear discriminant functions in a way that may be more useful in nonmetric spaces.
David Jacobs 0001, Daphna Weinshall, Yoram Gdalyahu
IEEE Trans. Pattern Anal. Mach. Intell.2
1999 Stochastic Image Segmentation by Typical Cuts
abstract
We present a stochastic clustering algorithm which uses pairwise similarity of elements, based on a new graph theoretical algorithm for the sampling of cuts in graphs. The stochastic nature of our method makes it robust against noise, including accidental edges and small spurious clusters. We demonstrate the robustness and superiority of our method for image segmentation on a few synthetic examples where other recently proposed methods (such as normalized-cut) fail. In addition, the complexity of our method is lower. We describe experiments with real images showing good segmentation results.
Yoram Gdalyahu, Daphna Weinshall, Michael Werman
CVPR2
1999 Motion of disturbances: detection and tracking of multi-body non-rigid motion
Gilad Halevy, Daphna Weinshall
Mach. Vis. Appl.2
1999 Flexible Syntactic Matching of Curves and Its Application to Automatic Hierarchical Classification of Silhouettes
abstract
Curve matching is one instance of the fundamental correspondence problem. Our flexible algorithm is designed to match curves under substantial deformations and arbitrary large scaling and rigid transformations. A syntactic representation is constructed for both curves and an edit transformation which maps one curve to the other is found using dynamic programming. We present extensive experiments where we apply the algorithm to silhouette matching. In these experiments, we examine partial occlusion, viewpoint variation, articulation, and class matching (where silhouettes of similar objects are matched). Based on the qualitative syntactic matching, we define a dissimilarity measure and we compute it for every pair of images in a database of 121 images. We use this experiment to objectively evaluate our algorithm. First, we compare our results to those reported by others. Second, we use the dissimilarity values in order to organize the image database into shape categories. The veridical hierarchical organization stands as evidence to the quality of our matching and similarity estimation.
Yoram Gdalyahu, Daphna Weinshall
IEEE Trans. Pattern Anal. Mach. Intell.2
1998 Automatic Hierarchical Classification of Silhouettes of 3D Objects
abstract
The organization of image databases can rely upon different aspects of image similarity. Here we extract silhouettes from images of three dimensional objects, and rely upon curve similarity for image classification. Our scheme avoids the embedding of images in a vector space. Instead, we propose a curve dissimilarity measure which relies upon a novel curve matching syntactic algorithm, and use it to represent the database as a complete graph, with nodes representing the images and dissimilarity values assigning weights to the edges. A robust clustering algorithm, which is based on a physical ferromagnet model, is used to find the hierarchical structure underlying the collection of images. We tested our scheme with a database of 90 real images of 6 objects, some of them very different, others rather similar. We get a perfect hierarchical classification of these images into 6 classes of objects belonging to 3 different families.
Yoram Gdalyahu, Daphna Weinshall
CVPR2
1998 Flexible Syntactic Matching of Curves
Yoram Gdalyahu, Daphna Weinshall
ECCV (2)2
1998 From Reference Frames to Reference Planes: Multi-View Parallax Geometry and Applications
Michal Irani, P. Anandan 0001, Daphna Weinshall
ECCV (2)3
1998 Condensing Image Databases when Retrieval is Based on Non-Metric Distances
abstract
One of the key problems in appearance-based vision is understanding how to use a set of labeled images to classify new images. Classification systems that can model human performance, or that use robust image matching methods, often make use of similarity judgments that are non-metric but when the triangle inequality is not obeyed, most existing pattern recognition techniques are not applicable. We note that exemplar-based (or nearest-neighbor) methods can be applied naturally when using a wide class of non-metric similarity functions. The key issue, however, is to find methods for choosing good representatives of a class that accurately characterize it. We note that existing condensing techniques for finding class representatives are ill-suited to deal with non-metric dataspaces. We then focus on developing techniques for solving this problem, emphasizing two points: First, we show that the distance between two images is not a good measure of how well one image can represent another in non-metric spaces. Instead, we use the vector correlation between the distances from each image to other previously seen images. Second, we show that in non-metric spaces, boundary points are less significant for capturing the structure of a class than they are in Euclidean spaces. We suggest that atypical points may be more important in describing classes. We demonstrate the importance of these ideas to learning that generalizes from experience by improving performance using both synthetic and real images.
David Jacobs 0001, Daphna Weinshall, Yoram Gdalyahu
ICCV2
1998 A Randomized Algorithm for Pairwise Clustering
Yoram Gdalyahu, Daphna Weinshall, Michael Werman
NIPS2
1998 Mechanisms of Generalization in Perceptual Learning
Daphna Weinshall
NIPS2
1998 Classification in Non-Metric Spaces
Daphna Weinshall, David Jacobs 0001, Yoram Gdalyahu
NIPS1
1998 Dual Computation of Projective Shape and Camera Positions from Multiple Images
Stefan Carlsson, Daphna Weinshall
Int. J. Comput. Vis.2
1997 Motion of Disturbances: Detection and Tracking of multi-Body non-Rigid Motion
abstract
We present a new approach to the tracking of very non rigid patterns of motion, such as water flowing down a stream. The algorithm is based on a "disturbance map", which is obtained by linearly subtracting the temporal average of the previous frames from the new frame. Every local motion creates a disturbance having the form of a wave, with a "head" at the present position of the motion and a historical "tail" that indicates the previous locations of that motion. These disturbances serve as loci of attraction for "tracking particles" that are scattered throughout the image. The algorithm is very fast and can be performed in real time. We provide excellent tracking results on various complex sequences, using both stabilized and moving cameras, showing: a busy ant column, waterfalls, rapids and, flowing streams, shoppers in a mall, and cars in a traffic intersection.
Gilad Halevy, Daphna Weinshall
CVPR2
1997 Using Bilateral Symmetry to Improve 3D Reconstruction from Image Sequences
Hagit Hel-Or, Daphna Weinshall
Comput. Vis. Image Underst.2
1997 On View Likelihood and Stability
abstract
We define two measures on views: view likelihood and view stability. View likelihood measures the probability that a certain view of a given 3D object is observed; it may be used to identify typical, or "characteristic" views. View stability measures how little the-image changes as the viewpoint is slightly perturbed; it may be used to identify "generic" views. Both definitions are shown to be identical up to the prior probability of camera orientations, and determined by the 2D metric used to compare images. We analytically derive the stability and likelihood measures for two feature-based 2D metrics, where the most stable and most likely view is shown to be the flattest view of the 3D shape. Incorporating view likelihood or stability in 3D object recognition and 3D reconstruction increases the chance of robust performance. In particular, we propose to use these measures to enhance 3D object recognition and 3D reconstruction algorithms, by adding a second step where the most likely solution is selected among all feasible solutions. These applications are demonstrated using simulated and real images.
Daphna Weinshall, Michael Werman
IEEE Trans. Pattern Anal. Mach. Intell.1
1996 Measures for Silhouettes Resemblance and Representative Silhouettes of Curved Objects
Yoram Gdalyahu, Daphna Weinshall
ECCV (2)2
1996 Duality of Multi-Point and Multi-Frame Geometry: Fundamental Shape Matrices and Tensors
Daphna Weinshall, Michael Werman, Amnon Shashua
ECCV (2)1
1996 Complexity of Indexing: Efficient and Learnable Large Database Indexing
Michael Werman, Daphna Weinshall
ECCV (1)2
1996 Distance Metric Between 3D Models and 2D Images for Recognition and Classification
abstract
Similarity measurements between 3D objects and 2D images are useful for the tasks of object recognition and classification. The authors distinguish between two types of similarity metrics: metrics computed in image-space (image metrics) and metrics computed in transformation-space (transformation metrics). Existing methods typically use image metrics; namely, metrics that measure the difference in the image between the observed image and the nearest view of the object. Example for such a measure is the Euclidean distance between feature points in the image and their corresponding points in the nearest view. (This measure can be computed by solving the exterior orientation calibration problem.) In this paper the authors introduce a different type of metrics: transformation metrics. These metrics penalize for the deformations applied to the object to produce the observed image. In particular, the authors define a transformation metric that optimally penalizes for "affine deformations" under weak-perspective. A closed-form solution, together with the nearest view according to this metric, are derived. The metric is shown to be equivalent to the Euclidean image metric, in the sense that they bound each other from both above and below. It therefore provides an easy-to-use closed-form approximation for the commonly-used least-squares distance between models and images. The authors demonstrate an image understanding application, where the true dimensions of a photographed battery charger are estimated by minimizing the transformation metric.
Ronen Basri, Daphna Weinshall
IEEE Trans. Pattern Anal. Mach. Intell.2
1995 Linear and Incremental Acquisition of Invariant Shape Models From Image Sequences
abstract
We show how to automatically acquire Euclidian shape representations of objects from noisy image sequences under weak perspective. The proposed method is linear and incremental, requiring no more than pseudoinverse. A nonlinear, but numerically sound preprocessing stage is added to improve the accuracy of the results even further. Experiments show that attention to noise and computational techniques improves the shape results substantially with respect to previous methods proposed for ideal images.>
Daphna Weinshall, Carlo Tomasi
IEEE Trans. Pattern Anal. Mach. Intell.1
1995 Similarity and Affine Invariant Distances Between 2D Point Sets
abstract
We develop expressions for measuring the distance between 2D point sets, which are invariant to either 2D affine transformations or 2D similarity transformations of the sets, and assuming a known correspondence between the point sets. We discuss the image normalization to be applied to the images before their comparison so that the computed distance is symmetric with respect to the two images. We then give a general (metric) definition of the distance between images, which leads to the same expressions for the similarity and affine cases. This definition avoids ad hoc decisions about normalization. Moreover, it makes it possible to compute the distance between images under different conditions, including cases where the images are treated asymmetrically. We demonstrate these results with real and simulated images.>
Michael Werman, Daphna Weinshall
IEEE Trans. Pattern Anal. Mach. Intell.2
1994 Shape from motion algorithms: a comparative analysis of scaled orthography and perspective
Boubakeur Boufama, Daphna Weinshall, Michael Werman
ECCV (1)2
1994 Stability and Likelihood of Views of Three Dimensional Objects
Daphna Weinshall, Michael Werman, Naftali Tishby
ECCV (1)1
1994 Utilizing symmetry in the reconstruction of three-dimensional shape from noisy images
Hagit Hel-Or, Daphna Weinshall
ECCV (1)2
1994 Similarity and affine distance between 2D point sets
abstract
We develop expressions for measuring the distance between 2D point sets, which are invariant to either 2D affine transformations or 2D similarity transformations of the sets, and assuming a known correspondence between the point sets. Moreover, it makes it possible to compute the distance between images under different conditions, including cases where the images are treated asymmetrically. We demonstrate these results with real images.
Michael Werman, Daphna Weinshall
ICPR (1)2
1993 Model-based invariants for 3D vision
abstract
A hierarchical representation of the 3D shape of objects that is invariant to affine and rigid 3D transformations is described. Model-based invariant functions of general 3D objects are defined and constructed using this representation. A linear algorithm for invariant structure from motion from a sequence of images is outlined. The model-based invariant functions are used as model-based object recognition operators, operating on raw image features. An efficient invariant implementation of alignment and geometric hashing is discussed.>
Daphna Weinshall
CVPR1
1993 Distance metric between 3D models and 2D images for recognition and classification
abstract
A transformation metric to measure the similarity between 3-D models and 2-D images is proposed. The transformation metric measures the amount of affine deformation applied to the object to produce the given image. A simple, closed-form solution for this metric is presented. This solution is optimal in transformation space, and it is used to bound the image metric from both above and below. The transformation metric can be used in several different ways in recognition and classification tasks.>
Daphna Weinshall, Ronen Basri
CVPR1
1993 3D object recognition by indexing structural invariants from multiple views
abstract
The authors present a method for 3-D object recognition from 2-D image sequences. The system uses feature points tracked over three or more views to compute structural invariants, which serve as 3-D shape representations. Object recognition is performed by using these Euclidean invariants as indices into a high-dimensional shape table. The use of indexing eliminates any need for matching models to images. In addition, the representation of 3-D objects is extracted from 2-D views, eliminating the cumbersome burden of having to obtain 3-D models. The proposed scheme was implemented using a mixed database of real and simulated objects. Experiments are outlined that show good recognition results on real objects and simulated objects corrupted with noise.>
Rakesh Mohan, Daphna Weinshall, Ramesh R. Sarukkai
ICCV2
1993 Linear and incremental acquisition of invariant shape models from image sequences
abstract
The authors show how to automatically acquire similarity-invariant shape representations of objects from noisy image sequences under a weak perspective. The incremental nature of the method makes it possible to process images one at a time, moving away from the storage-intensive batch methods of the past. It is based on the observation that the trajectories that points on the object form in weak-perspective image sequences are linear combinations of three of the trajectories themselves, and that the coefficients of the linear combinations represent shape in an affine-invariant basis. A nonlinear but numerically sound preprocessing state is added to improve the accuracy of the results even further. Experiments showed that attention to noise and computational techniques improved the shape results substantially with respect to previous methods.>
Daphna Weinshall, Carlo Tomasi
ICCV1
1993 Model-based invariants for 3-D vision
Daphna Weinshall
Int. J. Comput. Vis.1
1992 Local shape approximation from shading
abstract
Exact shape cannot be inferred from a local analysis of shading, but for shape interpolation, a crude local approximation may be sufficient. The author explores the limits of such local approximations that are easy to compute. In particular, the shape of shading is used to approximate the surface in areas of monotonic change of intensity. This analysis is complemented by a method to compute the direction of a single point light source from the shading on occluding contours.>
Daphna Weinshall
CVPR1
1992 Shortcuts in shape classification from two images
Daphna Weinshall
CVGIP Image Underst.1
1991 Direct Computation of Qualitative 3-D Shape and Motion Invariants
abstract
Structure from motion often refers to the computation of three-dimensional structure from a matched sequence of images. However, a depth map of a surface is difficult to compute and may not be a good representation for storage and recognition. Given matched images it is shown that the sign of the normal curvature in a given direction at a given point in the image can be computed from a simple difference of slopes of line segments in one image. Using this result, local surface patches can be classified as convex, concave, cylindrical, hyperbolic (saddle point), or planar. At the same time, the translational component of the optical flow, from which the focus of expansion can be computed, is obtained.>
Daphna Weinshall
IEEE Trans. Pattern Anal. Mach. Intell.1
1990 Direct computation of qualitative 3D shape and motion invariants
abstract
The author shows that the sign of the normal curvature in a given direction at a given point in the image can be computed from a simple difference of slopes of line-segments in one image. Using this result, local surface patches can be classified as convex, concave, parabolic (cylindrical), hyperbolic (saddle point) or planar. At the same time the translational component of the optical flow is obtained, from which the focus of expansion can be computed. In addition, the axes of principal curvature and the axes of zero curvature are obtained.>
Daphna Weinshall
ICCV1
1990 Qualitative Structure From Motion
Daphna Weinshall
NIPS1
1990 Qualitative Depth from Stereo, with Applications
Daphna Weinshall
Comput. Vis. Graph. Image Process.1
1989 A Self-Organizing Multiple-View Representations of 3D Objects
Daphna Weinshall, Shimon Edelman, Heinrich H. Bülthoff
NIPS1
1989 Integration of vision modules and labeling of surface discontinuities
abstract
It is assumed that a major goal of the early vision modules and their integration is to deliver a cartoon of the discontinuities in the scene and to label them in terms of their physical origin. The output of each of the vision modules is noisy, possibly sparse, and sometimes not unique. The authors suggest the use of a coupled Markov random field (MRF) at the output of each module (image cues)-stereo, motion, color, and texture-to achieve two goals: first, to counteract the noise and fill in sparse data, and secondly, to integrate the image within each MRF to find the module discontinuities and align them with the intensity edges. The authors outline a theory of how to label the discontinuities in terms of depth, orientation, albedo, illumination, and specular discontinuities. They present labeling results using a simple linear classifier operating on the output of the MRF associated with each vision module and coupled to the image data. The classifier has been trained on a small set of a mixture of synthetic and real data.>
Edward B. Gamble, Davi Geiger, Tomaso A. Poggio, Daphna Weinshall
IEEE Trans. Syst. Man Cybern.4
1988 Qualitative depth from vertical and horizontal binocular disparities, in agreement with psychophysical evidence
abstract
The author concentrates on the problem of obtaining depth information from binocular disparities. It is motivated by the fact that implementing registration algorithms and using the results for depth computations is hard in practice with real images due to noise and quantization errors. It is shown that qualitative depth information can be obtained from stereo disparities with almost no computations, and with no prior knowledge (or computation) of camera parameters. The only constraint is that the epipolar plane of the fixation point includes the X-axes of both cameras. Two expressions are derived that order all matched points in the images in two distinct depth-consistent ways from image coordinates only. One is a tilt-related order lambda , which depends only on the polar angles of the matched points, the other is a depth-related order chi . Using lambda for tilt estimation and point separation (in depth) demonstrates some anomalies and unusual characteristics that have been observed in psychophysical experiments, most notably the induced size effect.>
Daphna Weinshall
CVPR1
1988 Application Of Qualitative Depth And Shape From Stereo
abstract
Obtaining exact depth from binocular dis- parities is hard if camera calibration is needed. We will show that qualitative depth and shape information can be obtained from stereo disparities with almost no computations, and with no prior knowledge (or computation) of camera's parameters. First, we derive two expressions that order matched points in the images in two distinct depth-consistent ways from image coor- dinates only. Their use demonstrates some anomalies that have been observed in psychophysical experiments, most notably the induced size effect. Second, we apply the same approach to estimate some qualitative behavior of the normal to the surface of any object in the field of view. In a similar way we develop an algorithm to compute axes of zero-curvature from disparities alone. The algorithm is shown to be quite robust against viola- tions of the basic assumptions of the computation on synthetic images, with relatively large controlled deviations. It performs almost as well on real images, as demonstrated on an image of four cans at different orientations. In this example, the true zero- curvature axes have been found for three cans, and an estimate with some small error has been found for the fourth.
Daphna Weinshall
ICCV1
1988 Labeling edges with a linear network: An integration of low-level vision modules
Davi Geiger, Daphna Weinshall
Neural Networks2