VLDB 2026 Research / reviewers in the wild / expert
Shimon Ullman
dblp:93/2158
· DBLP profile ↗
65ranked-venue papers
3as first author
8since 2021 · last 2025
0000-0003-4331-298XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 64 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | From Infants to AI: Incorporating Infant-like Learning in Models Boosts Efficiency and Generalization in Learning Social Prediction Tasks
Shify Treger, Shimon Ullman |
CogSci | 2 |
| 2025 | Seeing More with Less: Human-like Representations in Vision ModelsabstractLarge multimodal models (LMMs) typically process visual inputs with uniform resolution across the entire field of view, leading to inefficiencies when non-critical image regions are processed as precisely as key areas. Inspired by the human visual system’s foveated approach, we apply a sampling method to leading architectures such as MDETR, BLIP2, InstructBLIP, LLaVA, and ViLT, and evaluate their performance with variable (foveated) resolution inputs. Results show that foveated sampling boosts accuracy in visual tasks like question answering and object detection under tight pixel budgets, improving performance by up to 2.7% on the GQA dataset, 2.1% on SEED-Bench, and 2.0% on VQAv2 compared to uniform sampling. Furthermore, we show that indiscriminate resolution increases yield diminishing returns, with models achieving up to 80% of their full capability using just 3% of the pixels, even on complex tasks. Foveated sampling prompts more human-like processing within models, such as neuronal selectivity and globally acting self-attention in vision transformers. This paper provides a foundational analysis of foveated sampling’s impact on existing models, suggesting that more efficient architectural adaptations, mimicking human visual processing, are a promising research venue for the community. Potential applications of our findings center low power minimal bandwidth devices (such as UAVs and edge devices), where compact and efficient vision is critical. Andrey Gizdov, Shimon Ullman, Daniel Harari |
CVPR | 2 |
| 2025 | Teaching VLMs to Localize Specific Objects from In-Context ExamplesabstractVision-Language Models (VLMs) have shown remarkable capabilities across diverse visual tasks, including image recognition, video understanding, and Visual Question Answering (VQA) when explicitly trained for these tasks. Despite these advances, we find that present-day VLMs (including the proprietary GPT-4o) lack a fundamental cognitive ability: learning to localize specific objects in a scene by taking into account the context. In this work, we focus on the task of few-shot personalized localization, where a model is given a small set of annotated images (in-context examples) -- each with a category label and bounding box -- and is tasked with localizing the same object type in a query image. Personalized localization can be particularly important in cases of ambiguity of several related objects that can respond to a text or an object that is hard to describe with words. To provoke personalized localization abilities in models, we present a data-centric solution that fine-tunes them using carefully curated data from video object tracking datasets. By leveraging sequences of frames tracking the same object across multiple shots, we simulate instruction-tuning dialogues that promote context awareness. To reinforce this, we introduce a novel regularization technique that replaces object labels with pseudo-names, ensuring the model relies on visual context rather than prior knowledge. Our method significantly enhances the few-shot localization performance of recent VLMs ranging from 7B to 72B in size, without sacrificing generalization, as demonstrated on several benchmarks tailored towards evaluating personalized localization abilities. This work is the first to explore and benchmark personalized few-shot localization for VLMs -- exposing critical weaknesses in present-day VLMs, and laying a foundation for future research in context-driven vision-language applications. Sivan Doveh, Nimrod Shabtay, Eli Schwartz, Hilde Kuehne, Raja Giryes, Rogério Feris, Leonid Karlinsky, James R. Glass, Assaf Arbelle, Shimon Ullman, Muhammad Jehanzeb Mirza |
ICCV | 10 |
| 2024 | Biologically Inspired Learning Model for Instructed VisionabstractAs part of the effort to understand how the brain learns, ongoing research seeks to combine biological knowledge with current artificial intelligence (AI) modeling in an attempt to find an efficient biologically plausible learning scheme. Current models often use a cortical-like combination of bottom-up (BU) and top-down (TD) processing, where the TD part carries feedback signals for learning. However, in the visual cortex, the TD pathway plays a second major role in visual attention, by guiding the visual process toward locations and tasks of interest. A biological model should therefore integrate both learning and visual guidance. We introduce a model that uses a cortical-like combination of BU and TD processing that naturally integrates the two major functions of the TD stream. This integration is achieved through an appropriate connectivity pattern between the BU and TD streams, a novel processing cycle that uses the TD stream twice, and a 'Counter-Hebb' learning mechanism that operates across both streams. We show that the 'Counter-Hebb' mechanism can provide an exact backpropagation synaptic modification. Additionally, our model can effectively guide the visual stream to perform a task of interest, achieving competitive performance on standard multi-task learning benchmarks compared to AI models. The successful combination of learning and visual guidance could provide a new view on combining BU and TD processing in human vision and suggests possible directions for both biologically plausible models and artificial instructed models, such as vision-language models (VLMs). Roy Abel, Shimon Ullman |
NeurIPS | 2 |
| 2023 | Teaching Structured Vision & Language Concepts to Vision & Language ModelsabstractVision and Language ($VL$) models have demonstrated remarkable zero-shot performance in a variety of tasks. However, some aspects of complex language understanding still remain a challenge. We introduce the collective notion of Structured Vision & Language Concepts (SVLC) which includes object attributes, relations, and states which are present in the text and visible in the image. Recent studies have shown that even the best$VL$models struggle with SVLC. A possible way of fixing this issue is by collecting dedicated datasets for teaching each SVLC type, yet this might be expensive and time-consuming. Instead, we propose a more elegant data-driven approach for enhancing$VL$models' understanding of SVLCs that makes more effective use of existing$VL$pre-training datasets and does not require any additional data. While automatic understanding of image structure still remains largely unsolved, language structure is much better modeled and understood, allowing for its effective utilization in teaching$VL$models. In this paper, we propose various techniques based on language structure understanding that can be used to manipulate the textual part of off-the-shelf paired$VL$datasets.$VL$models trained with the updated data exhibit a significant improvement of up to 15% in their SVLC understanding with only a mild degradation in their zero-shot capabilities both when training from scratch or fine-tuning a pre-trained model. Our code and pretrained models are available at: https://github.com/SivanDoveh/TSVLC Sivan Doveh, Assaf Arbelle, Sivan Harary, Eli Schwartz, Roei Herzig, Raja Giryes, Rogério Feris, Rameswar Panda, Shimon Ullman, Leonid Karlinsky |
CVPR | 9 |
| 2023 | Dense and Aligned Captions (DAC) Promote Compositional Reasoning in VL ModelsabstractVision and Language (VL) models offer an effective method for aligning representation spaces of images and text allowing for numerous applications such as cross-modal retrieval, visual and multi-hop question answering, captioning, and many more. However, the aligned image-text spaces learned by all the popular VL models are still suffering from the so-called 'object bias' - their representations behave as 'bags of nouns' mostly ignoring or downsizing the attributes, relations, and states of objects described/appearing in texts/images. Although some great attempts at fixing these `compositional reasoning' issues were proposed in the recent literature, the problem is still far from being solved. In this paper, we uncover two factors limiting the VL models' compositional reasoning performance. These two factors are properties of the paired VL dataset used for finetuning (or pre-training) the VL model: (i) the caption quality, or in other words 'image-alignment', of the texts; and (ii) the 'density' of the captions in the sense of mentioning all the details appearing on the image. We propose a fine-tuning approach for automatically treating these factors on a standard collection of paired VL data (CC3M). Applied to CLIP, we demonstrate its significant compositional reasoning performance increase of up to $\sim27$\% over the base model, up to $\sim20$\% over the strongest baseline, and by $6.7$\% on average. Our code is provided in the Supplementary and would be released upon acceptance. Sivan Doveh, Assaf Arbelle, Sivan Harary, Roei Herzig, Donghyun Kim 0006, Paola Cascante-Bonilla, Amit Alfassy, Rameswar Panda, Raja Giryes, Rogério Feris, Shimon Ullman, Leonid Karlinsky |
NeurIPS | 11 |
| 2021 | Detector-Free Weakly Supervised Grounding by SeparationabstractNowadays, there is an abundance of data involving images and surrounding free-form text weakly corresponding to those images. Weakly Supervised phrase-Grounding (WSG) deals with the task of using this data to learn to localize (or to ground) arbitrary text phrases in images without any additional annotations. However, most recent SotA methods for WSG assume an existence of a pre-trained object detector, relying on it to produce the ROIs for localization. In this work, we focus on the task of Detector-Free WSG (DF-WSG) to solve WSG without relying on a pre-trained detector. The key idea behind our proposed Grounding by Separation (GbS) method is synthesizing ‘text to image-regions’ associations by random alpha-blending of arbitrary image pairs and using the corresponding texts of the pair as conditions to recover the alpha map from the blended image via a segmentation network. At test time, this allows using the query phrase as a condition for a non-blended query image, thus interpreting the test image as a composition of a region corresponding to the phrase and the complement region. Our GbS shows an 8.5% accuracy improvement over previous DF-WSG SotA, for a range of benchmarks including Flickr30K, Visual Genome, and ReferIt, as well as a complementary improvement (above 7%) over the detector-based approaches for WSG. Assaf Arbelle, Sivan Doveh, Amit Alfassy, Joseph Shtok, Guy Lev, Eli Schwartz, Hilde Kuehne, Hila Levi, Prasanna Sattigeri, Rameswar Panda, Chun-Fu Chen 0001, Alexander M. Bronstein, Kate Saenko, Shimon Ullman, Raja Giryes, Rogério Feris, Leonid Karlinsky |
ICCV | 14 |
| 2021 | Multi-Task Learning By A Top-Down Control NetworkabstractAs the range of tasks performed by a general vision system expands, executing multiple tasks accurately and efficiently in a single network has become an important and still open problem. Recent computer vision approaches address this problem by branching networks, or by a channel-wise modulation of the network feature-maps with task specific vectors. We present a novel architecture that uses a dedicated top-down control network to modify the activation of all the units in the main recognition network in a manner that depends on the selected task, image content, and spatial location. We show the effectiveness of our scheme by achieving significantly better results than alternative state-of-the-art approaches on four datasets. We further demonstrate our advantages in terms of task selectivity, scaling the number of tasks and interpretability. Hila Levi, Shimon Ullman |
ICIP | 2 |
| 2020 | Cakewalk Sampling
Uri Patish, Shimon Ullman |
AAAI | 2 |
| 2020 | VQA With No Questions-Answers TrainingabstractMethods for teaching machines to answer visual questions have made significant progress in recent years, but current methods still lack important human capabilities, including integrating new visual classes and concepts in a modular manner, providing explanations for the answers and handling new domains without explicit examples. We propose a novel method that consists of two main parts: generating a question graph representation, and an answering procedure, guided by the abstract structure of the question graph to invoke an extendable set of visual estimators. Training is performed for the language part and the visual part on their own, but unlike existing schemes, the method does not require any training using images with associated questions and answers. This approach is able to handle novel domains (extended question types and new object classes, properties and relations) as long as corresponding visual estimators are available. In addition, it can provide explanations to its answers and suggest alternatives when questions are not grounded in the image. We demonstrate that this approach achieves both high performance and domain extensibility without any questions-answers training. Ben Zion Vatashsky, Shimon Ullman |
CVPR | 2 |
| 2019 | Efficient Coarse-to-Fine Non-Local Module for the Detection of Small Objects
Hila Levi, Shimon Ullman |
BMVC | 2 |
| 2018 | Action Classification via Concepts and AttributesabstractClasses in natural images tend to follow long tail distributions. This is problematic when there are insufficient training examples for rare classes. This effect is emphasized in compound classes, involving the conjunction of several concepts, such as those appearing in action-recognition datasets. In this paper, we propose to address this issue by learning how to utilize common visual concepts which are readily available. We detect the presence of prominent concepts in images and use them to infer the target labels instead of using visual features directly, combining tools from vision and natural-language processing. We validate our method on the recently introduced HICO dataset reaching a mAP of 31.54% and on the Stanford-40 Actions dataset, where the proposed method outperforms that obtained by direct visual features, obtaining an accuracy 83.12%. Moreover, the method provides for each class a semantically meaningful list of keywords and relevant image regions relating it to its constituent concepts. Amir Rosenfeld, Shimon Ullman |
ICPR | 2 |
| 2016 | Visual Concept Recognition and Localization via Iterative Introspection
Amir Rosenfeld, Shimon Ullman |
ACCV (5) | 2 |
| 2016 | Human Pose Estimation Using Deep Consensus Voting
Ita Lifshitz, Ethan Fetaya, Shimon Ullman |
ECCV (2) | 3 |
| 2015 | Graph Approximation and Clustering on a BudgetabstractWe consider the problem of learning from a similarity matrix (such as spectral clustering and low-dimensional embedding), when computing pairwise similarities are costly, and only a limited number of entries can be observed. We provide a theoretical analysis using standard notions of graph approximation, significantly generalizing previous results, which focused on spectral clustering with two clusters. We also propose a new algorithmic approach based on adaptive sampling, which experimentally matches or improves on previous methods, while being considerably more general and computationally cheaper. Ethan Fetaya, Ohad Shamir, Shimon Ullman |
AISTATS | 3 |
| 2015 | A model for full local image interpretation
Guy Ben-Yosef, Liav Assif, Daniel Harari, Shimon Ullman |
CogSci | 4 |
| 2015 | Do You See What I Mean? Visual Resolution of Linguistic AmbiguitiesabstractUnderstanding language goes hand in hand with the ability to integrate complex contextual information obtained via perception.In this work, we present a novel task for grounded language understanding: disambiguating a sentence given a visual scene which depicts one of the possible interpretations of that sentence.To this end, we introduce a new multimodal corpus containing ambiguous sentences, representing a wide range of syntactic, semantic and discourse ambiguities, coupled with videos that visualize the different interpretations for each sentence.We address this task by extending a vision model which determines if a sentence is depicted by a video.We demonstrate how such a model can be adjusted to recognize different interpretations of the same underlying sentence, allowing to disambiguate sentences in a unified fashion across the different ambiguity types. Yevgeni Berzak, Andrei Barbu, Daniel Harari, Boris Katz, Shimon Ullman |
EMNLP | 5 |
| 2015 | Learning Local Invariant Mahalanobis DistancesabstractFor many tasks and data types, there are natural transformations to which the data should be invariant or insensitive. For instance, in visual recognition, natural images should be insensitive to rotation and translation. This requirement and its implications have been important in many machine learning applications, and tolerance for image transformations was primarily achieved by using robust feature vectors. In this paper we propose a novel and computationally efficient way to learn a local Mahalanobis metric per datum, and show how we can learn a local invariant metric to any transformation in order to improve performance. Ethan Fetaya, Shimon Ullman |
ICML | 2 |
| 2013 | Learning to perceive coherent objects
Nimrod Dorfman, Daniel Harari, Shimon Ullman |
CogSci | 3 |
| 2013 | Minimal Nativism: How does cognitive development get off the ground?
Tomer D. Ullman, Josh Tenenbaum, Noah D. Goodman, Shimon Ullman, Elizabeth S. Spelke |
CogSci | 4 |
| 2012 | Modeling the Perception of Intentions
Barbara Tversky, Shimon Ullman, Dare A. Baldwin, Frank E. Pollick, Josh Tenenbaum, Tao Gao 0004, Peter C. Pantelis, David Pautler |
CogSci | 2 |
| 2012 | Using Linking Features in Learning Non-parametric Part Models
Leonid Karlinsky, Shimon Ullman |
ECCV (3) | 2 |
| 2010 | The chains model for detecting parts by their contextabstractDetecting an object part relies on two sources of information - the appearance of the part itself and the context supplied by surrounding parts. In this paper we consider problems in which a target part cannot be recognized reliably using its own appearance, such as detecting low-resolution hands, and must be recognized using the context of surrounding parts. We develop the `chains model' which can locate parts of interest in a robust and precise manner, even when the surrounding context is highly variable and deformable. In the proposed model, the relation between context features and the target part is modeled in a non-parametric manner using an ensemble of feature chains leading from parts in the context to the detection target. The method uses the configuration of the features in the image directly rather than through fitting an articulated 3-D model of the object. In addition, the chains are composable, meaning that new chains observed in the test image can be composed of sub-chains seen during training. Consequently, the model is capable of handling object poses which are infrequent, even non-existent, during training. We test the approach in different settings, including object parts detection, as well as complete object detection. The results show the advantages of the chains model for detecting and localizing parts of complex deformable objects. Leonid Karlinsky, Michael Dinerstein, Daniel Harari, Shimon Ullman |
CVPR | 4 |
| 2010 | Using body-anchored priors for identifying actions in single imagesabstractThis paper presents an approach to the visual recognition of human actions using only single images as input. The task is easy for humans but difficult for current approaches to object recognition, because action instances may be similar in terms of body pose, and often require detailed examination of relations between participating objects and body parts in order to be recognized. The proposed approach applies a two-stage interpretation procedure to each training and test image. The first stage produces accurate detection of the relevant body parts of the actor, forming a prior for the local evidence needed to be considered for identifying the action. The second stage extracts features that are ‘anchored’ to the detected body parts, and uses these features and their feature-to-part relations in order to recognize the action. The body anchored priors we propose apply to a large range of human actions. These priors allow focusing on the relevant regions and relations, thereby significantly simplifying the learning process and increasing recognition performance. Leonid Karlinsky, Michael Dinerstein, Shimon Ullman |
NIPS | 3 |
| 2010 | Learning to classify by ongoing feature selection
Dan Levi, Shimon Ullman |
Image Vis. Comput. | 2 |
| 2009 | Unsupervised feature optimization (UFO): Simultaneous selection of multiple features with their detection parametersabstractClass learning, both supervised and unsupervised, requires feature selection, which includes two main components. The first is the selection of a discriminative subset of features from a larger pool. The second is the selection of detection parameters for each feature to optimize classification performance. In this paper we present a method for the discovery of multiple classification features, their detection parameters and their consistent configurations, in the fully unsupervised setting. This is achieved by a global optimization of joint consistency between the features as a function of the detection parameters, without assuming any prior parametric model. We demonstrate how the proposed framework can be applied for learning different types of feature parameters, such as detection thresholds and geometric relations, resulting in the unsupervised discovery of informative configurations of objects parts. We test our approach on a wide range of classes and show good results. We also demonstrate how the approach can be used to unsupervisedly separate and learn visually similar sub-classes of a single category, such as facial views or hand poses. We use the approach to compare various criteria for feature consistency, including Mutual Information, Suspicious Coincidence, L2 and Jaccard index. Finally, we compare our approach to aparametric consistency optimization technique such as pLSA and show significantly better performance. Leonid Karlinsky, Michael Dinerstein, Shimon Ullman |
CVPR | 3 |
| 2009 | A hierarchical non-parametric method for capturing non-rigid deformations
Ady Ecker, Shimon Ullman |
Image Vis. Comput. | 2 |
| 2009 | Cortical Circuitry Implementing Graphical ModelsabstractIn this letter, we develop and simulate a large-scale network of spiking neurons that approximates the inference computations performed by graphical models. Unlike previous related schemes, which used sum and product operations in either the log or linear domains, the current model uses an inference scheme based on the sum and maximization operations in the log domain. Simulations show that using these operations, a large-scale circuit, which combines populations of spiking neurons as basic building blocks, is capable of finding close approximations to the full mathematical computations performed by graphical models within a few hundred milliseconds. The circuit is general in the sense that it can be wired for any graph structure, it supports multistate variables, and it uses standard leaky integrate-and-fire neuronal units. Following previous work, which proposed relations between graphical models and the large-scale cortical anatomy, we focus on the cortical microcircuitry and propose how anatomical and physiological aspects of the local circuitry may map onto elements of the graphical model implementation. We discuss in particular the roles of three major types of inhibitory neurons (small fast-spiking basket cells, large layer 2/3 basket cells, and double-bouquet neurons), subpopulations of strongly interconnected neurons with their unique connectivity patterns in different cortical layers, and the possible role of minicolumns in the realization of the population-based maximum operation. Shai Litvak, Shimon Ullman |
Neural Comput. | 2 |
| 2008 | Unsupervised Classification and Part Localization by Consistency Amplification
Leonid Karlinsky, Michael Dinerstein, Dan Levi, Shimon Ullman |
ECCV (2) | 4 |
| 2008 | From Aardvark to Zorro: A Benchmark for Mammal Image Classification
Michael Fink 0002, Shimon Ullman |
Int. J. Comput. Vis. | 2 |
| 2008 | Distinctive and compact features
Ayelet Akselrod-Ballin, Shimon Ullman |
Image Vis. Comput. | 2 |
| 2008 | Class-Based Feature Matching Across Unrestricted TransformationsabstractWe develop a novel method for class-based feature matching across large changes in viewing conditions. The method is based on the property that when objects share a similar part, the similarity is preserved across viewing conditions. Given a feature and a training set of object images, we first identify the subset of objects that share this feature. The transformation of the feature's appearance across viewing conditions is determined mainly by properties of the feature, rather than of the object in which it is embedded. Therefore, the transformed feature will be shared by approximately the same set of objects. Based on this consistency requirement, corresponding features can be reliably identified from a set of candidate matches. Unlike previous approaches, the proposed scheme compares feature appearances only in similar viewing conditions, rather than across different viewing conditions. As a result, the scheme is not restricted to locally planar objects or affine transformations. The approach also does not require examples of correct matches. We show that by using the proposed method, a dense set of accurate correspondences can be obtained. Experimental comparisons demonstrate that matching accuracy is significantly improved over previous schemes. Finally, we show that the scheme can be successfully used for invariant object recognition. Evgeniy Bart, Shimon Ullman |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2008 | Combined Top-Down/Bottom-Up SegmentationabstractWe construct a segmentation scheme that combines top-down with bottom-up processing. In the proposed scheme, segmentation and recognition are intertwined rather than proceeding in a serial manner. The top-down part applies stored knowledge about object shapes acquired through learning, whereas the bottom-up part creates a hierarchy of segmented regions based on uniformity criteria. Beginning with unsegmented training examples of class and non-class images, the algorithm constructs a bank of class-specific fragments and determines their figure-ground segmentation. This bank is then used to segment novel images in a top-down manner: the fragments are first used to recognize images containing class objects, and then to create a complete cover that best approximates these objects. The resulting segmentation is then integrated with bottom-up multi-scale grouping to better delineate the object boundaries. Our experiments, applied to a large set of four classes (horses, pedestrians, cars, faces), demonstrate segmentation results that surpass those achieved by previous top-down or bottom-up schemes. The main novel aspects of this work are the fragment learning phase, which efficiently learns the figure-ground labeling of segmentation fragments, even in training sets with high object and background variability; combining the top-down segmentation with bottom-up criteria to draw on their relative merits; and the use of segmentation to improve recognition. Eran Borenstein, Shimon Ullman |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Semantic Hierarchies for Recognizing Objects and PartsabstractThis paper describes the construction and use of a novel representation for the recognition of objects and their parts, the semantic hierarchy. Its advantages include improved classification performance, accurate detection and localization of object parts and sub-parts, and explicitly identifying the different appearances of each object part. The semantic hierarchy algorithm starts by constructing a minimal feature hierarchy and proceeds by adding semantically equivalent representatives to each node, using the entire hierarchy as a context for determining the identity and locations of added features. Part detection is obtained by a bottom-up top-down cycle. Unlike previous approaches, the semantic hierarchy learns to represent the set of possible appearances of object parts at all levels, and their statistical dependencies. The algorithm is fully automatic and is shown experimentally to substantially improve the recognition of objects and their parts. Boris Epshtein, Shimon Ullman |
CVPR | 2 |
| 2007 | Uncovering shared structures in multiclass classificationabstractThis paper suggests a method for multiclass learning with many classes by simultaneously learning shared characteristics common to the classes, and predictors for the classes in terms of these characteristics. We cast this as a convex optimization problem, using trace-norm regularization and study gradient-based optimization both for the linear case and the kernelized setting. Yonatan Amit, Michael Fink 0002, Nathan Srebro, Shimon Ullman |
ICML | 4 |
| 2006 | Satellite Features for the Classification of Visually Similar ClassesabstractWe show that the discrimination between visually similar classes often depends on the detection of socalled ‘satellite features’. These are local features which are not informative by themselves, and can only be detected reliably at locations specified relative to other features. This makes satellite features difficult to extract by current classification methods. We describe a novel scheme which can extract discriminative satellite features and use them to distinguish between visually similar classes. The algorithm first searches for a set of features ("anchor features") that can be found in all the similar classes. Such features can be detected because the classes are visually similar. The anchors are used to determine the locations of satellite features, which are extracted during learning and used in classification to distinguish between the similar classes. The algorithm is fully automatic, and is shown to work well for many categories of visually similar classes. Boris Epshtein, Shimon Ullman |
CVPR (2) | 2 |
| 2006 | Online multiclass learning by interclass hypothesis sharingabstractWe describe a general framework for online multiclass learning based on the notion of hypothesis sharing. In our framework sets of classes are associated with hypotheses. Thus, all classes within a given set share the same hypothesis. This framework includes as special cases commonly used constructions for multiclass categorization such as allocating a unique hypothesis for each class and allocating a single common hypothesis for all classes. We generalize the multiclass Perceptron to our framework and derive a unifying mistake bound analysis. Our construction naturally extends to settings where the number of classes is not known in advance but rather is revealed along the online learning process. We demonstrate the merits of our approach by comparing it to previous methods on both synthetic and natural datasets. 1. Michael Fink 0002, Shai Shalev-Shwartz, Yoram Singer, Shimon Ullman |
ICML | 4 |
| 2005 | Single-example Learning of Novel Classes using Representation by SimilarityabstractWe describe an object classification method that can learn from a single training example. In this method, a novel class is characterized by its similarity to a number of previously learned, familiar classes. We demonstrate that this similarity is well-preserved across different class instances. As a result, it generalizes well to new instances of the novel class. A simple comparison of the similarity patterns is therefore sufficient to obtain useful classification performance from a single training example. The similarity between the novel class and the familiar classes in the proposed method can be evaluated using a wide variety of existing classification schemes. It can therefore combine the merits of many different classification methods. Experiments on a database of 107 widely varying object classes demonstrate that the proposed method significantly improves the performance of the baseline algorithm. Evgeniy Bart, Shimon Ullman |
BMVC | 2 |
| 2005 | Cross-Generalization: Learning Novel Classes from a Single Example by Feature ReplacementabstractWe develop an object classification method that can learn a novel class from a single training example. In this method, experience with already learned classes is used to facilitate the learning of novel classes. Our classification scheme employs features that discriminate between class and non-class images. For a novel class, new features are derived by selecting features that proved useful for already learned classification tasks, and adapting these features to the new classification task. This adaptation is performed by replacing the features from already learned classes with similar features taken from the novel class. A single example of a novel class is sufficient to perform feature adaptation and achieve useful classification performance. Experiments demonstrate that the proposed algorithm can learn a novel class from a single training example, using 10 additional familiar classes. The performance is significantly improved compared to using no feature adaptation. The robustness of the proposed feature adaptation concept is demonstrated by similar performance gains across 107 widely varying object categories. Evgeniy Bart, Shimon Ullman |
CVPR (1) | 2 |
| 2005 | Identifying Semantically Equivalent Object FragmentsabstractWe describe a novel technique for identifying semantically equivalent parts in images belonging to the same object class, (e.g. eyes, license plates, aircraft wings etc.). The visual appearance of such object parts can differ substantially, and therefore, traditional image similarity-based methods are inappropriate for this task. The technique we propose is based on the use of common context. We first retrieve context fragments, which consistently appear together with a given input fragment in a stable geometric relation. We then use the context fragments in new images to infer the most likely position of equivalent parts. Given a set of image examples of objects in a class, the method can automatically learn the part structure of the domain - identify the main parts, and how their appearance changes across objects in the class. Two applications of the proposed algorithm are shown: the detection and identification of object parts and object recognition. Boris Epshtein, Shimon Ullman |
CVPR (1) | 2 |
| 2005 | Feature Hierarchies for Object ClassificationabstractThe paper describes a method for automatically extracting informative feature hierarchies for object classification, and shows the advantage of the features constructed hierarchically over previous methods. The extraction process proceeds in a top-down manner: informative top-level fragments are extracted first, and by a repeated application of the same feature extraction process the classification fragments are broken down successively into their own optimal components. The hierarchical decomposition terminates with atomic features that cannot be usefully decomposed into simpler features. The entire hierarchy, the different features and sub-features, and their optimal parameters, are learned during a training phase using training examples. Experimental comparisons show that these feature hierarchies are significantly more informative and better for classification compared with similar nonhierarchical features as well as previous methods for using feature hierarchies. Boris Epshtein, Shimon Ullman |
ICCV | 2 |
| 2004 | Image normalization by mutual informationabstractImage normalization refers to eliminating image variations (such as noise, illumination, or occlusion) that are related to conditions of image acquisition and are irrelevant to object identity. Image normalization can be used as a preprocessing stage to assist computer or human object perception. In this paper, a class-based image normalization method is proposed. Objects in this method are represented in the PCA basis, and mutual information is used to identify irrelevant principal components. These components are then discarded to obtain a normalized image which is not affected by the specific conditions of image acquisition. The method is demonstrated to produce visually pleasing results and to improve significantly the accuracy of known recognition algorithms. The use of mutual information is a significant advantage over the standard method of discarding components according to the eigenvalues, since eigenvalues correspond to variance and have no direct relation to the relevance of components to representation. An additional advantage of the proposed algorithm is that many types of image variations are handled in a unified framework. 1 Evgeniy Bart, Shimon Ullman |
BMVC | 2 |
| 2004 | View-Invariant Recognition Using Corresponding Object Fragments
Evgeniy Bart, Evgeny Byvatov, Shimon Ullman |
ECCV (2) | 3 |
| 2004 | Learning to Segment
Eran Borenstein, Shimon Ullman |
ECCV (3) | 2 |
| 2004 | Recognition invariance obtained by extended and invariant features
Shimon Ullman, Evgeniy Bart |
Neural Networks | 1 |
| 2003 | Object Recognition with Informative Features and Linear ClassificationabstractWe show that efficient object recognition can be obtained by combining informative features with linear classification. The results demonstrate the superiority of informative class-specific features, as compared with generic type features such as wavelets, for the task of object recognition. We show that information rich features can reach optimal performance with simple linear separation rules, while generic feature based classifiers require more complex classification schemes. This is significant because efficient and optimal methods have been developed for spaces that allow linear separation. To compare different strategies for feature extraction, we trained and compared classifiers working in feature spaces of the same low dimensionality, using two feature types (image fragments vs. wavelets) and two classification rules (linear hyperplane and a Bayesian network). The results show that by maximizing the individual information of the features, it is possible to obtain efficient classification by a simple linear separating rule, as well as more efficient learning. Michel Vidal-Naquet, Shimon Ullman |
ICCV | 2 |
| 2002 | Class-Specific, Top-Down Segmentation
Eran Borenstein, Shimon Ullman |
ECCV (2) | 2 |
| 1999 | Combining Class-Specific Fragments for Object ClassificationabstractWe describe an approach to object classification based on the conjunction of multiple classspecific object fragments detected in the image. The method represents members of a given class (such as a face, or a car) using combinations of common sub-structures, termed fragments. These fragments are partial 2-D patterns extracted from examples views of objects belonging to the class in question. An object view is covered by multiple, overlapping fragments of several types, and at multiple levels of complexity. We describe the detection of the individual fragments, and the combination of the fragments to detect complete objects. The combination of fragments to form a consistent overall arrangement is based in this scheme on a number of simple mechanisms: the use of overlapping fragments, a spatial `voting' scheme, and by imposing some constraints on the tolerated location of the fragments within the overall object view. We present experimental results of the application of the method to the detection of face and car views in cluttered scenes and to partially occluded objects. We present evidence that by combining fragments from different objects the method can deal successfully with intra-object variability within a class. The method is more economical and more resistant to occlusion and deformations than methods relying on global object views. Introduction In this paper we study the challenging task of detecting different objects from a given class (such as a face or a car) in an image. In addition to the unknown location and illumination of the object, the method must deal with intra-class variability between objects from the same class. The detection process must therefore cover a range of possible shapes, missing parts and additional clutter. To deal with the problem of shape variability and detect novel shapes of a given class, Turk & Pentland (1991) used the principal components of registered face views. Views of novel objects can be approximated by the superposition of several basis functions, or 'eigenfaces'. Poggio & Sung (1994) used a distribution-based modeling scheme for detecting faces in cluttered scenes. They represented a face view as a gray level vector with 283 components, and trained a multi-layered preceptron network to classify such data as face/non-face vectors. The generalization to novel shapes within the class is obtained in these schemes by the inherent generalization capacity of the neural network mechanism. Rowley, Baluja & Kanade (1995) and Lin, Kung & Lin (1996) BMVC99 Erez Sali, Shimon Ullman |
BMVC | 2 |
| 1999 | Computation of pattern invariance in brain-like structures
Shimon Ullman, Sergei Soloviev 0002 |
Neural Networks | 1 |
| 1998 | Recognizing Novel 3-D Objects Under New Illumination and Viewing Position Using a Small Number of ExamplesabstractA method is presented for class-based recognition using a small number of example views taken under several different viewing conditions. The main emphasis is on using a small number of examples. Previous work assumed that the set of examples is sufficient to span the entire space of possible objects, and that in generalizing to a new viewing conditions a sufficient number of previous examples under the new conditions will be available to the recognition system. Here we have considerably relaxed these assumptions and consequently obtained good class-based generalization from a small number of examples, even a single example view, for both viewing position and illumination changes. In addition, previous class-based approaches only focused on viewing position changes and did not deal with illumination changes. Here we used a class-based approach that can generalize for both illumination and viewing position changes. The method was applied to face and car model images. New views under viewing position and illumination changes were synthesized from a small number of examples. Erez Sali, Shimon Ullman |
ICCV | 2 |
| 1998 | Generalization to Novel Views: Universal, Class-based, and Model-based Processing
Yael Moses, Shimon Ullman |
Int. J. Comput. Vis. | 2 |
| 1997 | Face Recognition: The Problem of Compensating for Changes in Illumination DirectionabstractA face recognition system must recognize a face from a novel image despite the variations between images of the same face. A common approach to overcoming image variations because of changes in the illumination conditions is to use image representations that are relatively insensitive to these variations. Examples of such representations are edge maps, image intensity derivatives, and images convolved with 2D Gabor-like filters. Here we present an empirical study that evaluates the sensitivity of these representations to changes in illumination, as well as viewpoint and facial expression. Our findings indicated that none of the representations considered is sufficient by itself to overcome image variations because of a change in the direction of illumination. Similar results were obtained for changes due to viewpoint and expression. Image representations that emphasized the horizontal features were found to be less sensitive to changes in the direction of illumination. However, systems based only on such representations failed to recognize up to 20 percent of the faces in our database. Humans performed considerably better under the same conditions. We discuss possible reasons for this superiority and alternative methods for overcoming illumination effects in recognition. Yael Adini, Yael Moses, Shimon Ullman |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 1996 | Learning class regions by the union of ellipsoidsabstractIn many classification schemes objects are represented as points in multi-dimensional feature spaces. The classification scheme then attempts to discriminate between regions in the space occupied by objects of different classes. The performance of the classification method often depends on the shape of the class regions, e.g., whether or not they are linearly separable. In many practical cases, class regions have the structure of smooth low-dimensional manifolds. We develop a novel classification scheme that covers each class region by a set of ellipsoids that are oriented along the local orientation of the manifold. The scheme learns the class regions from sequential presentation of samples, and the ellipsoids are created and modified incrementally during the learning. In high dimensional feature spaces the ellipsoids cover can become significantly more efficient than alternative classification schemes. Michael Kositsky, Shimon Ullman |
ICPR | 2 |
| 1994 | Face Recognition: the Problem of Compensating for Changes in Illumination Direction
Yael Moses, Yael Adini, Shimon Ullman |
ECCV (1) | 3 |
| 1992 | Contour matching using local affine transformationsabstractA noniterative scheme for determining contour matches using locally affine transformations is proposed. The method assumes that contours are approximated by the orthographic projection of planar patches within oriented neighborhoods of varying sizes. For degenerate cases, a minimal matching solution is chosen closest to the minimal pure translation. Performance on noisy synthetic and natural contour imagery is reported.> Ivan A. Bachelder, Shimon Ullman |
CVPR | 2 |
| 1992 | Limitations of Non Model-Based Recognition Schemes
Yael Moses, Shimon Ullman |
ECCV | 2 |
| 1991 | A Pictorial Approach to Object Classification
Yerucham Shapira, Shimon Ullman |
IJCAI | 2 |
| 1991 | Linear Operator for Object Recognition
Ronen Basri, Shimon Ullman |
NIPS | 2 |
| 1991 | Recognition by Linear Combinations of ModelsabstractAn approach to visual object recognition in which a 3D object is represented by the linear combination of 2D images of the object is proposed. It is shown that for objects with sharp edges as well as with smooth bounding contours, the set of possible images of a given object is embedded in a linear space spanned by a small number of views. For objects with sharp edges, the linear combination representation is exact. For objects with smooth boundaries, it is an approximation that often holds over a wide range of viewing angles. Rigid transformations (with or without scaling) can be distinguished from more general linear transformations of the object by testing certain constraints placed on the coefficients of the linear combinations. Three alternative methods of determining the transformation that matches a model to a given image are proposed.> Shimon Ullman, Ronen Basri |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1990 | Grouping Contours by Iterated Pairing Networks
Amnon Shashua, Shimon Ullman |
NIPS | 2 |
| 1990 | Reading cursive handwriting by alignment of letter prototypes
Shimon Edelman, Tamar Flash, Shimon Ullman |
Int. J. Comput. Vis. | 3 |
| 1990 | Recognizing solid objects by alignment with an image
Daniel P. Huttenlocher, Shimon Ullman |
Int. J. Comput. Vis. | 2 |
| 1988 | The Alignment Of Objects With Smooth SurfacesabstractAbstract : This paper examines the recognition of rigid objects bounded by smooth surfaces, using an alignment approach. The projected image of such an object changes during rotation in a manner that is generally difficult to predict. An approach to this problem is suggested, using the 3-D surface curvature at the points along the silhouette. The curvature information requires a single number for each point along the object's silhouette, the magnitude of the curvature vector at the point. We have implemented this method, and tested it on images of complex 3-D objects. Models of the viewed objects were acquired using three images of each object. The implemented scheme was found to give accurate predictions of the objects' appearance for large transformations. Using this method, a small number of (viewer-centered) models can be used to predict the new appearance of an object from any given viewpoint. (JHD) Ronen Basri, Shimon Ullman |
ICCV | 2 |
| 1988 | Structural Saliency: The Detection Of Globally Salient Structures using A Locally Connected NetworkabstractCertain salient structures in images attract our immediate attention without requiring a systematic scan. We present a method for computing saliency by a simple iterative scheme, using a uniform network of locally connected processing elements. The network uses an optimization approach to produce a "saliency map," a representation of the image emphasizing salient locations. The main properties of the network are: (i) the computations are simple and local, (ii) globally salient structures emerge with a small number of iterations, and (iii) as a by-product of the computations, contours are smoothed and gaps are filled in. Amnon Shashua, Shimon Ullman |
ICCV | 2 |
| 1988 | Aligning A Model To An Image Using Minimal InformationabstractThe ability to determine the geometrical transformation which aligns a 3D object with its 2D image plays an important role in recognition and other visual tasks. We suggest that since usually the possible. transformations are constrained, they can be determined by using a small number of features. We prove that combinations of three co-planar points and lines can determine the alignment transformation uniquely or almost uniquely. Doron Shoham, Shimon Ullman |
ICCV | 2 |