EDBT 2026 Demo / reviewers in the wild / expert
Svetlana Lazebnik
dblp:55/5180
· DBLP profile ↗
76ranked-venue papers
13as first author
12since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 74 · 13 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 49 · 9 first-author · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | UnZipLoRA: Separating Content and Style from a Single ImageabstractThis paper introduces UnZipLoRA, a method for decomposing an image into its constituent subject and style, represented as two distinct LoRAs (Low-Rank Adaptations). Unlike existing personalization techniques that focus on either subject or style in isolation, or require separate training sets for each, UnZipLoRA disentangles these elements from a single image by training both the LoRAs simultaneously. UnZipLoRA ensures that the resulting LoRAs are compatible, i.e., they can be seamlessly combined using direct addition. UnZipLoRA enables independent manipulation and recontextualization of subject and style, including generating variations of each, applying the extracted style to new subjects, and recombining them to reconstruct the original image or create novel variations. To address the challenge of subject and style entanglement, UnZipLoRA employs a novel prompt separation technique, as well as column and block separation strategies to accurately preserve the characteristics of subject and style, and ensure compatibility between the learned LoRAs. Evaluation with human studies and quantitative metrics demonstrates UnZipLoRA's effectiveness compared to other state-of-the-art methods, including DreamBooth-LoRA, Inspiration Tree, and B-LoRA. Viraj Shah, Aiyu Cui, Svetlana Lazebnik |
ICCV | 4 |
| 2025 | A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint RewardsabstractTask specification for robotic manipulation in open-world environments is challenging, requiring flexible and adaptive objectives that align with human intentions and can evolve through iterative feedback. We introduce Iterative Keypoint Reward (IKER), a visually grounded, Python-based reward function that serves as a dynamic task specification. Our framework leverages VLMs to generate and refine these reward functions for multi-step manipulation tasks. Given RGB-D observations and free-form language instructions, we sample keypoints in the scene and generate a reward function conditioned on these keypoints. IKER operates on the spatial relationships between keypoints, leveraging commonsense priors about the desired behaviors, and enabling precise SE(3) control. We reconstruct real-world scenes in simulation and use the generated rewards to train reinforcement learning (RL) policies, which are then deployed into the real world-forming a real-to-sim-to-real loop. Our approach demonstrates notable capabilities across diverse scenarios, including both prehensile and non-prehensile tasks, showcasing multi-step task execution, spontaneous error recovery, and on-the-fly strategy adjustments. The results highlight IKER's effectiveness in enabling robots to perform multi-step tasks in dynamic environments through iterative reward shaping. Project Page: https://iker-robot.github.io/ Shivansh Patel, Xinchen Yin, Wenlong Huang, Shubham Garg, Hooshang Nayyeri, Li Fei-Fei 0001, Svetlana Lazebnik, Yunzhu Li |
ICRA | 7 |
| 2025 | Street TryOn: Learning In-the-Wild Virtual Try-On from Unpaired Person ImagesabstractMost virtual try-on research is motivated to serve the fashion business by generating images to demonstrate garments on studio models at a lower cost. However, virtual try-on should be a broader application that also allows customers to visualize garments on themselves using their own casual photos, known as in-the-wild try-on. Unfortunately, the existing methods, which achieve plausible results for studio try-on settings, perform poorly in the in-the-wild context. This is because these methods often require paired images (garment images paired with images of people wearing the same garment) for training. While such paired data is easy to collect from shopping websites for studio settings, it is difficult to obtain for in-the-wild scenes. In this work, we fill the gap by (1) introducing a Street-TryOn benchmark to support in-the-wild virtual try-on applications and (2) proposing a novel method to learn virtual try-on from a set of in-the-wild person images directly without requiring paired data. We tackle the unique challenges, including warping garments to more diverse human poses and rendering more complex backgrounds faithfully, by a novel DensePose warping correction method combined with diffusion-based conditional inpainting. Our experiments show competitive performance for standard studio try-on tasks and SOTA performance for street try-on and cross-domain try-on tasks. Aiyu Cui, Jay Mahajan, Viraj Shah, Preeti Gomathinayagam, Svetlana Lazebnik |
WACV | 6 |
| 2024 | Shadows Don't Lie and Lines Can't Bend! Generative Models Don't know Projective Geometry...for NowabstractGenerative models can produce impressively realistic images. This paper demonstrates that generated images have geometric features different from those of real images. We build a set of collections of generated images, prequalified to fool simple, signal-based classifiers into believing they are real. We then show that prequalified generated images can be identified reliably by classifiers that only look at geometric properties. We use three such classifiers. All three classifiers are denied access to image pixels, and look only at derived geometric features. The first classifier looks at the perspective field of the image, the second looks at lines detected in the image, and the third looks at relations between detected objects and shadows. Our procedure detects generated images more reliably than SOTA local signal based detectors, for images from a number of distinct generators. Saliency maps suggest that the classifiers can identify geometric problems reliably. We conclude that current generators cannot reliably reproduce geometric properties of real images. Ayush Sarkar, Hanlin Mai, Amitabh Mahapatra, Svetlana Lazebnik, David A. Forsyth, Anand Bhattad |
CVPR | 4 |
| 2024 | ZipLoRA: Any Subject in Any Style by Effectively Merging LoRAs
Viraj Shah, Nataniel Ruiz, Forrester Cole, Erika Lu, Svetlana Lazebnik, Yuanzhen Li, Varun Jampani |
ECCV (1) | 5 |
| 2024 | In Memoriam: Xiaoou Tang
Yasuyuki Matsushita, Svetlana Lazebnik, Jiri Matas |
Int. J. Comput. Vis. | 2 |
| 2022 | Revisiting Image-Language Networks for Open-Ended Phrase DetectionabstractMost existing work that grounds natural language phrases in images starts with the assumption that the phrase in question is relevant to the image. In this paper we address a more realistic version of the natural language grounding task where we must both identify whether the phrase is relevant to an image and localize the phrase. This can also be viewed as a generalization of object detection to an open-ended vocabulary, introducing elements of few- and zero-shot detection. We propose an approach for this task that extends Faster R-CNN to relate image regions and phrases. By carefully initializing the classification layers of our network using canonical correlation analysis (CCA), we encourage a solution that is more discerning when reasoning between similar phrases, resulting in over double the performance compared to a naive adaptation on three popular phrase grounding datasets, Flickr30K Entities, ReferIt Game, and Visual Genome, with test-time phrase vocabulary sizes of 5K, 32K, and 159K, respectively. Bryan A. Plummer, Kevin J. Shih, Svetlana Lazebnik, Stan Sclaroff, Kate Saenko |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2021 | Dressing in Order: Recurrent Person Image Generation for Pose Transfer, Virtual Try-on and Outfit EditingabstractWe proposes a flexible person generation framework called Dressing in Order (DiOr), which supports 2D pose transfer, virtual try-on, and several fashion editing tasks. The key to DiOr is a novel recurrent generation pipeline to sequentially put garments on a person, so that trying on the same garments in different orders will result in different looks. Our system can produce dressing effects not achievable by existing work, including different interactions of garments (e.g., wearing a top tucked into the bottom or over it), as well as layering of multiple garments of the same type (e.g., jacket over shirt over t-shirt). DiOr explicitly encodes the shape and texture of each garment, enabling these elements to be edited separately. Joint training on pose transfer and inpainting helps with detail preservation and coherence of generated garments. Extensive evaluations show that DiOr outperforms other recent methods like ADGAN [28] in terms of output quality, and handles a wide range of editing functions for which there is no direct supervision. Aiyu Cui, Daniel McKee, Svetlana Lazebnik |
ICCV | 3 |
| 2021 | GridToPix: Training Embodied Agents with Minimal SupervisionabstractWhile deep reinforcement learning (RL) promises freedom from hand-labeled data, great successes, especially for Embodied AI, require significant work to create supervision via carefully shaped rewards. Indeed, without shaped rewards, i.e., with only terminal rewards, present-day Embodied AI results degrade significantly across Embodied AI problems from single-agent Habitat-based PointGoal Navigation (SPL drops from 55 to 0) and two-agent AI2-THOR-based Furniture Moving (success drops from 58% to 1%) to three-agent Google Football-based 3 vs. 1 with Keeper (game score drops from 0.6 to 0.1). As training from shaped rewards doesn’t scale to more realistic tasks, the community needs to improve the success of training with terminal rewards. For this we propose GRIDTOPIX: 1) train agents with terminal rewards in gridworlds that generically mirror Embodied AI environments, i.e., they are independent of the task; 2) distill the learned policy into agents that reside in complex visual worlds. Despite learning from only terminal rewards with identical models and RL algorithms, GRIDTOPIX significantly improves results across tasks: from PointGoal Navigation (SPL improves from 0 to 64) and Furniture Moving (success improves from 1% to 25%) to football gameplay (game score improves from 0.1 to 0.6). GRIDTOPIX even helps to improve the results of shaped reward training. Unnat Jain, Iou-Jen Liu, Svetlana Lazebnik, Aniruddha Kembhavi, Luca Weihs, Alexander G. Schwing |
ICCV | 3 |
| 2021 | Interpretation of Emergent Communication in Heterogeneous Collaborative Embodied AgentsabstractCommunication between embodied AI agents has received increasing attention in recent years. Despite its use, it is still unclear whether the learned communication is interpretable and grounded in perception. To study the grounding of emergent forms of communication, we first introduce the collaborative multi-object navigation task ‘CoMON.' In this task, an ‘oracle agent' has detailed environment information in the form of a map. It communicates with a ‘navigator agent' that perceives the environment visually and is tasked to find a sequence of goals. To succeed at the task, effective communication is essential. CoMON hence serves as a basis to study different communication mechanisms between heterogeneous agents, that is, agents with different capabilities and roles. We study two common communication mechanisms and analyze their communication patterns through an egocentric and spatial lens. We show that the emergent communication can be grounded to the agent observations and the spatial structure of the 3D environment. Shivansh Patel, Saim Wani, Unnat Jain, Alexander G. Schwing, Svetlana Lazebnik, Manolis Savva, Angel X. Chang |
ICCV | 5 |
| 2021 | Bridging the Imitation Gap by Adaptive InsubordinationabstractIn practice, imitation learning is preferred over pure reinforcement learning whenever it is possible to design a teaching agent to provide expert supervision. However, we show that when the teaching agent makes decisions with access to privileged information that is unavailable to the student, this information is marginalized during imitation learning, resulting in an "imitation gap" and, potentially, poor results. Prior work bridges this gap via a progression from imitation learning to reinforcement learning. While often successful, gradual progression fails for tasks that require frequent switches between exploration and memorization. To better address these tasks and alleviate the imitation gap we propose 'Adaptive Insubordination' (ADVISOR). ADVISOR dynamically weights imitation and reward-based reinforcement learning losses during training, enabling on-the-fly switching between imitation and exploration. On a suite of challenging tasks set within gridworlds, multi-agent particle environments, and high-fidelity 3D simulators, we show that on-the-fly switching with ADVISOR outperforms pure imitation, pure reinforcement learning, as well as their sequential and parallel combinations. Luca Weihs, Unnat Jain, Iou-Jen Liu, Jordi Salvador, Svetlana Lazebnik, Aniruddha Kembhavi, Alexander G. Schwing |
NeurIPS | 5 |
| 2021 | Contextual Translation Embedding for Visual Relationship Detection and Scene Graph GenerationabstractRelations amongst entities play a central role in image understanding. Due to the complexity of modeling (subject, predicate, object) relation triplets, it is crucial to develop a method that can not only recognize seen relations, but also generalize to unseen cases. Inspired by a previously proposed visual translation embedding model, or VTransE [1] , we propose a context-augmented translation embedding model that can capture both common and rare relations. The previous VTransE model maps entities and predicates into a low-dimensional embedding vector space where the predicate is interpreted as a translation vector between the embedded features of the bounding box regions of the subject and the object. Our model additionally incorporates the contextual information captured by the bounding box of the union of the subject and the object, and learns the embeddings guided by the constraint predicate ≈ union (subject, object) - subject - object. In a comprehensive evaluation on multiple challenging benchmarks, our approach outperforms previous translation-based models and comes close to or exceeds the state of the art across a range of settings, from small-scale to large-scale datasets, from common to previously unseen relations. It also achieves promising results for the recently introduced task of scene graph generation. Zih-Siou Hung, Arun Mallya, Svetlana Lazebnik |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2020 | Memory-Efficient Incremental Learning Through Feature Adaptation
Ahmet Iscen, Jeffrey Zhang 0004, Svetlana Lazebnik, Cordelia Schmid |
ECCV (16) | 3 |
| 2020 | A Cordial Sync: Going Beyond Marginal Policies for Multi-agent Embodied Tasks
Unnat Jain, Luca Weihs, Eric Kolve, Ali Farhadi, Svetlana Lazebnik, Aniruddha Kembhavi, Alexander G. Schwing |
ECCV (5) | 5 |
| 2019 | Two Body Problem: Collaborative Visual Task CompletionabstractCollaboration is a necessary skill to perform tasks that are beyond one agent's capabilities. Addressed extensively in both conventional and modern AI, multi-agent collaboration has often been studied in the context of simple grid worlds. We argue that there are inherently visual aspects to collaboration which should be studied in visually rich environments. A key element in collaboration is communication that can be either explicit, through messages, or implicit, through perception of the other agents and the visual world. Learning to collaborate in a visual environment entails learning (1) to perform the task, (2) when and what to communicate, and (3) how to act based on these communications and the perception of the visual world. In this paper we study the problem of learning to collaborate directly from pixels in AI2-THOR and demonstrate the benefits of explicit and implicit modes of communication to perform visual tasks. Refer to our project page for more details: https://prior.allenai.org/projects/two-body-problem. Unnat Jain, Luca Weihs, Eric Kolve, Mohammad Rastegari, Svetlana Lazebnik, Ali Farhadi, Alexander G. Schwing, Aniruddha Kembhavi |
CVPR | 5 |
| 2019 | Combining Multiple Cues for Visual Madlibs Question Answering
Tatiana Tommasi, Arun Mallya, Bryan A. Plummer, Svetlana Lazebnik, Alexander C. Berg, Tamara L. Berg |
Int. J. Comput. Vis. | 4 |
| 2019 | Learning Two-Branch Neural Networks for Image-Text Matching TasksabstractImage-language matching tasks have recently attracted a lot of attention in the computer vision field. These tasks include image-sentence matching, i.e., given an image query, retrieving relevant sentences and vice versa, and region-phrase matching or visual grounding, i.e., matching a phrase to relevant regions. This paper investigates two-branch neural networks for learning the similarity between these two data modalities. We propose two network structures that produce different output representations. The first one, referred to as an embedding network, learns an explicit shared latent embedding space with a maximum-margin ranking loss and novel neighborhood constraints. Compared to standard triplet sampling, we perform improved neighborhood sampling that takes neighborhood information into consideration while constructing mini-batches. The second network structure, referred to as a similarity network, fuses the two branches via element-wise product and is trained with regression loss to directly predict a similarity score. Extensive experiments show that our networks achieve high accuracies for phrase localization on the Flickr30K Entities dataset and for bi-directional image-sentence retrieval on Flickr30K and MSCOCO datasets. Liwei Wang 0009, Yin Li 0003, Jing Huang 0014, Svetlana Lazebnik |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2018 | Two Can Play This Game: Visual Dialog With Discriminative Question Generation and AnsweringabstractHuman conversation is a complex mechanism with subtle nuances. It is hence an ambitious goal to develop artificial intelligence agents that can participate fluently in a conversation. While we are still far from achieving this goal, recent progress in visual question answering, image captioning, and visual question generation shows that dialog systems may be realizable in the not too distant future. To this end, a novel dataset was introduced recently and encouraging results were demonstrated, particularly for question answering. In this paper, we demonstrate a simple symmetric discriminative baseline, that can be applied to both predicting an answer as well as predicting a question. We show that this method performs on par with the state of the art, even memory net based methods. In addition, for the first time on the visual dialog dataset, we assess the performance of a system asking questions, and demonstrate how visual dialog can be generated from discriminative question generation and question answering. Unnat Jain, Svetlana Lazebnik, Alexander G. Schwing |
CVPR | 2 |
| 2018 | PackNet: Adding Multiple Tasks to a Single Network by Iterative PruningabstractThis paper presents a method for adding multiple tasks to a single deep neural network while avoiding catastrophic forgetting. Inspired by network pruning techniques, we exploit redundancies in large deep networks to free up parameters that can then be employed to learn new tasks. By performing iterative pruning and network re-training, we are able to sequentially "pack" multiple tasks into a single network while ensuring minimal drop in performance and minimal storage overhead. Unlike prior work that uses proxy losses to maintain accuracy on older tasks, we always optimize for the task at hand. We perform extensive experiments on a variety of network architectures and large-scale datasets, and observe much better robustness against catastrophic forgetting than prior work. In particular, we are able to add three fine-grained classification tasks to a single ImageNet-trained VGG-16 network and achieve accuracies close to those of separately trained networks for each task. Arun Mallya, Svetlana Lazebnik |
CVPR | 2 |
| 2018 | Piggyback: Adapting a Single Network to Multiple Tasks by Learning to Mask Weights
Arun Mallya, Dillon Davis, Svetlana Lazebnik |
ECCV (4) | 3 |
| 2018 | Conditional Image-Text Embedding Networks
Bryan A. Plummer, Paige Kordas, M. Hadi Kiapour, Shuai Zheng 0004, Robinson Piramuthu, Svetlana Lazebnik |
ECCV (12) | 6 |
| 2018 | Out of the Box: Reasoning with Graph Convolution Nets for Factual Visual Question AnsweringabstractAccurately answering a question about a given image requires combining observations with general knowledge. While this is effortless for humans, reasoning with general knowledge remains an algorithmic challenge. To advance research in this direction a novel fact-based' visual question answering (FVQA) task has been introduced recently along with a large set of curated facts which link two entities, i.e., two possible answers, via a relation. Given a question-image pair, deep network techniques have been employed to successively reduce the large set of facts until one of the two entities of the final remaining fact is predicted as the answer. We observe that a successive process which considers one fact at a time to form a local decision is sub-optimal. Instead, we develop an entity graph and use a graph convolutional network toreason' about the correct answer by jointly considering all entities. We show on the challenging FVQA dataset that this leads to an improvement in accuracy of around 7% compared to the state-of-the-art. Medhini Narasimhan, Svetlana Lazebnik, Alexander G. Schwing |
NeurIPS | 2 |
| 2017 | Enhancing Video Summarization via Vision-Language EmbeddingabstractThis paper addresses video summarization, or the problem of distilling a raw video into a shorter form while still capturing the original story. We show that visual representations supervised by freeform language make a good fit for this application by extending a recent submodular summarization approach [9] with representativeness and interestingness objectives computed on features from a joint vision-language embedding space. We perform an evaluation on two diverse datasets, UT Egocentric [18] and TV Episodes [45], and show that our new objectives give improved summarization ability compared to standard visual features alone. Our experiments also show that the vision-language embedding need not be trained on domainspecific data, but can be learned from standard still image vision-language datasets and transferred to video. A further benefit of our model is the ability to guide a summary using freeform text input at test time, allowing user customization. Bryan A. Plummer, Matthew Brown 0001, Svetlana Lazebnik |
CVPR | 3 |
| 2017 | Recurrent Models for Situation RecognitionabstractThis work proposes Recurrent Neural Network (RNN) models to predict structured `image situations' - actions and noun entities fulfilling semantic roles related to the action. In contrast to prior work relying on Conditional Random Fields (CRFs), we use a specialized action prediction network followed by an RNN for noun prediction. Our system obtains state-of-the-art accuracy on the challenging recent imSitu dataset, beating CRF-based models, including ones trained with additional data. Further, we show that specialized features learned from situation prediction can be transferred to the task of image captioning to more accurately describe human-object interactions. Arun Mallya, Svetlana Lazebnik |
ICCV | 2 |
| 2017 | Phrase Localization and Visual Relationship Detection with Comprehensive Image-Language CuesabstractThis paper presents a framework for localization or grounding of phrases in images using a large collection of linguistic and visual cues. We model the appearance, size, and position of entity bounding boxes, adjectives that contain attribute information, and spatial relationships between pairs of entities connected by verbs or prepositions. Special attention is given to relationships between people and clothing or body part mentions, as they are useful for distinguishing individuals. We automatically learn weights for combining these cues and at test time, perform joint inference over all phrases in a caption. The resulting system produces state of the art performance on phrase localization on the Flickr30k Entities dataset [33] and visual relationship detection on the Stanford VRD dataset [27]. Bryan A. Plummer, Arun Mallya, Chris M. Cervantes, Julia Hockenmaier, Svetlana Lazebnik |
ICCV | 5 |
| 2017 | Diverse and Accurate Image Description Using a Variational Auto-Encoder with an Additive Gaussian Encoding SpaceabstractThis paper explores image caption generation using conditional variational auto-encoders (CVAEs). Standard CVAEs with a fixed Gaussian prior yield descriptions with too little variability. Instead, we propose two models that explicitly structure the latent space around K components corresponding to different types of image content, and combine components to create priors for images that contain multiple types of content simultaneously (e.g., several kinds of objects). Our first model uses a Gaussian Mixture model (GMM) prior, while the second one defines a novel Additive Gaussian (AG) prior that linearly combines component means. We show that both models produce captions that are more diverse and more accurate than a strong LSTM baseline or a “vanilla” CVAE with a fixed Gaussian prior, with AG-CVAE showing particular promise. Liwei Wang 0009, Alexander G. Schwing, Svetlana Lazebnik |
NIPS | 3 |
| 2017 | Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence Models
Bryan A. Plummer, Liwei Wang 0009, Chris M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, Svetlana Lazebnik |
Int. J. Comput. Vis. | 6 |
| 2016 | Solving VIsual Madlibs with Multiple Cues
Tatiana Tommasi, Arun Mallya, Bryan A. Plummer, Svetlana Lazebnik, Alexander C. Berg, Tamara L. Berg |
BMVC | 4 |
| 2016 | Adaptive Object Detection Using Adjacency and Zoom PredictionabstractState-of-the-art object detection systems rely on an accurate set of region proposals. Several recent methods use a neural network architecture to hypothesize promising object locations. While these approaches are computationally efficient, they rely on fixed image regions as anchors for predictions. In this paper we propose to use a search strategy that adaptively directs computational resources to sub-regions likely to contain objects. Compared to methods based on fixed anchor locations, our approach naturally adapts to cases where object instances are sparse and small. Our approach is comparable in terms of accuracy to the state-of-the-art Faster R-CNN approach while using two orders of magnitude fewer anchors on average. Code is publicly available. Yongxi Lu, Tara Javidi, Svetlana Lazebnik |
CVPR | 3 |
| 2016 | Learning Deep Structure-Preserving Image-Text EmbeddingsabstractThis paper proposes a method for learning joint embeddings of images and text using a two-branch neural network with multiple layers of linear projections followed by nonlinearities. The network is trained using a largemargin objective that combines cross-view ranking constraints with within-view neighborhood structure preservation constraints inspired by metric learning literature. Extensive experiments show that our approach gains significant improvements in accuracy for image-to-text and textto-image retrieval. Our method achieves new state-of-theart results on the Flickr30K and MSCOCO image-sentence datasets and shows promise on the new task of phrase localization on the Flickr30K Entities dataset. Liwei Wang 0009, Yin Li 0003, Svetlana Lazebnik |
CVPR | 3 |
| 2016 | Learning Models for Actions and Person-Object Interactions with Transfer to Question Answering
Arun Mallya, Svetlana Lazebnik |
ECCV (1) | 2 |
| 2015 | Active Object Localization with Deep Reinforcement LearningabstractWe present an active detection model for localizing objects in scenes. The model is class-specific and allows an agent to focus attention on candidate regions for identifying the correct location of a target object. This agent learns to deform a bounding box using simple transformation actions, with the goal of determining the most specific location of target objects following top-down reasoning. The proposed localization agent is trained using deep reinforcement learning, and evaluated on the Pascal VOC 2007 dataset. We show that agents guided by the proposed model are able to localize a single instance of an object after analyzing only between 11 and 25 regions in an image, and obtain the best detection results among systems that do not use object proposals for object localization. Juan C. Caicedo, Svetlana Lazebnik |
ICCV | 2 |
| 2015 | Where to Buy It: Matching Street Clothing Photos in Online ShopsabstractIn this paper, we define a new task, Exact Street to Shop, where our goal is to match a real-world example of a garment item to the same item in an online shop. This is an extremely challenging task due to visual differences between street photos (pictures of people wearing clothing in everyday uncontrolled settings) and online shop photos (pictures of clothing items on people, mannequins, or in isolation, captured by professionals in more controlled settings). We collect a new dataset for this application containing 404,683 shop photos collected from 25 different online retailers and 20,357 street photos, providing a total of 39,479 clothing item matches between street and shop photos. We develop three different methods for Exact Street to Shop retrieval, including two deep learning baseline methods, and a method to learn a similarity measure between the street and shop domains. Experiments demonstrate that our learned similarity significantly outperforms our baselines that use existing deep learning based representations. M. Hadi Kiapour, Xufeng Han, Svetlana Lazebnik, Alexander C. Berg, Tamara L. Berg |
ICCV | 3 |
| 2015 | Learning Informative Edge Maps for Indoor Scene Layout PredictionabstractIn this paper, we introduce new edge-based features for the task of recovering the 3D layout of an indoor scene from a single image. Indoor scenes have certain edges that are very informative about the spatial layout of the room, namely, the edges formed by the pairwise intersections of room faces (two walls, wall and ceiling, wall and floor). In contrast with previous approaches that rely on area-based features like geometric context and orientation maps, our method attempts to directly detect these informative edges. We learn to predict 'informative edge' probability maps using two recent methods that exploit local and global context, respectively: structured edge detection forests, and a fully convolutional network for pixelwise labeling. We show that the fully convolutional network is quite successful at predicting the informative edges even when they lack contrast or are occluded, and that the accuracy can be further improved by training the network to jointly predict the edges and the geometric context. Using features derived from the 'informative edge' maps, we learn a maximum margin structured classifier that achieves state-of-the-art performance on layout prediction. Arun Mallya, Svetlana Lazebnik |
ICCV | 2 |
| 2015 | Flickr30k Entities: Collecting Region-to-Phrase Correspondences for Richer Image-to-Sentence ModelsabstractThe Flickr30k dataset has become a standard benchmark for sentence-based image description. This paper presents Flickr30k Entities, which augments the 158k captions from Flickr30k with 244k coreference chains linking mentions of the same entities in images, as well as 276k manually annotated bounding boxes corresponding to each entity. Such annotation is essential for continued progress in automatic image description and grounded language understanding. We present experiments demonstrating the usefulness of our annotations for text-to-image reference resolution, or the task of localizing textual entity mentions in an image, and for bidirectional image-sentence retrieval. These experiments confirm that we can further improve the accuracy of state-of-the-art retrieval methods by training with explicit region-to-phrase correspondence, but at the same time, they show that accurately inferring this correspondence given an image and a caption remains really challenging. Bryan A. Plummer, Liwei Wang 0009, Chris M. Cervantes, Juan C. Caicedo, Julia Hockenmaier, Svetlana Lazebnik |
ICCV | 6 |
| 2015 | Scene Parsing with Object Instance Inference Using Regions and Per-exemplar Detectors
Joseph Tighe, Marc Niethammer, Svetlana Lazebnik |
Int. J. Comput. Vis. | 3 |
| 2014 | Scene Parsing with Object Instances and Occlusion OrderingabstractThis work proposes a method to interpret a scene by assigning a semantic label at every pixel and inferring the spatial extent of individual object instances together with their occlusion relationships. Starting with an initial pixel labeling and a set of candidate object masks for a given test image, we select a subset of objects that explain the image well and have valid overlap relationships and occlusion ordering. This is done by minimizing an integer quadratic program either using a greedy method or a standard solver. Then we alternate between using the object predictions to refine the pixel labels and vice versa. The proposed system obtains promising results on two challenging subsets of the LabelMe and SUN datasets, the largest of which contains 45, 676 images and 232 classes. Joseph Tighe, Marc Niethammer, Svetlana Lazebnik |
CVPR | 3 |
| 2014 | Multi-scale Orderless Pooling of Deep Convolutional Activation Features
Yunchao Gong, Liwei Wang 0009, Svetlana Lazebnik |
ECCV (7) | 4 |
| 2014 | Improving Image-Sentence Embeddings Using Large Weakly Annotated Photo Collections
Yunchao Gong, Liwei Wang 0009, Micah Hodosh, Julia Hockenmaier, Svetlana Lazebnik |
ECCV (4) | 5 |
| 2014 | A Multi-View Embedding Space for Modeling Internet Images, Tags, and Their Semantics
Yunchao Gong, Qifa Ke, Michael Isard, Svetlana Lazebnik |
Int. J. Comput. Vis. | 4 |
| 2014 | Asymmetric Distances for Binary EmbeddingsabstractIn large-scale query-by-example retrieval, embedding image signatures in a binary space offers two benefits: data compression and search efficiency. While most embedding algorithms binarize both query and database signatures, it has been noted that this is not strictly a requirement. Indeed, asymmetric schemes that binarize the database signatures but not the query still enjoy the same two benefits but may provide superior accuracy. In this work, we propose two general asymmetric distances that are applicable to a wide variety of embedding techniques including locality sensitive hashing (LSH), locality sensitive binary codes (LSBC), spectral hashing (SH), PCA embedding (PCAE), PCAE with random rotations (PCAE-RR), and PCAE with iterative quantization (PCAE-ITQ). We experiment on four public benchmarks containing up to 1M images and show that the proposed asymmetric distances consistently lead to large improvements over the symmetric Hamming distance for all binary embedding techniques. Albert Gordo, Florent Perronnin, Yunchao Gong, Svetlana Lazebnik |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2013 | Learning Binary Codes for High-Dimensional Data Using Bilinear ProjectionsabstractRecent advances in visual recognition indicate that to achieve good retrieval and classification accuracy on large-scale datasets like Image Net, extremely high-dimensional visual descriptors, e.g., Fisher Vectors, are needed. We present a novel method for converting such descriptors to compact similarity-preserving binary codes that exploits their natural matrix structure to reduce their dimensionality using compact bilinear projections instead of a single large projection matrix. This method achieves comparable retrieval and classification accuracy to the original descriptors and to the state-of-the-art Product Quantization approach while having orders of magnitude faster code generation time and smaller memory footprint. Yunchao Gong, Sanjiv Kumar, Henry A. Rowley, Svetlana Lazebnik |
CVPR | 4 |
| 2013 | Finding Things: Image Parsing with Regions and Per-Exemplar DetectorsabstractThis paper presents a system for image parsing, or labeling each pixel in an image with its semantic category, aimed at achieving broad coverage across hundreds of object categories, many of them sparsely sampled. The system combines region-level features with per-exemplar sliding window detectors. Per-exemplar detectors are better suited for our parsing task than traditional bounding box detectors: they perform well on classes with little training data and high intra-class variation, and they allow object masks to be transferred into the test image for pixel-level segmentation. The proposed system achieves state-of-the-art accuracy on three challenging datasets, the largest of which contains 45,676 images and 232 labels. Joseph Tighe, Svetlana Lazebnik |
CVPR | 2 |
| 2013 | Superparsing - Scalable Nonparametric Image Parsing with Superpixels
Joseph Tighe, Svetlana Lazebnik |
Int. J. Comput. Vis. | 2 |
| 2013 | Iterative Quantization: A Procrustean Approach to Learning Binary Codes for Large-Scale Image RetrievalabstractThis paper addresses the problem of learning similarity-preserving binary codes for efficient similarity search in large-scale image collections. We formulate this problem in terms of finding a rotation of zero-centered data so as to minimize the quantization error of mapping this data to the vertices of a zero-centered binary hypercube, and propose a simple and efficient alternating minimization algorithm to accomplish this task. This algorithm, dubbed iterative quantization (ITQ), has connections to multiclass spectral clustering and to the orthogonal Procrustes problem, and it can be used both with unsupervised data embeddings such as PCA and supervised embeddings such as canonical correlation analysis (CCA). The resulting binary codes significantly outperform several other state-of-the-art methods. We also show that further performance improvements can result from transforming the data with a nonlinear kernel mapping prior to PCA or CCA. Finally, we demonstrate an application of ITQ to learning binary attributes or "classemes" on the ImageNet data set. Yunchao Gong, Svetlana Lazebnik, Albert Gordo, Florent Perronnin |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2012 | Angular Quantization-based Binary Codes for Fast Similarity SearchabstractThis paper focuses on the problem of learning binary embeddings for efficient retrieval of high-dimensional non-negative data. Such data typically arises in a large number of vision and text applications where counts or frequencies are used as features. Also, cosine distance is commonly used as a measure of dissimilarity between such vectors. In this work, we introduce a novel spherical quantization scheme to generate binary embedding of such data and analyze its properties. The number of quantization landmarks in this scheme grows exponentially with data dimensionality resulting in low-distortion quantization. We propose a very efficient method for computing the binary embedding using such large number of landmarks. Further, a linear transformation is learned to minimize the quantization error by adapting the method to the input data resulting in improved embedding. Experiments on image and text retrieval applications show superior performance of the proposed method over other existing state-of-the-art methods. Yunchao Gong, Sanjiv Kumar, Vishal Verma, Svetlana Lazebnik |
NIPS | 4 |
| 2011 | Iterative quantization: A procrustean approach to learning binary codesabstractThis paper addresses the problem of learning similarity-preserving binary codes for efficient retrieval in large-scale image collections. We propose a simple and efficient alternating minimization scheme for finding a rotation of zero-centered data so as to minimize the quantization error of mapping this data to the vertices of a zero-centered binary hypercube. This method, dubbed iterative quantization (ITQ), has connections to multi-class spectral clustering and to the orthogonal Procrustes problem, and it can be used both with unsupervised data embeddings such as PCA and supervised embeddings such as canonical correlation analysis (CCA). Our experiments show that the resulting binary coding schemes decisively outperform several other state-of-the-art methods. Yunchao Gong, Svetlana Lazebnik |
CVPR | 2 |
| 2011 | Comparing data-dependent and data-independent embeddings for classification and ranking of Internet imagesabstractThis paper presents a comparative evaluation of feature embeddings for classification and ranking in large-scale Internet image datasets. We follow a popular framework for scalable visual learning, in which the data is first transformed by a nonlinear embedding and then an efficient linear classifier is trained in the resulting space. Our study includes data-dependent embeddings inspired by the semi-supervised learning literature, and data-independent ones based on approximating specific kernels (such as the Gaussian kernel for GIST features and the histogram intersection kernel for bags of words). Perhaps surprisingly, we find that data-dependent embeddings, despite being computed from large amounts of unlabeled data, do not have any advantage over data-independent ones in the regime of scarce labeled data. On the other hand, we find that several data-dependent embeddings are competitive with popular data-independent choices for large-scale classification. Yunchao Gong, Svetlana Lazebnik |
CVPR | 2 |
| 2011 | Scene recognition and weakly supervised object localization with deformable part-based modelsabstractWeakly supervised discovery of common visual structure in highly variable, cluttered images is a key problem in recognition. We address this problem using deformable part-based models (DPM's) with latent SVM training [6]. These models have been introduced for fully supervised training of object detectors, but we demonstrate that they are also capable of more open-ended learning of latent structure for such tasks as scene recognition and weakly supervised object localization. For scene recognition, DPM's can capture recurring visual elements and salient objects; in combination with standard global image features, they obtain state-of-the-art results on the MIT 67-category indoor scene dataset. For weakly supervised object localization, optimization over latent DPM parameters can discover the spatial extent of objects in cluttered training images without ground-truth bounding boxes. The resulting method outperforms a recent state-of-the-art weakly supervised object localization approach on the PASCAL-07 dataset. Megha Pandey, Svetlana Lazebnik |
ICCV | 2 |
| 2011 | Understanding scenes on many levelsabstractThis paper presents a framework for image parsing with multiple label sets. For example, we may want to simultaneously label every image region according to its basic-level object category (car, building, road, tree, etc.), superordinate category (animal, vehicle, manmade object, natural object, etc.), geometric orientation (horizontal, vertical, etc.), and material (metal, glass, wood, etc.). Some object regions may also be given part names (a car can have wheels, doors, windshield, etc.). We compute co-occurrence statistics between different label types of the same region to capture relationships such as “roads are horizontal,” “cars are made of metal,” “cars have wheels” but “horses have legs,” and so on. By incorporating these constraints into a Markov Random Field inference framework and jointly solving for all the label sets, we are able to improve the classification accuracy for all the label sets at once, achieving a richer form of image understanding. Joseph Tighe, Svetlana Lazebnik |
ICCV | 2 |
| 2011 | Modeling and Recognition of Landmark Image Collections Using Iconic Scene Graphs
Rahul Raguram, Changchang Wu, Jan-Michael Frahm, Svetlana Lazebnik |
Int. J. Comput. Vis. | 4 |
| 2010 | Building Rome on a Cloudless Day
Jan-Michael Frahm, Pierre Fite Georgel, David Gallup, Tim Johnson, Rahul Raguram, Changchang Wu, Yi-Hung Jen, Enrique Dunn, Brian Clipp, Svetlana Lazebnik |
ECCV (4) | 10 |
| 2010 | SuperParsing: Scalable Nonparametric Image Parsing with Superpixels
Joseph Tighe, Svetlana Lazebnik |
ECCV (5) | 2 |
| 2009 | An empirical Bayes approach to contextual region classificationabstractThis paper presents a nonparametric approach to labeling of local image regions that is inspired by recent developments in information-theoretic denoising. The chief novelty of this approach rests in its ability to derive an unsupervised contextual prior over image classes from unlabeled test data. Labeled training data is needed only to learn a local appearance model for image patches (although additional supervisory information can optionally be incorporated when it is available). Instead of assuming a parametric prior such as a Markov random field for the class labels, the proposed approach uses the empirical Bayes technique of statistical inversion to recover a contextual model directly from the test data, either as a spatially varying or as a globally constant prior distribution over the classes in the image. Results on two challenging datasets convincingly demonstrate that useful contextual information can indeed be learned from unlabeled data. Svetlana Lazebnik, Maxim Raginsky |
CVPR | 1 |
| 2009 | Locality-sensitive binary codes from shift-invariant kernelsabstractThis paper addresses the problem of designing binary codes for high-dimensional data such that vectors that are similar in the original space map to similar binary strings. We introduce a simple distribution-free encoding scheme based on random projections, such that the expected Hamming distance between the binary codes of two vectors is related to the value of a shift-invariant kernel (e.g., a Gaussian kernel) between the vectors. We present a full theoretical analysis of the convergence properties of the proposed scheme, and report favorable experimental performance as compared to a recent state-of-the-art method, spectral hashing. Maxim Raginsky, Svetlana Lazebnik |
NIPS | 2 |
| 2009 | Supervised Learning of Quantizer Codebooks by Information Loss MinimizationabstractThis paper proposes a technique for jointly quantizing continuous features and the posterior distributions of their class labels based on minimizing empirical information loss such that the quantizer index of a given feature vector approximates a sufficient statistic for its class label. Informally, the quantized representation retains as much information as possible for classifying the feature vector correctly. We derive an alternating minimization procedure for simultaneously learning codebooks in the euclidean feature space and in the simplex of posterior class distributions. The resulting quantizer can be used to encode unlabeled points outside the training set and to predict their posterior class distributions, and has an elegant interpretation in terms of lossless source coding. The proposed method is validated on synthetic and real data sets and is applied to two diverse problems: learning discriminative visual vocabularies for bag-of-features image classification and image segmentation. Svetlana Lazebnik, Maxim Raginsky |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2008 | Modeling and Recognition of Landmark Image Collections Using Iconic Scene Graphs
Xiaowei Li 0007, Changchang Wu, Christopher Zach, Svetlana Lazebnik, Jan-Michael Frahm |
ECCV (1) | 4 |
| 2008 | Analysis of human attractiveness using manifold kernel regressionabstractThis paper uses a recently introduced manifold kernel regression technique to explore the relationship between facial shape and attractiveness on a heterogeneous dataset of over three thousand images gathered from the Web. Using the concept of the Frechet mean of images under a diffeomorphic transformation model, we evolve the average face as a function of attractiveness ratings. Examining these averages and associated deformation maps enables us to discern aggregate shape change trends for male and female faces. Bradley C. Davis, Svetlana Lazebnik |
ICIP | 2 |
| 2008 | Near-minimax recursive density estimation on the binary hypercubeabstractThis paper describes a recursive estimation procedure for multivariate binary densities using orthogonal expansions. For $d$ covariates, there are $2^d$ basis coefficients to estimate, which renders conventional approaches computationally prohibitive when $d$ is large. However, for a wide class of densities that satisfy a certain sparsity condition, our estimator runs in probabilistic polynomial time and adapts to the unknown sparsity of the underlying density in two key ways: (1) it attains near-minimax mean-squared error, and (2) the computational complexity is lower for sparser densities. Our method also allows for flexible control of the trade-off between mean-squared error and computational complexity. Maxim Raginsky, Svetlana Lazebnik, Rebecca Willett, Jorge G. Silva |
NIPS | 2 |
| 2007 | Projective Visual Hulls
Svetlana Lazebnik, Yasutaka Furukawa, Jean Ponce |
Int. J. Comput. Vis. | 1 |
| 2007 | Local Features and Kernels for Classification of Texture and Object Categories: A Comprehensive Study
Jianguo Zhang 0001, Marcin Marszalek, Svetlana Lazebnik, Cordelia Schmid |
Int. J. Comput. Vis. | 3 |
| 2007 | Segmenting, Modeling, and Matching Video Clips Containing Multiple Moving ObjectsabstractThis paper presents a novel representation for dynamic scenes composed of multiple rigid objects that may undergo different motions and are observed by a moving camera. Multiview constraints associated with groups of affine-covariant scene patches and a normalized description of their appearance are used to segment a scene into its rigid components, construct three-dimensional models of these components, and match instances of models recovered from different image sequences. The proposed approach has been applied to the detection and matching of moving objects in video sequences and to shot matching, i.e., the identification of shots that depict the same scene in a video clip. Fred Rothganger, Svetlana Lazebnik, Cordelia Schmid, Jean Ponce |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2006 | Beyond Bags of Features: Spatial Pyramid Matching for Recognizing Natural Scene CategoriesabstractThis paper presents a method for recognizing scene categories based on approximate global geometric correspondence. This technique works by partitioning the image into increasingly fine sub-regions and computing histograms of local features found inside each sub-region. The resulting "spatial pyramid" is a simple and computationally efficient extension of an orderless bag-of-features image representation, and it shows significantly improved performance on challenging scene categorization tasks. Specifically, our proposed method exceeds the state of the art on the Caltech-101 database and achieves high accuracy on a large database of fifteen natural scene categories. The spatial pyramid framework also offers insights into the success of several recently proposed image descriptions, including Torralba’s "gist" and Lowe’s SIFT descriptors. Svetlana Lazebnik, Cordelia Schmid, Jean Ponce |
CVPR (2) | 1 |
| 2006 | 3D Object Modeling and Recognition Using Local Affine-Invariant Image Descriptors and Multi-View Spatial Constraints
Fred Rothganger, Svetlana Lazebnik, Cordelia Schmid, Jean Ponce |
Int. J. Comput. Vis. | 2 |
| 2005 | A Maximum Entropy Framework for Part-Based Texture and Object RecognitionabstractThis paper presents a probabilistic part-based approach for texture and object recognition. Textures are represented using a part dictionary found by quantizing the appearance of scale- or affine- invariant keypoints. Object classes are represented using a dictionary of composite semi-local parts, or groups of neighboring keypoints with stable and distinctive appearance and geometric layout. A discriminative maximum entropy framework is used to learn the posterior distribution of the class label given the occurrences of parts from the dictionary in the training set. Experiments on two texture and two object databases demonstrate the effectiveness of this framework for visual classification. Svetlana Lazebnik, Cordelia Schmid, Jean Ponce |
ICCV | 1 |
| 2005 | Estimation of Intrinsic Dimensionality Using High-Rate Vector QuantizationabstractWe introduce a technique for dimensionality estimation based on the notion of quantization dimension, which connects the asymptotic optimal quantization error for a probability distribution on a manifold to its intrinsic dimension. The definition of quantization dimension yields a family of estimation algorithms, whose limiting case is equivalent to a recent method based on packing numbers. Using the formalism of high-rate vector quantization, we address issues of statistical consistency and analyze the behavior of our scheme in the presence of noise. Maxim Raginsky, Svetlana Lazebnik |
NIPS | 2 |
| 2005 | The Local Projective Shape of Smooth Surfaces and Their Outlines
Svetlana Lazebnik, Jean Ponce |
Int. J. Comput. Vis. | 1 |
| 2005 | A Sparse Texture Representation Using Local Affine RegionsabstractThis paper introduces a texture representation suitable for recognizing images of textured surfaces under a wide range of transformations, including viewpoint changes and nonrigid deformations. At the feature extraction stage, a sparse set of affine Harris and Laplacian regions is found in the image. Each of these regions can be thought of as a texture element having a characteristic elliptic shape and a distinctive appearance pattern. This pattern is captured in an affine-invariant fashion via a process of shape normalization followed by the computation of two novel descriptors, the spin image and the RIFT descriptor. When affine invariance is not required, the original elliptical shape servee as an additional discriminative feature for texture recognition. The proposed approach is evaluated in retrieval and classification tasks using the entire Brodatz database and a publicly available collection of 1,000 photographs of textured surfaces taken from different viewpoints. Svetlana Lazebnik, Cordelia Schmid, Jean Ponce |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2004 | Semi-Local Affine Parts for Object RecognitionabstractHAL is a multi-disciplinary open access archive for the deposit and dissemination of sci-entific research documents, whether they are pub-lished or not. The documents may come from teaching and research institutions in France or abroad, or from public or private research centers. L’archive ouverte pluridisciplinaire HAL, est destinée au dépôt et a ̀ la diffusion de documents scientifiques de niveau recherche, publiés ou non, émanant des établissements d’enseignement et de recherche français ou étrangers, des laboratoires publics ou privés. Svetlana Lazebnik, Cordelia Schmid, Jean Ponce |
BMVC | 1 |
| 2004 | Segmenting, Modeling, and Matching Video Clips Containing Multiple Moving Objects
Fred Rothganger, Svetlana Lazebnik, Cordelia Schmid, Jean Ponce |
CVPR (2) | 2 |
| 2003 | A Sparse Texture Representation Using Affine-Invariant RegionsabstractThis paper introduces a texture representation suitable for recognizing images of textured surfaces under a wide range of transformations, including viewpoint changes and nonrigid deformations. At the feature extraction stage, a sparse set of affine-invariant local patches is extracted from the image. This spatial selection process permits the computation of characteristic scale and neighborhood shape for every texture element. The proposed texture representation is evaluated in retrieval and classification tasks using the entire Brodatz database and a collection of photographs of textured surfaces taken from different viewpoints. Svetlana Lazebnik, Cordelia Schmid, Jean Ponce |
CVPR (2) | 1 |
| 2003 | 3D Object Modeling and Recognition Using Affine-Invariant Patches and Multi-View Spatial ConstraintsabstractThis paper presents a representation for three-dimensional objects in terms of affine-invariant image patches and their spatial relationships. Multi-view constraints associated with groups of patches are combined with a normalized representation of their appearance to guide matching and reconstruction, allowing the acquisition of true three-dimensional affine and Euclidean models from multiple images and their recognition in a single photograph taken from an arbitrary viewpoint. The proposed approach does not require a separate segmentation stage and is applicable to cluttered scenes. Preliminary modeling and recognition results are presented. Fred Rothganger, Svetlana Lazebnik, Cordelia Schmid, Jean Ponce |
CVPR (2) | 2 |
| 2003 | The Local Projective Shape of Smooth Surfaces and their OutlinesabstractWe examine projectively invariant local properties of smooth curves and surfaces. Oriented projective differential geometry is proposed as a theoretical framework for establishing such invariants and describing the local shape of surfaces and their outlines. This framework is applied to two problems: a projective proof of Koenderink's famous characterization of convexities, concavities, and inflections of apparent contours; and the determination of the relative orientation of rim tangents at frontier points. Svetlana Lazebnik, Jean Ponce |
ICCV | 1 |
| 2003 | Affine-Invariant Local Descriptors and Neighborhood Statistics for Texture RecognitionabstractWe present a framework for texture recognition based on local affine-invariant descriptors and their spatial layout. At modelling time, a generative model of local descriptors is learned from sample images using the EM algorithm. The EM framework allows the incorporation of unsegmented multitexture images into the training set. The second modelling step consists of gathering co-occurrence statistics of neighboring descriptors. At recognition time, initial probabilities computed from the generative model are refined using a relaxation step that incorporates co-occurrence statistics. Performance is evaluated on images of an indoor scene and pictures of wild animals. Svetlana Lazebnik, Cordelia Schmid, Jean Ponce |
ICCV | 1 |
| 2002 | On Pencils of Tangent Planes and the Recognition of Smooth 3D Shapes from Silhouettes
Svetlana Lazebnik, Amit Sethi, Cordelia Schmid, David J. Kriegman, Jean Ponce, Martial Hebert |
ECCV (3) | 1 |
| 2001 | On Computing Exact Visual Hulls of Solids Bounded by Smooth SurfacesabstractThis paper presents a method for computing the visual hull that is based on two novel representations: the rim mesh, which describes the connectivity of contour generators on the object surface; and the visual hull mesh, which describes the exact structure of the surface of the solid formed by intersecting a finite number of visual cones. We describe the topological features of these meshes and show how they can be identified in the image using epipolar constraints. These constraints are used to derive an image-based practical reconstruction algorithm that works with weakly calibrated cameras. Experiments on synthetic and real data validate the proposed approach. Svetlana Lazebnik, Edmond Boyer, Jean Ponce |
CVPR (1) | 1 |