VLDB 2026 Research / reviewers in the wild / expert
Lior Wolf
dblp:83/4103
· DBLP profile ↗
268ranked-venue papers
36as first author
113since 2021 · last 2026
0000-0001-5578-8892ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 212 · 33 first-author · 86 since 2021Graphics, computer vision, multimedia, augmented reality and games · 161 · 26 first-author · 60 since 2021Databases, data management, data science and information retrieval · 11 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorSystems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AlignTree: Efficient Defense Against LLM Jailbreak AttacksabstractLarge Language Models (LLMs) are vulnerable to adversarial attacks that bypass safety guidelines and generate harmful content. Mitigating these vulnerabilities requires defense mechanisms that are both robust and computationally efficient. However, existing approaches either incur high computational costs or rely on lightweight defenses that can be easily circumvented, rendering them impractical for real-world LLM-based systems. In this work, we introduce the AlignTree defense, which enhances model alignment while maintaining minimal computational overhead. AlignTree monitors LLM activations during generation and detects misaligned behavior using an efficient random forest classifier. This classifier operates on two signals: (i) the refusal direction - a linear representation that activates on misaligned prompts, and (ii) an SVM-based signal that captures non-linear features associated with harmful content. Unlike previous methods, AlignTree does not require additional prompts or auxiliary guard models. Through extensive experiments, we demonstrate the efficiency and robustness of AlignTree across multiple LLMs and benchmarks. Gil Goren, Shahar Katz, Lior Wolf |
AAAI | 3 |
| 2026 | Faithful Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution GuidanceabstractLarge language models (LLMs) achieve strong performance and have revolutionized NLP, but their lack of explainability keeps them treated as black boxes, limiting their use in domains that demand transparency and trust.A promising direction to address this issue is post-hoc text-based explanations, which aim to explain model decisions in natural language.Prior work has focused on generating convincing rationales that appear to be subjectively faithful, but it remains unclear whether these explanations are epistemically faithful, whether they reflect the internal evidence the model actually relied on for its decision.In this paper, we first assess the epistemic faithfulness of LLMgenerated explanations via counterfactuals and show that they are often unfaithful.We then introduce a training-free method that enhances faithfulness by guiding explanation generation through attention-level interventions, informed by token-level heatmaps extracted via a faithful attribution method.This method significantly improves epistemic faithfulness across multiple models, benchmarks, and prompts. Bar Alon 0001, Itamar Zimerman, Lior Wolf |
ACL (1) | 3 |
| 2026 | TensorLens: End-to-End Transformer Analysis via High-Order Attention TensorsabstractAttention matrices are fundamental to transformer research, supporting a broad range of applications including interpretability, visualization, manipulation, and distillation.Yet, most existing analyses focus on individual attention heads or layers, failing to account for the model's global behavior.While prior efforts have extended attention formulations across multiple heads via averaging and matrix multiplications or incorporated components such as normalization and FFNs, a unified and complete representation that encapsulates all transformer blocks is still lacking.We address this gap by introducing TensorLens, a novel formulation that captures the entire transformer as a single, input-dependent linear operator expressed through a high-order attentioninteraction tensor.This tensor jointly encodes attention, FFNs, activations, normalizations, and residual connections, offering a theoretically coherent and expressive linear representation of the model's computation.TensorLens is theoretically grounded and our empirical validation shows that it yields richer representations than previous attention-aggregation methods.Our experiments demonstrate that the attention tensor can serve as a powerful foundation for developing tools aimed at interpretability and model understanding. Ido Atad, Itamar Zimerman, Shahar Katz, Lior Wolf |
ACL (1) | 4 |
| 2026 | Detection-Driven Object Count Optimization for Text-to-Image Diffusion ModelsabstractAccurately controlling object count in text-to-image generation remains a key challenge. Supervised methods often fail, as training data rarely covers all count variations. Methods that manipulate the denoising process to add or remove objects can help; however, they still require labeled data, limit robustness and image quality, and rely on a slow, iterative process.Pre-trained differentiable counting models that rely on soft object density summation exist and could steer generation, but employing them presents three main challenges: (i) they are pre-trained on clean images, making them less effective during denoising steps that operate on noisy inputs; (ii) they are not robust to viewpoint changes; and (iii) optimization is computationally expensive, requiring repeated model evaluations per image.We propose a new framework that uses pre-trained object counting techniques and object detectors to guide generation. First, we optimize a counting token using an outer-loop loss computed on fully generated images. Second, we introduce a detection-driven scaling term that corrects errors caused by viewpoint and proportion shifts, etc., without requiring backpropagation through the detection model. Third, we show that the optimized parameters can be reused for new prompts, removing the need for repeated optimization. Our method provides efficiency through token reuse, flexibility via compatibility with various detectors, and accuracy with improved counting across diverse object categories. Oz Zafar, Yuval Cohen, Lior Wolf, Idan Schwartz |
WACV | 3 |
| 2026 | Kernel Reboot: Breaking the Boundaries of Neural Tangent Kernels for Neural FieldsabstractNeural fields (NFs) map continuous coordinates to signals such as color or density, but fast high-quality reconstruction from sparse observations remains difficult. Classical Neural Tangent Kernel (NTK) regression gives closed-form fits, yet it is fundamentally linear and cannot accumulate reusable task priors. We develop three algorithms that address these gaps. NTK-KIP learns a distilled support set of coordinates (and optional labels) so that a finite NTK can inpaint large missing regions from little observed data, yielding a compact non-linear representation instead of a raw kernel solve. MetaQuill meta-learns a shared initialization for an INR so that new scenes can be adapted by updating only a small task-specific weight offset, which provides true feature learning and a reusable prior. Finally, MetaQuill-KIP fuses both ideas: it seeds the task with a KIP-style non-linear warm start, then refines only that small offset around the meta-learned initialization. MetaQuill-KIP achieves high-PSNR reconstructions and semantically plausible inpainting under very sparse observations, while requiring only lightweight per-instance adaptation, whereas diffusion-style baselines typically depend on large pretrained generative priors and costly per-image tuning. This shows that NTK-driven neural fields can be made both non-linear and meta-learnable, narrowing the gap between analytic kernels and practical few-shot reconstruction. Amir Mallak, Alaa Maalouf, Lior Wolf, Daniela Rus, Dan Rosenbaum |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | The Hidden Attention of Mamba ModelsabstractThe Mamba layer offers an efficient selective state-space model (SSM) that is highly effective in modeling multiple domains, includingNLP, long-range sequence processing, and computer vision. Selective SSMs are viewed as dual models, in which one trains in parallel on the entire sequence via an IO-aware parallel scan, and deploys in an autoregressive manner. We add a third view and show that such models can be viewed as attention-driven models. This new perspective enables us to empirically and theoretically compare the underlying mechanisms to that of the attention in transformers and allows us to peer inside the inner workings of the Mamba model with explainability methods. Our code is publicly available. Ameen Ali, Itamar Zimerman, Lior Wolf |
ACL (1) | 3 |
| 2025 | Segment-Based Attention Masking for GPTsabstractCausal masking is a fundamental component in Generative Pre-Trained Transformer (GPT) models, playing a crucial role during training.Although GPTs can process the entire user prompt at once, the causal masking is applied to all input tokens step-by-step, mimicking the generation process.This imposes an unnecessary constraint during the initial "prefill" phase when the model processes the input prompt and generates the internal representations before producing any output tokens.In this work, attention is masked based on the known block structure at the prefill phase, followed by the conventional token-by-token autoregressive process after that.For example, in a typical chat prompt, the system prompt is treated as one block, and the user prompt as the next one.Each of these is treated as a unit for the purpose of masking, such that the first tokens in each block can access the subsequent tokens in a non-causal manner.Then, the model answer is generated in the conventional causal manner.The Segment-by-Segment scheme entails no additional computational overhead.When integrated using a lightweight fine-tuning into already trained models such as Llama and Qwen, MAS quickly increases models' performances.Our code will be available at: https://github.com/shacharKZ/ MAS-Segment- Shahar Katz, Liran Ringel, Yaniv Romano, Lior Wolf |
ACL (1) | 4 |
| 2025 | SphereUFormer: A U-Shaped Transformer for Spherical 360 PerceptionabstractThis paper proposes a novel method for omnidirectional 360° perception. Most common previous methods relied on equirectangular projection. This representation is easily applicable to 2D operation layers but introduces distortions into the image. Other methods attempted to remove the distortions by maintaining a sphere representation but relied on complicated convolution kernels that failed to show competitive results. In this work, we introduce a transformer-based architecture that, by incorporating a novel "Spherical Local Self-Attention" and other spherically-oriented modules, successfully operates in the spherical domain and outperforms the state-of-the-art in 360° perception benchmarks for depth estimation and semantic segmentation. Our code is available at https://github.com/yanivbenny/sphere_uformer. Yaniv Benny, Lior Wolf |
CVPR | 2 |
| 2025 | Adapting to the Unknown: Training-Free Audio-Visual Event Perception with Dynamic ThresholdsabstractIn the domain of audio-visual event perception, which focuses on the temporal localization and classification of events across distinct modalities (audio and visual), existing approaches are constrained by the vocabulary available in their training data. This limitation significantly impedes their capacity to generalize to novel, unseen event categories. Furthermore, the annotation process for this task is labor-intensive, requiring extensive manual labeling across modalities and temporal segments, limiting the scalability of current methods. Current state-of-the-art models ignore the shifts in event distributions over time, reducing their ability to adjust to changing video dynamics. Additionally, previous methods rely on late fusion to combine audio and visual information. While straightforward, this approach results in a significant loss of multimodal interactions. To address these challenges, we propose Audio-Visual Adaptive Video Analysis (AV2A), a model-agnostic approach that requires no further training and integrates a score-level fusion technique to retain richer multimodal interactions. AV2A also includes a within-video label shift algorithm, leveraging input video data and predictions from prior frames to dynamically adjust event distributions for subsequent frames. Moreover, we present the first training-free, open-vocabulary baseline for audio-visual event perception, demonstrating that AV2A achieves substantial improvements over naive training-free baselines. We demonstrate the effectiveness of AV2A on both zero-shot and weakly-supervised state-of-the-art methods, achieving notable improvements in performance metrics over existing approaches. Our code is available on Github. Eitan Shaar, Ariel Shaulov, Gal Chechik, Lior Wolf |
CVPR | 4 |
| 2025 | IlluSign: Illustrating Sign Language Videos by Leveraging the Attention MechanismabstractSign languages are dynamic visual languages that involve hand gestures, in combination with non-manual elements such as facial expressions. While video recordings of sign language are commonly used for education and documentation, the dynamic nature of signs can make it challenging to study them in detail, especially for new learners. This work aims to convert sign language video footage into static illustrations, which serve as an additional educational resource to complement video content. This process is usually done by an artist, and is therefore quite costly. We propose a method to illustrate sign language videos by leveraging generative models to capture both semantic and geometric image features. Our approach transfers a sketch-like style to keyframes and combines the start and end poses into a single image, enhanced with arrows to indicate hand motion.While many style transfer methods address domain adaptation at various levels of abstraction, applying a sketch-like style to sign language, particularly to detailed hand gestures, remains a significant challenge. To tackle this, we intervene in the denoising process of a diffusion model, injecting style as keys and values into high-resolution attention layers, and fusing geometric information from the image and edges as queries. For the final illustration, we use the attention mechanism to combine the attention weights from both the start and end illustrations, resulting in a soft combination. Our method offers a costeffective solution for generating sign language illustrations at inference time, addressing the lack of such resources in educational materials.11Watermarks, such as the SGB-FSS mark in the SignSuisse dataset [35], are ignored in this work. Janna Bruner, Amit Moryossef, Lior Wolf |
FG | 3 |
| 2025 | Predicting local fMRI activations from EEG: a Feasibility Study Using Both Classical and Modern Machine Learning PipelinesabstractfMRI’s clinical use is limited by cost, while EEG is more accessible but lacks spatial detail and deep brain coverage. Research aims to predict deep brain activations from combined fMRI and EEG data. We compare classical machine learning and a CNN-transformer pipeline for this mapping across multiple brain regions. As we show, in the first dataset, which is heavily tilted toward visual perception, the activations in the Hippocampus cannot be recovered reliably from EEG, using either pipeline. However, in other regions, predictability is much higher, and in those cases, the deep learning pipeline obtains better predictions. In a second dataset that is based on musical feedback while the visual is blocked, both pipelines yield improved results in the Hippocampus. Tomer Amit, Taly Markovits, Guy Gurevitch, Talma Hendler, Lior Wolf |
ICASSP | 5 |
| 2025 | Classifier-Guided Captioning Across ModalitiesabstractMost current captioning systems use language models trained on data from specific settings, such as image-based captioning via Amazon Mechanical Turk, limiting their ability to generalize to other modality distributions and contexts. This limitation hinders performance in tasks like audio or video captioning, where different semantic cues are needed. Addressing this challenge is crucial for creating more adaptable and versatile captioning frameworks applicable across diverse real-world contexts. In this work, we introduce a method to adapt captioning networks to the semantics of alternative settings, such as capturing audibility in audio captioning, where it is crucial to describe sounds and their sources. Our framework consists of two main components: (i) a frozen captioning system incorporating a language model (LM), and (ii) a text classifier that guides the captioning system. The classifier is trained on a dataset automatically generated by GPT-4, using tailored prompts specifically designed to enhance key aspects of the generated captions. Importantly, the framework operates solely during inference, eliminating the need for further training of the underlying captioning model. We evaluated the framework on various models and modalities, with a focus on audio captioning, and report promising results. Notably, when combined with an existing zero-shot audio captioning system, our framework improves its quality and sets state-of-the-art performance in zero-shot audio captioning. Ariel Shaulov, Tal Shaharabany, Eitan Shaar, Gal Chechik, Lior Wolf |
ICASSP | 5 |
| 2025 | DeciMamba: Exploring the Length Extrapolation Potential of MambaabstractLong-range sequence processing poses a significant challenge for Transformers due to their quadratic complexity in input length. A promising alternative is Mamba, which demonstrates high performance and achieves Transformer-level capabilities while requiring substantially fewer computational resources. In this paper we explore the length-generalization capabilities of Mamba, which we find to be relatively limited. Through a series of visualizations and analyses we identify that the limitations arise from a restricted effective receptive field, dictated by the sequence length used during training. To address this constraint, we introduce DeciMamba, a context-extension method specifically designed for Mamba. This mechanism, built on top of a hidden filtering mechanism embedded within the S6 layer, enables the trained model to extrapolate well even without additional training. Empirical experiments over real-world long-range NLP tasks show that DeciMamba can extrapolate to context lengths that are significantly longer than the ones seen during training, while enjoying faster inference. We will release our code and models. Assaf Ben-Kish, Itamar Zimerman, Shady Abu-Hussein, Nadav Cohen 0001, Amir Globerson, Lior Wolf, Raja Giryes |
ICLR | 6 |
| 2025 | Add-it: Training-Free Object Insertion in Images With Pretrained Diffusion ModelsabstractAdding Object into images based on text instructions is a challenging task in semantic image editing, requiring a balance between preserving the original scene and seamlessly integrating the new object in a fitting location. Despite extensive efforts, existing models often struggle with this balance, particularly with finding a natural location for adding an object in complex scenes. We introduce Add-it, a training-free approach that extends diffusion models' attention mechanisms to incorporate information from three key sources: the scene image, the text prompt, and the generated image itself. Our weighted extended-attention mechanism maintains structural consistency and fine details while ensuring natural object placement. Without task-specific fine-tuning, Add-it achieves state-of-the-art results on both real and generated image insertion benchmarks, including our newly constructed "Additing Affordance Benchmark" for evaluating object placement plausibility, outperforming supervised methods. Human evaluations show that Add-it is preferred in over 80% of cases, and it also demonstrates improvements in various automated metrics. Yoad Tewel, Rinon Gal, Dvir Samuel, Yuval Atzmon, Lior Wolf, Gal Chechik |
ICLR | 5 |
| 2025 | Explaining Modern Gated-Linear RNNs via a Unified Implicit Attention FormulationabstractRecent advances in efficient sequence modeling have led to attention-free layers, such as Mamba, RWKV, and various gated RNNs, all featuring sub-quadratic complexity in sequence length and excellent scaling properties, enabling the construction of a new type of foundation models. In this paper, we present a unified view of these models, formulating such layers as implicit causal self-attention layers. The formulation includes most of their sub-components and is not limited to a specific part of the architecture. The framework compares the underlying mechanisms on similar grounds for different layers and provides a direct means for applying explainability methods. Our experiments show that our attention matrices and attribution method outperform an alternative and a more limited formulation that was recently proposed for Mamba. For the other architectures for which our method is the first to provide such a view, our method is effective and competitive in the relevant metrics compared to the results obtained by state-of-the-art Transformer explainability methods. Our code is publicly available. Itamar Zimerman, Ameen Ali, Lior Wolf |
ICLR | 3 |
| 2025 | VideoJAM: Joint Appearance-Motion Representations for Enhanced Motion Generation in Video ModelsabstractDespite tremendous recent progress, generative video models still struggle to capture real-world motion, dynamics, and physics. We show that this limitation arises from the conventional pixel reconstruction objective, which biases models toward appearance fidelity at the expense of motion coherence. To address this, we introduce VideoJAM, a novel framework that instills an effective motion prior to video generators, by encouraging the model to learn a joint appearance-motion representation. VideoJAM is composed of two complementary units. During training, we extend the objective to predict both the generated pixels and their corresponding motion from a single learned representation. During inference, we introduce Inner-Guidance, a mechanism that steers the generation toward coherent motion by leveraging the model’s own evolving motion prediction as a dynamic guidance signal. Notably, our framework can be applied to any video model with minimal adaptations, requiring no modifications to the training data or scaling of the model. VideoJAM achieves state-of-the-art performance in motion coherence, surpassing highly competitive proprietary models while also enhancing the perceived visual quality of the generations. These findings emphasize that appearance and motion can be complementary and, when effectively integrated, enhance both the visual quality and the coherence of video generation. Hila Chefer, Uriel Singer, Amit Zohar, Yuval Kirstain, Adam Polyak, Yaniv Taigman, Lior Wolf, Shelly Sheynin |
ICML | 7 |
| 2025 | Discovering Directions of Uncertainty in Speech Inpainting
Kfir Cohen, Lior Wolf, Bracha Laufer-Goldshtein |
INTERSPEECH | 2 |
| 2025 | Reversed Attention: On The Gradient Descent Of Attention Layers In GPTabstractShahar Katz, Lior Wolf. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Shahar Katz, Lior Wolf |
NAACL (Long Papers) | 2 |
| 2025 | Revisiting LRP: Positional Attribution as the Missing Ingredient for Transformer ExplainabilityabstractThe development of effective explainability tools for Transformers is a crucial pursuit in deep learning research. One of the most promising approaches in this domain is Layer-wise Relevance Propagation (LRP), which propagates relevance scores backward through the network to the input space by redistributing activation values based on predefined rules. However, existing LRP-based methods for Transformer explainability entirely overlook a critical component of the Transformer architecture: its positional encoding (PE), resulting in violations of conservation, and the loss of an important and unique type of relevance, which is also associated with structural and positional features. To address this limitation, we reformulate the input space for Transformer explainability as a set of position-token pairs, rather than relying solely on the standard vocabulary space. This allows us to propose specialized theoretically-grounded LRP rules designed to propagate attributions across various positional encoding methods, including Rotary, Learned, and Absolute PE. Extensive experiments with both fine-tuned classifiers and zero-shot foundation models, such as LLaMA 3, demonstrate that our method significantly outperforms the SoTA in both vision and NLP explainability tasks. Our code is provided as a supplement. Yarden Bakish, Itamar Zimerman, Hila Chefer, Lior Wolf |
NeurIPS | 4 |
| 2025 | Execution Guided Line-by-Line Code GenerationabstractWe present a novel approach to neural code generation that incorporates real-time execution signals into the language model generation process. While large language models (LLMs) have demonstrated impressive code generation capabilities, they typically do not utilize execution feedback during inference, a critical signal that human programmers regularly leverage. Our method, Execution-Guided Classifier-Free Guidance EG-CFG, dynamically incorporates execution signals as the model generates code, providing line-by-line feedback that guides the generation process toward executable solutions.
EG-CFG employs a multi-stage process: first, we conduct beam search to sample candidate program completions for each line; second, we extract execution signals by executing these candidates against test cases; and finally, we incorporate these signals into the prompt during generation. By maintaining consistent signals across tokens within the same line and refreshing signals at line boundaries, our approach provides coherent guidance while preserving syntactic structure. Moreover, the method naturally supports native parallelism at the task level in which multiple agents operate in parallel, exploring diverse reasoning paths and collectively generating a broad set of candidate solutions.
Our experiments across diverse coding tasks demonstrate that EG-CFG significantly improves code generation performance compared to standard approaches, achieving state-of-the-art results across various levels of complexity, from foundational problems to challenging competitive programming and data science tasks. Boaz Lavon, Shahar Katz, Lior Wolf |
NeurIPS | 3 |
| 2025 | FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video GenerationabstractText-to-video diffusion models are notoriously limited in their ability to model temporal aspects such as motion, physics, and dynamic interactions. Existing approaches address this limitation by retraining the model or introducing external conditioning signals to enforce temporal consistency. In this work, we explore whether a meaningful temporal representation can be extracted directly from the predictions of a pre-trained model without any additional training or auxiliary inputs. We introduce __FlowMo__, a novel training-free guidance method that enhances motion coherence using only the model's own predictions in each diffusion step.
FlowMo first derives an appearance-debiased temporal representation by measuring the distance between latents corresponding to consecutive frames. This highlights the implicit temporal structure predicted by the model.
It then estimates motion coherence by measuring the patch-wise variance across the temporal dimension, and guides the model to reduce this variance dynamically during sampling. Extensive experiments across multiple text-to-video models demonstrate that FlowMo significantly improves motion coherence without sacrificing visual quality or prompt alignment, offering an effective plug-and-play solution for enhancing the temporal fidelity of pre-trained video diffusion models. Ariel Shaulov, Itay Hazan 0001, Lior Wolf, Hila Chefer |
NeurIPS | 3 |
| 2025 | ConsiStyle: Style Diversity in Training-Free Consistent T2I GenerationabstractIn text-to-image models, consistent character generation is the task of achieving text alignment while maintaining the subject's appearance across different prompts. However, since style and appearance are often entangled, the existing methods struggle to preserve consistent subject characteristics while adhering to varying style prompts. Current approaches for consistent text-to-image generation typically rely on large-scale fine-tuning on curated image sets or per-subject optimization, which either fail to generalize across prompts or do not align well with textual descriptions. Meanwhile, training-free methods often fail to maintain subject consistency across different styles. In this work, we introduce a training-free method that, for the first time, jointly achieves style preservation and subject consistency across varied styles. The attention matrices are manipulated such that Queries and Keys are obtained from the anchor image(s) that are used to define the subject, while the Values are imported from a parallel copy that is not subject-anchored. Additionally, cross-image components are added to the self-attention mechanism by expanding the Key and Value matrices. To do without shifting from the target style, we align the statistics of the Value matrices. As is demonstrated in a comprehensive battery of qualitative and quantitative experiments, our method effectively decouples style from subject appearance and enables faithful generation of text-aligned images with consistent characters across diverse styles. Code will be available at our project page: jbruner23.github.io/consistyle. Yohai Mazuz, Janna Bruner, Lior Wolf |
ACM Trans. Graph. | 3 |
| 2024 | Deep Quantum Error CorrectionabstractQuantum error correction codes (QECC) are a key component for realizing the potential of quantum computing. QECC, as its classical counterpart (ECC), enables the reduction of error rates, by distributing quantum logical information across redundant physical qubits, such that errors can be detected and corrected. In this work, we efficiently train novel end-to-end deep quantum error decoders. We resolve the quantum measurement collapse by augmenting syndrome decoding to predict an initial estimate of the system noise, which is then refined iteratively through a deep neural network. The logical error rates calculated over finite fields are directly optimized via a differentiable objective, enabling efficient decoding under the constraints imposed by the code. Finally, our architecture is extended to support faulty syndrome measurement, by efficient decoding of repeated syndrome sampling. The proposed method demonstrates the power of neural decoders for QECC by achieving state-of-the-art accuracy, outperforming for small distance topological codes, the existing end-to-end neural and classical decoders, which are often computationally prohibitive. Yoni Choukroun, Lior Wolf |
AAAI | 2 |
| 2024 | Diverse and Aligned Audio-to-Video Generation via Text-to-Video Model AdaptationabstractWe consider the task of generating diverse and realistic videos guided by natural audio samples from a wide variety of semantic classes. For this task, the videos are required to be aligned both globally and temporally with the input audio: globally, the input audio is semantically associated with the entire output video, and temporally, each segment of the input audio is associated with a corresponding segment of that video. We utilize an existing text-conditioned video generation model and a pre-trained audio encoder model. The proposed method is based on a lightweight adaptor network, which learns to map the audio-based representation to the input representation expected by the text-to-video generation model. As such, it also enables video generation conditioned on text, audio, and, for the first time as far as we can ascertain, on both text and audio. We validate our method extensively on three datasets demonstrating significant semantic diversity of audio-video samples and further propose a novel evaluation metric (AV-Align) to assess the alignment of generated videos with input audio samples. AV-Align is based on the detection and comparison of energy peaks in both modalities. In comparison to recent state-of-the-art approaches, our method generates videos that are better aligned with the input sound, both with respect to content and temporal axis. We also show that videos produced by our method present higher visual quality and are more diverse. Code and samples are available at: https://pages.cs.huji.ac.il/adiyoss-lab/TempoTokens/. Guy Yariv, Itai Gat, Sagie Benaim, Lior Wolf, Idan Schwartz, Yossi Adi |
AAAI | 4 |
| 2024 | Revisiting the Noise Model of Stochastic Gradient DescentabstractThe effectiveness of stochastic gradient descent (SGD) in neural network optimization is significantly influenced by stochastic gradient noise (SGN). Following the central limit theorem, SGN was initially described as Gaussian, but recently Simsekli et al (2019) demonstrated that the $S\alpha S$ Lévy distribution provides a better fit for the SGN. This assertion was purportedly debunked and rebounded to the Gaussian noise model that had been previously proposed. This study provides robust, comprehensive empirical evidence that SGN is heavy-tailed and is better represented by the $S\alpha S$ distribution. Our experiments include several datasets and multiple models, both discriminative and generative. Furthermore, we argue that different network parameters preserve distinct SGN properties. We develop a novel framework based on a Lévy-driven stochastic differential equation (SDE), where one-dimensional Lévy processes describe each parameter. This leads to a more accurate characterization of the dynamics of SGD around local minima. We use our framework to study SGD properties near local minima; these include the mean escape time and preferable exit directions. Barak Battash, Lior Wolf, Ofir Lindenbaum |
AISTATS | 2 |
| 2024 | Multi-Dimensional Hyena for Spatial Inductive BiasabstractThe advantage of Vision Transformers over CNNs is only fully manifested when trained over a large dataset, mainly due to the reduced inductive bias towards spatial locality within the transformer’s self-attention mechanism. In this work, we present a data-efficient vision transformer that does not rely on self-attention. Instead, it employs a novel generalization to multiple axes of the very recent Hyena layer. We propose several alternative approaches for obtaining this generalization and delve into their unique distinctions and considerations from both empirical and theoretical perspectives. The proposed Hyena N-D layer boosts the performance of various Vision Transformer architectures, such as ViT, Swin, and DeiT across multiple datasets. Furthermore, in the small dataset regime, our Hyena-based ViT is favorable to ViT variants from the recent literature that are specifically designed for solving the same challenge. Finally, we show that a hybrid approach that is based on Hyena N-D for the first layers in ViT, followed by layers that incorporate conventional attention, consistently boosts the performance of various vision transformer architectures. Our code is attached as supplementary. Itamar Zimerman, Lior Wolf |
AISTATS | 2 |
| 2024 | Fine-Tuning CLIP via Explainability Map Propagation for Boosting Image and Video Retrieval
Yoav Shalev, Lior Wolf |
ECIR (1) | 2 |
| 2024 | Backward Lens: Projecting Language Model Gradients into the Vocabulary SpaceabstractUnderstanding how Transformer-based Language Models (LMs) learn and recall information is a key goal of the deep learning community.Recent interpretability methods project weights and hidden states obtained from the forward pass to the models' vocabularies, helping to uncover how information flows within LMs.In this work, we extend this methodology to LMs' backward pass and gradients.We first prove that a gradient matrix can be cast as a low-rank linear combination of its forward and backward passes' inputs.We then develop methods to project these gradients into vocabulary items and explore the mechanics of how new information is stored in the LMs' neurons.Our code is available Shahar Katz, Yonatan Belinkov, Mor Geva, Lior Wolf |
EMNLP | 4 |
| 2024 | Efficient Verification-Based Face IdentificationabstractWe study the problem of performing face verification with an efficient neural model$f$. The efficiency of$f$stems from simplifying the face verification problem from an embedding nearest neighbor search into a binary problem; each user has its own neural network$f$. To allow information sharing between different individuals in the training set, we do not train$f$directly but instead generate the model weights using a hypernetwork$h$. This leads to the generation of a compact personalized model for face identification that can be deployed on edge devices. Key to the method's success is a novel way of generating hard negatives and carefully scheduling the training objectives. Our model leads to a substantially small$f$requiring only 23k parameters and 5M floating point operations (FLOPS). We use six face verification datasets to demonstrate that our method is on par or better than state-of-the-art models, with a significantly reduced number of parameters and computational burden. Furthermore, we perform an extensive ablation study to demonstrate the importance of each element in our method. Amit Rozner, Barak Battash, Ofir Lindenbaum, Lior Wolf |
FG | 4 |
| 2024 | A 2-Dimensional State Space Layer for Spatial Inductive BiasabstractA central objective in computer vision is to design models with appropriate 2-D inductive bias. Desiderata for 2-D inductive bias include two-dimensional position awareness, dynamic spatial locality, and translation and permutation invariance. To address these goals, we leverage an expressive variation of the multidimensional State Space Model (SSM). Our approach introduces efficient parameterization, accelerated computation, and a suitable normalization scheme. Empirically, we observe that incorporating our layer at the beginning of each transformer block of Vision Transformers (ViT), as well as when replacing the Conv2D filters of ConvNeXT with our proposed layers significantly enhances performance for multiple backbones and across multiple datasets. The new layer is effective even with a negligible amount of additional parameters and inference time. Ablation studies and visualizations demonstrate that the layer has a strong 2-D inductive bias. For example, vision transformers equipped with our layer exhibit effective performance even without positional encoding. Our code is attached as supplementary. Ethan Baron 0002, Itamar Zimerman, Lior Wolf |
ICLR | 3 |
| 2024 | The Hidden Language of Diffusion ModelsabstractText-to-image diffusion models have demonstrated an unparalleled ability to generate high-quality, diverse images from a textual prompt. However, the internal representations learned by these models remain an enigma. In this work, we present Conceptor, a novel method to interpret the internal representation of a textual concept by a diffusion model. This interpretation is obtained by decomposing the concept into a small set of human-interpretable textual elements. Applied over the state-of-the-art Stable Diffusion model, Conceptor reveals non-trivial structures in the representations of concepts. For example, we find surprising visual connections between concepts, that transcend their textual semantics. We additionally discover concepts that rely on mixtures of exemplars, biases, renowned artistic styles, or a simultaneous fusion of multiple meanings of the concept.
Through a large battery of experiments, we demonstrate Conceptor's ability to provide meaningful, robust, and faithful decompositions for a wide variety of abstract, concrete, and complex textual concepts, while allowing to naturally connect each decomposition element to its corresponding visual impact on the generated images. Hila Chefer, Oran Lang, Mor Geva, Volodymyr Polosukhin, Assaf Shocher, Michal Irani, Inbar Mosseri, Lior Wolf |
ICLR | 8 |
| 2024 | A Foundation Model for Error Correction CodesabstractIn recent years, Artificial Intelligence has undergone a paradigm shift with the rise of foundation models, which are trained on large amounts of data, typically in a self-supervised way, and can then be adapted to a wide range of downstream tasks. In this work, we propose the first foundation model for Error Correction Codes. This model is trained on multiple codes and can then be applied to an unseen code. To enable this, we extend the Transformer architecture in multiple ways: (1) a code-invariant initial embedding, which is also position- and length-invariant, (2) a learned modulation of the attention maps that is conditioned on the Tanner graph, and (3) a length-invariant code-aware noise prediction module that is based on the parity-check matrix. The proposed architecture is trained on multiple short- and medium-length codes and is able to generalize to unseen codes. Its performance on these codes matches and even outperforms the state of the art, despite having a smaller capacity than the leading code-specific transformers. The suggested framework therefore demonstrates, for the first time, the benefits of learning a universal decoder rather than a neural decoder optimized for a given code. Yoni Choukroun, Lior Wolf |
ICLR | 2 |
| 2024 | Dynamic Layer Tying for Parameter-Efficient TransformersabstractIn the pursuit of reducing the number of trainable parameters in deep transformer networks, we employ Reinforcement Learning to dynamically select layers during training and tie them together. Every few iterations, the RL agent is asked whether to train each layer $i$ independently or to copy the weights of a previous layer $j<i$. This facilitates weight sharing, reduces the number of trainable parameters, and also serves as an effective regularization technique. Experimental evaluations validate that our model modestly outperforms the baseline transformer model with regard to perplexity and drastically reduces the number of trainable parameters. In particular, the memory consumption during training is up to one order of magnitude less than the conventional training method. Tamir David Hay, Lior Wolf |
ICLR | 2 |
| 2024 | Separate and Diffuse: Using a Pretrained Diffusion Model for Better Source SeparationabstractThe problem of speech separation, also known as the cocktail party problem,
refers to the task of isolating a single speech signal from a mixture of speech
signals. Previous work on source separation derived an upper bound for the
source separation task in the domain of human speech. This bound is derived for
deterministic models. Recent advancements in generative models challenge this
bound. We show how the upper bound can be generalized to the case of random
generative models. Applying a diffusion model Vocoder that was pretrained to
model single-speaker voices on the output of a deterministic separation model leads
to state-of-the-art separation results. It is shown that this requires one to combine
the output of the separation model with that of the diffusion model. In our method,
a linear combination is performed, in the frequency domain, using weights that are
inferred by a learned model. We show state-of-the-art results on 2, 3, 5, 10, and 20
speakers on multiple benchmarks. In particular, for two speakers, our method is
able to surpass what was previously considered the upper performance bound. Shahar Lutati, Eliya Nachmani, Lior Wolf |
ICLR | 3 |
| 2024 | Learning Linear Block Error Correction CodesabstractError correction codes are a crucial part of the physical communication layer, ensuring the reliable transfer of data over noisy channels. The design of optimal linear block codes capable of being efficiently decoded is of major concern, especially for short block lengths. While neural decoders have recently demonstrated their advantage over classical decoding techniques, the neural design of the codes remains a challenge. In this work, we propose for the first time a unified encoder-decoder training of binary linear block codes. To this end, we adapt the coding setting to support efficient and differentiable training of the code for end-to-end optimization over the order two Galois field. We also propose a novel Transformer model in which the self-attention masking is performed in a differentiable fashion for the efficient backpropagation of the code gradient. Our results show that (i) the proposed decoder outperforms existing neural decoding on conventional codes, (ii) the suggested framework generates codes that outperform the analogous conventional codes, and (iii) the codes we developed not only excel with our decoder but also show enhanced performance with traditional decoding techniques. Yoni Choukroun, Lior Wolf |
ICML | 2 |
| 2024 | Converting Transformers to Polynomial Form for Secure Inference Over Homomorphic EncryptionabstractDesigning privacy-preserving DL solutions is a major challenge within the AI community. Homomorphic Encryption (HE) has emerged as one of the most promising approaches in this realm, enabling the decoupling of knowledge between a model owner and a data owner. Despite extensive research and application of this technology, primarily in CNNs, applying HE on transformer models has been challenging because of the difficulties in converting these models into a polynomial form. We break new ground by introducing the first polynomial transformer, providing the first demonstration of secure inference over HE with full transformers. This includes a transformer architecture tailored for HE, alongside a novel method for converting operators to their polynomial equivalent. This innovation enables us to perform secure inference on LMs and ViTs with several datasts and tasks. Our techniques yield results comparable to traditional models, bridging the performance gap with transformers of similar scale and underscoring the viability of HE for state-of-the-art applications. Finally, we assess the stability of our models and conduct a series of ablations to quantify the contribution of each model component. Our code is publicly available. Itamar Zimerman, Moran Baruch, Nir Drucker, Gilad Ezov, Omri Soceanu, Lior Wolf |
ICML | 6 |
| 2024 | Viewing Transformers Through the Lens of Long Convolutions LayersabstractDespite their dominance in modern DL and, especially, NLP domains, transformer architectures exhibit sub-optimal performance on long-range tasks compared to recent layers that are specifically designed for this purpose. In this work, drawing inspiration from key attributes of longrange layers, such as state-space layers, linear RNN layers, and global convolution layers, we demonstrate that minimal modifications to the transformer architecture can significantly enhance performance on the Long Range Arena (LRA) benchmark, thus narrowing the gap with these specialized layers. We identify that two key principles for long-range tasks are (i) incorporating an inductive bias towards smoothness, and (ii) locality. As we show, integrating these ideas into the attention mechanism improves results with a negligible amount of additional computation and without any additional trainable parameters. Our theory and experiments also shed light on the reasons for the inferior performance of transformers on long-range tasks and identify critical properties that are essential for successfully capturing long-range dependencies. Itamar Zimerman, Lior Wolf |
ICML | 2 |
| 2024 | Anomaly Detection with Variance Stabilized Density EstimationabstractWe propose a modified density estimation problem that is highly effective for detecting anomalies in tabular data. Our approach assumes that the density function is relatively stable (with lower variance) around normal samples. We have verified this hypothesis empirically using a wide range of real-world data. Then, we present a variance-stabilized density estimation problem for maximizing the likelihood of the observed samples while minimizing the variance of the density around normal samples. To obtain a reliable anomaly detector, we introduce a spectral ensemble of autoregressive models for learning the variance-stabilized distribution. We have conducted an extensive benchmark with 52 datasets, demonstrating that our method leads to state-of-the-art results while alleviating the need for data-specific hyperparameter tuning. Finally, we have used an ablation study to demonstrate the importance of each of the proposed components, followed by a stability analysis evaluating the robustness of our model. Amit Rozner, Barak Battash, Henry Li, Lior Wolf, Ofir Lindenbaum |
UAI | 4 |
| 2024 | Harnessing the flexibility of neural networks to predict dynamic theoretical parameters underlying human choice behaviorabstractReinforcement learning (RL) models are used extensively to study human behavior. These rely on normative models of behavior and stress interpretability over predictive capabilities. More recently, neural network models have emerged as a descriptive modeling paradigm that is capable of high predictive power yet with limited interpretability. Here, we seek to augment the expressiveness of theoretical RL models with the high flexibility and predictive power of neural networks. We introduce a novel framework, which we term theoretical-RNN (t-RNN), whereby a recurrent neural network is trained to predict trial-by-trial behavior and to infer theoretical RL parameters using artificial data of RL agents performing a two-armed bandit task. In three studies, we then examined the use of our approach to dynamically predict unseen behavior along with time-varying theoretical RL parameters. We first validate our approach using synthetic data with known RL parameters. Next, as a proof-of-concept, we applied our framework to two independent datasets of humans performing the same task. In the first dataset, we describe differences in theoretical RL parameters dynamic among clinical psychiatric vs. healthy controls. In the second dataset, we show that the exploration strategies of humans varied dynamically in response to task phase and difficulty. For all analyses, we found better performance in the prediction of actions for t-RNN compared to the stationary maximum-likelihood RL method. We discuss the use of neural networks to facilitate the estimation of latent RL parameters underlying choice behavior. Yoav Ger, Eliya Nachmani, Lior Wolf, Nitzan Shahar |
PLoS Comput. Biol. | 3 |
| 2024 | Still-Moving: Customized Video Generation without Customized Video DataabstractCustomizing text-to-image (T2I) models has seen tremendous progress recently, particularly in areas such as personalization, stylization, and conditional generation. However, expanding this progress to video generation is still in its infancy, primarily due to the lack of customized video data. In this work, we introduce Still-Moving, a novel generic framework for customizing a text-to-video (T2V) model, without requiring any customized video data. The framework applies to the prominent T2V design where the video model is built over a T2I model (e.g., via inflation). We assume access to a customized version of the T2I model, trained only on still image data (e.g., using DreamBooth). Naively plugging in the weights of the customized T2I model into the T2V model often leads to significant artifacts or insufficient adherence to the customization data. To overcome this issue, we train lightweight Spatial Adapters that adjust the features produced by the injected T2I layers. Importantly, our adapters are trained on "frozen videos" (i.e., repeated images), constructed from image samples generated by the customized T2I model. This training is facilitated by a novel Motion Adapter module, which allows us to train on such static videos while preserving the motion prior of the video model. At test time, we remove the Motion Adapter modules and leave in only the trained Spatial Adapters. This restores the motion prior of the T2V model while adhering to the spatial prior of the customized T2I model. We demonstrate the effectiveness of our approach on diverse tasks including personalized, stylized, and conditional generation. In all evaluated scenarios, our method seamlessly integrates the spatial prior of the customized T2I model with a motion prior supplied by the T2V model. Hila Chefer, Shiran Zada, Roni Paiss, Ariel Ephrat, Omer Tov, Michael Rubinstein, Lior Wolf, Tali Dekel, Tomer Michaeli, Inbar Mosseri |
ACM Trans. Graph. | 7 |
| 2024 | Training-Free Consistent Text-to-Image GenerationabstractText-to-image models offer a new level of creative flexibility by allowing users to guide the image generation process through natural language. However, using these models to consistently portraythe samesubject across diverse prompts remains challenging. Existing approaches fine-tune the model to teach it new words that describe specific user-provided subjects or add image conditioning to the model. These methods require lengthy persubject optimization or large-scale pre-training. Moreover, they struggle to align generated images with text prompts and face difficulties in portraying multiple subjects. Here, we presentConsiStory, atraining-freeapproach that enables consistent subject generation by sharing the internal activations of the pretrained model. We introduce a subject-driven shared attention block and correspondence-based feature injection to promote subject consistency between images. Additionally, we develop strategies to encourage layout diversity while maintaining subject consistency. We compareConsiStoryto a range of baselines, and demonstrate state-of-the-art performance on subject consistency and text alignment, without requiring a single optimization step. Finally,ConsiStorycan naturally extend to multi-subject scenarios, and even enable training-freepersonalizationfor common objects. Yoad Tewel, Omri Kaduri, Rinon Gal, Yoni Kasten, Lior Wolf, Gal Chechik, Yuval Atzmon |
ACM Trans. Graph. | 5 |
| 2023 | Degree-based stratification of nodes in Graph Neural Networks
Ameen Ali, Lior Wolf, Hakan Çevikalp |
ACML | 2 |
| 2023 | Cross-Domain Relation Adaptation
Ido Kessler, Omri Lifshitz, Sagie Benaim, Lior Wolf |
ACML | 4 |
| 2023 | AutoSAM: Adapting SAM to Medical Images by Overloading the Prompt Encoder
Tal Shaharabany, Aviad Dahan, Raja Giryes, Lior Wolf |
BMVC | 4 |
| 2023 | Zero-Shot Video Captioning by Evolving Pseudo-tokens
Yoad Tewel, Yoav Shalev, Roy Nadler, Idan Schwartz, Lior Wolf |
BMVC | 5 |
| 2023 | Similarity Maps for Self-Training Weakly-Supervised Phrase GroundingabstractA phrase grounding model receives an input image and a text phrase and outputs a suitable localization map. We present an effective way to refine a phrase ground model by considering self-similarity maps extracted from the latent representation of the model's image encoder. Our main insights are that these maps resemble localization maps and that by combining such maps, one can obtain useful pseudo-labels for performing self-training. Our results surpass, by a large margin, the state of the art in weakly supervised phrase grounding. A similar gap in performance is obtained for a recently proposed downstream task called WWbL, in which only the image is input, without any text. Our code is available at https://github.com/talshaharabany/Similarity-Maps-for-Self-Training-Weakly-Supervised-Phrase-Grounding Tal Shaharabany, Lior Wolf |
CVPR | 2 |
| 2023 | Focus Your Attention (with Adaptive IIR Filters)abstractWe present a new layer in which dynamic (i.e., input-dependent) Infinite Impulse Response (IIR) filters of order two are used to process the input sequence prior to applying conventional attention.The input is split into chunks, and the coefficients of these filters are determined based on previous chunks to maintain causality.Despite their relatively low order, the causal adaptive filters are shown to focus attention on the relevant sequence elements.The new layer is grounded in control theory, and is shown to generalize diagonal state-space layers.The layer performs on-par with state-of-the-art networks, with a fraction of their parameters and with time complexity that is sub-quadratic with input size.The obtained layer is favorable to layers such as Heyna, GPT2, and Mega, both with respect to the number of parameters and the obtained level of performance on multiple long-range sequence problems. Shahar Lutati, Itamar Zimerman, Lior Wolf |
EMNLP | 3 |
| 2023 | Energy Regularized RNNS for solving non-stationary Bandit problemsabstractWe consider a Multi-Armed Bandit problem in which the re- wards are non-stationary and are dependent on past actions and potentially on past contexts. At the heart of our method, we employ a recurrent neural network, which models these sequences. In order to balance between exploration and exploitation, we present an energy minimization term that pre- vents the neural network from becoming too confident in support of a certain action. This term provably limits the gap between the maximal and minimal probabilities assigned by the network. In a diverse set of experiments, we demonstrate that our method is at least as effective as methods suggested to solve the sub-problem of Rotting Bandits, and can solve intuitive extensions of various benchmark problems. We share our implementation at https://github.com/rotmanmi/Energy-Regularized-RNN. Michael Rotman, Lior Wolf |
ICASSP | 2 |
| 2023 | Learning a Weight Map for Weakly-Supervised LocalizationabstractIn the weakly supervised localization setting, supervision is given as an image-level label. We propose employing an image classifier f and training a generative network g that outputs, given the input image, a per-pixel weight map that indicates the location of the object within the image. Network g is trained by minimizing the discrepancy between the output of the classifier f on the original image and its output given the same image weighted by the output of g. Our results indicate that the method outperforms existing localization methods on the challenging fine-grained classification datasets. Tal Shaharabany, Lior Wolf |
ICASSP | 2 |
| 2023 | Box-based Refinement for Weakly Supervised and Unsupervised Localization TasksabstractIt has been established that training a box-based detector network can enhance the localization performance of weakly supervised and unsupervised methods. Moreover, we extend this understanding by demonstrating that these detectors can be utilized to improve the original network, paving the way for further advancements. To accomplish this, we train the detectors on top of the network output instead of the image data and apply suitable loss backpropagation. Our findings reveal a significant improvement in phrase grounding for the "what is where by looking" task, as well as various methods of unsupervised object discovery. Our code is available at https://github.com/eyalgomel/box-based-refinement. Eyal Gomel, Tal Shaharabany, Lior Wolf |
ICCV | 3 |
| 2023 | Discriminative Class Tokens for Text-to-Image Diffusion ModelsabstractRecent advances in text-to-image diffusion models have enabled the generation of diverse and high-quality images. While impressive, the images often fall short of depicting subtle details and are susceptible to errors due to ambiguity in the input text. One way of alleviating these issues is to train diffusion models on class-labeled datasets. This approach has two disadvantages: (i) supervised datasets are generally small compared to large-scale scraped text-image datasets on which text-to-image models are trained, affecting the quality and diversity of the generated images, or (ii) the input is a hard-coded label, as opposed to free-form text, limiting the control over the generated images.In this work, we propose a non-invasive fine-tuning technique that capitalizes on the expressive potential of freeform text while achieving high accuracy through discriminative signals from a pretrained classifier. This is done by iteratively modifying the embedding of an added input token of a text-to-image diffusion model, by steering generated images toward a given target class according to a classifier. Our method is fast compared to prior fine-tuning methods and does not require a collection of in-class images or retraining of a noise-tolerant classifier. We evaluate our method extensively, showing that the generated images are: (i) more accurate and of higher quality than standard diffusion models, (ii) can be used to augment training data in a low-resource setting, and (iii) reveal information about the data used to train the guiding classifier. The code is available at https://github.com/idansc/discriminative_class_tokens. Idan Schwartz, Vésteinn Snæbjarnarson, Hila Chefer, Serge J. Belongie, Lior Wolf, Sagie Benaim |
ICCV | 5 |
| 2023 | Decision S4: Efficient Sequence-Based RL via State Spaces Layers
Shmuel Bar-David, Itamar Zimerman, Eliya Nachmani, Lior Wolf |
ICLR | 4 |
| 2023 | Denoising Diffusion Error Correction Codes
Yoni Choukroun, Lior Wolf |
ICLR | 2 |
| 2023 | OCD: Learning to Overfit with Conditional Diffusion ModelsabstractWe present a dynamic model in which the weights are conditioned on an input sample x and are learned to match those that would be obtained by finetuning a base model on x and its label y. This mapping between an input sample and network weights is approximated by a denoising diffusion model. The diffusion model we employ focuses on modifying a single layer of the base model and is conditioned on the input, activations, and output of this layer. Since the diffusion model is stochastic in nature, multiple initializations generate different networks, forming an ensemble, which leads to further improvements. Our experiments demonstrate the wide applicability of the method for image classification, 3D reconstruction, tabular data, speech separation, and natural language processing. Shahar Lutati, Lior Wolf |
ICML | 2 |
| 2023 | Adaptation of Text-Conditioned Diffusion Models for Audio-to-Image Generation
Guy Yariv, Itai Gat, Lior Wolf, Yossi Adi, Idan Schwartz |
INTERSPEECH | 3 |
| 2023 | Annotator Consensus Prediction for Medical Image Segmentation with Diffusion Models
Tomer Amit, Shmuel Shichrur, Tal Shaharabany, Lior Wolf |
MICCAI (4) | 4 |
| 2023 | Reconstructing the Hemodynamic Response Function via a Bimodal Transformer
Yoni Choukroun, Lior Golgher, Pablo Blinder, Lior Wolf |
MICCAI (2) | 4 |
| 2023 | Dynamically-Scaled Deep Canonical Correlation Analysis
Tomer Friedlander, Lior Wolf |
PAKDD (3) | 2 |
| 2023 | Semi-supervised learning of partial differential operators and dynamical flowsabstractThe evolution of many dynamical systems is generically governed by nonlinear partial differential equations (PDEs), whose solution, in a simulation framework, requires vast amounts of computational resources. In this work, we present a novel method that combines a hyper-network solver with a Fourier Neural Operator architecture. Our method treats time and space separately and as a result, it successfully propagates initial conditions in continuous time steps by employing the general composition properties of the partial differential operators. Following previous works, supervision is provided at a specific time point. We test our method on various time evolution PDEs, including nonlinear fluid flows in one, two, or three spatial dimensions. The results show that the new method improves the learning accuracy at the time of the supervision point, and can interpolate the solutions to any intermediate time. Michael Rotman, Amit Dekel, Ran Ilan Ber, Lior Wolf, Yaron Oz |
UAI | 4 |
| 2023 | Attend-and-Excite: Attention-Based Semantic Guidance for Text-to-Image Diffusion ModelsabstractRecent text-to-image generative models have demonstrated an unparalleled ability to generate diverse and creative imagery guided by a target text prompt. While revolutionary, current state-of-the-art diffusion models may still fail in generating images that fully convey the semantics in the given text prompt. We analyze the publicly available Stable Diffusion model and assess the existence of catastrophic neglect , where the model fails to generate one or more of the subjects from the input prompt. Moreover, we find that in some cases the model also fails to correctly bind attributes ( e.g. , colors) to their corresponding subjects. To help mitigate these failure cases, we introduce the concept of Generative Semantic Nursing (GSN) , where we seek to intervene in the generative process on the fly during inference time to improve the faithfulness of the generated images. Using an attention-based formulation of GSN, dubbed Attend-and-Excite , we guide the model to refine the cross-attention units to attend to all subject tokens in the text prompt and strengthen --- or excite --- their activations, encouraging the model to generate all subjects described in the text prompt. We compare our approach to alternative approaches and demonstrate that it conveys the desired concepts more faithfully across a range of text prompts. Code is available at our project page: https://attendandexcite.github.io/Attend-and-Excite/. Hila Chefer, Yuval Alaluf, Yael Vinker, Lior Wolf, Daniel Cohen-Or |
ACM Trans. Graph. | 4 |
| 2022 | Dynamic Dual-Output Diffusion ModelsabstractIterative denoising-based generation, also known as denoising diffusion models, has recently been shown to be comparable in quality to other classes of generative models, and even surpass them. Including, in particular, Generative Adversarial Networks, which are currently the state of the art in many subtasks of image generation. However, a major drawback of this method is that it requires hundreds of iterations to produce a competitive result. Recent works have proposed solutions that allow for faster generation with fewer iterations, but the image quality gradually deteriorates with increasingly fewer iterations being applied during generation. In this paper, we reveal some of the causes that affect the generation quality of diffusion models, especially when sampling with few iterations, and come up with a simple, yet effective, solution to mitigate them. We consider two opposite equations for the iterative denoising, the first predicts the applied noise, and the second predicts the image directly. Our solution takes the two options and learns to dynamically alternate between them through the denoising process. Our proposed solution is general and can be applied to any existing diffusion model. As we show, when applied to various SOTA architectures, our solution immediately improves their generation quality, with negligible added complexity and parameters. We experiment on multiple datasets and configurations and run an extensive ablation study to support these findings. Yaniv Benny, Lior Wolf |
CVPR | 2 |
| 2022 | Image Animation with Perturbed MasksabstractWe present a novel approach for image-animation of a source image by a driving video, both depicting the same type of object. We do not assume the existence of pose models and our method is able to animate arbitrary objects without the knowledge of the object's structure. Furthermore, both, the driving video and the source image are only seen during test-time. Our method is based on a shared mask generator, which separates the foreground object from its background, and captures the object's general pose and shape. To control the source of the identity of the output frame, we employ perturbations to interrupt the unwanted identity information on the driver's mask. A mask-refinement module then replaces the identity of the driver with the identity of the source. Conditioned on the source image, the transformed mask is then decoded by a multi-scale generator that renders a realistic image, in which the content of the source frame is animated by the pose in the driving video. Due to the lack of fully supervised data, we train on the task of reconstructing frames from the same video the source image is taken from. Our method is shown to greatly outperform the state-of-the-art methods on multiple benchmarks. Our code and samples are available at https://github.com/itsyoavshalevlImage-Animation-with-Perturbed-Masks. Yoav Shalev, Lior Wolf |
CVPR | 2 |
| 2022 | ZeroCap: Zero-Shot Image-to-Text Generation for Visual-Semantic ArithmeticabstractRecent text-to-image matching models apply contrastive learning to large corpora of uncurated pairs of images and sentences. While such models can provide a powerful score for matching and subsequent zero-shot tasks, they are not capable of generating caption given an image. In this work, we repurpose such models to generate a descriptive text given an image at inference time, without any further training or tuning step. This is done by combining the visual-semantic model with a large language model, benefiting from the knowledge in both web-scale models. The resulting captions are much less restrictive than those obtained by supervised captioning methods. Moreover, as a zero-shot learning method, it is extremely flexible and we demonstrate its ability to perform image arithmetic in which the inputs can be either images or text and the output is a sentence. This enables novel high-level vision capabilities such as comparing two images or solving visual analogy tests. Our code is available at: https://github.com/YoadTew/zero-shot-image-to-text. Yoad Tewel, Yoav Shalev, Idan Schwartz, Lior Wolf |
CVPR | 4 |
| 2022 | Image-Based CLIP-Guided Essence Transfer
Hila Chefer, Sagie Benaim, Roni Paiss, Lior Wolf |
ECCV (13) | 4 |
| 2022 | No Token Left Behind: Explainability-Aided Image Classification and Generation
Roni Paiss, Hila Chefer, Lior Wolf |
ECCV (12) | 3 |
| 2022 | FewGAN: Generating from the Joint Distribution of a Few ImagesabstractWe introduce FewGAN, a generative model for generating novel, high-quality and diverse images whose patch distribution lies in the joint patch distribution of a small number of N > 1 training samples. The method is, in essence, a hierarchical patch-GAN that applies quantization at the first coarse scale, in a similar fashion to VQ-GAN, followed by a pyramid of residual fully convolutional GANs at finer scales. Our key idea is to first use quantization to learn a fixed set of patch embeddings for training images. We then use a separate set of side images to model the structure of generated images using an autoregressive model trained on the learned patch embeddings of training images. Using quantization at the coarsest scale allows the model to generate both conditional and unconditional novel images. Subsequently, a patch-GAN renders the fine details, resulting in high-quality images. In an extensive set of experiments, it is shown that FewGAN outperforms baselines both quantitatively and qualitatively. Lior Ben-Moshe, Sagie Benaim, Lior Wolf |
ICIP | 3 |
| 2022 | Unsupervised Disentanglement with Tensor Product Representations on the Torus
Michael Rotman, Amit Dekel, Shir Gur, Yaron Oz, Lior Wolf |
ICLR | 5 |
| 2022 | Anomaly Detection for Tabular Data with Internal Contrastive Learning
Tom Shenkar, Lior Wolf |
ICLR | 2 |
| 2022 | XAI for Transformers: Better Explanations through Conservative PropagationabstractTransformers have become an important workhorse of machine learning, with numerous applications. This necessitates the development of reliable methods for increasing their transparency. Multiple interpretability methods, often based on gradient information, have been proposed. We show that the gradient in a Transformer reflects the function only locally, and thus fails to reliably identify the contribution of input features to the prediction. We identify Attention Heads and LayerNorm as main reasons for such unreliable explanations and propose a more stable way for propagation through these layers. Our proposal, which can be seen as a proper extension of the well-established LRP method to Transformers, is shown both theoretically and empirically to overcome the deficiency of a simple gradient-based approach, and achieves state-of-the-art explanation performance on a broad range of Transformer models and datasets. Ameen Ali, Thomas Schnake, Oliver Eberle, Grégoire Montavon, Klaus-Robert Müller, Lior Wolf |
ICML | 6 |
| 2022 | Neural Inverse KinematicabstractInverse kinematic (IK) methods recover the parameters of the joints, given the desired position of selected elements in the kinematic chain. While the problem is well-defined and low-dimensional, it has to be solved rapidly, accounting for multiple possible solutions. In this work, we propose a neural IK method that employs the hierarchical structure of the problem to sequentially sample valid joint angles conditioned on the desired position and on the preceding joints along the chain. In our solution, a hypernetwork $f$ recovers the parameters of multiple primary networks {$g_1,g_2,…,g_N$, where $N$ is the number of joints}, such that each $g_i$ outputs a distribution of possible joint angles, and is conditioned on the sampled values obtained from the previous primary networks $g_j, j Cite this Paper BibTeX @InProceedings{pmlr-v162-bensadoun22a, title = {Neural Inverse Kinematic}, author = {Bensadoun, Raphael and Gur, Shir and Blau, Nitsan and Wolf, Lior}, booktitle = {Proceedings of the 39th International Conference on Machine Learning}, pages = {1787--1797}, year = {2022}, editor = {Chaudhuri, Kamalika and Jegelka, Stefanie and Song, Le and Szepesvari, Csaba and Niu, Gang and Sabato, Sivan}, volume = {162}, series = {Proceedings of Machine Learning Research}, month = {17--23 Jul}, publisher = {PMLR}, pdf = {https://proceedings.mlr.press/v162/bensadoun22a/bensadoun22a.pdf}, url = {https://proceedings.mlr.press/v162/bensadoun22a.html}, abstract = {Inverse kinematic (IK) methods recover the parameters of the joints, given the desired position of selected elements in the kinematic chain. While the problem is well-defined and low-dimensional, it has to be solved rapidly, accounting for multiple possible solutions. In this work, we propose a neural IK method that employs the hierarchical structure of the problem to sequentially sample valid joint angles conditioned on the desired position and on the preceding joints along the chain. In our solution, a hypernetwork $f$ recovers the parameters of multiple primary networks {$g_1,g_2,…,g_N$, where $N$ is the number of joints}, such that each $g_i$ outputs a distribution of possible joint angles, and is conditioned on the sampled values obtained from the previous primary networks $g_j, j Copy to Clipboard Download Endnote %0 Conference Paper %T Neural Inverse Kinematic %A Raphael Bensadoun %A Shir Gur %A Nitsan Blau %A Lior Wolf %B Proceedings of the 39th International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2022 %E Kamalika Chaudhuri %E Stefanie Jegelka %E Le Song %E Csaba Szepesvari %E Gang Niu %E Sivan Sabato %F pmlr-v162-bensadoun22a %I PMLR %P 1787--1797 %U https://proceedings.mlr.press/v162/bensadoun22a.html %V 162 %X Inverse kinematic (IK) methods recover the parameters of the joints, given the desired position of selected elements in the kinematic chain. While the problem is well-defined and low-dimensional, it has to be solved rapidly, accounting for multiple possible solutions. In this work, we propose a neural IK method that employs the hierarchical structure of the problem to sequentially sample valid joint angles conditioned on the desired position and on the preceding joints along the chain. In our solution, a hypernetwork $f$ recovers the parameters of multiple primary networks {$g_1,g_2,…,g_N$, where $N$ is the number of joints}, such that each $g_i$ outputs a distribution of possible joint angles, and is conditioned on the sampled values obtained from the previous primary networks $g_j, j Copy to Clipboard Download APA Bensadoun, R., Gur, S., Blau, N. & Wolf, L.. (2022). Neural Inverse Kinematic. Proceedings of the 39th International Conference on Machine Learning, in Proceedings of Machine Learning Research 162:1787-1797 Available from https://proceedings.mlr.press/v162/bensadoun22a.html. Copy to Clipboard Download Related Material Download PDF This site last compiled Sun, 05 Jul 2026 14:54:17 +0000 Github Account Copyright © The authors and PMLR 2026. MLResearchPress Raphael Bensadoun, Shir Gur, Nitsan Blau, Lior Wolf |
ICML | 4 |
| 2022 | Geometric Transformer for End-to-End Molecule Properties PredictionabstractTransformers have become methods of choice in many applications thanks to their ability to represent complex interactions between elements. However, extending the Transformer architecture to non-sequential data such as molecules and enabling its training on small datasets remains a challenge. In this work, we introduce a Transformer-based architecture for molecule property prediction, which is able to capture the geometry of the molecule. We modify the classical positional encoder by an initial encoding of the molecule geometry, as well as a learned gated self-attention mechanism. We further suggest an augmentation scheme for molecular data capable of avoiding the overfitting induced by the overparameterized architecture. The proposed framework outperforms the state-of-the-art methods while being based on pure machine learning solely, i.e. the method does not incorporate domain knowledge from quantum chemistry and does not use extended geometric inputs besides the pairwise atomic distances. Yoni Choukroun, Lior Wolf |
IJCAI | 2 |
| 2022 | Zero-Shot Voice Conditioning for Denoising Diffusion TTS ModelsabstractWe present a novel way of conditioning a pretrained denoising diffusion speech model to produce speech in the voice of a novel person unseen during training.The method requires a short (∼ 3 seconds) sample from the target person, and generation is steered at inference time, without any training steps.At the heart of the method lies a sampling process that combines the estimation of the denoising model with a low-pass version of the new speaker's sample.The objective and subjective evaluations show that our sampling method can generate a voice similar to that of the target speaker in terms of frequency, with an accuracy comparable to state-of-the-art methods, and without training. Alon Levkovitch, Eliya Nachmani, Lior Wolf |
INTERSPEECH | 3 |
| 2022 | SepIt: Approaching a Single Channel Speech Separation BoundabstractWe present an upper bound for the Single Channel Speech Separation task, which is based on an assumption regarding the nature of short segments of speech.Using the bound, we are able to show that while the recent methods have made great progress for a few speakers, there is room for improvement for five and ten speakers.We then introduce a Deep neural network, SepIt, that iteratively improves the different speakers' estimation.At test time, SpeIt has a varying number of iterations per test sample, based on a mutual information criterion that arises from our analysis.In an extensive set of experiments, SepIt outperforms the state of the art neural networks for 2, 3, 5, and 10 speakers. Shahar Lutati, Eliya Nachmani, Lior Wolf |
INTERSPEECH | 3 |
| 2022 | fMRI Neurofeedback Learning Patterns are Predictive of Personal and Clinical Traits
Rotem Leibovitz, Jhonathan Osin, Lior Wolf, Guy Gurevitch, Talma Hendler |
MICCAI (1) | 3 |
| 2022 | End-to-End Segmentation of Medical Images via Patch-Wise Polygons Prediction
Tal Shaharabany, Lior Wolf |
MICCAI (5) | 2 |
| 2022 | A-Muze-Net: Music Generation by Composing the Harmony Based on the Generated Melody
Or Goren, Eliya Nachmani, Lior Wolf |
MMM (1) | 3 |
| 2022 | Optimizing Relevance Maps of Vision Transformers Improves RobustnessabstractIt has been observed that visual classification models often rely mostly on spurious cues such as the image background, which hurts their robustness to distribution changes. To alleviate this shortcoming, we propose to monitor the model's relevancy signal and direct the model to base its prediction on the foreground object.This is done as a finetuning step, involving relatively few samples consisting of pairs of images and their associated foreground masks. Specifically, we encourage the model's relevancy map (i) to assign lower relevance to background regions, (ii) to consider as much information as possible from the foreground, and (iii) we encourage the decisions to have high confidence. When applied to Vision Transformer (ViT) models, a marked improvement in robustness to domain-shifts is observed. Moreover, the foreground masks can be obtained automatically, from a self-supervised variant of the ViT model itself; therefore no additional supervision is required. Our code is available at: https://github.com/hila-chefer/RobustViT. Hila Chefer, Idan Schwartz, Lior Wolf |
NeurIPS | 3 |
| 2022 | Error Correction Code TransformerabstractError correction code is a major part of the physical communication layer, ensuring the reliable transfer of data over noisy channels.Recently, neural decoders were shown to outperform classical decoding techniques.However, the existing neural approaches present strong overfitting, due to the exponential training complexity, or a restrictive inductive bias, due to reliance on Belief Propagation.Recently, Transformers have become methods of choice in many applications, thanks to their ability to represent complex interactions between elements.In this work, we propose to extend for the first time the Transformer architecture to the soft decoding of linear codes at arbitrary block lengths.We encode each channel's output dimension to a high dimension for a better representation of the bits' information to be processed separately.The element-wise processing allows the analysis of channel output reliability, while the algebraic code and the interaction between the bits are inserted into the model via an adapted masked self-attention module.The proposed approach demonstrates the power and flexibility of Transformers and outperforms existing state-of-the-art neural decoders by large margins, at a fraction of their time complexity. Yoni Choukroun, Lior Wolf |
NeurIPS | 2 |
| 2022 | What is Where by Looking: Weakly-Supervised Open-World Phrase-Grounding without Text InputsabstractGiven an input image, and nothing else, our method returns the bounding boxes of objects in the image and phrases that describe the objects. This is achieved within an open world paradigm, in which the objects in the input image may not have been encountered during the training of the localization mechanism. Moreover, training takes place in a weakly supervised setting, where no bounding boxes are provided. To achieve this, our method combines two pre-trained networks: the CLIP image-to-text matching score and the BLIP image captioning tool. Training takes place on COCO images and their captions and is based on CLIP. Then, during inference, BLIP is used to generate a hypothesis regarding various regions of the current image. Our work generalizes weakly supervised segmentation and phrase grounding and is shown empirically to outperform the state of the art in both domains. It also shows very convincing results in the novel task of weakly-supervised open-world purely visual phrase-grounding presented in our work.For example, on the datasets used for benchmarking phrase-grounding, our method results in a very modest degradation in comparison to methods that employ human captions as an additional input. Tal Shaharabany, Yoad Tewel, Lior Wolf |
NeurIPS | 3 |
| 2022 | Video and Text Matching with Conditioned EmbeddingsabstractWe present a method for matching a text sentence from a given corpus to a given video clip and vice versa. Traditionally video and text matching is done by learning a shared embedding space and the encoding of one modality is independent of the other. In this work, we encode the dataset data in a way that takes into account the query’s relevant information. The power of the method is demonstrated to arise from pooling the interaction data between words and frames. Since the encoding of the video clip depends on the sentence compared to it, the representation needs to be recomputed for each potential match. To this end, we propose an efficient shallow neural network. Its training employs a hierarchical triplet loss that is extendable to paragraph/video matching. The method is simple, provides explainability, and achieves state-of-the-art results for both sentence-clip and video-text by a sizable margin across five different datasets: ActivityNet, DiDeMo, YouCook2, MSR-VTT, and LSMDC. We also show that our conditioned representation can be transferred to video-guided machine translation, where we improved the current results on VATEX. Source code is available at https://github.com/AmeenAli/VideoMatch. Ameen Ali, Idan Schwartz, Tamir Hazan, Lior Wolf |
WACV | 4 |
| 2022 | DeepFake Detection Based on Discrepancies Between Faces and Their ContextabstractWe propose a method for detecting face swapping and other identity manipulations in single images. Face swapping methods, such as DeepFake, manipulate the face region, aiming to adjust the face to the appearance of its context, while leaving the context unchanged. We show that this modus operandi produces discrepancies between the two regions (e.g., Fig. 1). These discrepancies offer exploitable telltale signs of manipulation. Our approach involves two networks: (i) a face identification network that considers the face region bounded by a tight semantic segmentation, and (ii) a context recognition network that considers the face context (e.g., hair, ears, neck). We describe a method which uses the recognition signals from our two networks to detect such discrepancies, providing a complementary detection signal that improves conventional real versus fake classifiers commonly used for detecting fake images. Our method achieves state of the art results on the FaceForensics++ and Celeb-DF-v2 benchmarks for face manipulation detection, and even generalizes to detect fakes produced by unseen methods. Yuval Nirkin, Lior Wolf, Yosi Keller, Tal Hassner |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Visualization of Supervised and Self-Supervised Neural Networks via Attribution Guided FactorizationabstractNeural network visualization techniques mark image locations by their relevancy to the network's classification. Existing methods are effective in highlighting the regions that affect the resulting classification the most. However, as we show, these methods are limited in their ability to identify the support for alternative classifications, an effect we name the saliency bias hypothesis. In this work, we integrate two lines of research: gradient-based methods and attribution-based methods, and develop an algorithm that provides per-class explainability. The algorithm back-projects the per pixel local influence, in a manner that is guided by the local attributions, while correcting for salient features that would otherwise bias the explanation. In an extensive battery of experiments, we demonstrate the ability of our methods to class-specific visualization, and not just the predicted label. Remarkably, the method obtains state of the art results in benchmarks that are commonly applied to gradient-based methods as well as in those that are employed mostly for evaluating attribution methods. Using a new unsupervised procedure, our method is also successful in demonstrating that self-supervised methods learn semantic information. Our code is available at: https://github.com/shirgur/AGFVisualization. Shir Gur, Ameen Ali, Lior Wolf |
AAAI | 3 |
| 2021 | Computational Visual Ceramicology: Matching Image Outlines to Catalog SketchesabstractField archeologists are called upon to identify potsherds, for which they rely on their professional experience and on reference works. We have developed a recognition method starting from images captured on site, which relies on the shape of the sherd's fracture outline. The method sets up a new target for deep-learning, integrating information from points along inner and outer surfaces to learn about shapes. Training the classifiers required tackling multiple challenges that arose on account of our working with real-world archeological data: paucity of labeled data; extreme imbalance between instances of different categories; and the need to avoid neglecting rare classes and to take note of minute distinguishing features of some classes. The scarcity of training data was overcome by using synthetically-produced virtual potsherds and by employing multiple data-augmentation techniques. A novel form of training loss allowed us to overcome classification problems caused by under-populated classes and inhomogeneous distribution of discriminative features. Barak Itkin, Lior Wolf, Nachum Dershowitz |
AAAI | 2 |
| 2021 | Sample Selection for Universal Domain AdaptationabstractThis paper studies the problem of unsupervised domain adaption in the universal scenario, in which only some of the classes are shared between the source and target domains. We present a scoring scheme that is effective in identifying the samples of the shared classes. The score is used to select samples in the target domain for which to apply specific losses during training; pseudo-labels for high scoring samples and confidence regularization for low scoring samples. Taken together, our method is shown to outperform, by a sizeable margin, the current state of the art on the literature benchmarks. Omri Lifshitz, Lior Wolf |
AAAI | 2 |
| 2021 | Shuffling Recurrent Neural NetworksabstractWe propose a novel recurrent neural network model, where the hidden state hₜ is obtained by permuting the vector elements of the previous hidden state hₜ₋₁ and adding the output of a learned function β(xₜ) of the input xₜ at time t. In our model, the prediction is given by a second learned function, which is applied to the hidden state s(hₜ). The method is easy to implement, extremely efficient, and does not suffer from vanishing nor exploding gradients. In an extensive set of experiments, the method shows competitive results, in comparison to the leading literature baselines. We share our implementation at https://github.com/rotmanmi/SRNN. Michael Rotman, Lior Wolf |
AAAI | 2 |
| 2021 | Learning Query Expansion over the Nearest Neighbor Graph
Benjamin Eliot Klein, Lior Wolf |
BMVC | 2 |
| 2021 | Scale-Localized Abstract ReasoningabstractWe consider the abstract relational reasoning task, which is commonly used as an intelligence test. Since some patterns have spatial rationales, while others are only semantic, we propose a multi-scale architecture that processes each query in multiple resolutions. We show that indeed different rules are solved by different resolutions and a combined multi-scale approach outperforms the existing state of the art in this task on all benchmarks by 5-54%. The success of our method is shown to arise from multiple novelties. First, it searches for relational patterns in multiple resolutions, which allows it to readily detect visual relations, such as location, in higher resolution, while allowing the lower resolution module to focus on semantic relations, such as shape type. Second, we optimize the reasoning network of each resolution proportionally to its performance, hereby we motivate each resolution to specialize on the rules for which it performs better than the others and ignore cases that are already solved by the other resolutions. Third, we propose a new way to pool information along the rows and the columns of the illustration-grid of the query. Our work also analyses the existing benchmarks, demonstrating that the RAVEN dataset selects the negative examples in a way that is easily exploited. We, therefore, propose a modified version of the RAVEN dataset, named RAVEN-FAIR. Our code and pretrained models are available at https://github.com/yanivbenny/MRNet. Yaniv Benny, Niv Pekar, Lior Wolf |
CVPR | 3 |
| 2021 | Transformer Interpretability Beyond Attention VisualizationabstractSelf-attention techniques, and specifically Transformers, are dominating the field of text processing and are becoming increasingly popular in computer vision classification tasks. In order to visualize the parts of the image that led to a certain classification, existing methods either rely on the obtained attention maps or employ heuristic propagation along the attention graph. In this work, we propose a novel way to compute relevancy for Transformer networks. The method assigns local relevance based on the Deep Taylor Decomposition principle and then propagates these relevancy scores through the layers. This propagation involves attention layers and skip connections, which challenge existing methods. Our solution is based on a specific formulation that is shown to maintain the total relevancy across layers. We benchmark our method on very recent visual Transformer networks, as well as on a text classification problem, and demonstrate a clear advantage over the existing explainability methods. Our code is available at: https://github.com/hila-chefer/Transformer-Explainability. Hila Chefer, Shir Gur, Lior Wolf |
CVPR | 3 |
| 2021 | Single-Shot Freestyle Dance ReenactmentabstractThe task of motion transfer between a source dancer and a target person is a special case of the pose transfer problem, in which the target person changes their pose in accordance with the motions of the dancer. In this work, we propose a novel method that can reanimate a single image by arbitrary video sequences, unseen during training. The method combines three networks: (i) a segmentation-mapping network, (ii) a realistic frame-rendering network, and (iii) a face refinement network. By separating this task into three stages, we are able to attain a novel sequence of realistic frames, capturing natural motion and appearance. Our method obtains significantly better visual quality than previous methods and is able to animate diverse body types and appearances, which are captured in challenging poses. Oran Gafni, Oron Ashual, Lior Wolf |
CVPR | 3 |
| 2021 | HyperSeg: Patch-Wise Hypernetwork for Real-Time Semantic SegmentationabstractWe present a novel, real-time, semantic segmentation network in which the encoder both encodes and generates the parameters (weights) of the decoder. Furthermore, to allow maximal adaptivity, the weights at each decoder block vary spatially. For this purpose, we design a new type of hypernetwork, composed of a nested U-Net for drawing higher level context features, a multi-headed weight generating module which generates the weights of each block in the decoder immediately before they are consumed, for efficient memory utilization, and a primary network that is composed of novel dynamic patch-wise convolutions. Despite the usage of less-conventional blocks, our architecture obtains real-time performance. In terms of the runtime vs. accuracy trade-off, we surpass state of the art (SotA) results on popular semantic segmentation benchmarks: PASCAL VOC 2012 (val. set) and real-time semantic segmentation on Cityscapes, and CamVid. The code is available: https://nirkin.com/hyperseg. Yuval Nirkin, Lior Wolf, Tal Hassner |
CVPR | 2 |
| 2021 | Permuted AdaIN: Reducing the Bias Towards Global Statistics in Image ClassificationabstractRecent work has shown that convolutional neural network classifiers overly rely on texture at the expense of shape cues. We make a similar but different distinction between shape and local image cues, on the one hand, and global image statistics, on the other. Our method, called Permuted Adaptive Instance Normalization (pAdaIN), reduces the representation of global statistics in the hidden layers of image classifiers. pAdaIN samples a random per-mutation π that rearranges the samples in a given batch. Adaptive Instance Normalization (AdaIN) is then applied between the activations of each (non-permuted) sample i and the corresponding activations of the sample π(i), thus swapping statistics between the samples of the batch. Since the global image statistics are distorted, this swapping procedure causes the network to rely on cues, such as shape or texture. By choosing the random permutation with probability p and the identity permutation otherwise, one can control the effect’s strength.With the correct choice of p, fixed apriori for all experiments and selected without considering test data, our method consistently outperforms baselines in multiple settings. In image classification, our method improves on both CIFAR100 and ImageNet using multiple architectures. In the setting of robustness, our method improves on both ImageNet-C and Cifar-100-C for multiple architectures. In the setting of domain adaptation and domain generalization, our method achieves state of the art results on the transfer learning task from GTAV to Cityscapes and on the PACS benchmark. Oren Nuriel, Sagie Benaim, Lior Wolf |
CVPR | 3 |
| 2021 | Maximal Multiverse Learning for Promoting Cross-Task Generalization of Fine-Tuned Language ModelsabstractLanguage modeling with BERT consists of two phases of (i) unsupervised pre-training on unlabeled text, and (ii) fine-tuning for a specific supervised task.We present a method that leverages the second phase to its fullest, by applying an extensive number of parallel classifier heads, which are enforced to be orthogonal, while adaptively eliminating the weaker heads during training.We conduct an extensive inter-and intradataset evaluation, showing that our method improves the generalization ability of BERT, sometimes leading to a +9% gain in accuracy.These results highlight the importance of a proper fine-tuning procedure, especially for relatively smaller-sized datasets.Our code is attached as supplementary. Itzik Malkiel, Lior Wolf |
EACL | 2 |
| 2021 | Caption Enriched Samples for Improving Hateful Memes DetectionabstractThe recently introduced hateful meme challenge demonstrates the difficulty of determining whether a meme is hateful or not.Specifically, both unimodal language models and multimodal vision-language models cannot reach the human level of performance.Motivated by the need to model the contrast between the image content and the overlayed text, we suggest applying an off-the-shelf image captioning tool in order to capture the first.We demonstrate that the incorporation of such automatic captions during fine-tuning improves the results for various unimodal and multimodal models.Moreover, in the unimodal case, continuing the pre-training of language models on augmented and original caption pairs, is highly beneficial to the classification accuracy.Our code is publicly available 1 . Efrat Blaier, Itzik Malkiel, Lior Wolf |
EMNLP (1) | 3 |
| 2021 | MTAdam: Automatic Balancing of Multiple Training Loss TermsabstractWhen training neural models, it is common to combine multiple loss terms.The balancing of these terms requires considerable human effort and is computationally demanding.Moreover, the optimal trade-off between the loss terms can change as training progresses, e.g., for adversarial terms.In this work, we generalize the Adam optimization algorithm to handle multiple loss terms.The guiding principle is that for every layer, the gradient magnitude of the terms should be balanced.To this end, the Multi-Term Adam (MTAdam) computes the derivative of each loss term separately, infers the first and second moments per parameter and loss term, and calculates a first moment for the magnitude per layer of the gradients arising from each loss.This magnitude is used to continuously balance the gradients across all layers, in a manner that both varies from one layer to the next and dynamically changes over time.Our results show that training with the new method leads to fast recovery from suboptimal initial loss weighting and to training outcomes that match or improve conventional training with the prescribed hyperparameters of each method. Itzik Malkiel, Lior Wolf |
EMNLP (1) | 2 |
| 2021 | Generating Master Faces for Dictionary Attacks with a Network-Assisted Latent Space EvolutionabstractA master face is a face image that passes face-based identity-authentication for a large portion of the population. These faces can be used to impersonate, with a high probability of success, any user, without having access to any user-information. We optimize these faces, by using an evolutionary algorithm in the latent embedding space of the StyleGAN face generator. Multiple evolutionary strategies are compared, and we propose a novel approach that employs a neural network in order to direct the search in the direction of promising samples, without adding fitness evaluations. The results we present demonstrate that it is possible to obtain a high coverage of the LFW identities (over 40%) with less than 10 master faces, for three leading deep face recognition systems. Ron Shmelkin, Tomer Friedlander, Lior Wolf |
FG | 3 |
| 2021 | Single Channel Voice Separation for Unknown Number of Speakers Under Reverberant and Noisy SettingsabstractWe present a unified network for voice separation of an unknown number of speakers. The proposed approach is composed of several separation heads optimized together with a speaker classification branch. The separation is carried out in the time domain, together with parameter sharing between all separation heads. The classification branch estimates the number of speakers while each head is specialized in separating a different number of speakers. We evaluate the proposed model under both clean and noisy reverberant settings. Results suggest that the proposed approach is superior to the baseline model by a significant margin. Additionally, we present a new noisy and reverberant dataset of up to five different speakers speaking simultaneously. Shlomo E. Chazan, Lior Wolf, Eliya Nachmani, Yossi Adi |
ICASSP | 2 |
| 2021 | High Fidelity Speech Regeneration with Application to Speech EnhancementabstractSpeech enhancement has seen great improvement in recent years mainly through contributions in denoising, speaker separation, and dereverberation methods that mostly deal with environmental effects on vocal audio. To enhance speech beyond the limitations of the original signal, we take a regeneration approach, in which we recreate the speech from its essence, including the semi-recognized speech, prosody features, and identity. We propose a wav-to-wav generative model for speech that can generate 24khz speech in a real-time manner and which utilizes a compact speech representation, composed of ASR and identity features, to achieve a higher level of intelligibility. Inspired by voice conversion methods, we train to augment the speech characteristics while preserving the identity of the source using an auxiliary identity network. Perceptual acoustic metrics and subjective tests show that the method obtains valuable improvements over recent baselines. Adam Polyak, Lior Wolf, Yossi Adi, Ori Kabeli, Yaniv Taigman |
ICASSP | 2 |
| 2021 | Adaptive Gradient Balancing for Undersampled MRI Reconstruction and Image-to-Image TranslationabstractRecent accelerated MRI reconstruction models have used Deep Neural Networks (DNNs) to reconstruct relatively high-quality images from highly undersampled k-space data, enabling much faster MRI scanning. However, these techniques sometimes struggle to reconstruct sharp images that preserve fine detail while maintaining a natural appearance. In this work, we enhance the image quality by using a Conditional Wasserstein Generative Adversarial Network combined with a novel Adaptive Gradient Balancing (AGB) technique that automates the process of combining the adversarial and pixel-wise terms and streamlines hyperparameter tuning. In addition, we introduce a Densely Connected Iterative Network, which is an undersampled MRI reconstruction network that utilizes dense connections. In MRI, our method minimizes artifacts, while maintaining a high-quality reconstruction that produces sharper images than other techniques. To demonstrate the general nature of our method, it is further evaluated on a battery of image-to-image translation experiments, demonstrating an ability to recover from sub-optimal weighting in multi-term adversarial training. Itzik Malkiel, Sangtae Ahn, Valentina Taviani, Anne Menini, Lior Wolf, Christopher J. Hardy |
ICCP | 5 |
| 2021 | Generic Attention-model Explainability for Interpreting Bi-Modal and Encoder-Decoder TransformersabstractTransformers are increasingly dominating multi-modal reasoning tasks, such as visual question answering, achieving state-of-the-art results thanks to their ability to contextualize information using the self-attention and co-attention mechanisms. These attention modules also play a role in other computer vision tasks including object detection and image segmentation. Unlike Transformers that only use self-attention, Transformers with co-attention require to consider multiple attention maps in parallel in order to highlight the information that is relevant to the prediction in the model’s input. In this work, we propose the first method to explain prediction by any Transformer-based architecture, including bi-modal Transformers and Transformers with co-attentions. We provide generic solutions and apply these to the three most commonly used of these architectures: (i) pure self-attention, (ii) self-attention combined with co-attention, and (iii) encoder-decoder attention. We show that our method is superior to all existing methods which are adapted from single modality explainability. Our code is available at: https://github.com/hila-chefer/Transformer-MM-Explainability. Hila Chefer, Shir Gur, Lior Wolf |
ICCV | 3 |
| 2021 | A Hierarchical Transformation-Discriminating Generative Model for Few Shot Anomaly DetectionabstractAnomaly detection, the task of identifying unusual samples in data, often relies on a large set of training samples. In this work, we consider the setting of few-shot anomaly detection in images, where only a few images are given at training. We devise a hierarchical generative model that captures the multi-scale patch distribution of each training image. We further enhance the representation of our model by using image transformations and optimize scale-specific patch-discriminators to distinguish between real and fake patches of the image, as well as between different transformations applied to those patches. The anomaly score is obtained by aggregating the patch-based votes of the correct transformation across scales and image regions. We demonstrate the superiority of our method on both the one-shot and few-shot settings, on the datasets of Paris, CIFAR10, MNIST and FashionMNIST as well as in the setting of defect detection on MVTec. In all cases, our method outperforms the recent baseline methods. Shelly Sheynin, Sagie Benaim, Lior Wolf |
ICCV | 3 |
| 2021 | Identity and Attribute Preserving Thumbnail UpscalingabstractWe consider the task of upscaling a low resolution thumbnail image of a person, to a higher resolution image, which preserves the person’s identity and other attributes. Since the thumbnail image is of low resolution, many higher resolution versions exist. Previous approaches produce solutions where the person’s identity is not preserved, or biased solutions, such as predominantly Caucasian faces. We address the existing ambiguity by first augmenting the feature extractor to better capture facial identity, facial attributes (such as smiling or not) and race, and second, use this feature extractor to generate high-resolution images which are identity preserving as well as conditioned on race and facial attributes. Our results indicate an improvement in face similarity recognition and lookalike generation as well as in the ability to generate higher resolution images which preserve an input thumbnail identity and whose race and attributes are maintained. Noam Gat, Sagie Benaim, Lior Wolf |
ICIP | 3 |
| 2021 | Scene Graph tO Image Generation with Contextualized Object Layout RefinementabstractGenerating images from scene graphs is a challenging task that attracted substantial interest recently. Prior works have approached this task by generating an intermediate layout description of the target image. However, the representation of each object in the layout was generated independently, which resulted in high overlap, low coverage, and an overall blurry layout. We propose a novel method that alleviates these issues by generating the entire layout description gradually to improve inter-object dependency. We empirically show on the COCO-STUFF dataset that our approach improves the quality of both the intermediate layout and the final image. Our approach improves the layout coverage by almost 20 points, and drops object overlap to negligible amounts. Our code is available at github.com/yanivbenny/COLoR. Maor Ivgi, Yaniv Benny, Avichai Ben-David, Jonathan Berant, Lior Wolf |
ICIP | 5 |
| 2021 | Natural Statistics Of Network Activations And Implications For Knowledge DistillationabstractIn a matter that is analogous to the study of natural image statistics, we study the natural statistics of the deep neural network activations at various layers. As we show, these statistics, similar to image statistics, follow a power law. We also show, both analytically and empirically, that with depth the exponent of this power law increases at a linear rate.As a direct implication of our discoveries, we present a method for performing Knowledge Distillation (KD). While classical KD methods consider the logits of the teacher network, more recent methods obtain a leap in performance by considering the activation maps. This, however, uses metrics that are suitable for comparing images. We propose to employ two additional loss terms that are based on the spectral properties of the intermediate activation maps. The proposed method obtains state of the art results on multiple image recognition KD benchmarks. Michael Rotman, Lior Wolf |
ICIP | 2 |
| 2021 | Fidelity-based Deep Adiabatic Scheduling
Eli Ovits, Lior Wolf |
ICLR | 2 |
| 2021 | HyperHyperNetwork for the Design of Antenna ArraysabstractWe present deep learning methods for the design of arrays and single instances of small antennas. Each design instance is conditioned on a target radiation pattern and is required to conform to specific spatial dimensions and to include, as part of its metallic structure, a set of predetermined locations. The solution, in the case of a single antenna, is based on a composite neural network that combines a simulation network, a hypernetwork, and a refinement network. In the design of the antenna array, we add an additional design level and employ a hypernetwork within a hypernetwork. The learning objective is based on measuring the similarity of the obtained radiation pattern to the desired one. Our experiments demonstrate that our approach is able to design novel antennas and antenna arrays that are compliant with the design requirements, considerably better than the baseline methods. We compare the solutions obtained by our method to existing designs and demonstrate a high level of overlap. When designing the antenna array of a cellular phone, the obtained solution displays improved properties over the existing one. Shahar Lutati, Lior Wolf |
ICML | 2 |
| 2021 | Recovering AES Keys with a Deep Cold Boot AttackabstractCold boot attacks inspect the corrupted random access memory soon after the power has been shut down. While most of the bits have been corrupted, many bits, at random locations, have not. Since the keys in many encryption schemes are being expanded in memory into longer keys with fixed redundancies, the keys can often be restored. In this work we combine a deep error correcting code technique together with a modified SAT solver scheme in order to apply the attack to AES keys. Even though AES consists Rijndael SBOX elements, that are specifically designed to be resistant to linear and differential cryptanalysis, our method provides a novel formalization of the AES key scheduling as a computational graph, which is implemented by neural message passing network. Our results show that our methods outperform the state of the art attack methods by a very large gap. Itamar Zimerman, Eliya Nachmani, Lior Wolf |
ICML | 3 |
| 2021 | Many-Speakers Single Channel Speech Separation with Optimal Permutation TrainingabstractSingle channel speech separation has experienced great progress in the last few years. However, training neural speech separation for a large number of speakers (e.g., more than 10 speakers) is out of reach for the current methods, which rely on the Permutation Invariant Loss (PIT). In this work, we present a permutation invariant training that employs the Hungarian algorithm in order to train with an $O(C^3)$ time complexity, where $C$ is the number of speakers, in comparison to $O(C!)$ of PIT based methods. Furthermore, we present a modified architecture that can handle the increased number of speakers. Our approach separates up to $20$ speakers and improves the previous results for large $C$ by a wide margin. Shaked Dovrat, Eliya Nachmani, Lior Wolf |
Interspeech | 3 |
| 2021 | Meta Internal LearningabstractInternal learning for single-image generation is a framework, where a generator is trained to produce novel images based on a single image. Since these models are trained on a single image, they are limited in their scale and application. To overcome these issues, we propose a meta-learning approach that enables training over a collection of images, in order to model the internal statistics of the sample image more effectively.In the presented meta-learning approach, a single-image GAN model is generated given an input image, via a convolutional feedforward hypernetwork $f$. This network is trained over a dataset of images, allowing for feature sharing among different models, and for interpolation in the space of generative models. The generated single-image model contains a hierarchy of multiple generators and discriminators. It is therefore required to train the meta-learner in an adversarial manner, which requires careful design choices that we justify by a theoretical analysis. Our results show that the models obtained are as suitable as single-image GANs for many common image applications, {significantly reduce the training time per image without loss in performance}, and introduce novel capabilities, such as interpolation and feedforward modeling of novel images. Raphael Bensadoun, Shir Gur, Tomer Galanti, Lior Wolf |
NeurIPS | 4 |
| 2021 | Data Augmenting Contrastive Learning of Speech Representations in the Time DomainabstractContrastive Predictive Coding (CPC), based on predicting future segments of speech from past segments is emerging as a powerful algorithm for representation learning of speech signal. However, it still under-performs compared to other methods on unsupervised evaluation benchmarks. Here, we intro-duce WavAugment, a time-domain data augmentation library which we adapt and optimize for the specificities of CPC (raw waveform input, contrastive loss, past versus future structure). We find that applying augmentation only to the segments from which the CPC prediction is performed yields better results than applying it also to future segments from which the samples (both positive and negative) of the contrastive loss are drawn. After selecting the best combination of pitch modification, additive noise and reverberation on unsupervised metrics on LibriSpeech (with a gain of 18-22% relative on the ABX score), we apply this combination without any change to three new datasets in the Zero Resource Speech Benchmark 2017 and beat the state-of-the-art using out-of-domain training data. Finally, we show that the data-augmented pretrained features improve a downstream phone recognition task in the Libri-light semi-supervised setting (10 min, 1 h or 10 h of labelled data) reducing the PER by 15% relative. Eugene Kharitonov, Morgane Rivière, Gabriel Synnaeve, Lior Wolf, Pierre-Emmanuel Mazaré, Matthijs Douze, Emmanuel Dupoux |
SLT | 4 |
| 2021 | On random kernels of residual architecturesabstractWe analyze the finite corrections to the neural tangent kernel (NTK) of residual and densely connected networks, as a function of both depth and width. Surprisingly, our analysis reveals that given a fixed depth, residual networks provide the best tradeoff between the parameter complexity and the coefficient of variation (normalized variance), followed by densely connected networks and vanilla MLPs. While in networks that do not use skip connections, convergence to the NTK requires one to fix the depth, while increasing the layers’ width. Our findings show that in ResNets, convergence to the NTK may occur when depth and width simultaneously tend to infinity, provided with a proper initialization. In DenseNets, however, the convergence of the NTK to its limit as the width tends to infinity is guaranteed, at a rate that is independent of both the depth and scale of the weights. Our experiments validate the theoretical results and demonstrate the advantage of deep ResNets and DenseNets for kernel regression with random gradient features. Etai Littwin, Tomer Galanti, Lior Wolf |
UAI | 3 |
| 2021 | Structural Analogy from a Single Image PairabstractAbstract The task of unsupervised image‐to‐image translation has seen substantial advancements in recent years through the use of deep neural networks. Typically, the proposed solutions learn the characterizing distribution of two large, unpaired collections of images, and are able to alter the appearance of a given image, while keeping its geometry intact. In this paper, we explore the capabilities of neural networks to understand image structure given only a single pair of images, and . We seek to generate images that are structurally aligned: that is, to generate an image that keeps the appearance and style of , but has a structural arrangement that corresponds to . The key idea is to map between image patches at different scales. This enables controlling the granularity at which analogies are produced, which determines the conceptual distinction between style and content. In addition to structural alignment, our method can be used to generate high quality imagery in other conditional generation tasks utilizing images and only: guided image synthesis, style and texture transfer, text translation as well as video translation. Our code and additional results are available in https://github.com/rmokady/structural-analogy/ Sagie Benaim, Ron Mokady, Amit Bermano, Lior Wolf |
Comput. Graph. Forum | 4 |
| 2021 | Evaluation Metrics for Conditional Image GenerationabstractAbstract We present two new metrics for evaluating generative models in the class-conditional image generation setting. These metrics are obtained by generalizing the two most popular unconditional metrics: the Inception Score (IS) and the Fréchet Inception Distance (FID). A theoretical analysis shows the motivation behind each proposed metric and links the novel metrics to their unconditional counterparts. The link takes the form of a product in the case of IS or an upper bound in the FID case. We provide an extensive empirical evaluation, comparing the metrics to their unconditional variants and to other metrics, and utilize them to analyze existing generative models, thus providing additional insights about their performance, from unlearned classes to mode collapse. Yaniv Benny, Tomer Galanti, Sagie Benaim, Lior Wolf |
Int. J. Comput. Vis. | 4 |
| 2021 | Risk Bounds for Unsupervised Cross-Domain Mapping with IPMsabstractThe recent empirical success of unsupervised cross-domain mapping algorithms, in mapping between two domains that share common characteristics, is not well-supported by theoretical justifications. This lacuna is especially troubling, given the clear ambiguity in such mappings. We work with adversarial training methods based on integral probability metrics (IPMs) and derive a novel risk bound, which upper bounds the risk between the learned mapping $h$ and the target mapping $y$, by a sum of three terms: (i) the risk between $h$ and the most distant alternative mapping that was learned by the same cross-domain mapping algorithm, (ii) the minimal discrepancy between the target domain and the domain obtained by applying a hypothesis $h^*$ on the samples of the source domain, where $h^*$ is a hypothesis selectable by the same algorithm, and (iii) an approximation error term that decreases as the capacity of the class of discriminators increases and is empirically shown to be small. The bound is directly related to Occam's razor and encourages the selection of the minimal architecture that supports a small mapping discrepancy. The bound leads to multiple algorithmic consequences, including a method for hyperparameter selection and early stopping in cross-domain mapping. Tomer Galanti, Sagie Benaim, Lior Wolf |
J. Mach. Learn. Res. | 3 |
| 2020 | Interactive Scene Generation via Scene Graphs with AttributesabstractWe introduce a simple yet expressive image generation method. On the one hand, it does not require the user to paint the masks or define a bounding box of the various objects, since the model does it by itself. On the other hand, it supports defining a coarse location and size of each object. Based on this, we offer a simple, interactive GUI, that allows a layman user to generate diverse images effortlessly.From a technical perspective, we introduce a dual embedding of layout and appearance. In this scheme, the location, size, and appearance of an object can change independently of each other. This way, the model is able to generate innumerable images per scene graph, to better express the intention of the user.In comparison to previous work, we also offer better quality and higher resolution outputs. This is due to a superior architecture, which is based on a novel set of discriminators. Those discriminators better constrain the shape of the generated mask, as well as capturing the appearance encoding in a counterfactual way.Our code is publicly available at https://www.github.com/ashual/scene_generation. Oron Ashual, Lior Wolf |
AAAI | 2 |
| 2020 | Relative Attributing Propagation: Interpreting the Comparative Contributions of Individual Units in Deep Neural NetworksabstractAs Deep Neural Networks (DNNs) have demonstrated superhuman performance in a variety of fields, there is an increasing interest in understanding the complex internal mechanisms of DNNs. In this paper, we propose Relative Attributing Propagation (RAP), which decomposes the output predictions of DNNs with a new perspective of separating the relevant (positive) and irrelevant (negative) attributions according to the relative influence between the layers. The relevance of each neuron is identified with respect to its degree of contribution, separated into positive and negative, while preserving the conservation rule. Considering the relevance assigned to neurons in terms of relative priority, RAP allows each neuron to be assigned with a bi-polar importance score concerning the output: from highly relevant to highly irrelevant. Therefore, our method makes it possible to interpret DNNs with much clearer and attentive visualizations of the separated attributions than the conventional explaining methods. To verify that the attributions propagated by RAP correctly account for each meaning, we utilize the evaluation metrics: (i) Outside-inside relevance ratio, (ii) Segmentation mIOU and (iii) Region perturbation. In all experiments and metrics, we present a sizable gap in comparison to the existing literature. Woo-Jeoung Nam, Shir Gur, Jaesik Choi, Lior Wolf, Seong-Whan Lee |
AAAI | 4 |
| 2020 | Meta Decision Trees for Explainable Recommendation SystemsabstractWe tackle the problem of building explainable recommendation systems that are based on a per-user decision tree, with decision rules that are based on single attribute values. We build the trees by applying learned regression functions to obtain the decision rules as well as the values at the leaf nodes. The regression functions receive as input the embedding of the user's training set, as well as the embedding of the samples that arrive at the current node. The embedding and the regressors are learned end-to-end with a loss that encourages the decision rules to be sparse. By applying our method, we obtain a collaborative filtering solution that provides a direct explanation to every rating it provides. With regards to accuracy, it is competitive with other algorithms. However, as expected, explainability comes at a cost and the accuracy is typically slightly lower than the state of the art result reported in the literature. Our code is available at \urlhttps://github.com/shulmaneyal/metatrees. Eyal Shulman, Lior Wolf |
AIES | 2 |
| 2020 | ScopeFlow: Dynamic Scene Scoping for Optical FlowabstractWe propose to modify the common training protocols of optical flow, leading to sizable accuracy improvements without adding to the computational complexity of the training process. The improvement is based on observing the bias in sampling challenging data that exists in the current training protocol, and improving the sampling process. In addition, we find that both regularization and augmentation should decrease during the training protocol. Using an existing low parameters architecture, the method is ranked first on the MPI Sintel benchmark among all other methods, improving the best two frames method accuracy by more than 10%. The method also surpasses all similar architecture variants by more than 12% and 19.7% on the KITTI benchmarks, achieving the lowest Average End-Point Error on KITTI2012 among two-frame methods, without using extra datasets. Aviram Bar-Haim, Lior Wolf |
CVPR | 2 |
| 2020 | Wish You Were Here: Context-Aware Human Generation
Oran Gafni, Lior Wolf |
CVPR | 2 |
| 2020 | OneGAN: Simultaneous Unsupervised Learning of Conditional Image Generation, Foreground Segmentation, and Fine-Grained Clustering
Yaniv Benny, Lior Wolf |
ECCV (26) | 2 |
| 2020 | A Gated Hypernet Decoder for Polar CodesabstractHypernetworks were recently shown to improve the performance of message passing algorithms for decoding error correcting codes. In this work, we demonstrate how hypernet-works can be applied to decode polar codes by employing a new formalization of the polar belief propagation decoding scheme. We demonstrate that our method improves the previous results of neural polar decoders and achieves, for large SNRs, the same bit-error-rate performances as the successive list cancellation method, which is known to be better than any belief propagation decoders and very close to the maximum likelihood decoder. Eliya Nachmani, Lior Wolf |
ICASSP | 2 |
| 2020 | Electric Analog Circuit Design with Hypernetworks And A Differential SimulatorabstractThe manual design of analog circuits is a tedious task of parameter tuning that requires hours of work by human experts. In this work, we make a significant step towards a fully automatic design method that is based on deep learning. The method selects the components and their configuration, as well as their numerical parameters. By contrast, the current literature methods are limited to the parameter fitting part only. A two-stage network is used, which first generates a chain of circuit components and then predicts their parameters. A hypernetwork scheme is used in which a weight generating network, which is conditioned on the circuit's power spectrum, produces the parameters of a primal RNN network that places the components. A differential simulator is used for refining the numerical values of the components. We show that our model provides an efficient design solution, and is superior to alternative solutions. Michael Rotman, Lior Wolf |
ICASSP | 2 |
| 2020 | Comparing Vision-based to Sonar-based 3D ReconstructionabstractOur understanding of sonar based sensing is very limited in comparison to light based imaging. In this work, we synthesize a ShapeNet variant in which echolocation replaces the role of vision. A new hypernetwork method is presented for 3D reconstruction from a single echolocation view. The success of the method demonstrates the ability to reconstruct a 3D shape from bat-like sonar, and not just obtain the relative position of the bat with respect to obstacles. In addition, it is shown that integrating information from multiple orientations around the same view point helps performance. The sonar-based method we develop is analog to the state-of-the-art single image reconstruction method, which allows us to directly compare the two imaging modalities. Based on this analysis, we learn that while 3D can be reliably reconstructed form sonar, as far as the current technology shows, the accuracy is lower than the one obtained based on vision, that the performance in sonar and in vision are highly correlated, that both modalities favor shapes that are not round, and that while the current vision method is able to better reconstruct the 3D shape, its advantage with respect to estimating the normal's direction is much lower. Netanel Frank, Lior Wolf, Danny Olshansky, Arjan Boonman, Yossi Yovel |
ICCP | 2 |
| 2020 | Vid2Game: Controllable Characters Extracted from Real-World Videos
Oran Gafni, Lior Wolf, Yaniv Taigman |
ICLR | 2 |
| 2020 | End to End Trainable Active Contours via Differentiable Rendering
Shir Gur, Tal Shaharabany, Lior Wolf |
ICLR | 3 |
| 2020 | Masked Based Unsupervised Content Transfer
Ron Mokady, Sagie Benaim, Lior Wolf, Amit Bermano |
ICLR | 3 |
| 2020 | Voice Separation with an Unknown Number of Multiple SpeakersabstractWe present a new method for separating a mixed audio sequence, in which multiple voices speak simultaneously. The new method employs gated neural networks that are trained to separate the voices at multiple processing steps, while maintaining the speaker in each output channel fixed. A different model is trained for every number of possible speakers, and the model with the largest number of speakers is employed to select the actual number of speakers in a given sample. Our method greatly outperforms the current state of the art, which, as we show, is not competitive for more than two speakers. Eliya Nachmani, Yossi Adi, Lior Wolf |
ICML | 3 |
| 2020 | Unsupervised Cross-Domain Singing Voice ConversionabstractWe present a wav-to-wav generative model for the task of singing voice conversion from any identity. Our method utilizes both an acoustic model, trained for the task of automatic speech recognition, together with melody extracted features to drive a waveform-based generator. The proposed generative architecture is invariant to the speaker's identity and can be trained to generate target singers from unlabeled training data, using either speech or singing sources. The model is optimized in an end-to-end fashion without any manual supervision, such as lyrics, musical notes or parallel samples. The proposed approach is fully-convolutional and can generate audio in real-time. Experiments show that our method significantly outperforms the baseline methods while generating convincingly better audio samples than alternative attempts. Adam Polyak, Lior Wolf, Yossi Adi, Yaniv Taigman |
INTERSPEECH | 2 |
| 2020 | TTS Skins: Speaker Conversion via ASRabstractWe present a fully convolutional wav-to-wav network for converting between speakers' voices, without relying on text. Our network is based on an encoder-decoder architecture, where the encoder is pre-trained for the task of Automatic Speech Recognition, and a multi-speaker waveform decoder is trained to reconstruct the original signal in an autoregressive manner. We train the network on narrated audiobooks, and demonstrate multi-voice TTS in those voices, by converting the voice of a TTS robot. Adam Polyak, Lior Wolf, Yaniv Taigman |
INTERSPEECH | 2 |
| 2020 | Learning Personal Representations from fMRI by Predicting Neurofeedback Performance
Jhonathan Osin, Lior Wolf, Guy Gurevitch, Nimrod Jakob Keynan, Tom Fruchtman-Steinbok, Ayelet Or-Borichov, Talma Hendler |
MICCAI (7) | 2 |
| 2020 | On the Modularity of HypernetworksabstractIn the context of learning to map an input $I$ to a function $h_I:\mathcal{X}\to \mathbb{R}$, two alternative methods are compared: (i) an embedding-based method, which learns a fixed function in which $I$ is encoded as a conditioning signal $e(I)$ and the learned function takes the form $h_I(x) = q(x,e(I))$, and (ii) hypernetworks, in which the weights $\theta_I$ of the function $h_I(x) = g(x;\theta_I)$ are given by a hypernetwork $f$ as $\theta_I=f(I)$. In this paper, we define the property of modularity as the ability to effectively learn a different function for each input instance $I$. For this purpose, we adopt an expressivity perspective of this property and extend the theory of~\cite{devore} and provide a lower bound on the complexity (number of trainable parameters) of neural networks as function approximators, by eliminating the requirements for the approximation method to be robust. Our results are then used to compare the complexities of $q$ and $g$, showing that under certain conditions and when letting the functions $e$ and $f$ be as large as we wish, $g$ can be smaller than $q$ by orders of magnitude. This sheds light on the modularity of hypernetworks in comparison with the embedding-based method. Besides, we show that for a structured target function, the overall number of trainable parameters in a hypernetwork is smaller by orders of magnitude than the number of trainable parameters of a standard neural network and an embedding method. Tomer Galanti, Lior Wolf |
NeurIPS | 2 |
| 2020 | Hierarchical Patch VAE-GAN: Generating Diverse Videos from a Single SampleabstractWe consider the task of generating diverse and novel videos from a single video sample. Recently, new hierarchical patch-GAN based approaches were proposed for generating diverse images, given only a single sample at training time. Moving to videos, these approaches fail to generate diverse samples, and often collapse into generating samples similar to the training video. We introduce a novel patch-based variational autoencoder (VAE) which allows for a much greater diversity in generation. Using this tool, a new hierarchical video generation scheme is constructed: at coarse scales, our patch-VAE is employed, ensuring samples are of high diversity. Subsequently, at finer scales, a patch-GAN renders the fine details, resulting in high quality videos. Our experiments show that the proposed method produces diverse samples in both the image domain, and the more challenging video domain. Our code and supplementary material (SM) with additional samples are available at https://shirgur.github.io/hp-vae-gan Shir Gur, Sagie Benaim, Lior Wolf |
NeurIPS | 3 |
| 2020 | On Infinite-Width Hypernetworksabstract{\em Hypernetworks} are architectures that produce the weights of a task-specific {\em primary network}. A notable application of hypernetworks in the recent literature involves learning to output functional representations. In these scenarios, the hypernetwork learns a representation corresponding to the weights of a shallow MLP, which typically encodes shape or image information. While such representations have seen considerable success in practice, they remain lacking in the theoretical guarantees in the wide regime of the standard architectures. In this work, we study wide over-parameterized hypernetworks. We show that unlike typical architectures, infinitely wide hypernetworks do not guarantee convergence to a global minima under gradient descent. We further show that convexity can be achieved by increasing the dimensionality of the hypernetwork's output, to represent wide MLPs. In the dually infinite-width regime, we identify the functional priors of these architectures by deriving their corresponding GP and NTK kernels, the latter of which we refer to as the {\em hyperkernel}. As part of this study, we make a mathematical contribution by deriving tight bounds on high order Taylor expansion terms of standard fully connected ReLU networks. Etai Littwin, Tomer Galanti, Lior Wolf, Greg Yang |
NeurIPS | 3 |
| 2020 | Generating Correct Answers for Progressive Matrices Intelligence TestsabstractRaven’s Progressive Matrices are multiple-choice intelligence tests, where one tries to complete the missing location in a 3x3 grid of abstract images. Previous attempts to address this test have focused solely on selecting the right answer out of the multiple choices. In this work, we focus, instead, on generating a correct answer given the grid, which is a harder task, by definition. The proposed neural model combines multiple advances in generative models, including employing multiple pathways through the same network, using the reparameterization trick along two pathways to make their encoding compatible, a selective application of variational losses, and a complex perceptual loss that is coupled with a selective backpropagation procedure. Our algorithm is able not only to generate a set of plausible answers but also to be competitive to the state of the art methods in multiple-choice tests. Niv Pekar, Yaniv Benny, Lior Wolf |
NeurIPS | 3 |
| 2020 | Supervised and Unsupervised Learning of Parameterized Color EnhancementabstractWe treat the problem of color enhancement as an image translation task, which we tackle using both supervised and unsupervised learning. Unlike traditional image to image generators, our translation is performed using a global parameterized color transformation instead of learning to directly map image information. In the supervised case, every training image is paired with a desired target image and a convolutional neural network (CNN) learns from the expert retouched images the parameters of the transformation. In the unpaired case, we employ two-way generative adversarial networks (GANs) to learn these parameters and apply a circularity constraint. We achieve state-of-the-art results compared to both supervised (paired data) and unsupervised (unpaired data) image enhancement methods on the MIT-Adobe FiveK benchmark. Moreover, we show the generalization capability of our method, by applying it on photos from the early 20th century and to dark video frames. Yoav Chai, Raja Giryes, Lior Wolf |
WACV | 3 |
| 2020 | End to End Lip Synchronization with a Temporal AutoEncoderabstractWe study the problem of syncing the lip movement in a video with the audio stream. Our solution finds an optimal alignment using a dual-domain recurrent neural network that is trained on synthetic data we generate by dropping and duplicating video frames. Once the alignment is found, we modify the video in order to sync the two sources. Our method is shown to greatly outperform the literature methods on a variety of existing and new benchmarks. As an application, we demonstrate our ability to robustly align text-to-speech generated audio with an existing video stream. Our code is attached as supplementary. Yoav Shalev, Lior Wolf |
WACV | 2 |
| 2020 | Unsupervised Generation of Free-Form and Parameterized AvatarsabstractWe study two problems involving the task of mapping images between different domains. The first problem, transfers an image in one domain to an analog image in another domain. The second problem, extends the previous one by mapping an input image to a tied pair, consisting of a vector of parameters and an image that is created using a graphical engine from this vector of parameters. Similar to the first problem, the mapping's objective is to have the output image as similar as possible to the input image. In both cases, no supervision is given during training in the form of matching inputs and outputs. We compare the two unsupervised learning problems to the problem of unsupervised domain adaptation, define generalization bounds that are based on discrepancy, and employ a GAN to implement network solutions that correspond to these bounds. Experimentally, our methods are shown to solve the problem of automatically creating avatars. Adam Polyak, Yaniv Taigman, Lior Wolf |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2019 | A Formal Approach to ExplainabilityabstractWe regard explanations as a blending of the input sample and the model's output and offer a few definitions that capture various desired properties of the function that generates these explanations. We study the links between these properties and between explanation-generating functions and intermediate representations of learned models and are able to show, for example, that if the activations of a given layer are consistent with an explanation, then so do all other subsequent layers. In addition, we study the intersection and union of explanations as a way to construct new explanations. Lior Wolf, Tomer Galanti, Tamir Hazan |
AIES | 1 |
| 2019 | Single Image Depth Estimation Trained via Depth From Defocus CuesabstractEstimating depth from a single RGB images is a fundamental task in computer vision, which is most directly solved using supervised deep learning. In the field of unsupervised learning of depth from a single RGB image, depth is not given explicitly. Existing work in the field receives either a stereo pair, a monocular video, or multiple views, and, using losses that are based on structure-from-motion, trains a depth estimation network. In this work, we rely, instead of different views, on depth from focus cues. Learning is based on a novel Point Spread Function convolutional layer, which applies location specific kernels that arise from the Circle-Of-Confusion in each image location. We evaluate our method on data derived from five common datasets for depth estimation and lightfield images, and present results that are on par with supervised methods on KITTI and Make3D datasets and outperform unsupervised learning approaches. Since the phenomenon of depth from defocus is not dataset specific, we hypothesize that learning based on it would overfit less to the specific content in each dataset. Our experiments show that this is indeed the case, and an estimator learned on one dataset using our method provides better results on other datasets, than the directly supervised methods. Shir Gur, Lior Wolf |
CVPR | 2 |
| 2019 | End-To-End Supervised Product Quantization for Image Search and RetrievalabstractProduct Quantization, a dictionary based hashing method, is one of the leading unsupervised hashing techniques. While it ignores the labels, it harnesses the features to construct look up tables that can approximate the feature space. In recent years, several works have achieved state of the art results on hashing benchmarks by learning binary representations in a supervised manner. This work presents Deep Product Quantization (DPQ), a technique that leads to more accurate retrieval and classification than the latest state of the art methods, while having similar computational complexity and memory footprint as the Product Quantization method. To our knowledge, this is the first work to introduce a dictionary-based representation that is inspired by Product Quantization and which is learned end-to-end, and thus benefits from the supervised signal. DPQ explicitly learns soft and hard representations to enable an efficient and accurate asymmetric search, by using a straight-through estimator. Our method obtains state of the art results on an extensive array of retrieval and classification experiments. Benjamin Eliot Klein, Lior Wolf |
CVPR | 2 |
| 2019 | Semi-supervised Monaural Singing Voice Separation with a Masking Network Trained on Synthetic MixturesabstractWe study the problem of semi-supervised singing voice separation, in which the training data contains a set of samples of mixed music (singing and instrumental) and an unmatched set of instrumental music. Our solution employs a single mapping function g, which, applied to a mixed sample, recovers the underlying instrumental music, and, applied to an instrumental sample, returns the same sample. The network g is trained using purely instrumental samples, as well as on synthetic mixed samples that are created by mixing reconstructed singing voices with random instrumental samples. Our results indicate that we are on a par with or better than fully supervised methods, which are also provided with training samples of unmixed singing voices, and are better than other recent semi-supervised methods. Michael Michelashvili, Sagie Benaim, Lior Wolf |
ICASSP | 3 |
| 2019 | Unsupervised Polyglot Text-to-speechabstractWe present a TTS neural network that is able to produce speech in multiple languages. The proposed network is able to transfer a voice, which was presented as a sample in a source language, into one of several target languages. Training is done without using matching or parallel data, i.e., without samples of the same speaker in multiple languages, making the method much more applicable. The conversion is based on learning a polyglot network that has multiple per-language sub-networks and adding loss terms that preserve the speaker's identity in multiple languages. We evaluate the proposed polyglot neural network for three languages with a total of more than 400 speakers and demonstrate convincing conversion capabilities. Eliya Nachmani, Lior Wolf |
ICASSP | 2 |
| 2019 | Attention-based Wavenet Autoencoder for Universal Voice ConversionabstractWe present a method for converting any voice to a target voice. The method is based on a WaveNet autoencoder, with the addition of a novel attention component that supports the modification of timing between the input and the output samples. Training the attention is done in an unsupervised way, by teaching the neural network to recover the original timing from an artificially modified one. Adding a generic voice robot, which we convert to the target voice, we present a robust Text To Speech pipeline that is able to train without any transcript. Our experiments show that the proposed method is able to recover the timing of the speaker and that the proposed pipeline provides a competitive Text To Speech method. Adam Polyak, Lior Wolf |
ICASSP | 2 |
| 2019 | Specifying Object Attributes and Relations in Interactive Scene GenerationabstractWe introduce a method for the generation of images from an input scene graph. The method separates between a layout embedding and an appearance embedding. The dual embedding leads to generated images that better match the scene graph, have higher visual quality, and support more complex scene graphs. In addition, the embedding scheme supports multiple and diverse output images per scene graph, which can be further controlled by the user. We demonstrate two modes of per-object control: (i) importing elements from other images, and (ii) navigation in the object space by selecting an appearance archetype. Our code is publicly available at https://www.github.com/ashual/scene_generation. Oron Ashual, Lior Wolf |
ICCV | 2 |
| 2019 | Domain Intersection and Domain DifferenceabstractWe present a method for recovering the shared content between two visual domains as well as the content that is unique to each domain. This allows us to map from one domain to the other, in a way in which the content that is specific for the first domain is removed and the content that is specific for the second is imported from any image in the second domain. In addition, our method enables generation of images from the intersection of the two domains as well as their union, despite having no such samples during training. The method is shown analytically to contain all the sufficient and necessary constraints. It also outperforms the literature methods in an extensive set of experiments. Sagie Benaim, Michael Khaitov, Tomer Galanti, Lior Wolf |
ICCV | 4 |
| 2019 | Bidirectional One-Shot Unsupervised Domain MappingabstractWe study the problem of mapping between a domain A, in which there is a single training sample and a domain B, for which we have a richer training set. The method we present is able to perform this mapping in both directions. For example, we can transfer all MNIST images to the visual domain captured by a single SVHN image and transform the SVHN image to the domain of the MNIST images. Our method is based on employing one encoder and one decoder for each domain, without utilizing weight sharing. The autoencoder of the single sample domain is trained to match both this sample and the latent space of domain B. Our results demonstrate convincing mapping between domains, where either the source or the target domain are defined by a single sample, far surpassing existing solutions. Our code is made publicly available at https://github.com/tomercohen11/BiOST. Tomer Cohen, Lior Wolf |
ICCV | 2 |
| 2019 | Live Face De-Identification in VideoabstractWe propose a method for face de-identification that enables fully automatic video modification at high frame rates. The goal is to maximally decorrelate the identity, while having the perception (pose, illumination and expression) fixed. We achieve this by a novel feed-forward encoder-decoder network architecture that is conditioned on the high-level representation of a person's facial image. The network is global, in the sense that it does not need to be retrained for a given video or for a given identity, and it creates natural looking image sequences with little distortion in time. Oran Gafni, Lior Wolf, Yaniv Taigman |
ICCV | 2 |
| 2019 | Unsupervised Microvascular Image Segmentation Using an Active Contours Mimicking Neural NetworkabstractThe task of blood vessel segmentation in microscopy images is crucial for many diagnostic and research applications. However, vessels can look vastly different, depending on the transient imaging conditions, and collecting data for supervised training is laborious. We present a novel deep learning method for unsupervised segmentation of blood vessels. The method is inspired by the field of active contours and we introduce a new loss term, which is based on the morphological Active Contours Without Edges (ACWE) optimization method. The role of the morphological operators is played by novel pooling layers that are incorporated to the network's architecture. We demonstrate the challenges that are faced by previous supervised learning solutions, when the imaging conditions shift. Our unsupervised method is able to outperform such previous methods in both the labeled dataset, and when applied to similar but different datasets. Our code, as well as efficient pytorch reimplementations of the baseline methods VesselNN and DeepVess are attached as supplementary. Shir Gur, Lior Wolf, Lior Golgher, Pablo Blinder |
ICCV | 2 |
| 2019 | Deep Meta Functionals for Shape RepresentationabstractWe present a new method for 3D shape reconstruction from a single image, in which a deep neural network directly maps an image to a vector of network weights. The network parametrized by these weights represents a 3D shape by classifying every point in the volume as either within or outside the shape. The new representation has virtually unlimited capacity and resolution, and can have an arbitrary topology. Our experiments show that it leads to more accurate shape inference from a 2D projection than the existing methods, including voxel-, silhouette-, and mesh-based methods. The code will be available at: https: //github.com/gidilittwin/Deep-Meta. Gidi Littwin, Lior Wolf |
ICCV | 2 |
| 2019 | Transductive Learning for Reading Handwritten Tibetan ManuscriptsabstractWe examine the use case of performing handwritten character recognition (HCR) on a newly compiled collection of Tibetan historical documents, which presents multiple challenges, including inherent challenges such as image quality and the lack of word separation, and dataset challenges such as a lack of supervised training data. To tackle these challenges, we introduce an end-to-end unsupervised full-document HCR approach composed of unsupervised line segmentation and a convolutional recurrent neural network, trained using solely synthetic data. Various augmentations are applied to these synthesized images, and we compare the effect of each augmentation on the HCR results. Since we work on a collection of historical manuscripts, we can fit the model to the available test data. During training, our network has access to both the labeled synthetic training data and the unlabeled images of the test set, and we adapt and evaluate four different semi-supervised learning and domain adaptation approaches for transductive learning in HCR. We test our approach on a set of 167 images from the "Kadam" collection, containing 829 lines. We show that correct data augmentation is crucial for the success of HCR trained solely on synthetic data and that using an effective transductive learning approach drastically improves results. Sivan Keret, Lior Wolf, Nachum Dershowitz, Eric Werner, Orna Almogi, Dorji Wangchuk |
ICDAR | 2 |
| 2019 | A Universal Music Translation Network
Noam Mor, Lior Wolf, Adam Polyak, Yaniv Taigman |
ICLR (Poster) | 2 |
| 2019 | Emerging Disentanglement in Auto-Encoder Based Unsupervised Image Content Transfer
Ori Press, Tomer Galanti, Sagie Benaim, Lior Wolf |
ICLR (Poster) | 4 |
| 2019 | Unsupervised Learning of the Set of Local Maxima
Lior Wolf, Sagie Benaim, Tomer Galanti |
ICLR (Poster) | 1 |
| 2019 | Unsupervised Singing Voice ConversionabstractWe present a deep learning method for singing voice conversion. The proposed network is not conditioned on the text or on the notes, and it directly converts the audio of one singer to the voice of another. Training is performed without any form of supervision: no lyrics or any kind of phonetic features, no notes, and no matching samples between singers. The proposed network employs a single CNN encoder for all singers, a single WaveNet decoder, and a classifier that enforces the latent representation to be singer-agnostic. Each singer is represented by one embedding vector, which the decoder is conditioned on. In order to deal with relatively small datasets, we propose a new data augmentation scheme, as well as new training losses and protocols that are based on backtranslation. Our evaluation presents evidence that the conversion produces natural signing voices that are highly recognizable as the target singer. Eliya Nachmani, Lior Wolf |
INTERSPEECH | 2 |
| 2019 | Hyper-Graph-Network Decoders for Block CodesabstractNeural decoders were shown to outperform classical message passing techniques for short BCH codes. In this work, we extend these results to much larger families of algebraic block codes, by performing message passing with graph neural networks. The parameters of the sub-network at each variable-node in the Tanner graph are obtained from a hypernetwork that receives the absolute values of the current message as input. To add stability, we employ a simplified version of the arctanh activation that is based on a high order Taylor approximation of this activation function. Our results show that for a large number of algebraic block codes, from diverse families of codes (BCH, LDPC, Polar), the decoding obtained with our method outperforms the vanilla belief propagation method as well as other learning techniques from the literature. Eliya Nachmani, Lior Wolf |
NeurIPS | 2 |
| 2018 | A Two-Step Disentanglement MethodabstractWe address the problem of disentanglement of factors that generate a given data into those that are correlated with the labeling and those that are not. Our solution is simpler than previous solutions and employs adversarial training. First, the part of the data that is correlated with the labels is extracted by training a classifier. Then, the other part is extracted such that it enables the reconstruction of the original data but does not contain label information. The utility of the new method is demonstrated on visual datasets as well as on financial data. Our code is available at https://github.com/naamahadad/A-Two-Step-Disentanglement-Method. Naama Hadad, Lior Wolf, Shimon Shahar |
CVPR | 2 |
| 2018 | Unsupervised Correlation AnalysisabstractLinking between two data sources is a basic building block in numerous computer vision problems. In this paper, we set to answer a fundamental cognitive question: are prior correspondences necessary for linking between different domains? One of the most popular methods for linking between domains is Canonical Correlation Analysis (CCA). All current CCA algorithms require correspondences between the views. We introduce a new method Unsupervised Correlation Analysis (UCA), which requires no prior correspondences between the two domains. The correlation maximization term in CCA is replaced by a combination of a reconstruction term (similar to autoencoders), full cycle loss, orthogonality and multiple domain confusion terms. Due to lack of supervision, the optimization leads to multiple alternative solutions with similar scores and we therefore introduce a consensus-based mechanism that is often able to recover the desired solution. Remarkably, this suffices in order to link remote domains such as text and images. We also present results on well accepted CCA benchmarks, showing that performance far exceeds other unsupervised baselines, and approaches supervised performance in some cases. Yedid Hoshen, Lior Wolf |
CVPR | 2 |
| 2018 | Estimating the Success of Unsupervised Image to Image Translation
Sagie Benaim, Tomer Galanti, Lior Wolf |
ECCV (5) | 3 |
| 2018 | NAM: Non-Adversarial Unsupervised Domain Mapping
Yedid Hoshen, Lior Wolf |
ECCV (14) | 2 |
| 2018 | Non-Adversarial Unsupervised Word TranslationabstractUnsupervised word translation from nonparallel inter-lingual corpora has attracted much research interest.Very recently, neural network methods trained with adversarial loss functions achieved high accuracy on this task.Despite the impressive success of the recent techniques, they suffer from the typical drawbacks of generative adversarial models: sensitivity to hyper-parameters, long training time and lack of interpretability.In this paper, we make the observation that two sufficiently similar distributions can be aligned correctly with iterative matching methods.We present a novel method that first aligns the second moment of the word distributions of the two languages and then iteratively refines the alignment.Extensive experiments on word translation of European and Non-European languages show that our method achieves better performance than recent state-of-the-art deep adversarial approaches and is competitive with the supervised baseline.It is also efficient, easy to parallelize on CPU and interpretable. Yedid Hoshen, Lior Wolf |
EMNLP | 2 |
| 2018 | Deep learning for the design of nano-photonic structuresabstractOur visual perception of our surroundings is ultimately limited by the diffraction-limit, which stipulates that optical information smaller than roughly half the illumination wavelength is not retrievable. Over the past decades, many breakthroughs have led to unprecedented imaging capabilities beyond the diffraction-limit, with applications in biology and nanotechnology. In this context, nano-photonics has had a profound impact on the field of optics by enabling the manipulation of light-matter interaction with subwave-length structures [1, 2, 3]. However, despite the many advances in this field, its impact and penetration in our daily life has been hindered by a convoluted and iterative process, cycling through modeling, nanofabrication and nano-characterization. The fundamental reason is the fact that not only the prediction of the optical response is very time consuming and requires solving Maxwell's equations with dedicated numerical packages [4, 5, 6]. But, more significantly, the inverse problem, i.e. designing a nanostructure with an on-demand optical response, is currently a prohibitive task even with the most advanced numerical tools due to the high non-linearity of the problem [7, 8]. Here, we harness the power of Deep Learning and show its ability to predict the geometry of nanostructures based solely on their far-field response. This approach addresses in a direct way the currently inaccessible inverse problem breaking the ground for on-demand design of optical response with applications such as sensing, imaging and also for Plasmons mediated cancer thermotherapy. Itzik Malkiel, Michael Mrejen, Achiya Nagler, Uri Arieli, Lior Wolf, Haim Suchowski |
ICCP | 5 |
| 2018 | Toward a Dataset-Agnostic Word Segmentation MethodabstractWord segmentation in documents is a critical stage towards word and character recognition, as well as word spotting. Despite recent advancements in word segmentation and object detection, detecting instances of words in a cluttered handwritten document remains a non-trivial task that requires a large amount of labeled documents for training. We present a flexible and general framework for word segmentation in handwritten documents, which incorporates techniques from the recent object detection literature as well as document analysis tools. Our method utilizes information that is relevant for word segmentation and ignores other highly variable information contained in a handwritten text, thus allowing for efficient transfer learning between datasets and alleviating the need for labeled training data. Our approach efficiently detects words in a variety of scanned document images, including historical handwritten documents and modern day handwritten documents, presenting excellent results on existing benchmarks. In addition, we demonstrate the usefulness of our approach by achieving state-of-the-art results for segmentation-free word spotting tasks. Gregory Axler, Lior Wolf |
ICIP | 2 |
| 2018 | The Role of Minimal Complexity Functions in Unsupervised Learning of Semantic Mappings
Tomer Galanti, Lior Wolf, Sagie Benaim |
ICLR (Poster) | 2 |
| 2018 | Identifying Analogies Across Domains
Yedid Hoshen, Lior Wolf |
ICLR (Poster) | 2 |
| 2018 | VoiceLoop: Voice Fitting and Synthesis via a Phonological Loop
Yaniv Taigman, Lior Wolf, Adam Polyak, Eliya Nachmani |
ICLR (Poster) | 2 |
| 2018 | Fitting New Speakers Based on a Short Untranscribed SampleabstractLearning-based Text To Speech systems have the potential to generalize from one speaker to the next and thus require a relatively short sample of any new voice. However, this promise is currently largely unrealized. We present a method that is designed to capture a new speaker from a short untranscribed audio sample. This is done by employing an additional network that given an audio sample, places the speaker in the embedding space. This network is trained as part of the speech synthesis system using various consistency losses. Our results demonstrate a greatly improved performance on both the dataset speakers, and, more importantly, when fitting new voices, even from very short samples. Eliya Nachmani, Adam Polyak, Yaniv Taigman, Lior Wolf |
ICML | 4 |
| 2018 | One-Shot Unsupervised Cross Domain TranslationabstractGiven a single image $x$ from domain $A$ and a set of images from domain $B$, our task is to generate the analogous of $x$ in $B$. We argue that this task could be a key AI capability that underlines the ability of cognitive agents to act in the world and present empirical evidence that the existing unsupervised domain translation methods fail on this task. Our method follows a two step process. First, a variational autoencoder for domain $B$ is trained. Then, given the new sample $x$, we create a variational autoencoder for domain $A$ by adapting the layers that are close to the image in order to directly fit $x$, and only indirectly adapt the other layers. Our experiments indicate that the new method does as well, when trained on one sample $x$, as the existing domain transfer methods, when these enjoy a multitude of training samples from domain $A$. Our code is made publicly available at https://github.com/sagiebenaim/OneShotTranslation Sagie Benaim, Lior Wolf |
NeurIPS | 2 |
| 2018 | Regularizing by the Variance of the Activations' Sample-VariancesabstractNormalization techniques play an important role in supporting efficient and often more effective training of deep neural networks. While conventional methods explicitly normalize the activations, we suggest to add a loss term instead. This new loss term encourages the variance of the activations to be stable and not vary from one random mini-batch to the next. As we prove, this encourages the activations to be distributed around a few distinct modes. We also show that if the inputs are from a mixture of two Gaussians, the new loss would either join the two together, or separate between them optimally in the LDA sense, depending on the prior probabilities. Finally, we are able to link the new regularization term to the batchnorm method, which provides it with a regularization perspective. Our experiments demonstrate an improvement in accuracy over the batchnorm technique for both CNNs and fully connected networks. Etai Littwin, Lior Wolf |
NeurIPS | 2 |
| 2018 | Automatic Program Synthesis of Long Programs with a Learned Garbage CollectorabstractWe consider the problem of generating automatic code given sample input-output pairs. We train a neural network to map from the current state and the outputs to the program's next statement. The neural network optimizes multiple tasks concurrently: the next operation out of a set of high level commands, the operands of the next statement, and which variables can be dropped from memory. Using our method we are able to create programs that are more than twice as long as existing state-of-the-art solutions, while improving the success rate for comparable lengths, and cutting the run-time by two orders of magnitude. Our code, including an implementation of various literature baselines, is publicly available at https://github.com/amitz25/PCCoder Amit Zohar, Lior Wolf |
NeurIPS | 2 |
| 2018 | A Method for Segmentation, Matching and Alignment of Dead Sea ScrollsabstractThe Dead Sea Scrolls are of great historical significance. Lamentably, in the decades since their discovery, many fragments have deteriorated. Fortunately, low-resolution grayscale infrared images of the Palestinian Archaeological Museum plates holding the scrolls in their discovered state are extant, along with recent high-quality multispectral images by the Israel Antiquities Authority. However, the necessary task of identifying each fragment in the new images on the old plates is tedious and time consuming to perform manually, and is often problematic when fragments have been moved from the original plate. We describe an automated system that segments the new and old images of fragments from the background on which they were imaged, finds their matches on the old plates and aligns and superimposes them. To this end, we developed a deep-learning based segmentation method and a cascade approach for template matching, based on scale, shape analysis and dense matching. We have tested the proposed method on five plates, comprising about 120 fragments. We present both quantitative and qualitative analyses of the results and perform an ablation study to evaluate the importance of each component of our system. Gil Levi, Pinhas Nisnevich, Adiel Ben-Shalom, Nachum Dershowitz, Lior Wolf |
WACV | 5 |
| 2018 | Confidence Prediction for Lexicon-Free OCRabstractHaving a reliable accuracy score is crucial for real world applications of OCR, since such systems are judged by the number of false readings. Lexicon-based OCR systems, which deal with what is essentially a multi-class classification problem, often employ methods explicitly taking into account the lexicon, in order to improve accuracy. However, in lexicon-free scenarios, filtering errors requires an explicit confidence calculation. In this work we show two explicit confidence measurement techniques, and show that they are able to achieve a significant reduction in misreads on both standard benchmarks and a proprietary dataset. Noam Mor, Lior Wolf |
WACV | 2 |
| 2018 | Structured GANsabstractWe present Generative Adversarial Networks (GANs), in which the symmetric property of the generated images is controlled. This is obtained through the generator network's architecture, while the training procedure and the loss remain the same. The symmetric GANs are applied to face image synthesis in order to generate novel faces with a varying amount of symmetry. We also present an unsupervised face rotation capability, which is based on the novel notion of one-shot fine tuning. Irad Peleg, Lior Wolf |
WACV | 2 |
| 2017 | Linking Image and Text with 2-Way NetsabstractLinking two data sources is a basic building block in numerous computer vision problems. Canonical Correlation Analysis (CCA) achieves this by utilizing a linear optimizer in order to maximize the correlation between the two views. Recent work makes use of non-linear models, including deep learning techniques, that optimize the CCA loss in some feature space. In this paper, we introduce a novel, bi-directional neural network architecture for the task of matching vectors from two data sources. Our approach employs two tied neural network channels that project the two views into a common, maximally correlated space using the Euclidean loss. We show a direct link between the correlation-based loss and Euclidean loss, enabling the use of Euclidean loss for correlation maximization. To overcome common Euclidean regression optimization problems, we modify well-known techniques to our problem, including batch normalization and dropout. We show state of the art results on a number of computer vision matching tasks including MNIST image matching and sentence-image matching on the Flickr8k, Flickr30k and COCO datasets. Aviv Eisenschtat, Lior Wolf |
CVPR | 2 |
| 2017 | Optical Flow Requires Multiple Strategies (but Only One Network)abstractWe show that the matching problem that underlies optical flow requires multiple strategies, depending on the amount of image motion and other factors. We then study the implications of this observation on training a deep neural network for representing image patches in the context of descriptor based optical flow. We propose a metric learning method, which selects suitable negative samples based on the nature of the true match. This type of training produces a network that displays multiple strategies depending on the input and leads to state of the art results on the KITTI 2012 and KITTI 2015 optical flow benchmarks. Tal Schuster, Lior Wolf, David Gadot |
CVPR | 2 |
| 2017 | Improved Stereo Matching with Constant Highway Networks and Reflective Confidence LearningabstractWe present an improved three-step pipeline for the stereo matching problem and introduce multiple novelties at each stage. We propose a new highway network architecture for computing the matching cost at each possible disparity, based on multilevel weighted residual shortcuts, trained with a hybrid loss that supports multilevel comparison of image patches. A novel post-processing step is then introduced, which employs a second deep convolutional neural network for pooling global information from multiple disparities. This network outputs both the image disparity map, which replaces the conventional winner takes all strategy, and a confidence in the prediction. The confidence score is achieved by training the network with a new technique that we call the reflective loss. Lastly, the learned confidence is employed in order to better detect outliers in the refinement step. The proposed pipeline achieves state of the art accuracy on the largest and most competitive stereo benchmarks, and the learned confidence is shown to outperform all existing alternatives. Amit Shaked, Lior Wolf |
CVPR | 2 |
| 2017 | InterpoNet, a Brain Inspired Neural Network for Optical Flow Dense InterpolationabstractSparse-to-dense interpolation for optical flow is a fundamental phase in the pipeline of most of the leading optical flow estimation algorithms. The current state-of-the-art method for interpolation, EpicFlow, is a local average method based on an edge aware geodesic distance. We propose a new data-driven sparse-to-dense interpolation algorithm based on a fully convolutional network. We draw inspiration from the filling-in process in the visual cortex and introduce lateral dependencies between neurons and multi-layer supervision into our learning process. We also show the importance of the image contour to the learning process. Our method is robust and outperforms EpicFlow on competitive optical flow benchmarks with several underlying matching algorithms. This leads to state-of-the-art performance on the Sintel and KITTI 2012 benchmarks. Shay Zweig, Lior Wolf |
CVPR | 2 |
| 2017 | Temporal Tessellation: A Unified Approach for Video AnalysisabstractWe present a general approach to video understanding, inspired by semantic transfer techniques that have been successfully used for 2D image analysis. Our method considers a video to be a 1D sequence of clips, each one associated with its own semantics. The nature of these semantics - natural language captions or other labels - depends on the task at hand. A test video is processed by forming correspondences between its clips and the clips of reference videos with known semantics, following which, reference semantics can be transferred to the test video. We describe two matching methods, both designed to ensure that (a) reference clips appear similar to test clips and (b), taken together, the semantics of the selected reference clips is consistent and maintains temporal coherence. We use our method for video captioning on the LSMDC'16 benchmark, video summarization on the SumMe and TV-Sum benchmarks, Temporal Action Detection on the Thumos2014 benchmark, and sound prediction on the Greatest Hits benchmark. Our method not only surpasses the state of the art, in four out of five benchmarks, but importantly, it is the only single method we know of that was successfully applied to such a diverse range of tasks. Dotan Kaufman, Gil Levi, Tal Hassner, Lior Wolf |
ICCV | 4 |
| 2017 | Unsupervised Creation of Parameterized AvatarsabstractWe study the problem of mapping an input image to a tied pair consisting of a vector of parameters and an image that is created using a graphical engine from the vector of parameters. The mapping's objective is to have the output image as similar as possible to the input image. During training, no supervision is given in the form of matching inputs and outputs. This learning problem extends two literature problems: unsupervised domain adaptation and cross domain transfer. We define a generalization bound that is based on discrepancy, and employ a GAN to implement a network solution that corresponds to this bound. Experimentally, our method is shown to solve the problem of automatically creating avatars. Lior Wolf, Yaniv Taigman, Adam Polyak |
ICCV | 1 |
| 2017 | VASESKETCH: Automatic 3D Representation of Pottery from Paper Catalog DrawingsabstractWe describe an automated pipeline for digitization of catalog drawings of pottery types. This work is aimed at extracting a structured description of the main geometric features and a 3D representation of each class. The pipeline includes methods for understanding a 2D drawing and using it for constructing a 3D model of the pottery. These will be used to populate a reference database for classification of potsherds. Furthermore, we extend the pipeline with methods for breaking the 3D model to obtain synthetic sherds and methods for capturing images of these sherds in a way that matches the imaging methodology of archaeologists. These will serve to build a massive set of synthetic sherd images that will help train and test future automated classification systems. Francesco Banterle, Barak Itkin, Matteo Dellepiane, Lior Wolf, Marco Callieri, Nachum Dershowitz, Roberto Scopigno |
ICDAR | 4 |
| 2017 | Relating Articles Textually and VisuallyabstractHistorical documents have been undergoing large-scale digitization over the past years, placing massive image collections online. Optical character recognition (OCR) often performs poorly on such material, which makes searching within these resources problematic and textual analysis of such documents difficult. We present two approaches to overcome this obstacle, one textual and one visual. We show that, for tasks like finding newspaper articles related by topic, poor-quality OCR text suffices. An ordinary vector-space model is used to represent articles. Additional improvements obtain by adding words with similar distributional representations. As an alternative to OCR-based methods, one can perform image-based search, using word spotting. Synthetic images are generated for every word in a lexicon, and word-spotting is used to compile vectors of their occurrences. Retrieval is by means of a usual nearest-neighbor search. The results of this visual approach are comparable to those obtained using noisy OCR. We report on experiments applying both methods, separately and together, on historical Hebrew newspapers, with their added problem of rich morphology. Nachum Dershowitz, Daniel Labenski, Adi Silberpfennig, Lior Wolf, Yaron Tsur |
ICDAR | 4 |
| 2017 | Qumran Letter Restoration by Rotation and Reflection Modified PixelCNNabstractThe task of restoring fragmentary letters is fundamental to the reading of ancient manuscripts. We present a method to complete broken letters in the Dead Sea Scrolls, which is based on PixelCNN++. Since the generation of the broken letters is conditioned on the extant scroll, we modify the original method to allow reconstructions in multiple directions. Results on both simulated data and real scrolls demonstrate the advantage of our method over the baseline. The implementation may be found at https://github.com/ghostcow/pixel-cnn-qumran. Lior Uzan, Nachum Dershowitz, Lior Wolf |
ICDAR | 3 |
| 2017 | Unsupervised Cross-Domain Image Generation
Yaniv Taigman, Adam Polyak, Lior Wolf |
ICLR (Poster) | 3 |
| 2017 | Learning to Align the Source Code to the Compiled Object CodeabstractWe propose a new neural network architecture and use it for the task of statement-by-statement alignment of source code and its compiled object code. Our architecture learns the alignment between the two sequences – one being the translation of the other – by mapping each statement to a context-dependent representation vector and aligning such vectors using a grid of the two sequence domains. Our experiments include short C functions, both artificial and human-written, and show that our neural network architecture is able to predict the alignment with high accuracy, outperforming known baselines. We also demonstrate that our model is general and can learn to solve graph problems such as the Traveling Salesman Problem. Dor Levy, Lior Wolf |
ICML | 2 |
| 2017 | One-Sided Unsupervised Domain MappingabstractIn unsupervised domain mapping, the learner is given two unmatched datasets $A$ and $B$. The goal is to learn a mapping $G_{AB}$ that translates a sample in $A$ to the analog sample in $B$. Recent approaches have shown that when learning simultaneously both $G_{AB}$ and the inverse mapping $G_{BA}$, convincing mappings are obtained. In this work, we present a method of learning $G_{AB}$ without learning $G_{BA}$. This is done by learning a mapping that maintains the distance between a pair of samples. Moreover, good mappings are obtained, even by maintaining the distance between different parts of the same sample before and after mapping. We present experimental results that the new method not only allows for one sided mapping learning, but also leads to preferable numerical results over the existing circularity-based constraint. Our entire code is made publicly available at~\url{https://github.com/sagiebenaim/DistanceGAN}. Sagie Benaim, Lior Wolf |
NIPS | 2 |
| 2016 | Stemming and Segmentation for Classical Tibetan
Orna Almogi, Lena Dankin, Nachum Dershowitz, Yair Hoffman, Dimitri Pauls, Dorji Wangchuk, Lior Wolf |
CICLing (1) | 7 |
| 2016 | PatchBatch: A Batch Augmented Loss for Optical FlowabstractWe propose a new pipeline for optical flow computation, based on Deep Learning techniques. We suggest using a Siamese CNN to independently, and in parallel, compute the descriptors of both images. The learned descriptors are then compared efficiently using the L2 norm and do not require network processing of patch pairs. The success of the method is based on an innovative loss function that computes higher moments of the loss distributions for each training batch. Combined with an Approximate Nearest Neighbor patch matching method and a flow interpolation technique, state of the art performance is obtained on the most challenging and competitive optical flow benchmarks. David Gadot, Lior Wolf |
CVPR | 2 |
| 2016 | The Multiverse Loss for Robust Transfer LearningabstractDeep learning techniques are renowned for supporting effective transfer learning. However, as we demonstrate, the transferred representations support only a few modes of separation and much of its dimensionality is unutilized. In this work, we suggest to learn, in the source domain, multiple orthogonal classifiers. We prove that this leads to a reduced rank representation, which, however, supports more discriminative directions. Interestingly, the softmax probabilities produced by the multiple classifiers are likely to be identical. Experimental results, on CIFAR-100 and LFW, further demonstrate the effectiveness of our method. Etai Littwin, Lior Wolf |
CVPR | 2 |
| 2016 | CNN-N-Gram for HandwritingWord RecognitionabstractGiven an image of a handwritten word, a CNN is employed to estimate its n-gram frequency profile, which is the set of n-grams contained in the word. Frequencies for unigrams, bigrams and trigrams are estimated for the entire word and for parts of it. Canonical Correlation Analysis is then used to match the estimated profile to the true profiles of all words in a large dictionary. The CNN that is used employs several novelties such as the use of multiple fully connected branches. Applied to all commonly used handwriting recognition benchmarks, our method outperforms, by a very large margin, all existing methods. Arik Poznanski, Lior Wolf |
CVPR | 2 |
| 2016 | RNN Fisher Vectors for Action Recognition and Image Annotation
Guy Lev, Gil Sadeh, Benjamin Eliot Klein, Lior Wolf |
ECCV (6) | 4 |
| 2016 | Learning to Count with CNN Boosting
Elad Walach, Lior Wolf |
ECCV (2) | 2 |
| 2016 | DeepChess: End-to-End Deep Neural Network for Automatic Learning in Chess
Eli David, Nathan S. Netanyahu, Lior Wolf |
ICANN (2) | 3 |
| 2016 | Complexity of multiverse networks and their multilayer generalizationabstractMultiverse networks were recently proposed as a method for promoting more effective transfer learning. While an extensive analysis was proposed, this analysis failed to capture two main aspects of these networks: (i) the rank of the representation is much lower than the rank predicted by the analysis; and (ii) the contribution of increased multiplicity in such networks diminishes quickly. In this work, we propose additional analysis of multiverse networks which addresses both deficits. A major contribution of our work is quantifying the Rademacher complexity of the multiverse network. It is shown that the complexity upper bound of multiverse networks is significantly lower than that of conventional networks, and diminishes by a factor of √k, k being the multiplicity. In addition, we generalize the notion of multiverse networks to multilayer multiverse networks. We derive the Rademacher complexity formula to such networks and present experimental results. Etai Littwin, Lior Wolf |
ICPR | 2 |
| 2016 | Softbot: Software-based lead-through for rigid servo robotsabstractIn this paper, we present an interactive control method for rigid robotics. The core of the method is a neural network classifier that maps position and torque readings to a force direction in 3D space. We show that running our method online allows a human to move the robot along a desired path by performing intuitive pushes and pulls on the robot's joints. The setup is sensorless: no additional sensors, other than those already integral to the rigid joint itself, are added to the manipulator or used in the process. Yackov Lubarsky, Amit Wolf, Lior Wolf, Curime Batliner, Jake Newsum |
ICRA | 3 |
| 2016 | Texture instance similarity via dense correspondencesabstractThis paper concerns the task of evaluating the similarity of textures instances: Rather than discriminating between different texture classes, our goal is to identify when two images display the same texture instance. To address this problem, we propose an approach inspired by alignment based recognition theories. We offer a pixel-based method, employing a robust, dense correspondence estimation engine, applied to an efficient, novel representation, to match the pixels of two texture photos. We describe means for quantifying the quality of these matches, considering in particular the quality of the flow established between the two images. These quality measures are effectively combined into similarity scores by using standard linear SVM classifiers. By relying on a general, alignment based approach our method can be applied to different problem domains (different texture classes) with little modification. We demonstrate this by reporting state-of-the-art results on benchmarks for fingerprint recognition and two new benchmarks for texture-based animal identification. Tal Hassner, Gilad Saban, Lior Wolf |
WACV | 3 |
| 2015 | Associating neural word embeddings with deep image representations using Fisher VectorsabstractIn recent years, the problem of associating a sentence with an image has gained a lot of attention. This work continues to push the envelope and makes further progress in the performance of image annotation and image search by a sentence tasks. In this work, we are using the Fisher Vector as a sentence representation by pooling the word2vec embedding of each word in the sentence. The Fisher Vector is typically taken as the gradients of the log-likelihood of descriptors, with respect to the parameters of a Gaussian Mixture Model (GMM). In this work we present two other Mixture Models and derive their Expectation-Maximization and Fisher Vector expressions. The first is a Laplacian Mixture Model (LMM), which is based on the Laplacian distribution. The second Mixture Model presented is a Hybrid Gaussian-Laplacian Mixture Model (HGLMM) which is based on a weighted geometric mean of the Gaussian and Laplacian distribution. Finally, by using the new Fisher Vectors derived from HGLMMs to represent sentences, we achieve state-of-the-art results for both the image annotation and the image search by a sentence tasks on four benchmarks: Pascal1K, Flickr8K, Flickr30K, and COCO. Benjamin Eliot Klein, Guy Lev, Gil Sadeh, Lior Wolf |
CVPR | 4 |
| 2015 | A Dynamic Convolutional Layer for short rangeweather predictionabstractWe present a new deep network layer called “Dynamic Convolutional Layer” which is a generalization of the convolutional layer. The conventional convolutional layer uses filters that are learned during training and are held constant during testing. In contrast, the dynamic convolutional layer uses filters that will vary from input to input during testing. This is achieved by learning a function that maps the input to the filters. We apply the dynamic convolutional layer to the application of short range weather prediction and show performance improvements compared to other baselines. Benjamin Eliot Klein, Lior Wolf, Yehuda Afek |
CVPR | 2 |
| 2015 | Web-scale training for face identificationabstractScaling machine learning methods to very large datasets has attracted considerable attention in recent years, thanks to easy access to ubiquitous sensing and data from the web. We study face recognition and show that three distinct properties have surprising effects on the transferability of deep convolutional networks (CNN): (1) The bottleneck of the network serves as an important transfer learning regularizer, and (2) in contrast to the common wisdom, performance saturation may exist in CNN's (as the number of training samples grows); we propose a solution for alleviating this by replacing the naive random subsampling of the training set with a bootstrapping process. Moreover, (3) we find a link between the representation norm and the ability to discriminate in a target domain, which sheds lights on how such networks represent faces. Based on these discoveries, we are able to improve face recognition accuracy on the widely used LFW benchmark, both in the verification (1:1) and identification (1:N) protocols, and directly compare, for the first time, with the state of the art Commercially-Off-The-Shelf system and show a sizable leap in performance. Yaniv Taigman, Ming Yang 0007, Marc'Aurelio Ranzato, Lior Wolf |
CVPR | 4 |
| 2015 | Live Repetition CountingabstractThe task of counting the number of repetitions of approximately the same action in an input video sequence is addressed. The proposed method runs online and not on the complete pre-captured video. It analyzes sequentially blocks of 20 non-consecutive frames. The cycle length within each block is evaluated using a convolutional neural network and the information is then integrated over time. The entropy of the network's predictions is used in order to automatically start and stop the repetition counter and to select the appropriate time scale. Coupled with a region of interest detection mechanism, the method is robust enough to handle real world videos, even when the camera is moving. A unique property of our method is that it is shown to successfully train on entirely unrealistic data created by synthesizing moving random patches. Ofir Levy, Lior Wolf |
ICCV | 2 |
| 2015 | Viral transcript alignmentabstractWe present an end-to-end system for aligning transcript letters to their coordinates in a manuscript image. An intuitive GUI and an automatic line detection method enable the user to perform an exact alignment of parts of document pages. In order to bridge large regions in between annotation, and augment the manual effort, the system employs an optical-flow engine for directly matching at the pixel level the image of a line of a historical text with a synthetic image created from the transcript's matching line. Meanwhile, by accumulating aligned letters, and performing letter spotting, the system is able to bootstrap a rapid semi-automatic transcription of the remaining text. Thus, the amount of manual work is greatly diminished and the transcript alignment task becomes practical regardless of the corpus size. Gil Sadeh, Lior Wolf, Tal Hassner, Nachum Dershowitz, Daniel Stökl Ben Ezra |
ICDAR | 2 |
| 2015 | Improving OCR for an under-resourced script using unsupervised word-spottingabstractOptical character recognition (OCR) quality, especially for under-resourced scripts like Bangla, as well as for documents printed in old typefaces, is a major concern. An efficient and effective pipeline for OCR betterment is proposed here. The method is unsupervised. It employs a baseline OCR engine as a black box plus a dataset of unlabeled document images. That engine is applied to the images, followed by a visual encoding designed to support efficient word spotting. Given a new document to be analyzed, the black-box recognition engine is first applied. Then, for each result, word spotting is carried out within the dataset. The unreliable OCR outputs of the retrieved word spotting results are then considered. The word that is the centroid of the set of OCR words, measured by edit distance, is deemed a candidate reading. Adi Silberpfennig, Lior Wolf, Nachum Dershowitz, Bhagesh Seraogi, Bidyut B. Chaudhuri |
ICDAR | 2 |
| 2015 | In Defense of Word Embedding for Generic Text Representation
Guy Lev, Benjamin Eliot Klein, Lior Wolf |
NLDB | 3 |
| 2014 | Congruency-Based RerankingabstractWe present a tool for re-ranking the results of a specific query by considering the (n+1) × (n+1) matrix of pairwise similarities among the elements of the set of n retrieved results and the query itself. The re-ranking thus makes use of the similarities between the various results and does not employ additional sources of information. The tool is based on graphical Bayesian models, which reinforce retrieved items strongly linked to other retrievals, and on repeated clustering to measure the stability of the obtained associations. The utility of the tool is demonstrated within the context of visual search of documents from the Cairo Genizah and for retrieval of paintings by the same artist and in the same style. Itai Ben-Shalom, Noga Levy, Lior Wolf, Nachum Dershowitz, Adiel Ben-Shalom, Roni Shweka, Yaacov Choueka, Tamir Hazan, Yaniv Bar |
CVPR | 3 |
| 2014 | DeepFace: Closing the Gap to Human-Level Performance in Face VerificationabstractIn modern face recognition, the conventional pipeline consists of four stages: detect => align => represent => classify. We revisit both the alignment step and the representation step by employing explicit 3D face modeling in order to apply a piecewise affine transformation, and derive a face representation from a nine-layer deep neural network. This deep network involves more than 120 million parameters using several locally connected layers without weight sharing, rather than the standard convolutional layers. Thus we trained it on the largest facial dataset to-date, an identity labeled dataset of four million facial images belonging to more than 4, 000 identities. The learned representations coupling the accurate model-based alignment with the large facial database generalize remarkably well to faces in unconstrained environments, even with a simple classifier. Our method reaches an accuracy of 97.35% on the Labeled Faces in the Wild (LFW) dataset, reducing the error of the current state of the art by more than 27%, closely approaching human-level performance. Yaniv Taigman, Ming Yang 0007, Marc'Aurelio Ranzato, Lior Wolf |
CVPR | 4 |
| 2014 | A Simple and Fast Word Spotting MethodabstractA simple and efficient pipeline for word spotting in handwritten documents is proposed. The method allows for extremely rapid querying, while still maintaining high accuracy. The dataset images that are to be queried are preprocessed by a simple binarization operation, followed by the extraction of multiple overlapping candidate targets. Each binary target, as well as the binarized query, is resized to fit a fixed-size rectangle and represented by conventional image descriptors. Then, a cosine similarity operator -- followed by maximum pooling over random groups -- is used to represent each target or query as a concise 250D vector. Retrieval is performed in a fraction of a second by nearest-neighbor search within that space, followed by a simple suppression of extra overlapping candidates. Alon Kovalchuk, Lior Wolf, Nachum Dershowitz |
ICFHR | 2 |
| 2014 | When standard RANSAC is not enough: cross-media visual matching with hypothesis relevancy
Tal Hassner, Liav Assif, Lior Wolf |
Mach. Vis. Appl. | 3 |
| 2014 | Propagating Waves of Directionality and Coordination Orchestrate Collective Cell MigrationabstractThe ability of cells to coordinately migrate in groups is crucial to enable them to travel long distances during embryonic development, wound healing and tumorigenesis, but the fundamental mechanisms underlying intercellular coordination during collective cell migration remain elusive despite considerable research efforts. A novel analytical framework is introduced here to explicitly detect and quantify cell clusters that move coordinately in a monolayer. The analysis combines and associates vast amount of spatiotemporal data across multiple experiments into transparent quantitative measures to report the emergence of new modes of organized behavior during collective migration of tumor and epithelial cells in wound healing assays. First, we discovered the emergence of a wave of coordinated migration propagating backward from the wound front, which reflects formation of clusters of coordinately migrating cells that are generated further away from the wound edge and disintegrate close to the advancing front. This wave emerges in both normal and tumor cells, and is amplified by Met activation with hepatocyte growth factor/scatter factor. Second, Met activation was found to induce coinciding waves of cellular acceleration and stretching, which in turn trigger the emergence of a backward propagating wave of directional migration with about an hour phase lag. Assessments of the relations between the waves revealed that amplified coordinated migration is associated with the emergence of directional migration. Taken together, our data and simplified modeling-based assessments suggest that increased velocity leads to enhanced coordination: higher motility arises due to acceleration and stretching that seems to increase directionality by temporarily diminishing the velocity components orthogonal to the direction defined by the monolayer geometry. Spatial and temporal accumulation of directionality thus defines coordination. The findings offer new insight and suggest a basic cellular mechanism for long-term cell guidance and intercellular communication during collective cell migration. Assaf Zaritsky, Doron Kaplan, Inbal Hecht, Sari Natan, Lior Wolf, Nir S. Gov, Eshel Ben-Jacob, Ilan Tsarfaty |
PLoS Comput. Biol. | 5 |
| 2013 | The SVM-Minus Similarity Score for Video Face RecognitionabstractChallenge, but also an opportunity to eliminate spurious similarities. Luckily, a major source of confusion in visual similarity of faces is the 3D head orientation, for which image analysis tools provide an accurate estimation. The method we propose belongs to a family of classifier-based similarity scores. We present an effective way to discount pose induced similarities within such a framework, which is based on a newly introduced classifier called SVM-minus. The presented method is shown to outperform existing techniques on the most challenging and realistic publicly available video face recognition benchmark, both by itself, and in concert with other methods. Lior Wolf, Noga Levy |
CVPR | 1 |
| 2013 | Fast High Dimensional Vector Multiplication Face RecognitionabstractThis paper advances descriptor-based face recognition by suggesting a novel usage of descriptors to form an over-complete representation, and by proposing a new metric learning pipeline within the same/not-same framework. First, the Over-Complete Local Binary Patterns (OCLBP) face representation scheme is introduced as a multi-scale modified version of the Local Binary Patterns (LBP) scheme. Second, we propose an efficient matrix-vector multiplication-based recognition system. The system is based on Linear Discriminant Analysis (LDA) coupled with Within Class Covariance Normalization (WCCN). This is further extended to the unsupervised case by proposing an unsupervised variant of WCCN. Lastly, we introduce Diffusion Maps (DM) for non-linear dimensionality reduction as an alternative to the Whitened Principal Component Analysis (WPCA) method which is often used in face recognition. We evaluate the proposed framework on the LFW face recognition dataset under the restricted, unrestricted and unsupervised protocols. In all three cases we achieve very competitive results. Oren Barkan, Jonathan Weill, Lior Wolf, Hagai Aronowitz |
ICCV | 3 |
| 2013 | OCR-Free Transcript AlignmentabstractRecent large-scale digitization and preservation efforts have made images of original manuscripts, accompanied by transcripts, commonly available. An important challenge, for which no practical system exists, is that of aligning transcript letters to their coordinates in manuscript images. Here we propose a system that directly matches the image of a historical text with a synthetic image created from the transcript for the purpose. This, rather than attempting to recognize individual letters in the manuscript image using optical character recognition (OCR). Our method matches the pixels of the two images by employing a dedicated dense flow mechanism coupled with novel local image descriptors designed to spatially integrate local patch similarities. Matching these pixel representations is performed using a message passing algorithm. The various stages of our method make it robust with respect to document degradation, to variations between script styles and to non-linear image transformations. Robustness, as well as practicality of the system, are verified by comprehensive empirical experiments. Tal Hassner, Lior Wolf, Nachum Dershowitz |
ICDAR | 2 |
| 2013 | Integrating Copies Obtained from Old and New Preservation EffortsabstractThe Dead Sea Scrolls were discovered in the Qumran area and elsewhere in the Judean desert beginning in 1947 and were photographed in infrared in the 1950s. Recently, the Israel Antiquities Authority embarked on an ambitious project to digitize all the fragments using multi-spectral cameras. We describe a method that utilizes information from both of these image sets: the highly detailed multispectral images and the older infrared images, which preserve the state of the fragments as it was shortly after discovery. We use a two-step registration procedure to align the image sets. First, a coarse global transformation is applied to the whole image of the new set, producing a rough alignment, followed by a fine, local wrapping based on interest point matching. The aligned images can be used to improve image binarization and to identify and repair fragments that have degraded further over the years. Additionally, the fine alignment parameters can be used for coarse attribute classification, such as the period when written. Yoram Zarai, Tamar Lavee, Nachum Dershowitz, Lior Wolf |
ICDAR | 4 |
| 2013 | Benchmark for multi-cellular segmentation of bright field microscopy imagesabstractBACKGROUND: Multi-cellular segmentation of bright field microscopy images is an essential computational step when quantifying collective migration of cells in vitro. Despite the availability of various tools and algorithms, no publicly available benchmark has been proposed for evaluation and comparison between the different alternatives. DESCRIPTION: A uniform framework is presented to benchmark algorithms for multi-cellular segmentation in bright field microscopy images. A freely available set of 171 manually segmented images from diverse origins was partitioned into 8 datasets and evaluated on three leading designated tools. CONCLUSIONS: The presented benchmark resource for evaluating segmentation algorithms of bright field images is the first public annotated dataset for this purpose. This annotated dataset of diverse examples allows fair evaluations and comparisons of future segmentation methods. Scientists are encouraged to assess new algorithms on this benchmark, and to contribute additional annotated datasets. Assaf Zaritsky, Nathan Manor, Lior Wolf, Eshel Ben-Jacob, Ilan Tsarfaty |
BMC Bioinform. | 3 |
| 2012 | Motion Interchange Patterns for Action Recognition in Unconstrained Videos
Orit Kliper-Gross, Yaron Gurovich, Tal Hassner, Lior Wolf |
ECCV (6) | 4 |
| 2012 | Minimal Correlation Classification
Noga Levy, Lior Wolf |
ECCV (6) | 2 |
| 2012 | The Action Similarity Labeling ChallengeabstractRecognizing actions in videos is rapidly becoming a topic of much research. To facilitate the development of methods for action recognition, several video collections, along with benchmark protocols, have previously been proposed. In this paper, we present a novel video database, the "Action Similarity LAbeliNg" (ASLAN) database, along with benchmark protocols. The ASLAN set includes thousands of videos collected from the web, in over 400 complex action classes. Our benchmark protocols focus on action similarity (same/not-same), rather than action classification, and testing is performed on never-before-seen actions. We propose this data set and benchmark as a means for gaining a more principled understanding of what makes actions different or similar, rather than learning the properties of particular action classes. We present baseline results on our benchmark, and compare them to human performance. To promote further study of action similarity techniques, we make the ASLAN database, benchmarks, and descriptor encodings publicly available to the research community. Orit Kliper-Gross, Tal Hassner, Lior Wolf |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2011 | Face recognition in unconstrained videos with matched background similarityabstractRecognizing faces in unconstrained videos is a task of mounting importance. While obviously related to face recognition in still images, it has its own unique characteristics and algorithmic requirements. Over the years several methods have been suggested for this problem, and a few benchmark data sets have been assembled to facilitate its study. However, there is a sizable gap between the actual application needs and the current state of the art. In this paper we make the following contributions. (a) We present a comprehensive database of labeled videos of faces in challenging, uncontrolled conditions (i.e., `in the wild'), the `YouTube Faces' database, along with benchmark, pair-matching tests1. (b) We employ our benchmark to survey and compare the performance of a large variety of existing video face recognition techniques. Finally, (c) we describe a novel set-to-set similarity measure, the Matched Background Similarity (MBGS). This similarity is shown to considerably improve performance on the benchmark tests. Lior Wolf, Tal Hassner, Itay Maoz |
CVPR | 1 |
| 2011 | Active clustering of document fragments using information derived from both images and catalogsabstractMany significant historical corpora contain leaves that are mixed up and no longer bound in their original state as multi-page documents. The reconstruction of old manuscripts from a mix of disjoint leaves can therefore be of paramount importance to historians and literary scholars. Previously, it was shown that visual similarity provides meaningful pair-wise similarities between handwritten leaves. Here, we go a step further and suggest a semiautomatic clustering tool that helps reconstruct the original documents. The proposed solution is based on a graphical model that makes inferences based on catalog information provided for each leaf as well as on the pairwise similarities of handwriting. Several novel active clustering techniques are explored, and the solution is applied to a significant part of the Cairo Genizah, where the problem of joining leaves remains unsolved even after a century of extensive study by hundreds of scholars. Lior Wolf, Lior Litwak, Nachum Dershowitz, Roni Shweka, Yaacov Choueka |
ICCV | 1 |
| 2011 | Computerized paleography: Tools for historical manuscriptsabstractThe Digital Age has brought with it large-scale digitization of historical records. The modern scholar of history or of other disciplines is often faced today with hundreds of thousands of readily-available and potentially-relevant full or fragmentary documents, but without computer aids that would make it possible to find the sought-after needles in the proverbial haystack of online images. The problems are even more acute when documents are handwritten, since optical character recognition does not provide quality results. We consider two tools: (1) a handwriting matching tool that is used to join together fragments of the same scribe, and (2) a paleographic classification tool that matches a given document to a large set of paleographic samples. Both tools are carefully designed not only to provide a high level of accuracy, but also to provide a clean and concise justification of the inferred results. This last requirement engenders challenges, such as sparsity of the representation, for which existing solutions are inappropriate for document analysis. Lior Wolf, Liza Potikha, Nachum Dershowitz, Roni Shweka, Yaacov Choueka |
ICIP | 1 |
| 2011 | Protein stability: a single recorded mutation aids in predicting the effects of other mutations in the same amino acid siteabstractMOTIVATION: Accurate prediction of protein stability is important for understanding the molecular underpinnings of diseases and for the design of new proteins. We introduce a novel approach for the prediction of changes in protein stability that arise from a single-site amino acid substitution; the approach uses available data on mutations occurring in the same position and in other positions. Our algorithm, named Pro-Maya (Protein Mutant stAbilitY Analyzer), combines a collaborative filtering baseline model, Random Forests regression and a diverse set of features. Pro-Maya predicts the stability free energy difference of mutant versus wild type, denoted as ΔΔG. RESULTS: We evaluated our algorithm extensively using cross-validation on two previously utilized datasets of single amino acid mutations and a (third) validation set. The results indicate that using known ΔΔG values of mutations at the query position improves the accuracy of ΔΔG predictions for other mutations in that position. The accuracy of our predictions in such cases significantly surpasses that of similar methods, achieving, e.g. a Pearson's correlation coefficient of 0.79 and a root mean square error of 0.96 on the validation set. Because Pro-Maya uses a diverse set of features, including predictions using two other methods, it also performs slightly better than other methods in the absence of additional experimental data on the query positions. AVAILABILITY: Pro-Maya is freely available via web server at http://bental.tau.ac.il/ProMaya. CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Gilad Wainreb, Lior Wolf, Haim Ashkenazy, Yves Dehouck, Nir Ben-Tal |
Bioinform. | 2 |
| 2011 | Prior Knowledge for Part CorrespondenceabstractAbstract Classical approaches to shape correspondence base their computation purely on the properties, in particular geometric similarity, of the shapes in question. Their performance still falls far short of that of humans in challenging cases where corresponding shape parts may differ significantly in geometry or even topology. We stipulate that in these cases, shape correspondence by humans involves recognition of the shape parts where prior knowledge on the parts would play a more dominant role than geometric similarity. We introduce an approach to part correspondence which incorporates prior knowledge imparted by a training set of pre‐segmented, labeled models and combines the knowledge with content‐driven analysis based on geometric similarity between the matched shapes. First, the prior knowledge is learned from the training set in the form of per‐label classifiers. Next, given two query shapes to be matched, we apply the classifiers to assign a probabilistic label to each shape face. Finally, by means of a joint labeling scheme, the probabilistic labels are used synergistically with pairwise assignments derived from geometric similarity to provide the resulting part correspondence. We show that the incorporation of knowledge is especially effective in dealing with shapes exhibiting large intra‐class variations. We also show that combining knowledge and content analyses outperforms approaches guided by either attribute alone. Oliver van Kaick, Andrea Tagliasacchi, Oana Sidi, Hao (Richard) Zhang, Daniel Cohen-Or, Lior Wolf, Ghassan Hamarneh |
Comput. Graph. Forum | 6 |
| 2011 | Content aware video manipulation
Moshe Guttmann, Lior Wolf, Daniel Cohen-Or |
Comput. Vis. Image Underst. | 2 |
| 2011 | Identifying Join Candidates in the Cairo Genizah
Lior Wolf, Rotem Littman, Naama Mayer, Tanya German, Nachum Dershowitz, Roni Shweka, Yaacov Choueka |
Int. J. Comput. Vis. | 1 |
| 2011 | Effective Unconstrained Face Recognition by Combining Multiple Descriptors and Learned Background StatisticsabstractComputer vision systems have demonstrated considerable improvement in recognizing and verifying faces in digital images. Still, recognizing faces appearing in unconstrained, natural conditions remains a challenging task. In this paper, we present a face-image, pair-matching approach primarily developed and tested on the "Labeled Faces in the Wild" (LFW) benchmark that reflects the challenges of face recognition from unconstrained images. The approach we propose makes the following contributions. 1) We present a family of novel face-image descriptors designed to capture statistics of local patch similarities. 2) We demonstrate how unlabeled background samples may be used to better evaluate image similarities. To this end, we describe a number of novel, effective similarity measures. 3) We show how labeled background samples, when available, may further improve classification performance, by employing a unique pair-matching pipeline. We present state-of-the-art results on the LFW pair-matching benchmarks. In addition, we show our system to be well suited for multilabel face classification (recognition) problem, on both the LFW images and on images from the laboratory controlled multi-PIE database. Lior Wolf, Tal Hassner, Yaniv Taigman |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2011 | Gene Expression in the Rodent Brain is Associated with Its Regional ConnectivityabstractThe putative link between gene expression of brain regions and their neural connectivity patterns is a fundamental question in neuroscience. Here this question is addressed in the first large scale study of a prototypical mammalian rodent brain, using a combination of rat brain regional connectivity data with gene expression of the mouse brain. Remarkably, even though this study uses data from two different rodent species (due to the data limitations), we still find that the connectivity of the majority of brain regions is highly predictable from their gene expression levels-the outgoing (incoming) connectivity is successfully predicted for 73% (56%) of brain regions, with an overall fairly marked accuracy level of 0.79 (0.83). Many genes are found to play a part in predicting both the incoming and outgoing connectivity (241 out of the 500 top selected genes, p-value<1e-5). Reassuringly, the genes previously known from the literature to be involved in axon guidance do carry significant information about regional brain connectivity. Surveying the genes known to be associated with the pathogenesis of several brain disorders, we find that those associated with schizophrenia, autism and attention deficit disorder are the most highly enriched in the connectivity-related genes identified here. Finally, we find that the profile of functional annotation groups that are associated with regional connectivity in the rodent is significantly correlated with the annotation profile of genes previously found to determine neural connectivity in C. elegans (Pearson correlation of 0.24, p<1e-6 for the outgoing connections and 0.27, p<1e-5 for the incoming). Overall, the association between connectivity and gene expression in a specific extant rodent species' brain is likely to be even stronger than found here, given the limitations of current data. Lior Wolf, Chen Goldberg, Nathan Manor, Roded Sharan, Eytan Ruppin |
PLoS Comput. Biol. | 1 |
| 2010 | An eye for an eye: A single camera gaze-replacement methodabstractThe camera in video conference systems is typically positioned above, or below, the screen, causing the gaze of the users to appear misplaced. We propose an effective solution to this problem that is based on replacing the eyes of the user. This replacement, when done accurately, is enough to achieve a natural looking video. At an initialization stage the user is asked to look straight at the camera. We store these frames, then track the eyes accurately in the video sequence and replace the eyes, taking care of illumination and ghosting artifacts. We have tested the system on a large number of videos demonstrating the effectiveness of the proposed solution. Lior Wolf, Ziv Freund, Shai Avidan |
CVPR | 1 |
| 2010 | Visual recognition using mappings that replicate marginsabstractWe consider the problem of learning to map between two vector spaces given pairs of matching vectors, one from each space. This problem naturally arises in numerous vision problems, for example, when mapping between the images of two cameras, or when the annotations of each image is multidimensional. We focus on the common asymmetric case, where one vector space X is more informative than the other Y, and find a transformation from Y to X. We present a new optimization problem that aims to replicate in the transformed Y the margins that dominate the structure of X. This optimization problem is convex, and efficient algorithms are presented. Links to various existing methods such as CCA and SVM are drawn, and the effectiveness of the method is demonstrated in several visual domains. Lior Wolf, Nathan Manor |
CVPR | 1 |
| 2010 | Automatic Cephalometric Evaluation of Patients Suffering from Sleep-Disordered Breathing
Lior Wolf, Tamir Yedidya, Roy Ganor, Michael Chertok, Ariela Nachmani, Yehuda Finkelstein |
MICCAI (3) | 1 |
| 2010 | The Virtual Director: a Correlation-Based Online Viewing of Human MotionabstractAbstract Automatic camera control for scenes depicting human motion is an imperative topic in motion capture base animation, computer games, and other animation based fields. This challenging control problem is complex and combines both geometric constraints, visibility requirements, and aesthetic elements. Therefore, existing optimization‐based approaches for human action overview are often too demanding for online computation. In this paper, we introduce an effective automatic camera control which is extremely efficient and allows online performance. Rather than optimizing a complex quality measurement, at each time it selects one active camera from a multitude of cameras that render the dynamic scene. The selection is based on the correlation between each view stream and the human motion in the scene. Two factors allow for rapid selection among tens of candidate views in real‐time, even for complex multi‐character scenes: the efficient rendering of the multitude of view streams, and optimized calculations of the correlations using modified CCA. In addition to the method's simplicity and speed, it exhibits good agreement with both cinematic idioms and previous human motion camera control work. Our evaluations show that the method is able to cope with the challenges put forth by severe occlusions, multiple characters and complex scenes. Jackie Assa, Lior Wolf, Daniel Cohen-Or |
Comput. Graph. Forum | 2 |
| 2010 | Optimizing Photo CompositionabstractAbstract Aesthetic images evoke an emotional response that transcends mere visual appreciation. In this work we develop a novel computational means for evaluating the composition aesthetics of a given image based on measuring several well‐grounded composition guidelines. A compound operator of crop‐and‐retarget is employed to change the relative position of salient regions in the image and thus to modify the composition aesthetics of the image. We propose an optimization method for automatically producing a maximally‐aesthetic version of the input image. We validate the performance of the method and show its effectiveness in a variety of experiments. Ligang Liu 0001, Renjie Chen 0001, Lior Wolf, Daniel Cohen-Or |
Comput. Graph. Forum | 3 |
| 2009 | Similarity Scores Based on Background Samples
Lior Wolf, Tal Hassner, Yaniv Taigman |
ACCV (2) | 1 |
| 2009 | Multiple One-Shots for Utilizing Class Label InformationabstractThe One-Shot Similarity (OSS) kernel [3, 4] has recently been introduced as a means of boosting the performance of face recognition systems. Given two vectors, their One-Shot Similarity score (Fig. 1) reflects the likelihood of each vector belonging to the same class as the other vector and not in a class defined by a fixed set of “negative” examples. In this paper we explore how the One-Shot Similarity may nevertheless benefit from the availability of such labels. (a) we present a system utilizing identity and pose information to improve facial image pair-matching performance using multiple One-Shot scores; (b) we show how separating pose and identity may lead to better face recognition rates in unconstrained, “wild” facial images; (c) we explore how far we can get using a single descriptor with different similarity tests as opposed to the popular multiple descriptor approaches; and (d) we demonstrate the benefit of learned metrics for improved One-Shot performance. Yaniv Taigman, Lior Wolf, Tal Hassner |
BMVC | 2 |
| 2009 | One-sided object cutout using principal-channelsabstractWe introduce principal-channels for cutting out objects from an image by one-sided scribbles. We demonstrate that few scribbles, all from within the object of interest, are sufficient to mark it out. One-sided scribbles provide significantly less information than two-sided ones. Thus, it is required to maximize the use of image-information. Our approach is to first analyze the image with a large filter bank and generate a high-dimensional feature space. We then extract a set of principal-channels that discern one object from another. We show that by applying an iterative graph-cut optimization over the principal-channels, we can cut out the object of interest. Lior Gavish, Lior Shapira, Lior Wolf, Daniel Cohen-Or |
CAD/Graphics | 3 |
| 2009 | Semi-automatic stereo extraction from video footageabstractWe present a semi-automatic system that converts conventional video shots to stereoscopic video pairs. The system requires just a few user-scribbles in a sparse set of frames. The system combines a diffusion scheme, which takes into account the local saliency and the local motion at each video location, coupled with a classification scheme that assigns depth to image patches. The system tolerates both scene motion and camera motion. In typical shots, containing hundreds of frames, even in the face of significant motion, it is enough to mark scribbles on the first and last frames of the shot. Once marked, plausible stereo results are obtained in a matter of seconds, leading to a scalable video conversion system. Finally, we validate our results with ground truth stereo video. Moshe Guttmann, Lior Wolf, Daniel Cohen-Or |
ICCV | 2 |
| 2009 | The One-Shot similarity kernelabstractThe One-Shot similarity measure has recently been introduced in the context of face recognition where it was used to produce state-of-the-art results. Given two vectors, their One-Shot similarity score reflects the likelihood of each vector belonging in the same class as the other vector and not in a class defined by a fixed set of “negative” examples. The potential of this approach has thus far been largely unexplored. In this paper we analyze the One-Shot score and show that: (1) when using a version of LDA as the underlying classifier, this score is a Conditionally Positive Definite kernel and may be used within kernel-methods (e.g., SVM), (2) it can be efficiently computed, and (3) that it is effective as an underlying mechanism for image representation. We further demonstrate the effectiveness of the One-Shot similarity score in a number of applications including multiclass identification and descriptor generation. Lior Wolf, Tal Hassner, Yaniv Taigman |
ICCV | 1 |
| 2009 | Local Trinary Patterns for human action recognitionabstractWe present a novel action recognition method which is based on combining the effective description properties of Local Binary Patterns with the appearance invariance and adaptability of patch matching based methods. The resulting method is extremely efficient, and thus is suitable for real-time uses of simultaneous recovery of human action of several lengths and starting points. Tested on all publicity available datasets in the literature known to us, our system repeatedly achieves state of the art performance. Lastly, we present a new benchmark that focuses on uncut motion recognition in broadcast sports video. Lahav Yeffet, Lior Wolf |
ICCV | 2 |
| 2009 | Emerging imagesabstractEmergence refers to the unique human ability to aggregate information from seemingly meaningless pieces, and to perceive a whole that is meaningful. This special skill of humans can constitute an effective scheme to tell humans and machines apart. This paper presents a synthesis technique to generate images of 3D objects that are detectable by humans, but difficult for an automatic algorithm to recognize. The technique allows generating an infinite number of images with emerging figures. Our algorithm is designed so that locally the synthesized images divulge little useful information or cues to assist any segmentation or recognition procedure. Therefore, as we demonstrate, computer vision algorithms are incapable of effectively processing such images. However, when a human observer is presented with an emergence image, synthesized using an object she is familiar with, the figure emerges when observed as a whole. We can control the difficulty level of perceiving the emergence effect through a limited set of parameters. A procedure that synthesizes emergence images can be an effective tool for exploring and understanding the factors affecting computer vision techniques. Niloy J. Mitra, Hung-Kuo Chu, Tong-Yee Lee, Lior Wolf, Yehezkel Yeshurun, Daniel Cohen-Or |
ACM Trans. Graph. | 4 |
| 2008 | An experimental study of employing visual appearance as a phenotypeabstractVisual and non-visual data are often related through complex, indirect links, thus making the prediction of one from the other difficult. Examples include the partially- understood connections between firing of VI neurons and visual stimuli, the coupling between recorded speech and video of the corresponding lip movements, and the attempts to infer criminal intentions from surveillance videos. In this study, we explore the exploitation of the visual/non-visual relation between genetic sequences and visual appearance. This exploitation is currently considered infeasible due to the many hidden variables and unknown factors involved, the considerable variability and noise that exist in images and the high-dimensionality of the data. Despite the difficulties, we show convincing evidence that the application of correlations between genotype and visual phenotype for identification is feasible with current technologies. To this end, we employ sensitive forced- matching tests, that can accurately detect correlations between data sets. These tests are used to compare the performance of several existing algorithms, as well as novel ones that we have designed for the task. Lior Wolf, Yoni Donner |
CVPR | 1 |
| 2008 | Local Regularization for Multiclass Classification Facing Significant Intraclass Variations
Lior Wolf, Yoni Donner |
ECCV (4) | 1 |
| 2008 | Using Biologically Inspired Features for Face Processing
Ethan Meyers, Lior Wolf |
Int. J. Comput. Vis. | 2 |
| 2007 | Image representations beyond histograms of gradients: The role of Gestalt descriptorsabstractHistograms of orientations and the statistics derived from them have proven to be effective image representations for various recognition tasks. In this work we attempt to improve the accuracy of object detection systems by including new features that explicitly capture mid-level Gestalt concepts. Four new image features are proposed, inspired by the Gestalt principles of continuity, symmetry, closure and repetition. The resulting image representations are used jointly with existing state-of-the-art features and together enable better detectors for challenging real-world data sets. As baseline features, we use Riesenhuber and Poggio's C1 features and Dalan and Triggs' histogram of oriented gradients feature. Given that both of these baseline features have already shown state of the art performance in multiple object detection benchmarks, that our new mid-level representations can further improve detection results warrants special consideration. We evaluate the performance of these detection systems on the publicly available StreetScenes and Caltech101 databases among others. Stanley M. Bileschi, Lior Wolf |
CVPR | 2 |
| 2007 | Artificial Complex Cells via the Tropical SemiringabstractThe seminal work of Hubel and Wiesel (1965) and the vast amount of work that followed it prove that hierarchies of increasingly complex cells play a central role in cortical computations. Computational models, pioneered by Fukushima (1980), suggest that these hierarchies contain feature-building cells ("S-cells") and pooling cells ("C-cells"). More recently, Riesenhuber & Poggio have developed the HMAX model (1999), in which S-cells perform linear combinations, while C-cells perform a MAX operation. We note that methods for computing the connectivity of S-cells abound since many algorithms for suggesting informative linear combinations exist. There are, however, only few published methods that are suitable for the construction of C-cells. Here, we build a novel dimensionality reduction algorithm for learning the connectivity of C-cells, using the framework of the max-plus ("tropical") semiring. Lior Wolf, Moshe Guttmann |
CVPR | 1 |
| 2007 | Modeling Appearances with Low-Rank SVMabstractSeveral authors have noticed that the common representation of images as vectors is sub-optimal. The process of vectorization eliminates spatial relations between some of the nearby image measurements and produces a vector of a dimension which is the product of the measurements' dimensions. It seems that images may be better represented when taking into account their structure as a 2D (or multi-D) array. Our work bears similarities to recent work such as 2DPCA or Coupled Subspace Analysis in that we treat images as 2D arrays. The main difference, however, is that unlike previous work which separated representation from the discriminative learning stage, we achieve both by the same method. Our framework, "low-rank separators ", studies the use of a separating hyperplane which are constrained to have the structure of low-rank matrices. We first prove that the low-rank constraint provides preferable generalization properties. We then define two "low-rank SVM problems" and propose algorithms to solve these. Finally, we provide supporting experimental evidence for the framework. Lior Wolf, Hueihan Jhuang, Tamir Hazan |
CVPR | 1 |
| 2007 | A Biologically Inspired System for Action RecognitionabstractWe present a biologically-motivated system for the recognition of actions from video sequences. The approach builds on recent work on object recognition based on hierarchical feedforward architectures [25, 16, 20] and extends a neurobiological model of motion processing in the visual cortex [10]. The system consists of a hierarchy of spatio-temporal feature detectors of increasing complexity: an input sequence is first analyzed by an array of motion- direction sensitive units which, through a hierarchy of processing stages, lead to position-invariant spatio-temporal feature detectors. We experiment with different types of motion-direction sensitive units as well as different system architectures. As in [16], we find that sparse features in intermediate stages outperform dense ones and that using a simple feature selection approach leads to an efficient system that performs better with far fewer features. We test the approach on different publicly available action datasets, in all cases achieving the highest results reported to date. Hueihan Jhuang, Thomas Serre, Lior Wolf, Tomaso A. Poggio |
ICCV | 3 |
| 2007 | Non-homogeneous Content-driven Video-retargetingabstractVideo retargeting is the process of transforming an existing video to fit the dimensions of an arbitrary display. A compelling retargeting aims at preserving the viewers' experience by maintaining the information content of important regions in the frame, whilst keeping their aspect ratio. An efficient algorithm for video retargeting is introduced. It consists of two stages. First, the frame is analyzed to detect the importance of each region in the frame. Then, a transformation that respects the analysis shrinks less important regions more than important ones. Our analysis is fully automatic and based on local saliency, motion detection and object detectors. The performance of the proposed algorithm is demonstrated on a variety of video sequences, and compared to the state of the art in image retargeting. Lior Wolf, Moshe Guttmann, Daniel Cohen-Or |
ICCV | 1 |
| 2007 | Diorama Construction From a Single ImageabstractAbstract Diorama artists produce a spectacular 3D effect in a confined space by generating depth illusions that are faithful to the ordering of the objects in a large real or imaginary scene. Indeed, cognitive scientists have discovered that depth perception is mostly affected by depth order and precedence among objects. Motivated by these findings, we employ ordinal cues to construct a model from a single image that similarly to Dioramas, intensifies the depth perception. We demonstrate that such models are sufficient for the creation of realistic 3D visual experiences. The initial step of our technique extracts several relative depth cues that are well known to exist in the human visual system. Next, we integrate the resulting cues to create a coherent surface. We introduce wide slits in the surface, thus generalizing the concept of cardboard cutout layers. Lastly, the surface geometry and texture are extended alongside the slits, to allow small changes in the viewpoint which enriches the depth illusion. Jackie Assa, Lior Wolf |
Comput. Graph. Forum | 2 |
| 2007 | Robust Object Recognition with Cortex-Like MechanismsabstractWe introduce a new general framework for the recognition of complex visual scenes, which is motivated by biology: We describe a hierarchical system that closely follows the organization of visual cortex and builds an increasingly complex and invariant feature representation by alternating between a template matching and a maximum pooling operation. We demonstrate the strength of the approach on a range of recognition tasks: From invariant single object recognition in clutter to multiclass categorization problems and complex scene understanding tasks that rely on the recognition of both shape-based as well as texture-based objects. Given the biological constraints that the system had to satisfy, the approach performs surprisingly well: It has the capability of learning from only a few training examples and competes with state-of-the-art systems. We also discuss the existence of a universal, redundant dictionary of features that could handle the recognition of most object categories. In addition to its relevance for computer vision, the success of this approach suggests a plausibility proof for a class of feedforward models of object recognition in cortex. Thomas Serre, Lior Wolf, Stanley M. Bileschi, Maximilian Riesenhuber, Tomaso A. Poggio |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2006 | Perception Strategies in Hierarchical Vision SystemsabstractFlat appearance-based systems, which combine clever image representations with standard classifiers, might be the most effective way to recognize objects using current technologies. In the future, however, it seems probable that hierarchical representations might have better performance. In such systems, the image representation consists of a sequence of sets of features, where each subsequent set is computed based on the previous sets. The main contributions of this paper are to: (1) pose the question "what is the best way to employ discriminative methods for hierarchical image representations?"; (2) enumerate some of the alternative hierarchies while drawing connections to recent work by brain researchers; (3) study experimentally the different alternatives. As we will show, the strategy used can make a substantial difference. Lior Wolf, Stanley M. Bileschi, Ethan Meyers |
CVPR (2) | 1 |
| 2006 | Patch-Based Texture Edges and Segmentation
Lior Wolf, Sharon X. Huang, Ian Martin, Dimitris N. Metaxas |
ECCV (2) | 1 |
| 2006 | A Critical View of Context
Lior Wolf, Stanley M. Bileschi |
Int. J. Comput. Vis. | 1 |
| 2006 | Wide Baseline Matching between Unsynchronized Video Sequences
Lior Wolf, Assaf Zomet |
Int. J. Comput. Vis. | 1 |
| 2005 | A Unified System For Object Detection, Texture Recognition, and Context Analysis Based on the Standard Model Feature SetabstractRecently, a neuroscience inspired set of visual features was introduced. It was shown that this representation facilitates better performance than stateof-the-art vision systems for object recognition in cluttered and unsegmented images. In this paper, we investigate the utility of these features in other common scene-understanding tasks. We show that this outstanding performance extends to shape-based object detection in the usual windowing framework, to amorphous object detection as a texture classification task, and finally to context understanding These tasks are performed on a large set of images which were collected as a benchmark for the problem of scene understanding. The final system is able to reliably identify cars, pedestrians, bicycles, sky, road, buildings and trees in a diverse set of images. 1. Stanley M. Bileschi, Lior Wolf |
BMVC | 2 |
| 2005 | Object Recognition with Features Inspired by Visual CortexabstractWe introduce a novel set of features for robust object recognition. Each element of this set is a complex feature obtained by combining position- and scale-tolerant edge-detectors over neighboring positions and multiple orientations. Our system's architecture is motivated by a quantitative model of visual cortex. We show that our approach exhibits excellent recognition performance and outperforms several state-of-the-art systems on a variety of image datasets including many different object categories. We also demonstrate that our system is able to learn from very few examples. The performance of the approach constitutes a suggestive plausibility proof for a class of feedforward models of object recognition in cortex. Thomas Serre, Lior Wolf, Tomaso A. Poggio |
CVPR (2) | 2 |
| 2005 | Combining Variable Selection with Dimensionality ReductionabstractThis paper bridges the gap between variable selection methods (e.g., Pearson coefficients, KS test) and dimensionality reduction algorithms (e.g., PCA, LDA). Variable selection algorithms encounter difficulties dealing with highly correlated data, since many features are similar in quality. Dimensionality reduction algorithms tend to combine all variables and cannot select a subset of significant variables. Our approach combines both methodologies by applying variable selection followed by dimensionality reduction. This combination makes sense only when using the same utility function in both stages, which we do. The resulting algorithm benefits from complex features as variable selection algorithms do, and at the same time enjoys the benefits of dimensionality reduction. Lior Wolf, Stanley M. Bileschi |
CVPR (2) | 1 |
| 2005 | Robust Boosting for Learning from Few ExamplesabstractWe present and analyze a novel regularization technique based on enhancing our dataset with corrupted copies of our original data. The motivation is that since the learning algorithm lacks information about which parts of the data are reliable, it has to make more robust classification functions. Using this framework, we propose a simple addition to the gentle boosting algorithm which enables it to work with only a few examples. We test this new algorithm on a variety of datasets and show convincing results. Lior Wolf, Ian Martin |
CVPR (1) | 1 |
| 2005 | Feature Selection for Unsupervised and Supervised Inference: The Emergence of Sparsity in a Weight-Based ApproachabstractThe problem of selecting a subset of relevant features in a potentially overwhelming quantity of data is classic and found in many branches of science. Examples in computer vision, text processing and more recently bio-informatics are abundant. In text classification tasks, for example, it is not uncommon to have 104 to 107 features of the size of the vocabulary containing word frequency counts, with the expectation that only a small fraction of them are relevant. Typical examples include the automatic sorting of URLs into a web directory and the detection of spam email. In this work we present a definition of "relevancy" based on spectral properties of the Laplacian of the features' measurement matrix. The feature selection process is then based on a continuous ranking of the features defined by a least-squares optimization process. A remarkable property of the feature relevance function is that sparse solutions for the ranking values naturally emerge as a result of a "biased non-negativity" of a key matrix in the process. As a result, a simple least-squares optimization process converges onto a sparse solution, i.e., a selection of a subset of features which form a local maximum over the relevance function. The feature selection algorithm can be embedded in both unsupervised and supervised inference problems and empirical evidence show that the feature selections typically achieve high accuracy even when only a small fraction of the features are relevant. Lior Wolf, Amnon Shashua |
J. Mach. Learn. Res. | 1 |
| 2004 | Kernel Feature Selection with Side Data Using a Spectral Approach
Amnon Shashua, Lior Wolf |
ECCV (3) | 2 |
| 2003 | Kernel Principal Angles for Classification Machines with Applications to Image Sequence InterpretationabstractWe consider the problem of learning with instances defined over a space of sets of vectors. We derive a new positive definite kernel f(A, B) defined over pairs of matrices A, B based on the concept of principal angles between two linear subspaces. We show that the principal angles can be recovered using only inner-products between pairs of column vectors of the input matrices thereby allowing the original column vectors of A, B to be mapped onto arbitrarily high-dimensional feature spaces. We apply this technique to inference over image sequences applications of face recognition and irregular motion trajectory detection. Lior Wolf, Amnon Shashua |
CVPR (1) | 1 |
| 2003 | Feature Selection for Unsupervised and Supervised Inference: the Emergence of Sparsity in a Weighted-based ApproachabstractThe ability to reliably infer the nature of telephone conversations opens up a variety of applications, ranging from designing context-sensitive user interfaces on smartphones, to providing new tools for social psychologists and social scientists to study and understand social life of different subpopulations within different contexts. Using a unique corpus of everyday telephone conversations collected from eight residences over the duration of a year, we investigate the utility of popular features, extracted solely from the content, in classifying business-oriented calls from others. Through feature selection experiments, we find that the discrimination can be performed robustly for a majority of the calls using a small set of features. Remarkably, features learned from unsupervised methods, specifically latent Dirichlet allocation, perform almost as well as with as those from supervised methods. The unsupervised clusters learned in this task shows promise of finer grain inference of social nature of telephone conversations. Lior Wolf, Amnon Shashua |
ICCV | 1 |
| 2003 | Learning over Sets using Kernel Principal Angles
Lior Wolf, Amnon Shashua |
J. Mach. Learn. Res. | 1 |
| 2002 | Sequence-to-Sequence Self Calibration
Lior Wolf, Assaf Zomet |
ECCV (2) | 1 |
| 2002 | Video de-Abstraction or How to save money on your wedding videoabstractThere exist an increasing body of work dealing with video still abstraction, the extraction of representative still images from a video sequence. This work focuses in the other direction: given a video abstract and raw unedited video data, we produce an edited video. We focus on the application of generating wedding videos. We use the existing wedding photo album as an abstract, and produce an edited wedding video from it. The photo album serves us in determining importance of raw shots, as well as style and order. Aya Aner-Wolf, Lior Wolf |
WACV | 2 |
| 2002 | On Projection Matrices Pk-> P2k=3, ..., 6, and their Applications in Computer Vision
Lior Wolf, Amnon Shashua |
Int. J. Comput. Vis. | 1 |
| 2001 | Time-varying Shape Tensors for Scenes with Multiply Moving PointsabstractWe derive single view indexing functions for dynamic scenes - where dynamic is defined as a scene consisting of multiply moving points each moving independently with constant velocity. The indexing functions we derive are view independent and form a generalization of the "shape tensors" associated with rigid scenes by introducing a time-varying parameter We derive those indexing functions under full 3D projective, 3D affine, and various reduced configurations. The indexing functions were implemented and tested for matching against objects for which their non-rigid motion is an intrinsic part of their character - human gait recognition and hand gesture identification are the two chosen application examples. Anat Levin, Lior Wolf, Amnon Shashua |
CVPR (1) | 2 |
| 2001 | Two-body Segmentation from Two Perspective ViewsabstractWe consider a scene containing two independently and generally moving objects, viewed by two general perspective views. Using matching points arising from both objects simultaneously we derive a geometrical constraint, applicable to points from both objects, we call the segmentation matrix. We then use this constraint in order to recover the fundamental matrices associated with, each object, or simply to segment the scene into the two objects. Moreover, when the two bodies move in pure translation relative to each other we can both segment the scene and recover the affine calibration (homography at infinity) of the camera geometry. Unlike algorithms suggested in the past we need only two images, we work with general projective cameras (rather than affine or orthographic) and with general body motion, and no prior information beyond point matches is required. Lior Wolf, Amnon Shashua |
CVPR (1) | 1 |
| 2001 | On Projection Matrices and their Applications in Computer Vision
Lior Wolf, Amnon Shashua |
ICCV | 1 |
| 2001 | Affine 3-D Reconstruction from Two Projective Images of Independently Translating PlanesabstractConsider two views of a multi-body scene consisting of k planar bodies moving in pure translation one relative to the other. We show that the fundamental matrices, one per body, live in a 3-dimensional subspace, which when represented as a step-3 extensor is the common transversal on the collection of extensors defined by the homograph matrices H/sub 1/,...,H/sub k/ of the moving planes. We show that as much as five bodies are necessary for recovering the common transversal from the homograph matrices, from which we show how to recover the fundamental matrices and the affine calibration between the two cameras. Lior Wolf, Amnon Shashua |
ICCV | 1 |
| 2001 | Omni-Rig: Linear Self-Recalibration of a Rig with Varying Internal and External Parameters
Assaf Zomet, Lior Wolf, Amnon Shashua |
ICCV | 2 |
| 2000 | Homography Tensors: On Algebraic Entities that Represent Three Views of Static or Moving Planar Points
Amnon Shashua, Lior Wolf |
ECCV (1) | 2 |
| 2000 | On the Structure and Properties of the Quadrifocal Tensor
Amnon Shashua, Lior Wolf |
ECCV (1) | 2 |
| 2000 | Join Tensors: On 3D-to-3D Alignment of Dynamic SetsabstractIntroduces a family of 4/spl times/4/spl times/4 tensors, referred to as "join tensors" or Jtensors for short, which perform "3D to 3D" alignment between coordinate systems of sets of dynamic 3D points. 3D configurations of points are obtained by a 3D measuring device (such as a structured light or laser range sensor, or a stereo rig) at times t/sub 1/, t/sub 2/, t/sub 3/ from different viewing positions in addition to the motion of the sensor the points are also allowed to move in space; each point can move along an arbitrary straight-line path-we refer to this situation as "dynamic". The problem is to recover the motion of the sensor given the 3D correspondences of the points over time. We introduce Jtensors to capture the problem described above. Three observations P, P', P'' of a point measured at three time instants contribute a linear measurement to the Jtensor, regardless of whether the point has moved in space or has remained stationary while the sensor has changed position. Lior Wolf, Amnon Shashua, Yonatan Wexler |
ICPR | 1 |