Dov Danon

dblp:161/3307 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
2since 2021 · last 2021
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
3D vision · 91% Generative modeling · 9%
Computer graphics and multimedia
1 paper
Image and video processing · 100%
Human-computer interaction and pervasive computing
1 paper
Interaction techniques and input · 100%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 9 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Image and video processing › image fusion
multi-modal image fusion
0.512021
Single Pair Cross-Modality Super Resolution · CVPR 2021
Image and video processing
super-resolution
0.512021
Single Pair Cross-Modality Super Resolution · CVPR 2021
Computer vision › 3D vision › image registration
cross-modal registration
0.412020
Unsupervised Multi-Modal Image Registration via Geometry Preserving Image-to-Image Translation · CVPR 2020
Computer vision › 3D vision
image registration
0.412020
Unsupervised Multi-Modal Image Registration via Geometry Preserving Image-to-Image Translation · CVPR 2020
Computer vision › 3D vision
jigsaw puzzle solving
0.412020
Solving Jigsaw Puzzles With Eroded Boundaries · CVPR 2020
Interaction techniques and input › spatial interaction › navigation
image browsing
0.212015
DynamicMaps: Similarity-based Browsing through a Massive Set of Images · CHI 2015
Machine learning › Generative modeling › generative adversarial network
image-to-image translation
0.112020
Unsupervised Multi-Modal Image Registration via Geometry Preserving Image-to-Image Translation · CVPR 2020
Information retrieval
relevance feedback
0.112015
DynamicMaps: Similarity-based Browsing through a Massive Set of Images · CHI 2015
Information retrieval
retrieval models
0.112015
DynamicMaps: Similarity-based Browsing through a Massive Set of Images · CHI 2015

Methods — techniques the papers use, named apart from their topics

self-supervised learning · 0.5internal transformer · 0.5nearest neighbor search · 0.4spatial transformation network · 0.4inpainting · 0.4image-to-image translation network · 0.4geometry-preserving translation · 0.4generative adversarial network · 0.4multidimensional embedding · 0.2multi-dimensional embedding · 0.2
YearPublicationVenuePosition
2021 Single Pair Cross-Modality Super Resolution
abstract
Non-visual imaging sensors are widely used in the industry for different purposes. Those sensors are more expensive than visual (RGB) sensors, and usually produce images with lower resolution. To this end, Cross-Modality Super-Resolution methods were introduced, where an RGB image of a high-resolution assists in increasing the resolution of a low-resolution modality. However, fusing images from different modalities is not a trivial task, since each multi-modal pair varies greatly in its internal correlations. For this reason, traditional state-of-the-arts which are trained on external datasets often struggle with yielding an artifact-free result that is still loyal to the target modality characteristics.We present CMSR, a single-pair approach for Cross-Modality Super-Resolution. The network is internally trained on the two input images only, in a self-supervised manner, learns their internal statistics and correlations, and applies them to up-sample the target modality. CMSR contains an internal transformer which is trained on-the-fly together with the up-sampling process itself and without supervision, to allow dealing with pairs that are only weakly aligned. We show that CMSR produces state-of-the-art super resolved images, yet without introducing artifacts or irrelevant details that originate from the RGB image only.
Guy Shacht, Dov Danon, Sharon Fogel, Daniel Cohen-Or
CVPR2
2021 Image resizing by reconstruction from deep features
abstract
Traditional image resizing methods usually work in pixel space and use various saliency measures. The challenge is to adjust the image shape while trying to preserve important content. In this paper we perform image resizing in feature space using the deep layers of a neural network containing rich important semantic information. We directly adjust the image feature maps, extracted from a pre-trained classification network, and reconstruct the resized image using neural-network based optimization. This novel approach leverages the hierarchical encoding of the network, and in particular, the high-level discriminative power of its deeper layers, that can recognize semantic regions and objects, thereby allowing maintenance of their aspect ratios. Our use of reconstruction from deep features results in less noticeable artifacts than use of imagespace resizing operators. We evaluate our method on benchmarks, compare it to alternative approaches, and demonstrate its strengths on challenging images.
Dov Danon, Moab Arar, Daniel Cohen-Or, Ariel Shamir
Comput. Vis. Media1
2020 Unsupervised Multi-Modal Image Registration via Geometry Preserving Image-to-Image Translation
abstract
Many applications, such as autonomous driving, heavily rely on multi-modal data where spatial alignment between the modalities is required. Most multi-modal registration methods struggle computing the spatial correspondence between the images using prevalent cross-modality similarity measures. In this work, we bypass the difficulties of developing cross-modality similarity measures, by training an image-to-image translation network on the two input modalities. This learned translation allows training the registration network using simple and reliable mono-modality metrics. We perform multi-modal registration using two networks - a spatial transformation network and a translation network. We show that by encouraging our translation network to be geometry preserving, we manage to train an accurate spatial transformation network. Compared to state-of-the-art multi-modal methods our presented method is unsupervised, requiring no pairs of aligned modalities for training, and can be adapted to any pair of modalities. We evaluate our method quantitatively and qualitatively on commercial datasets, showing that it performs well on several modalities and achieves accurate alignment.
Moab Arar, Yiftach Ginger, Dov Danon, Amit Bermano, Daniel Cohen-Or
CVPR3
2020 Solving Jigsaw Puzzles With Eroded Boundaries
abstract
Jigsaw puzzle solving is an intriguing problem which has been explored in computer vision for decades. This paper focuses on a specific variant of the problem-solving puzzles with eroded boundaries. Such erosion makes the problem extremely difficult, since most existing solvers utilize solely the information at the boundaries. Nevertheless, this variant is important since erosion and missing data often occur at the boundaries. The key idea of our proposed approach is to inpaint the eroded boundaries between puzzle pieces and later leverage the quality of the inpainted area to classify a pair of pieces as ”neighbors or not”. An interesting feature of our architecture is that the same GAN discriminator is used for both inpainting and classification; training of the second task is simply a continuation of the training of the first, beginning from the point it left off. We show that our approach outperforms other SOTA methods.
Dov Bridger, Dov Danon, Ayellet Tal
CVPR2
2020 Implicit pairs for boosting unpaired image-to-image translation
abstract
In image-to-image translation the goal is to learn a mapping from one image domain to another. In the case of supervised approaches the mapping is learned from paired samples. However, collecting large sets of image pairs is often either prohibitively expensive or not possible. As a result, in recent years more attention has been given to techniques that learn the mapping from unpaired sets. In our work, we show that injecting implicit pairs into unpaired sets strengthens the mapping between the two domains, improves the compatibility of their distributions, and leads to performance boosting of unsupervised techniques by up to 12% across several measurements. The competence of the implicit pairs is further displayed with the use of pseudo-pairs, i.e., paired samples which only approximate a real pair. We demonstrate the effect of the approximated implicit samples on image-to-image translation problems, where such pseudo-pairs may be synthesized in one direction, but not in the other. We further show that pseudo-pairs are significantly more effective as implicit pairs in an unpaired setting, than directly using them explicitly in a paired setting.
Yiftach Ginger, Dov Danon, Hadar Averbuch-Elor, Daniel Cohen-Or
Vis. Informatics2
2019 Unsupervised natural image patch learning
abstract
A metric for natural image patches is an important tool for analyzing images. An efficient means of learning one is to train a deep network to map an image patch to a vector space, in which the Euclidean distance reflects patch similarity. Previous attempts learned such an embedding in a supervised manner, requiring the availability of many annotated images. In this paper, we present an unsupervised embedding of natural image patches, avoiding the need for annotated images. The key idea is that the similarity of two patches can be learned from the prevalence of their spatial proximity in natural images. Clearly, relying on this simple principle, many spatially nearby pairs are outliers. However, as we show, these outliers do not harm the convergence of the metric learning. We show that our unsupervised embedding approach is more effective than a supervised one or one that uses deep patch representations. Moreover, we show that it naturally lends itself to an efficient self-supervised domain adaptation technique onto a target domain that contains a common foreground object.
Dov Danon, Hadar Averbuch-Elor, Ohad Fried, Daniel Cohen-Or
Comput. Vis. Media1
2015 DynamicMaps: Similarity-based Browsing through a Massive Set of Images
abstract
We present a novel system for browsing through a very large set of images according to similarity. The images are dynamically placed on a 2D canvas next to their nearest neighbors in a high-dimensional feature space. The layout and choice of images is generated on-the-fly during user interaction, reflecting the user's navigation tendencies and interests. This intuitive solution for image browsing provides a continuous experience of navigating through an infinite 2D grid arranged by similarity. In contrast to common multidimensional embedding methods, our solution does not entail an upfront creation of a full global map. Image map generation is dynamic, fast and scalable, independent of the number of images in the dataset, and seamlessly supports online updates to the dataset. Thus, the technique is a viable solution for massive and constantly varying datasets consisting of millions of images. Evaluation of our approach shows that when using DynamicMaps, users viewed many more images per minute compared to a standard relevance feedback interface, suggesting that it supports more fluid and natural interaction that enables easier and faster movement in the image space. Most users preferred DynamicMaps, indicating it is more exploratory, better supports serendipitous browsing and more fun to use
Yanir Kleiman, Joel Lanir, Dov Danon, Yasmin Felberbaum, Daniel Cohen-Or
CHI3