VLDB 2026 Research / reviewers in the wild / expert
Daniel Glasner
dblp:28/1971
· DBLP profile ↗
14ranked-venue papers
6as first author
3since 2021 · last 2024
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 3 since 2021Artificial intelligence and machine learning · 10 · 5 first-author · 3 since 2021Theory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Generative modeling · 39% Probabilistic and Bayesian machine learning · 27% 3D vision · 10% | |
| Computer graphics and multimedia
7 papers |
Image and video processing · 52% Rendering · 26% Computational photography and imaging · 20% |
Topics — the 23 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling › diffusion model
text-to-image generation |
1.0 | 2 | 2024 | MarkovGen: Structured Prediction for Efficient Text-to-Image Generation · CVPR 2024 Rethinking FID: Towards a Better Evaluation Metric for Image Generation · CVPR 2024 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
distribution distance estimation |
0.8 | 1 | 2024 | Rethinking FID: Towards a Better Evaluation Metric for Image Generation · CVPR 2024 |
Machine learning › Generative modeling › generative model evaluation
image generation evaluation |
0.8 | 1 | 2024 | Rethinking FID: Towards a Better Evaluation Metric for Image Generation · CVPR 2024 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
markov random field |
0.8 | 1 | 2024 | MarkovGen: Structured Prediction for Efficient Text-to-Image Generation · CVPR 2024 |
Computer vision › Image recognition and object detection
image classification |
0.5 | 1 | 2021 | Understanding Robustness of Transformers for Image Classification · ICCV 2021 |
Machine learning › Trustworthy machine learning
robustness |
0.5 | 1 | 2021 | Understanding Robustness of Transformers for Image Classification · ICCV 2021 |
Computer vision › 3D vision › stereo vision
stereo matching |
0.3 | 1 | 2017 | Toward Perceptually-Consistent Stereo: A Scanline Study · ICCV 2017 |
Image and video processing
stereo vision |
0.3 | 1 | 2017 | Toward Perceptually-Consistent Stereo: A Scanline Study · ICCV 2017 |
Image and video processing › super-resolution › image super-resolution
single image super-resolution |
0.3 | 2 | 2013 | Accurate Blur Models vs. Image Priors in Single Image Super-resolution · ICCV 2013 Super-resolution from a single image · ICCV 2009 |
Machine learning › Generative modeling
diffusion model |
0.2 | 1 | 2024 | MarkovGen: Structured Prediction for Efficient Text-to-Image Generation · CVPR 2024 |
Machine learning › Generative modeling › diffusion model › diffusion model acceleration
sampling acceleration |
0.2 | 1 | 2024 | MarkovGen: Structured Prediction for Efficient Text-to-Image Generation · CVPR 2024 |
Image and video processing › super-resolution
image super-resolution |
0.2 | 1 | 2013 | Accurate Blur Models vs. Image Priors in Single Image Super-resolution · ICCV 2013 |
Rendering › appearance modeling
reflectance and appearance modeling |
0.2 | 1 | 2013 | Fabricating BRDFs at high spatial resolution using wave optics · ACM Trans. Graph. 2013 |
Rendering › bidirectional reflectance distribution function
spatially-varying BRDF |
0.2 | 1 | 2013 | Fabricating BRDFs at high spatial resolution using wave optics · ACM Trans. Graph. 2013 |
Machine learning › Deep learning architectures and training
transformer |
0.1 | 1 | 2021 | Understanding Robustness of Transformers for Image Classification · ICCV 2021 |
Machine learning › Deep learning architectures and training › transformer
vision transformer |
0.1 | 1 | 2021 | Understanding Robustness of Transformers for Image Classification · ICCV 2021 |
Computer vision › 3D vision › object pose estimation
object detection and pose estimation |
0.1 | 1 | 2011 | Viewpoint-aware object detection and pose estimation · ICCV 2011 |
Computer vision › 3D vision
object pose estimation |
0.1 | 1 | 2011 | Viewpoint-aware object detection and pose estimation · ICCV 2011 |
Image and video processing
image segmentation |
0.1 | 1 | 2011 | Contour-based joint clustering of multiple segmentations · CVPR 2011 |
Image and video processing
image restoration |
0.1 | 1 | 2009 | Super-resolution from a single image · ICCV 2009 |
Image and video processing
super-resolution |
0.1 | 1 | 2009 | Super-resolution from a single image · ICCV 2009 |
Computer vision › Segmentation and scene understanding
scene understanding |
0.1 | 1 | 2015 | Hot or Not: Exploring Correlations between Appearance and Temperature · ICCV 2015 |
Geometric modeling and processing › shape representation › 2d shape representation
contour representation |
0.0 | 1 | 2011 | Contour-based joint clustering of multiple segmentations · CVPR 2011 |
Methods — techniques the papers use, named apart from their topics
maximum mean discrepancy · 0.8markov random field · 0.8differentiable inference layer · 0.8backpropagation · 0.8Gaussian RBF kernel · 0.8CLIP embedding · 0.8energy minimization · 0.6empirical robustness study · 0.5statistical correlation analysis · 0.4convolutional neural network · 0.4wave optics · 0.2phase modulation · 0.2liquid crystal spatial light modulator · 0.2wave optics analysis · 0.2photolithography · 0.2gradient regularization · 0.2blur kernel estimation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | MarkovGen: Structured Prediction for Efficient Text-to-Image GenerationabstractModern text-to-image generation models produce high-quality images that are both photorealistic and faithful to the text prompts. However, this quality comes at significant computational cost: nearly all of these models are iterative and require running sampling multiple times with large models. This iterative process is needed to ensure that different regions of the image are not only aligned with the text prompt, but also compatible with each other. In this work, we propose a light-weight approach to achieving this compatibility between different regions of an image, using a Markov Random Field (MRF) model. We demonstrate the effectiveness of this method on top of the latent token-based Muse text-to-image model. The MRF richly encodes the compatibility among image tokens at different spatial locations to improve quality and significantly reduce the required number of Muse sampling steps. Inference with the MRF is significantly cheaper, and its parameters can be quickly learned through back-propagation by modeling MRF inference as a differentiable neural-network layer. Our full model, MarkovGen, uses this proposed MRF model to both speed up Muse by 1.5 x and produce higher quality images by decreasing undesirable image artifacts. Sadeep Jayasumana, Daniel Glasner, Srikumar Ramalingam, Andreas Veit, Ayan Chakrabarti, Sanjiv Kumar |
CVPR | 2 |
| 2024 | Rethinking FID: Towards a Better Evaluation Metric for Image GenerationabstractAs with many machine learning problems, the progress of image generation methods hinges on good evaluation metrics. One of the most popular is the Frechet Inception Distance (FID). FID estimates the distance between a distribution of Inception-v3 features of real images, and those of images generated by the algorithm. We highlight important drawbacks of FID: Inception's poor representation of the rich and varied content generated by modern text-to-image models, incorrect normality assumptions, and poor sample complexity. We call for a reevaluation of FID's use as the primary quality metric for generated images. We empirically demonstrate that FID contradicts human raters, it does not reflect gradual improvement of iterative text-to-image models, it does not capture distortion levels, and that it produces inconsistent results when varying the sample size. We also propose an alternative new metric, CMMD, based on richer CLIP embeddings and the maximum mean discrepancy distance with the Gaussian RBF kernel. It is an unbiased estimator that does not make any assumptions on the probability distribution of the embeddings and is sample efficient. Through extensive experiments and analysis, we demonstrate that FID-based evaluations of text-to-image models may be unreliable, and that CMMD of-fers a more robust and reliable assessment of image quality. A reference implementation of CMMD is available at: https://github.com/google-research/google-research/tree/master/cmmd. Sadeep Jayasumana, Srikumar Ramalingam, Andreas Veit, Daniel Glasner, Ayan Chakrabarti, Sanjiv Kumar |
CVPR | 4 |
| 2021 | Understanding Robustness of Transformers for Image ClassificationabstractDeep Convolutional Neural Networks (CNNs) have long been the architecture of choice for computer vision tasks. Recently, Transformer-based architectures like Vision Transformer (ViT) have matched or even surpassed ResNets for image classification. However, details of the Transformer architecture –such as the use of non-overlapping patches– lead one to wonder whether these networks are as robust. In this paper, we perform an extensive study of a variety of different measures of robustness of ViT models and compare the findings to ResNet baselines. We investigate robustness to input perturbations as well as robustness to model perturbations. We find that when pre-trained with a sufficient amount of data, ViT models are at least as robust as the ResNet counterparts on a broad range of perturbations. We also find that Transformers are robust to the removal of almost any single layer, and that while activations from later layers are highly correlated with each other, they nevertheless play an important role in classification. Srinadh Bhojanapalli, Ayan Chakrabarti, Daniel Glasner, Daliang Li, Thomas Unterthiner, Andreas Veit |
ICCV | 3 |
| 2017 | Toward Perceptually-Consistent Stereo: A Scanline StudyabstractTwo types of information exist in a stereo pair: correlation (matching) and decorrelation (half-occlusion). Vision science has shown that both types of information are used in the visual cortex, and that people can perceive depth even when correlation cues are absent or very weak, a capability that remains absent from most computational stereo systems. As a step toward stereo algorithms that are more consistent with these perceptual phenomena, we re-examine the topic of scanline stereo as energy minimization. We represent a disparity profile as a piecewise smooth function with explicit breakpoints between its smooth pieces, and we show this allows correlation and decorrelation to be integrated into an objective that requires only two types of local information: the correlation and its spatial gradient. Experimentally, we show the global optimum of this objective matches human perception on a broad collection of wellknown perceptual stimuli, and that it also provides reasonable piecewise-smooth interpretations of depth in natural images, even without exploiting monocular boundary cues. Jialiang Wang 0001, Daniel Glasner, Todd E. Zickler |
ICCV | 2 |
| 2015 | Hot or Not: Exploring Correlations between Appearance and TemperatureabstractIn this paper we explore interactions between the appearance of an outdoor scene and the ambient temperature. By studying statistical correlations between image sequences from outdoor cameras and temperature measurements we identify two interesting interactions. First, semantically meaningful regions such as foliage and reflective oriented surfaces are often highly indicative of the temperature. Second, small camera motions are correlated with the temperature in some scenes. We propose simple scene-specific temperature prediction algorithms which can be used to turn a camera into a crude temperature sensor. We find that for this task, simple features such as local pixel intensities outperform sophisticated, global features such as from a semantically-trained convolutional neural network. Daniel Glasner, Pascal Fua, Todd E. Zickler, Lihi Zelnik-Manor |
ICCV | 1 |
| 2015 | A Global Approach for Solving Edge-Matching PuzzlesabstractWe consider apictorial edge-matching puzzles, in which the goal is to arrange a collection of puzzle pieces with colored edges so that the colors match along the edges of adjacent pieces. We devise an algebraic representation for this problem and provide conditions under which it exactly characterizes a puzzle. Using the new representation, we recast the combinatorial, discrete problem of solving puzzles as a global, polynomial system of equations with continuous variables. We further propose new algorithms for generating approximate solutions to the continuous problem by solving a sequence of convex relaxations. Shahar Z. Kovalsky, Daniel Glasner, Ronen Basri |
SIAM J. Imaging Sci. | 2 |
| 2014 | A reflectance displayabstractWe present a reflectance display: a dynamic digital display capable of showing images and videos with spatially-varying, user-defined reflectance functions. Our display is passive: it operates by phase-modulation of reflected light. As such, it does not rely on any illumination recording sensors, nor does it require expensive on-the-fly rendering. It reacts to lighting changes instantaneously and consumes only a minimal amount of energy. Our work builds on the wave optics approach to BRDF fabrication of Levin et al. shortciteLevinBRDFFab13. We replace their expensive one-time hardware fabrication with a programable liquid crystal spatial light modulator, retaining high resolution of approximately 160 dpi. Our approach enables the display of a much wider family of angular reflectances, and it allows the display of dynamic content with time varying reflectance properties---"reflectance videos". To facilitate these new capabilities we develop novel reflectance design algorithms with improved resolution tradeoffs. We demonstrate the utility of our display with a diverse set of experiments including display of custom reflectance images and videos, interactive reflectance editing, display of 3D content reproducing lighting and depth variation, and simultaneous display of two independent channels on one screen. Daniel Glasner, Todd E. Zickler, Anat Levin |
ACM Trans. Graph. | 1 |
| 2013 | Accurate Blur Models vs. Image Priors in Single Image Super-resolutionabstractOver the past decade, single image Super-Resolution (SR) research has focused on developing sophisticated image priors, leading to significant advances. Estimating and incorporating the blur model, that relates the high-res and low-res images, has received much less attention, however. In particular, the reconstruction constraint, namely that the blurred and down sampled high-res output should approximately equal the low-res input image, has been either ignored or applied with default fixed blur models. In this work, we examine the relative importance of the image prior and the reconstruction constraint. First, we show that an accurate reconstruction constraint combined with a simple gradient regularization achieves SR results almost as good as those of state-of-the-art algorithms with sophisticated image priors. Second, we study both empirically and theoretically the sensitivity of SR algorithms to the blur model assumed in the reconstruction constraint. We find that an accurate blur model is more important than a sophisticated image prior. Finally, using real camera data, we demonstrate that the default blur models of various SR algorithms may differ from the camera blur, typically leading to over-smoothed results. Our findings highlight the importance of accurately estimating camera blur in reconstructing raw lowers images acquired by an actual camera. Netalee Efrat, Daniel Glasner, Alexander Apartsin, Boaz Nadler, Anat Levin |
ICCV | 2 |
| 2013 | Fabricating BRDFs at high spatial resolution using wave opticsabstractRecent attempts to fabricate surfaces with custom reflectance functions boast impressive angular resolution, yet their spatial resolution is limited. In this paper we present a method to construct spatially varying reflectance at a high resolution of up to 220dpi, orders of magnitude greater than previous attempts, albeit with a lower angular resolution. The resolution of previous approaches is limited by the machining, but more fundamentally, by the geometric optics model on which they are built. Beyond a certain scale geometric optics models break down and wave effects must be taken into account. We present an analysis of incoherent reflectance based on wave optics and gain important insights into reflectance design. We further suggest and demonstrate a practical method, which takes into account the limitations of existing micro-fabrication techniques such as photolithography to design and fabricate a range of reflection effects, based on wave interference. Anat Levin, Daniel Glasner, Frédo Durand, William T. Freeman, Wojciech Matusik, Todd E. Zickler |
ACM Trans. Graph. | 2 |
| 2012 | Viewpoint-aware object detection and continuous pose estimation
Daniel Glasner, Meirav Galun, Sharon Alpert, Ronen Basri, Gregory Shakhnarovich |
Image Vis. Comput. | 1 |
| 2011 | Contour-based joint clustering of multiple segmentationsabstractWe present an unsupervised, shape-based method for joint clustering of multiple image segmentations. Given two or more closely-related images, such as nearby frames in a video sequence or images of the same scene taken under different lighting conditions, our method generates a joint segmentation of the images. We introduce a novel contour-based representation that allows us to cast the shape-based joint clustering problem as a quadratic semi-assignment problem. Our score function is additive. We use complex-valued affinities to assess the quality of matching the edge elements at the exterior bounding contour of clusters, while ignoring the contributions of elements that fall in the interior of the clusters. We further combine this contour-based score with region information and use a linear programming relaxation to solve for the joint clusters. We evaluate our approach on the occlusion boundary data-set of Stein et al. Daniel Glasner, Shiv Vitaladevuni, Ronen Basri |
CVPR | 1 |
| 2011 | Viewpoint-aware object detection and pose estimationabstractWe describe an approach to category-level detection and viewpoint estimation for rigid 3D objects from single 2D images. In contrast to many existing methods, we directly integrate 3D reasoning with an appearance-based voting architecture. Our method relies on a nonparametric representation of a joint distribution of shape and appearance of the object class. Our voting method employs a novel parametrization of joint detection and viewpoint hypothesis space, allowing efficient accumulation of evidence. We combine this with a re-scoring and refinement mechanism, using an ensemble of view-specific Support Vector Machines. We evaluate the performance of our approach in detection and pose estimation of cars on a number of benchmark datasets. Daniel Glasner, Meirav Galun, Sharon Alpert, Ronen Basri, Gregory Shakhnarovich |
ICCV | 1 |
| 2010 | A Preemptive Algorithm for Maximizing Disjoint Paths on Trees
Yossi Azar, Uriel Feige, Daniel Glasner |
Algorithmica | 3 |
| 2009 | Super-resolution from a single imageabstractMethods for super-resolution can be broadly classified into two families of methods: (i) The classical multi-image super-resolution (combining images obtained at subpixel misalignments), and (ii) Example-Based super-resolution (learning correspondence between low and high resolution image patches from a database). In this paper we propose a unified framework for combining these two families of methods. We further show how this combined approach can be applied to obtain super resolution from as little as a single image (with no database or prior examples). Our approach is based on the observation that patches in a natural image tend to redundantly recur many times inside the image, both within the same scale, as well as across different scales. Recurrence of patches within the same image scale (at subpixel misalignments) gives rise to the classical super-resolution, whereas recurrence of patches across different scales of the same image gives rise to example-based super-resolution. Our approach attempts to recover at each pixel its best possible resolution increase based on its patch redundancy within and across scales. Daniel Glasner, Shai Bagon, Michal Irani |
ICCV | 1 |