EDBT 2026 Demo / reviewers in the wild / expert
Ivaylo Boyadzhiev
dblp:122/4817
· DBLP profile ↗
13ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0001-8555-2952ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | BADGR: Bundle Adjustment Diffusion Conditioned by Gradients for Wide-Baseline Floor Plan ReconstructionabstractReconstructing precise camera poses and floor plan layouts from wide-baseline RGB panoramas is a difficult and unsolved problem. We introduce BADGR, a novel diffusion model that jointly performs reconstruction and bundle adjustment (BA) to refine poses and layouts from a coarse state, using 1D floor boundary predictions from dozens of sparsely captured images. Unlike guided diffusion models, BADGR is conditioned on dense per-column outputs from a single-step Levenberg Marquardt (LM) optimizer and is trained to predict camera and wall positions, while minimizing reprojection errors for view consistency. The objective of layout generation from denoising diffusion process complements BA optimization by providing additional learned layout-structural constraints on top of the co-visible features across images. These constraints help BADGR make plausible guesses about spatial relationships, which constrain the pose graph, such as wall adjacency and collinearity, while also learning to mitigate errors from dense boundary observations using global context. BADGR trains exclusively on 2D floor plans, simplifying data acquisition, enabling robust augmentation, and supporting a variety of input densities. Our experiments validate our method, which significantly outperforms the state-of-the-art pose and floor plan layout reconstruction with different input densities. Visit project website at: https://badgr-diffusion.github.io. Yuguang Li, Ivaylo Boyadzhiev, Zixuan Liu 0001, Linda G. Shapiro, Alex Colburn |
CVPR | 2 |
| 2022 | PSMNet: Position-aware Stereo Merging Network for Room Layout EstimationabstractIn this paper, we propose a new deep learning-based method for estimating room layout given a pair of 360° panoramas. Our system, called Position-aware Stereo Merging Network or PSMNet, is an end-to-end joint layout-pose estimator. PSMNet consists of a Stereo Pano Pose (SP2) transformer and a novel Cross-Perspective Projection (CP2) layer. The stereo-view SP2 transformer is used to implicitly infer correspondences between views, and can handle noisy poses. The pose-aware CP2layer is designed to render features from the adjacent view to the anchor (reference) view, in order to perform view fusion and estimate the visible layout. Our experiments and analysis validate our method, which significantly outperforms the state-of-the-art layout estimators, especially for large and complex room spaces. Haiyan Wang 0019, Will Hutchcroft, Yuguang Li, Zhiqiang Wan, Ivaylo Boyadzhiev, Yingli Tian, Sing Bing Kang |
CVPR | 5 |
| 2022 | LASER: LAtent SpacE Rendering for 2D Visual LocalizationabstractWe present LASER, an image-based Monte Carlo Localization (MCL) framework for 2D floor maps. LASER introduces the concept of latent space rendering, where 2D pose hypotheses on the floor map are directly rendered into a geometrically-structured latent space by aggregating viewing ray features. Through a tightly coupled rendering codebook scheme, the viewing ray features are dynamically determined at rendering-time based on their geometries (i.e. length, incident-angle), endowing our representation with view-dependent fine-grain variability. Our codebook scheme effectively disentangles feature encoding from rendering, allowing the latent space rendering to run at speeds above 10KHz. Moreover, through metric learning, our geometrically-structured latent space is common to both pose hypotheses and query images with arbitrary field of views. As a result, LASER achieves state-of-the-art performance on large-scale indoor localization datasets (i. e. ZInD [5] and Structured3D [38]) for both panorama and perspective image queries, while significantly outperforming existing learning-based methods in speed. Zhixiang Min, Naji Khosravan, Zachary Bessinger, Manjunath Narayana, Sing Bing Kang, Enrique Dunn, Ivaylo Boyadzhiev |
CVPR | 7 |
| 2022 | CoVisPose: Co-visibility Pose Transformer for Wide-Baseline Relative Pose Estimation in 360$^\circ $ Indoor Panoramas
Will Hutchcroft, Yuguang Li, Ivaylo Boyadzhiev, Zhiqiang Wan, Haiyan Wang 0019, Sing Bing Kang |
ECCV (32) | 3 |
| 2022 | SALVe: Semantic Alignment Verification for Floorplan Reconstruction from Sparse Panoramas
John Lambert, Yuguang Li, Ivaylo Boyadzhiev, Lambert Wixson, Manjunath Narayana, Will Hutchcroft, James Hays, Frank Dellaert, Sing Bing Kang |
ECCV (31) | 3 |
| 2022 | Generating Topological Structure of Floorplans from Room AttributesabstractAnalysis of indoor spaces requires topological information. In this paper, we propose to extract topological information from room attributes using what we call Iterative and adaptive graph Topology Learning (ITL). ITL progressively predicts multiple relations between rooms; at each iteration, it improves node embeddings, which in turn facilitates the generation of a better topological graph structure. This notion of iterative improvement of node embeddings and topological graph structure is in the same spirit as [5]. However, while [5] computes the adjacency matrix based on node similarity, we learn the graph metric using a relational decoder to extract room correlations. Experiments using a new challenging indoor dataset validate our proposed method. Qualitative and quantitative evaluation for layout topology prediction and floorplan generation applications also demonstrate the effectiveness of ITL. Yu Yin 0001, Will Hutchcroft, Naji Khosravan, Ivaylo Boyadzhiev, Yun Fu 0001, Sing Bing Kang |
ICMR | 4 |
| 2022 | Semantically supervised appearance decomposition for virtual staging from a single panoramaabstractWe describe a novel approach to decompose a single panorama of an empty indoor environment into four appearance components: specular, direct sunlight, diffuse and diffuse ambient without direct sunlight. Our system is weakly supervised by automatically generated semantic maps (with floor, wall, ceiling, lamp, window and door labels) that have shown success on perspective views and are trained for panoramas using transfer learning without any further annotations. A GAN-based approach supervised by coarse information obtained from the semantic map extracts specular reflection and direct sunlight regions on the floor and walls. These lighting effects are removed via a similar GAN-based approach and a semantic-aware inpainting step. The appearance decomposition enables multiple applications including sun direction estimation, virtual furniture insertion, floor material replacement, and sun direction change, providing an effective tool for virtual home staging. We demonstrate the effectiveness of our approach on a large and recently released dataset of panoramas of empty homes. Tiancheng Zhi, Bowei Chen 0004, Ivaylo Boyadzhiev, Sing Bing Kang, Martial Hebert, Srinivasa G. Narasimhan |
ACM Trans. Graph. | 3 |
| 2021 | Zillow Indoor Dataset: Annotated Floor Plans With 360deg Panoramas and 3D Room LayoutsabstractWe present Zillow Indoor Dataset (ZInD): A large indoor dataset with 71,474 panoramas from 1,524 real unfurnished homes. ZInD provides annotations of 3D room layouts, 2D and 3D floor plans, panorama location in the floor plan, and locations of windows and doors. The ground truth construction took over 1,500 hours of annotation work. To the best of our knowledge, ZInD is the largest real dataset with layout annotations. A unique property is the room layout data, which follows a real world distribution (cuboid, more general Manhattan, and non-Manhattan layouts) as opposed to the mostly cuboid or Manhattan layouts in current publicly available datasets. Also, the scale and annotations provided are valuable for effective research related to room layout and floor plan analysis. To demonstrate ZInD’s benefits, we benchmark on room layout estimation from single panoramas and multi-view registration. Steve Cruz, Will Hutchcroft, Yuguang Li, Naji Khosravan, Ivaylo Boyadzhiev, Sing Bing Kang |
CVPR | 5 |
| 2016 | Do-it-yourself lighting design for product videographyabstractThe growth of online marketplaces for selling goods has increased the need for product photography by novice users and consumers. Additionally, the increased use of online media and large-screen billboards promotes the adoption of videos for advertising, going beyond just using still imagery. Lighting is a key distinction between professional and casual product videography. Professionals use specialized hardware setups, and bring expert skills to create good lighting that shows off the product's shape and material, while also producing aesthetically pleasing results. In this paper, we introduce a new do-it-yourself (DIY) approach to lighting design that lets novice users create studio quality product videography. We identify design principles to light products through emphasizing highlights, rim lighting, and contours. We devise a set of computational metrics to achieve these design goals. Our workflow is: the user acquires a video of the product by mounting a video camera on a tripod and using a tablet to light objects by waving the tablet around the object. We automatically analyze and split this acquired video into snippets that match our design principles. Finally, we present an interface that lets users easily select snippets with specific characteristics and then assembles them to produce a final pleasing video of the product. Alternatively, they can rely on our template mechanism to automatically assemble a video. Ivaylo Boyadzhiev, Jiawen Chen 0001, Sylvain Paris, Kavita Bala |
ICCP | 1 |
| 2015 | Talk abstract: Computational lighting design and band-sifting operatorsabstractIn this talk, I present two projects that are inspired by how photographers work. These projects are in collaboration with Ivo Boyadzhiev and Kavita Bala at Cornell University and Ted Adelson at MIT. Sylvain Paris, Ivaylo Boyadzhiev, Kavita Bala, Edward H. Adelson |
ICIP | 2 |
| 2015 | Band-Sifting Decomposition for Image-Based Material EditingabstractPhotographers often “prep” their subjects to achieve various effects; for example, toning down overly shiny skin, covering blotches, etc. Making such adjustments digitally after a shoot is possible, but difficult without good tools and good skills. Making such adjustments to video footage is harder still. We describe and study a set of 2D image operations, based on multiscale image analysis, that are easy and straightforward and that can consistently modify perceived material properties. These operators first build a subband decomposition of the image and then selectively modify the coefficients within the subbands. We call this selection process band sifting . We show that different siftings of the coefficients can be used to modify the appearance of properties such as gloss, smoothness, pigmentation, or weathering. The band-sifting operators have particularly striking effects when applied to faces; they can provide “knobs” to make a face look wetter or drier, younger or older, and with heavy or light variation in pigmentation. Through user studies, we identify a set of operators that yield consistent subjective effects for a variety of materials and scenes. We demonstrate that these operators are also useful for processing video sequences. Ivaylo Boyadzhiev, Kavita Bala, Sylvain Paris, Edward H. Adelson |
ACM Trans. Graph. | 1 |
| 2013 | User-assisted image compositing for photographic lightingabstractGood lighting is crucial in photography and can make the difference between a great picture and a discarded image. Traditionally, professional photographers work in a studio with many light sources carefully set up, with the goal of getting a near-final image at exposure time, with post-processing mostly focusing on aspects orthogonal to lighting. Recently, a new workflow has emerged for architectural and commercial photography, where photographers capture several photos from a fixed viewpoint with a moving light source. The objective is not to produce the final result immediately, but rather to capture useful data that are later processed, often significantly, in photo editing software to create the final well-lit image. This new workflow is flexible, requires less manual setup, and works well for time-constrained shots. But dealing with several tens of unorganized layers is painstaking, requiring hours to days of manual effort, as well as advanced photo editing skills. Our objective in this paper is to make the compositing step easier. We describe a set of optimizations to assemble the input images to create a fewbasis lightsthat correspond to common goals pursued by photographers, e.g., accentuating edges and curved regions. We also introducemodifiersthat capture standard photographic tasks, e.g., to alter the lights to soften highlights and shadows, akin to umbrellas and soft boxes. Our experiments with novice and professional users show that our approach allows them to quickly create satisfying results, whereas working with unorganized images requires considerably more time. Casual users particularly benefit from our approach since coping with a large number of layers is daunting for them and requires significant experience. Ivaylo Boyadzhiev, Sylvain Paris, Kavita Bala |
ACM Trans. Graph. | 1 |
| 2012 | User-guided white balance for mixed lighting conditionsabstractProper white balance is essential in photographs to eliminate color casts due to illumination. The single-light case is hard to solve automatically but relatively easy for humans. Unfortunately, many scenes contain multiple light sources such as an indoor scene with a window, or when a flash is used in a tungsten-lit room. The light color can then vary on a per-pixel basis and the problem becomes challenging at best, even with advanced image editing tools. We propose a solution to the ill-posed mixed light white balance problem, based on user guidance. Users scribble on a few regions that should have the same color, indicate one or more regions of neutral color, and select regions where the current color looks correct. We first expand the provided scribble groups to more regions using pixel similarity and a robust voting scheme. We formulate the spatially varying white balance problem as a sparse data interpolation problem in which the user scribbles and their extensions form constraints. We demonstrate that our approach can produce satisfying results on a variety of scenes with intuitive scribbles and without any knowledge about the lights. Ivaylo Boyadzhiev, Kavita Bala, Sylvain Paris, Frédo Durand |
ACM Trans. Graph. | 1 |