EDBT 2026 Demo / reviewers in the wild / expert
SaiKiran Kumar Tedla
dblp:376/7961
· DBLP profile ↗
8ranked-venue papers
6as first author
8since 2021 · last 2025
0000-0002-2679-2881ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Examining Joint Demosaicing and Denoising for Single-, Quad-, and Nona-Bayer PatternsabstractCamera sensors have color filters arranged in a mosaic layout, traditionally following the Bayer pattern. Demosaicing is a critical step camera hardware applies to obtain a full-channel RGB image. Many smartphones now have multiple sensors with different patterns, such as Quad-Bayer or Nona-Bayer. Most modern deep network-based models perform joint demosaicing and denoising with the strategy of training a separate network per pattern. Relying on individual models per pattern requires additional memory overhead and makes it challenging to switch quickly between cameras. In this work, we are interested in analyzing strategies for joint demosaicing and denoising for the three main mosaic layouts ($1 \times 1$ Single-Bayer, $2 \times 2$ Quad-Bayer, and $3 \times 3$ Nona-Bayer). We found concatenating a three-channel mosaic embedding to the input image and training a unified demosaicing architecture yields results that outperform existing Quad-Bayer and Nona-Bayer models and are comparable to Single-Bayer models. Additionally, we describe a maskout strategy that enhances the model performance and facilitates dead pixel correction-a step often overlooked by existing Al-based demosaicing models. As part of this effort, we captured a new demosaicing dataset of 638 RAW images that contain challenging scenes with patches annotated for training, validation, and testing. Code and data is available at https://github.com/SamsungLabs/unified-demosaicing. SaiKiran Kumar Tedla, Abhijith Punnappurath, Luxi Zhao 0002, Michael S. Brown |
ICCP | 1 |
| 2025 | Gain-MLP: Improving HDR Gain Map Encoding via a Lightweight MLP
Trevor D. Canham, SaiKiran Kumar Tedla, Michael Murdoch, Michael S. Brown |
ICCV | 2 |
| 2025 | Multispectral Demosaicing via Dual CamerasabstractMultispectral (MS) images capture detailed scene information across a wide range of spectral bands, making them invaluable for applications requiring rich spectral data. Integrating MS imaging into multi camera devices, such as smartphones, has the potential to enhance both spectral applications and RGB image quality. A critical step in processing MS data is demosaicing, which reconstructs color information from the mosaic MS images captured by the camera. This paper proposes a method for MS image demosaicing specifically designed for dual-camera setups where both RGB and MS cameras capture the same scene. Our approach leverages co-captured RGB images, which typically have higher spatial fidelity, to guide the demosaicing of lower-fidelity MS images. We introduce the Dual-camera RGB-MS Dataset - a large collection of paired RGB and MS mosaiced images with ground-truth demosaiced outputs - that enables training and evaluation of our method. Experimental results demonstrate that our method achieves state-of-the-art accuracy compared to existing techniques. SaiKiran Kumar Tedla, Junyong Lee 0001, Beixuan Yang, Mahmoud Afifi, Michael S. Brown |
ICCV | 1 |
| 2025 | Learning to Refocus with Video Diffusion ModelsabstractFocus is a cornerstone of photography, yet autofocus systems often fail to capture the intended subject, and users frequently wish to adjust focus after capture. We introduce a novel method for realistic post-capture refocusing using video diffusion models. From a single defocused image, our approach generates a perceptually accurate focal stack, represented as a video sequence, enabling interactive refocusing and unlocking a range of downstream applications. We release a large-scale focal stack dataset acquired under diverse real-world smartphone conditions to support this work and future research. Our method consistently outperforms existing approaches in both perceptual quality and robustness across challenging scenarios, paving the way for more advanced focus-editing capabilities in everyday photography. Code and data are available at www.learn2refocus.github.io SaiKiran Kumar Tedla, Zhoutong Zhang, Xuaner Cecilia Zhang, Shumian Xin |
SIGGRAPH Asia | 1 |
| 2025 | Generating the Past, Present and Future from a Motion-Blurred ImageabstractWe seek to answer the question: what can a motion-blurred image reveal about a scene's past, present, and future? Although motion blur obscures image details and degrades visual quality, it also encodes information about scene and camera motion during an exposure. Previous techniques leverage this information to estimate a sharp image from an input blurry one, or to predict a sequence of video frames showing what might have occurred at the moment of image capture. However, they rely on handcrafted priors or network architectures to resolve ambiguities in this inverse problem, and do not incorporate image and video priors on large-scale datasets. As such, existing methods struggle to reproduce complex scene dynamics and do not attempt to recover what occurred before or after an image was taken. Here, we introduce a new technique that repurposes a pre-trained video diffusion model trained on internet-scale datasets to recover videos revealing complex scene dynamics during the moment of capture and what might have occurred immediately into the past or future. Our approach is robust and versatile; it outperforms previous methods for this task, generalizes to challenging in-the-wild images, and supports downstream tasks such as recovering camera trajectories, object motion, and dynamic 3D scene structure. Code and data are available at blur2vid.github.io SaiKiran Kumar Tedla, Kelly Zhu, Trevor D. Canham, Felix Taubner, Michael S. Brown, Kiriakos N. Kutulakos, David B. Lindell |
ACM Trans. Graph. | 1 |
| 2024 | LookToFocus: Image Focus via Eye TrackingabstractWe present LookToFocus, a method to perform real-time manual camera focus based on eye tracking. LookToFocus and two alternative methods for manual focus1 photography tasks were compared in a user study. A novel manual focus camera simulation was used to test the methods. The first two methods were LookToFocus and LookToFocusNB (no bounding box). The third method, TapToFocus, used touch for manual focus and image capture, analogous to typical smartphone interaction. LookToFocusNB had the fastest mean capture time at 1429 ms; the mean capture times were 1431 ms for LookToFocusNB and 2416 ms for TapToFocus. When compared to TapToFocus, LookToFocus and LookToFocusNB had faster capture times because these algorithms start converging to the optimal focus as soon as the user spots the target. LookToFocus and LookToFocusNB also had no significant decrease in sharpness error compared to TapToFocus. Users preferred LookToFocus over both LookToFocusNB and TapToFocus. SaiKiran Kumar Tedla, I. Scott MacKenzie, Michael S. Brown |
ETRA | 1 |
| 2023 | Graphics2RAW: Mapping Computer Graphics Images to Sensor RAW ImagesabstractComputer graphics (CG) rendering platforms produce imagery with ever-increasing photo realism. The narrowing domain gap between real and synthetic imagery makes it possible to use CG images as training data for deep learning models targeting high-level computer vision tasks, such as autonomous driving and semantic segmentation. CG images, however, are currently not suitable for low-level vision tasks targeting RAW sensor images. This is because RAW images are encoded in sensor-specific color spaces and incur pre-white-balance color casts caused by the sensor’s response to scene illumination. CG images are rendered directly to a device-independent perceptual color space without needing white balancing. As a result, it is necessary to apply a mapping procedure to close the domain gap between graphics and RAW images. To this end, we introduce a framework to process graphics images to mimic RAW sensor images accurately. Our approach allows a one-to-many mapping, where a single graphics image can be transformed to match multiple sensors and multiple scene illuminations. In addition, our approach requires only a handful of example RAW-DNG files from the target sensor as parameters for the mapping process. We compare our method to alternative strategies and show that our approach produces more realistic RAW images and provides better results on three low-level vision tasks: RAW denoising, illumination estimation, and neural rendering for night photography. Finally, as part of this work, we provide a dataset of 292 realistic CG images for training low-light imaging models. Donghwan Seo, Abhijith Punnappurath, Luxi Zhao 0002, Abdelrahman Abdelhamed, SaiKiran Kumar Tedla, Jihwan Choe, Michael S. Brown |
ICCV | 5 |
| 2023 | Examining Autoexposure for Challenging ScenesabstractAutoexposure (AE) is a critical step applied by camera systems to ensure properly exposed images. While current AE algorithms are effective in well-lit environments with constant illumination, these algorithms still struggle in environments with bright light sources or scenes with abrupt changes in lighting. A significant hurdle in developing new AE algorithms for challenging environments, especially those with time-varying lighting, is the lack of suitable image datasets. To address this issue, we have captured a new 4D exposure dataset that provides a large solution space (i.e., shutter speed range from $\frac{1}{{1500}}$ to 15 seconds) over a temporal sequence with moving objects, bright lights, and varying lighting. In addition, we have designed a software platform to allow AE algorithms to be used in a plug-and-play manner with the dataset. Our dataset and associate platform enable repeatable evaluation of different AE algorithms and provide a much-needed starting point to develop better AE methods. We examine several existing AE strategies using our dataset and show that most users prefer a simple saliency method for challenging lighting conditions. SaiKiran Kumar Tedla, Beixuan Yang, Michael S. Brown |
ICCV | 1 |