EDBT 2026 Demo / reviewers in the wild / expert
Roni Sengupta
dblp:359/9890 · also Soumyadip Sengupta
· DBLP profile ↗
39ranked-venue papers
11as first author
24since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 6 first-author · 20 since 2021Artificial intelligence and machine learning · 25 · 9 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TalkingHeadBench: A Multi-Modal Benchmark & Analysis of Talking-Head DeepFake DetectionabstractThe rapid advancement of talking-head deepfake generation fueled by advanced generative models has elevated the realism of synthetic videos to a level that poses substantial risks in domains such as media, politics, and finance. However, current benchmarks for deepfake talking-head detection fail to reflect this progress, relying on outdated generators and offering limited insight into model robustness and generalization. We introduce TalkingHeadBench, a new benchmark designed to address this gap, featuring talking-head videos from six modern generators, with an additional two emerging generators used exclusively for testing generalization. The dataset is built on an expert-led curation process that filters over 60% of samples to remove videos with noticeable artifacts, presenting a more difficult challenge for detectors. Our evaluation protocols are designed to measure generalization across identity and generator shifts. Benchmarking seven state-of-the-art detectors reveals that models with high accuracy on older datasets like FaceForensics++ show a significant performance drop on our curated data, particularly at strict false positive rates (e.g., TPR@FPR=0.1%). In addition, we identify a trend where detectors focus on background cues instead of facial features using Grad-CAM visualization. Our benchmark aims to accelerate research towards more robust and generalizable detection models in the face of rapidly evolving generative techniques. We release our benchmark and dataset with all data splits and protocols at https://anaxqx.github.io/talkingheadbench.github.io. Xinqi Xiong, Prakrut Patel, Qingyuan Fan, Amisha Wadhwa, Sarathy Selvam, Luchao Qi, Roni Sengupta |
WACV | 9 |
| 2025 | ScribbleLight: Single Image Indoor Relighting with ScribblesabstractImage-based relighting of indoor rooms creates an immersive virtual understanding of the space, which is useful for interior design, virtual staging, and real estate. Relighting indoor rooms from a single image is especially challenging due to complex illumination interactions between multiple lights and cluttered objects featuring a large variety in geometrical and material complexity. Recently, generative models have been successfully applied to imagebased relighting conditioned on a target image or a latent code, albeit without detailed local lighting control. In this paper, we introduce ScribbleLight, a generative model that supports local fine-grained control of lighting effects through scribbles that describe changes in lighting. Our key technical novelty is an Albedo-conditioned Stable Image Diffusion model that preserves the intrinsic color and texture of the original image after relighting and an encoder-decoder-based ControlNet architecture that enables geometry-preserving lighting effects with normal map and scribble annotations. We demonstrate ScribbleLight’s ability to create different lighting effects (e.g., turning lights on/off, adding highlights, cast shadows, or indirect lighting from unseen lights) from sparse scribble annotations. Jun Myeong Choi, Annie Wang, Pieter Peers, Anand Bhattad, Roni Sengupta |
CVPR | 5 |
| 2025 | NFL-BA: Near-Field Light Bundle Adjustment for SLAM in Dynamic LightingabstractSimultaneous Localization and Mapping (SLAM) systems typically assume static, distant illumination; however, many real-world scenarios, such as endoscopy, subterranean robotics, and search & rescue in collapsed environments, require agents to operate with a co-located light and camera in the absence of external lighting. In such cases, dynamic near-field lighting introduces strong, view-dependent shading that significantly degrades SLAM performance. We introduce Near-Field Lighting Bundle Adjustment Loss (NFL-BA) which explicitly models near-field lighting as a part of Bundle Adjustment loss and enables better performance for scenes captured with dynamic lighting. NFL-BA can be integrated into neural rendering-based SLAM systems with implicit or explicit scene representations. Our evaluations mainly focus on endoscopy procedure where SLAM can enable autonomous navigation, guidance to unsurveyed regions, blindspot detections, and 3D visualizations, which can significantly improve patient outcomes and endoscopy experience for both physicians and patients. Replacing Photometric Bundle Adjustment loss of SLAM systems with NFL-BA leads to significant improvement in camera tracking, 37% for MonoGS and 14% for EndoGSLAM, and leads to state-of-the-art camera tracking and mapping performance on the C3VD colonoscopy dataset. Further evaluation on indoor scenes captured with phone camera with flashlight turned on, also demonstrate significant improvement in SLAM performance due to NFL-BA. Andrea Dunn Beltran, Daniel Rho, Marc Niethammer, Roni Sengupta |
NeurIPS | 4 |
| 2025 | The Aging Multiverse: Generating Condition-Aware Facial Aging Tree via Training-Free DiffusionabstractWe introduce the Aging Multiverse, a framework for generating multiple plausible facial aging trajectories from a single image, each conditioned on external factors such as environment, health, and lifestyle. Unlike prior methods that model aging as a single deterministic path, our approach creates an aging tree that visualizes diverse futures. To enable this, we propose a training-free diffusion-based method that balances identity preservation, age accuracy, and condition control. Our key contributions include attention mixing to modulate editing strength and a Simulated Aging Regularization strategy to stabilize edits. Extensive experiments and user studies demonstrate state-of-the-art performance across identity preservation, aging realism, and conditional alignment, outperforming existing editing and age-progression models, which often fail to account for one or more of the editing criteria. By transforming aging into a multi-dimensional, controllable, and interpretable process, our approach opens up new creative and practical avenues in digital storytelling, health education, and personalized visualization. Bang Gong, Luchao Qi, Jiaye Wu 0001, Zhicheng Fu, Chunbo Song, John W. Nicholson 0001, Roni Sengupta |
SIGGRAPH Asia | 7 |
| 2025 | My3DGen: A Scalable Personalized 3D Generative ModelabstractIn recent years, generative 3D face models (e.g., EG3D) have been developed to tackle the problem of synthesizing photo-realistic faces. However, these models are often unable to capture facial features unique to each individual, highlighting the importance of personalization. Some prior works have shown promise in personalizing generative face models, but these studies primarily focus on 2D settings. Also, these methods require both fine-tuning and storing a large number of parameters for each user, posing a hindrance to achieving scalable personalization. Another challenge of personalization is the limited number of training images available for each individual, which often leads to overfitting when using full fine-tuning methods. Our proposed approach, My3DGen, generates a personalized 3D prior of an individual using as few as 50 training images. My3DGen allows for novel view synthesis, semantic editing of a given face (e.g. adding a smile), and synthesizing novel appearances, all while preserving the original person's identity. We decouple the 3D facial features into global features and personalized features by freezing the pre-trained EG3D and training additional personalized weights through low-rank decomposition. As a result, My3DGen introduces only 240K personalized parameters per individual, leading to a 127× reduction in trainable parameters compared to the 30.6M required for fine-tuning the entire parameter space. Despite this significant reduction in storage, our model preserves identity features without compromising the quality of downstream applications, both quantitatively and qualitatively. Luchao Qi, Jiaye Wu 0001, Annie N. Wang, Shengze Wang 0002, Roni Sengupta |
WACV | 5 |
| 2025 | Continual Learning of Personalized Generative Face Models with Experience ReplayabstractWe introduce a novel continual learning problem: how to sequentially update the weights of a personalized 2D and 3D generative face model as new batches of photos in different appearances, styles, poses, and lighting are captured regularly. We observe that naive sequential fine-tuning of the model leads to catastrophic forgetting of past representations of the individual's face. We then demonstrate that a simple random sampling-based experience replay method is effective at mitigating catastrophic forgetting when a relatively large number of images can be stored and replayed. However, for long-term deployment of these models with relatively smaller storage, this simple random samplingbased replay technique also forgets past representations. Thus, we introduce a novel experience replay algorithm that combines random sampling with StyleGAN's latent space to represent the buffer as an optimal convex hull. We observe that our proposed convex hull-based experience replay is more effective in preventing forgetting than a random sampling baseline and the lower bound. We introduce continual learning datasets for five celebrities, along with the evaluation framework, metrics, and visualizations to examine this problem. See our project page for more details. Annie N. Wang, Luchao Qi, Roni Sengupta |
WACV | 3 |
| 2025 | MyTimeMachine: Personalized Facial Age TransformationabstractFacial aging is a complex process, highly dependent on multiple factors like gender, ethnicity, lifestyle, etc., making it extremely challenging to learn a global aging prior to predict aging for any individual accurately. Existing techniques often produce realistic and plausible aging results, but the re-aged images often do not resemble the person's appearance at the target age and thus need personalization. In many practical applications of virtual aging, e.g. VFX in movies and TV shows, access to a personal photo collection of the user depicting aging in a small time interval (20~40 years) is often available. However, naive attempts to personalize global aging techniques on personal photo collections often fail. Thus, we propose MyTimeMachine (MyTM), a method that combines a global aging prior with a personalized photo collection (ranging from as few as 10 images, ideally 50) to learn individualized age transformations. We introduce a novel Adapter Network that combines personalized aging features with global aging features and generates a re-aged image with StyleGAN2. We also introduce three loss functions to personalize the Adapter Network with personalized aging loss, extrapolation regularization, and adaptive w-norm regularization. Our method demonstrates strong performance on fair-use imagery of widely recognizable individuals, producing photorealistic and identity-consistent age transformations that generalize well across diverse appearances. It also extends naturally to video, delivering high-quality, temporally consistent results that closely resemble actual appearances at target ages—outperforming state-of-the-art approaches. Luchao Qi, Jiaye Wu 0001, Bang Gong, Annie N. Wang, David Jacobs 0001, Roni Sengupta |
ACM Trans. Graph. | 6 |
| 2025 | Learning View Synthesis for Desktop Telepresence With Few RGBD CamerasabstractRecent telepresence systems have shown significant improvements in quality compared to prior systems. However, they struggle to achieve both low cost and high quality at the same time. In this work, we envision a future where telepresence systems become a commodity and can be installed on typical desktops. To this end, we present a high-quality view synthesis method that uses a cost-effective capture system that consists of commodity hardware accessible to the general public. We propose a neural renderer that uses a few RGBD cameras as input to synthesize novel views of a user and their surroundings. At the core of the renderer is Multi-Layer Point Cloud (MPC), a novel 3D representation that improves reconstruction accuracy by removing non-linear biases in depth cameras. Our temporally-aware renderer further improves the stability of synthesized videos by conditioning on past information. Additionally, we propose Spatial Skip Connections (SSC) to improve image upsampling under limited GPU memory. Experimental results show that our renderer outperforms recent methods in terms of view synthesis quality. Our method generalizes to new users and challenging content (e.g. hand gestures and clothing deformation) without costly per-video optimization, object templates, or heavy pre-processing. The code and dataset will be made available. Shengze Wang 0002, Ryan Schmelzle, Liujie Zheng, Youngjoong Kwon, Roni Sengupta, Henry Fuchs |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2024 | Personalized Video Relighting With an At-Home Light Stage
Jun Myeong Choi, Max Christman, Roni Sengupta |
ECCV (40) | 3 |
| 2024 | Leveraging Near-Field Lighting for Monocular Depth Estimation from Endoscopy Videos
Akshay Paruchuri, Samuel Ehrenstein, Inbar Fried, Stephen M. Pizer, Marc Niethammer, Roni Sengupta |
ECCV (32) | 7 |
| 2024 | NePhi: Neural Deformation Fields for Approximately Diffeomorphic Medical Image Registration
Lin Tian 0001, Thomas Hastings Greer, Raúl San José Estépar, Roni Sengupta, Marc Niethammer |
ECCV (88) | 4 |
| 2024 | Universal Guidance for Diffusion ModelsabstractTypical diffusion models are trained to accept a particular form of conditioning, most commonly text, and cannot be conditioned on other modalities without retraining. In this work, we propose a universal guidance algorithm that enables diffusion models to be controlled by arbitrary guidance modalities without the need to retrain any use-specific components. We show that our algorithm successfully generates quality images with guidance functions including segmentation, face recognition, object detection, style guidance and classifier signals. Arpit Bansal, Hong-Min Chu, Avi Schwarzschild, Roni Sengupta, Micah Goldblum, Jonas Geiping, Tom Goldstein |
ICLR | 4 |
| 2024 | Structure-Preserving Image Translation for Depth Estimation in Colonoscopy
Akshay Paruchuri, Sarah McGill, Roni Sengupta |
MICCAI (11) | 5 |
| 2024 | Motion Matters: Neural Motion Transfer for Better Camera Physiological MeasurementabstractMachine learning models for camera-based physiological measurement can have weak generalization due to a lack of representative training data. Body motion is one of the most significant sources of noise when attempting to recover the subtle cardiac pulse from a video. We explore motion transfer as a form of data augmentation to introduce motion variation while preserving physiological changes of interest. We adapt a neural video synthesis approach to augment videos for the task of remote photoplethysmography (rPPG) and study the effects of motion augmentation with respect to 1) the magnitude and 2) the type of motion. After training on motion-augmented versions of publicly available datasets, we demonstrate a 47% improvement over existing inter-dataset results using various state-of-the-art methods on the PURE dataset. We also present inter-dataset results on five benchmark datasets to show improvements of up to 79% using TS-CAN, a neural rPPG estimation method. Our findings illustrate the usefulness of motion transfer as a data augmentation technique for improving the generalization of models for camera-based physiological sensing. We release our code for using motion transfer as a data augmentation technique on three publicly available datasets, UBFC-rPPG, PURE, and SCAMPS, and models pre-trained on motion-augmented data here: https://motion-matters.github.io/ Akshay Paruchuri, Xin Liu 0034, Yulu Pan, Shwetak N. Patel, Daniel McDuff, Roni Sengupta |
WACV | 6 |
| 2024 | Joint Depth Prediction and Semantic Segmentation with Multi-View SAMabstractMulti-task approaches to joint depth and segmentation prediction are well-studied for monocular images. Yet, predictions from a single-view are inherently limited, while multiple views are available in many robotics applications. On the other end of the spectrum, video-based and full 3D methods require numerous frames to perform reconstruction and segmentation. With this work we propose a Multi-View Stereo (MVS) technique for depth prediction that benefits from rich semantic features of the Segment Anything Model (SAM). This enhanced depth prediction, in turn, serves as a prompt to our Transformer-based semantic segmentation decoder. We report the mutual benefit that both tasks enjoy in our quantitative and qualitative studies on the ScanNet dataset. Our approach consistently outperforms single-task MVS and segmentation models, along with multi-task monocular methods. Mykhailo Shvets, Dongxu Zhao 0001, Marc Niethammer, Roni Sengupta, Alexander C. Berg |
WACV | 4 |
| 2023 | Measured Albedo in the Wild: Filling the Gap in Intrinsics EvaluationabstractIntrinsic image decomposition and inverse rendering are long-standing problems in computer vision. To evaluate albedo recovery, most algorithms report their quantitative performance with a mean Weighted Human Disagreement Rate (WHDR) metric on the IIW dataset. However, WHDR focuses only on relative albedo values and often fails to capture overall quality of the albedo. In order to comprehensively evaluate albedo, we collect a new dataset, Measured Albedo in the Wild (MAW), and propose three new metrics that complement WHDR: intensity, chromaticity and texture metrics. We show that existing algorithms often improve WHDR metric but perform poorly on other metrics. We then finetune different algorithms on our MAW dataset to significantly improve the quality of the reconstructed albedo both quantitatively and qualitatively. Since the proposed intensity, chromaticity, and texture metrics and the WHDR are all complementary we further introduce a relative performance measure that captures average performance. By analysing existing algorithms we show that there is significant room for improvement. Our dataset and evaluation metrics will enable researchers to develop algorithms that improve albedo reconstruction. Jiaye Wu 0001, Sanjoy Chowdhury, Hariharmano Shanmugaraja, David Jacobs 0001, Roni Sengupta |
ICCP | 5 |
| 2023 | MVPSNet: Fast Generalizable Multi-view Photometric StereoabstractWe propose a fast and generalizable solution to Multiview Photometric Stereo (MVPS), called MVPSNet. The key to our approach is a feature extraction network that effectively combines images from the same view captured under multiple lighting conditions to extract geometric features from shading cues for stereo matching. We demonstrate these features, termed ‘Light Aggregated Feature Maps’ (LAFM), are effective for feature matching even in textureless regions, where traditional multi-view stereo methods often fail. Our method produces similar reconstruction results to PS-NeRF, a state-of-the-art MVPS method that optimizes a neural network per-scene, while being 411× faster (105 seconds vs. 12 hours) in inference. Additionally, we introduce a new synthetic dataset for MVPS, sMVPS, which is shown to be effective for training a generalizable MVPS method. Dongxu Zhao 0001, Daniel Lichy, Pierre-Nicolas Perrin, Jan-Michael Frahm, Roni Sengupta |
ICCV | 5 |
| 2023 | rPPG-Toolbox: Deep Remote PPG ToolboxabstractCamera-based physiological measurement is a fast growing field of computer vision. Remote photoplethysmography (rPPG) utilizes imaging devices (e.g., cameras) to measure the peripheral blood volume pulse (BVP) via photoplethysmography, and enables cardiac measurement via webcams and smartphones. However, the task is non-trivial with important pre-processing, modeling and post-processing steps required to obtain state-of-the-art results. Replication of results and benchmarking of new models is critical for scientific progress; however, as with many other applications of deep learning, reliable codebases are not easy to find or use. We present a comprehensive toolbox, rPPG-Toolbox, unsupervised and supervised rPPG models with support for public benchmark datasets, data augmentation and systematic evaluation: https://github.com/ubicomplab/rPPG-Toolbox. Xin Liu 0034, Girish Narayanswamy, Akshay Paruchuri, Jiankai Tang, Roni Sengupta, Shwetak N. Patel, Yuntao Wang 0001, Daniel McDuff |
NeurIPS | 7 |
| 2022 | Fast Light-Weight Near-Field Photometric StereoabstractWe introduce the first end-to-end learning-based solution to near-field Photometric Stereo (PS), where the light sources are close to the object of interest. This setup is especially useful for reconstructing large immobile objects. Our method is fast, producing a mesh from 52 512x384 resolution images in about 1 second on a commodity GPU, thus potentially unlocking several AR/VR applications. Existing approaches rely on optimization coupled with a far-field PS network operating on pixels or small patches. Using optimization makes these approaches slow and memory intensive (requiring 17GB GPU and 27GB of CPU memory) while using only pixels or patches makes them highly sus-ceptible to noise and calibration errors. To address these issues, we develop a recursive multi-resolution scheme to estimate surface normal and depth maps of the whole image at each step. The predicted depth map at each scale is then used to estimate 'per-pixel lighting, for the next scale. This design makes our approach almost 45x faster and 2° more accurate (11.3° vs. 13.3° Mean Angular Error) than the state-of-the-art near-field PS reconstruction technique, which uses iterative optimization. Daniel Lichy, Roni Sengupta, David Jacobs 0001 |
CVPR | 2 |
| 2022 | Robust High-Resolution Video Matting with Temporal GuidanceabstractWe introduce a robust, real-time, high-resolution human video matting method that achieves new state-of-the-art performance. Our method is much lighter than previous approaches and can process 4K at 76 FPS and HD at 104 FPS on an Nvidia GTX 1080Ti GPU. Unlike most existing methods that perform video matting frame-by-frame as independent images, our method uses a recurrent architecture to exploit temporal information in videos and achieves significant improvements in temporal coherence and matting quality. Furthermore, we propose a novel training strategy that enforces our network on both matting and segmentation objectives. This significantly improves our model’s robustness. Our method does not require any auxiliary inputs such as a trimap or a pre-captured background image, so it can be widely applied to existing human matting applications. Our code is available at https://peterl1n.github.io/RobustVideoMatting/ Shanchuan Lin, Imran Saleemi, Roni Sengupta |
WACV | 4 |
| 2022 | SfSNet: Learning Shape, Reflectance and Illuminance of Faces in the WildabstractWe present SfSNet, an end-to-end learning framework for producing an accurate decomposition of an unconstrained human face image into shape, reflectance and illuminance. SfSNet is designed to reflect a physical lambertian rendering model. SfSNet learns from a mixture of labeled synthetic and unlabeled real-world images. This allows the network to capture low-frequency variations from synthetic and high-frequency details from real images through the photometric reconstruction loss. SfSNet consists of a new decomposition architecture with residual blocks that learns a complete separation of albedo and normal. This is used along with the original image to predict lighting. SfSNet produces significantly better quantitative and qualitative results than state-of-the-art methods for inverse rendering and independent normal and illumination estimation. We also introduce a companion network, SfSMesh, that utilizes normals estimated by SfSNet to reconstruct a 3D face mesh. We demonstrate that SfSMesh produces face meshes with greater accuracy than state-of-the-art methods on real-world images. Roni Sengupta, Daniel Lichy, Angjoo Kanazawa, Carlos Domingo Castillo, David Jacobs 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Shape and Material Capture at HomeabstractIn this paper, we present a technique for estimating the geometry and reflectance of objects using only a camera, flashlight, and optionally a tripod. We propose a simple data capture technique in which the user goes around the object, illuminating it with a flashlight and capturing only a few images. Our main technical contribution is the introduction of a recursive neural architecture, which can predict geometry and reflectance at 2k×2kresolution given an input image at 2k×2kand estimated geometry and reflectance from the previous step at 2k−1×2k−1. This recursive architecture, termed RecNet, is trained with 256×256 resolution but can easily operate on 1024×1024 images during inference. We show that our method produces more accurate surface normal and albedo, especially in regions of specular highlights and cast shadows, compared to previous approaches, given three or fewer input images. Daniel Lichy, Jiaye Wu 0001, Roni Sengupta, David Jacobs 0001 |
CVPR | 3 |
| 2021 | Real-Time High-Resolution Background MattingabstractWe introduce a real-time, high-resolution background replacement technique which operates at 30fps in 4K resolution, and 60fps for HD on a modern GPU. Our technique is based on background matting, where an additional frame of the background is captured and used in recovering the alpha matte and the foreground layer. The main challenge is to compute a high-quality alpha matte, preserving strand-level hair details, while processing high-resolution images in real-time. To achieve this goal, we employ two neural networks; a base network computes a low-resolution result which is refined by a second network operating at high-resolution on selective patches. We introduce two large-scale video and image matting datasets: VideoMatte240K and PhotoMatte13K/85. Our approach yields higher quality results compared to the previous state-of-the-art in background matting, while simultaneously yielding a dramatic boost in both speed and resolution. Shanchuan Lin, Andrey Ryabtsev, Roni Sengupta, Brian Curless, Steven M. Seitz, Ira Kemelmacher-Shlizerman |
CVPR | 3 |
| 2021 | A Light Stage on Every DeskabstractEvery time you sit in front of a TV or monitor, your face is actively illuminated by time-varying patterns of light. This paper proposes to use this time-varying illumination for synthetic relighting of your face with any new illumination condition. In doing so, we take inspiration from the light stage work of Debevec et al. [4], who first demonstrated the ability to relight people captured in a controlled lighting environment. Whereas existing light stages require expensive, room-scale spherical capture gantries and exist in only a few labs in the world, we demonstrate how to acquire useful data from a normal TV or desktop monitor. Instead of subjecting the user to uncomfortable rapidly flashing light patterns, we operate on images of the user watching a YouTube video or other standard content. We train a deep network on images plus monitor patterns of a given user and learn to predict images of that user under any target illumination (monitor pattern). Experimental evaluation shows that our method produces realistic relighting results. Video results are available at grail.cs.washington.edu/projects/Light_Stage_on_Every_Desk/. Roni Sengupta, Brian Curless, Ira Kemelmacher-Shlizerman, Steven M. Seitz |
ICCV | 1 |
| 2020 | Background Matting: The World Is Your Green ScreenabstractWe propose a method for creating a matte - the per-pixel foreground color and alpha - of a person by taking photos or videos in an everyday setting with a handheld camera. Most existing matting methods require a green screen background or a manually created trimap to produce a good matte. Automatic, trimap-free methods are appearing, but are not of comparable quality. In our trimap free approach, we ask the user to take an additional photo of the background without the subject at the time of capture. This step requires a small amount of foresight but is far less timeconsuming than creating a trimap. We train a deep network with an adversarial loss to predict the matte. We first train a matting network with a supervised loss on ground truth data with synthetic composites. To bridge the domain gap to real imagery with no labeling, we train another matting network guided by the first network and by a discriminator that judges the quality of composites. We demonstrate results on a wide variety of photos and videos and show significant improvement over the state of the art. Roni Sengupta, Vivek Jayaram, Brian Curless, Steven M. Seitz, Ira Kemelmacher-Shlizerman |
CVPR | 1 |
| 2020 | Lifespan Age Transformation Synthesis
Roy Or-El, Roni Sengupta, Ohad Fried, Eli Shechtman, Ira Kemelmacher-Shlizerman |
ECCV (6) | 2 |
| 2019 | Neural Inverse Rendering of an Indoor Scene From a Single ImageabstractInverse rendering aims to estimate physical attributes of a scene, e.g., reflectance, geometry, and lighting, from image(s). Inverse rendering has been studied primarily for single objects or with methods that solve for only one of the scene attributes. We propose the first learning based approach that jointly estimates albedo, normals, and lighting of an indoor scene from a single image. Our key contribution is the Residual Appearance Renderer (RAR), which can be trained to synthesize complex appearance effects (e.g., inter-reflection, cast shadows, near-field illumination, and realistic shading), which would be neglected otherwise. This enables us to perform self-supervised learning on real data using a reconstruction loss, based on re-synthesizing the input image from the estimated components. We finetune with real data after pretraining with synthetic data. To this end, we use physically-based rendering to create a large-scale synthetic dataset, named SUNCG-PBR, which is a significant improvement over prior datasets. Experimental results show that our approach outperforms state-of-the-art methods that estimate one or more scene attributes. Roni Sengupta, Jinwei Gu, Guilin Liu, David Jacobs 0001, Jan Kautz |
ICCV | 1 |
| 2018 | SfSNet: Learning Shape, Reflectance and Illuminance of Faces 'in the Wild'abstractWe present SfSNet, an end-to-end learning framework for producing an accurate decomposition of an unconstrained human face image into shape, reflectance and illuminance. SfSNet is designed to reflect a physical lambertian rendering model. SfSNet learns from a mixture of labeled synthetic and unlabeled real world images. This allows the network to capture low frequency variations from synthetic and high frequency details from real images through the photometric reconstruction loss. SfSNet consists of a new decomposition architecture with residual blocks that learns a complete separation of albedo and normal. This is used along with the original image to predict lighting. SfSNet produces significantly better quantitative and qualitative results than state-of-the-art methods for inverse rendering and independent normal and illumination estimation. Roni Sengupta, Angjoo Kanazawa, Carlos Domingo Castillo, David Jacobs 0001 |
CVPR | 1 |
| 2017 | A New Rank Constraint on Multi-view Fundamental Matrices, and Its Application to Camera Location RecoveryabstractAccurate estimation of camera matrices is an important step in structure from motion algorithms. In this paper we introduce a novel rank constraint on collections of fundamental matrices in multi-view settings. We show that in general, with the selection of proper scale factors, a matrix formed by stacking fundamental matrices between pairs of images has rank 6. Moreover, this matrix forms the symmetric part of a rank 3 matrix whose factors relate directly to the corresponding camera matrices. We use this new characterization to produce better estimations of fundamental matrices by optimizing an L1-cost function using Iterative Re-weighted Least Squares and Alternate Direction Method of Multiplier. We further show that this procedure can improve the recovery of camera locations, particularly in multi-view settings in which fewer images are available. Roni Sengupta, Tal Amir, Meirav Galun, Tom Goldstein, David Jacobs 0001, Amit Singer, Ronen Basri |
CVPR | 1 |
| 2016 | Frontal to profile face verification in the wildabstractWe have collected a new face data set that will facilitate research in the problem of frontal to profile face verification `in the wild'. The aim of this data set is to isolate the factor of pose variation in terms of extreme poses like profile, where many features are occluded, along with other `in the wild' variations. We call this data set the Celebrities in Frontal-Profile (CFP) data set. We find that human performance on Frontal-Profile verification in this data set is only slightly worse (94.57% accuracy) than that on Frontal-Frontal verification (96.24% accuracy). However we evaluated many state-of-the-art algorithms, including Fisher Vector, Sub-SML and a Deep learning algorithm. We observe that all of them degrade more than 10% from Frontal-Frontal to Frontal-Profile verification. The Deep learning implementation, which performs comparable to humans on Frontal-Frontal, performs significantly worse (84.91% accuracy) on Frontal-Profile. This suggests that there is a gap between human performance and automatic face recognition methods for large pose variation in unconstrained images. Roni Sengupta, Jun-Cheng Chen, Carlos Domingo Castillo, Vishal M. Patel, Rama Chellappa, David Jacobs 0001 |
WACV | 1 |
| 2013 | Evenly Spaced Pareto Front Approximations for Tricriteria Problems Based on Triangulation
Günter Rudolph, Heike Trautmann, Roni Sengupta, Oliver Schütze 0001 |
EMO | 3 |
| 2013 | Multi-objective node deployment in WSNs: In search of an optimal trade-off among coverage, lifetime, energy consumption, and connectivity
Roni Sengupta, Swagatam Das, Md. Nasir, Bijaya K. Panigrahi |
Eng. Appl. Artif. Intell. | 1 |
| 2013 | Risk minimization in biometric sensor networks: an evolutionary multi-objective optimization approach
Roni Sengupta, Swagatam Das, Md. Nasir, Ponnuthurai N. Suganthan |
Soft Comput. | 1 |
| 2012 | An improved multi-objective optimization algorithm based on fuzzy dominance for risk minimization in biometric sensor networkabstractBiometric system is very important for recognition in several security areas. In this paper we deal in designing biometric sensor manager by optimizing the risk. Risk is modeled as a multi-objective optimization with Global False Acceptance Rate and Global False Rejection Rate as two objectives. In practice, multiple biometric sensors are used and the decision is taken locally at each sensor and the data is passed to the sensor manager. At the sensor manager the data is fused using a fusion rule and the final decision is taken. The optimization involves designing the data fusion rule and setting the sensor thresholds. We have implemented a recent fuzzy dominance based decomposition technique for multi-objective optimization called MOEA/DFD and have compared its performance on other contemporary state-of-arts in multi-objective optimization field like MOEA/D, NSGAII. The algorithm introduces a fuzzy Pareto dominance concept to compare two solutions and uses the scalar decomposition method only when one of the solutions fails to dominate the other in terms of a fuzzy dominance level. We have simulated the algorithms on different number of sensor setups consisting of 3, 6, 8 sensors respectively. We have also varied the apriori probability of imposter from 0.1 to 0.9 to verify the performance of the system with varying threat. One of the most significant advantages of using multi-objective optimization is that with a single run just by changing the decision making logic applied to the obtained Pareto front one can find the required threshold and decision strategies for varying threat of imposter. But with single objective optimization one need to run the algorithms each time with change in threat of imposter. Thus multi-objective representation appears to be more useful and better than single objective one. In all the test instances MOEA/DFD performs better than all other algorithms. Md. Nasir, Roni Sengupta, Swagatam Das, Ponnuthurai N. Suganthan |
IEEE Congress on Evolutionary Computation | 2 |
| 2012 | A Multi-Objective Evolutionary approach for linear antenna array design and synthesisabstractThe linear antenna array design problem is one of the most important in electromagnetism. While designing a linear antenna array, the goal of the designer is to achieve the “minimum average side lobe level” and a “null control” in specific directions. In contrast to the existing methods that attempt to minimize a weighted sum of these two objectives considered here, in this paper our contribution is twofold. First, we have considered these as two distinct objectives which are optimized simultaneously in a multi-objective framework. Second, for directivity purposes, we have introduced another objective called the “maximum side lobe level” in the design formulation. The resulting multi-objective optimization problem is solved by using the recently-proposed decomposition-based Multi-Objective Particle Swarm Optimizer (dMOPSO). Our experimental results indicate that the proposed approach is able to obtain results which are better than those obtained by two other state-of-the-art Multi-Objective Evolutionary Algorithms (MOEAs). Additionally, the individual minima reached by dMOPSO outperform those achieved by two single-objective evolutionary algorithms. Subhrajit Roy, Saúl Zapotecas Martínez, Carlos A. Coello Coello, Roni Sengupta |
IEEE Congress on Evolutionary Computation | 4 |
| 2012 | Energy-efficient differentiated coverage of dynamic objects using an improved evolutionary multi-objective optimization algorithm with fuzzy-dominanceabstractWe present an energy efficient sensor manager for differentiated coverage of dynamic object group changing their positions with time. The information about the location of the object group is provided to the sensor manager. The manager invokes optimization algorithm whenever the obtained coverage falls below a threshold to sleep schedule the sensor network. Multi-objective Optimization (MO) algorithms help in finding a better trade-off among energy consumption, lifetime, and coverage. Here the motion of the particle is modeled to follow a polynomial variation and with a constant acceleration. We formulate the scheduling problem as a combinatorial, constrained and multi-objective optimization problem with energy and non-coverage as the two objectives to be minimized. The proposed scheme uses a recent variant of a powerful MO algorithm known as Decomposition based Multi-Objective Evolutionary Algorithm (MOEA/D). Systematic comparison with the original MOEA/D and another well-known MO algorithm, NSGA-II (Non-dominated Sorting Genetic Algorithm) quantifies the superiority of the proposed approach. Roni Sengupta, Swagatam Das, Md. Nasir, Athanasios V. Vasilakos, Witold Pedrycz |
IEEE Congress on Evolutionary Computation | 1 |
| 2012 | A dynamic neighborhood learning based particle swarm optimizer for global numerical optimization
Md. Nasir, Swagatam Das, Dipankar Maity, Roni Sengupta, Udit Halder, Ponnuthurai N. Suganthan |
Inf. Sci. | 4 |
| 2012 | An Evolutionary Multiobjective Sleep-Scheduling Scheme for Differentiated Coverage in Wireless Sensor NetworksabstractWe propose an online, multiobjective optimization (MO) algorithm to efficiently schedule the nodes of a wireless sensor network (WSN) and to achieve maximum lifetime. Instead of dealing with traditional grid or uniform coverage, we focus on the differentiated or probabilistic coverage where different regions require different levels of sensing. The MO algorithm helps to attain a better tradeoff among energy consumption, lifetime, and coverage. The algorithm can be run every time a node failure occurs due to power failure of the node battery so that it may reschedule the network. This scheduling is modeled as a combinatorial, multiobjective, and constrained optimization problem with energy and noncoverage as the two objectives. The basic evolutionary multiobjective optimizer used is known as decomposition-based multiobjective evolutionary algorithm (MOEA/D) which is modified by integrating the concept of fuzzy Pareto dominance. The performance of the resulting algorithm, which is called MOEA/DFD, is compared with the performance of the original MOEA/D, which is another very well known MO algorithm called nondominated sorting genetic algorithm (NSGA-II), and an IBM optimization software package called CPLEX. In all the tests, MOEA/DFD is observed to outperform all other algorithms. Roni Sengupta, Swagatam Das, Md. Nasir, Athanasios V. Vasilakos, Witold Pedrycz |
IEEE Trans. Syst. Man Cybern. Part C | 1 |
| 2011 | An improved Multiobjective Evolutionary Algorithm based on decomposition with fuzzy dominanceabstractThis paper presents a new Multiobjective Evolutionary Algorithm (MOEA) based on decomposition, with fuzzy dominance (MOEA/DFD). The algorithm introduces a fuzzy Pareto dominance concept to compare two solutions and uses the scalar decomposition method only when one of the solutions fails to dominate the other in terms of a fuzzy dominance level. The diversity is maintained through the uniformly distributed weight vectors. In addition, Dynamic Resource Allocation (DRA) is used to distribute the computational effort based on the utilities of the individuals. To assess the performance of the proposed algorithm, experiments were conducted on two general benchmarks and ten unconstrained benchmark problems taken from the competition on real parameter MOEAs held under the 2009 IEEE Congress on Evolutionary Computation (CEC). As per the IGD metric, MOEA/DFD outperforms other major MOEAs in most cases. Mohammed Nasir, Arnab Kumar Mondal, Roni Sengupta, Swagatam Das, Ajith Abraham |
IEEE Congress on Evolutionary Computation | 3 |