VLDB 2026 Research / reviewers in the wild / expert
Shahrukh Athar
dblp:79/9032 · also ShahRukh Athar
· DBLP profile ↗
21ranked-venue papers
12as first author
13since 2021 · last 2025
0000-0002-8871-669XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 10 first-author · 13 since 2021Artificial intelligence and machine learning · 8 · 5 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Rig3DGS: Creating Controllable Portraits From Casual Monocular VideosabstractWe present Rig3DGS, a novel technique for creating reanimatable 3D portraits from short monocular smartphone videos. Rig3DGS learns to reconstruct a set of controllable 3D Gaussians from a monocular video of a dynamic subject captured with varying head poses and facial expressions in an in-the-wild scene. In contrast to synchronized multi-view studio captures, this in-the-wild, single camera setup brings fresh challenges to learning high quality 3D Gaussians. We address these challenges by learning to deform 3D Gaussians from a fixed canonical space to the deformed space that is consistent with the target facial expression and headpose. Our key contribution is a carefully designed deformation model that is guided by a 3D face morphable model. This deformation not only enables control over facial expression and head-poses but also allows our method to generates high-quality photorealistic renders of the whole scene. Once trained, Rig3DGS is able to generate photorealistic renders of a subject and their scene for novel facial expression, head-poses, and viewing directions. Through extensive experiments we demonstrate that Rig3DGS significantly outperforms prior art while being orders of magnitude faster. Alfredo Rivero, Shahrukh Athar, Zhixin Shu, Dimitris Samaras |
3DV | 2 |
| 2025 | HeadCraft: Modeling High-Detail Shape Variations for Animated 3DMMsabstractCurrent advances in human head modeling allow to generate plausible-looking 3D head models via neural representations, such as NeRFs and SDFs. Nevertheless, constructing complete high-fidelity head models with explicitly controlled animation remains an issue. Furthermore, completing the head geometry based on a partial observation, e.g. coming from a depth sensor, while preserving a high level of detail is often problematic for the existing methods. We introduce a generative model for detailed 3D head meshes on top of an articulated 3DMM which allows explicit animation and high-detail preservation at the same time. Our method is trained in two stages. First, we register a parametric head model with vertex displacements to each mesh of the recently introduced NPHM dataset of accurate 3D head scans. The estimated displacements are baked into a hand-crafted UV layout. Second, we train a StyleGAN model in order to generalize over the UV maps of displacements, which we later refer to HeadCraft. The decomposition of the parametric model and high-quality vertex displacements allows us to animate the model and modify the regions semantically. We demonstrate the results of unconditional sampling, fitting to a scan and editing. The code and data are available at https://seva100.github.io/headcraft. Artem Sevastopolsky, Philip-William Grassal, Simon Giebenhain, Shahrukh Athar, Luisa Verdoliva, Matthias Nießner |
3DV | 4 |
| 2025 | Shadow Removal Refinement via Material-Consistent Shadow EdgesabstractShadow boundaries can be confused with material boundaries as both exhibit sharp changes in luminance or contrast within a scene. However, shadows do not modify the intrinsic color or texture of surfaces. Therefore, on both sides of shadow edges traversing regions with the same material, the original color and texture should be the same if the shadow is removed properly. These shadow/shadow-free pairs are very useful but difficult-to-collect supervision signals. The crucial contribution of this paper is to learn how to identify those shadow edges that traverse material-consistent regions and how to use them as self-supervision for shadow removal refinement during test time. To achieve this, we fine-tune SAM, an image segmentation foundation model, to produce a shadow-invariant segmentation and then extract material-consistent shadow edges by comparing the SAM segmentation with the shadow mask. Utilizing these shadow edges, we in-troduce color- and texture-consistency losses to enhance the shadow removal process. We demonstrate the effectiveness of our method in improving shadow removal results on more challenging, in-the-wild images, outperforming the state-of-the-art shadow removal methods. Addition-ally, we propose a new metric and an annotated dataset for evaluating the performance of shadow removal methods without the need for paired shadow/shadow-free data. Our code and dataset are available at: https://github.com/cvlab-stonybrook/ShadowRemovalRefine Shilin Hu, Hieu Le 0001, Shahrukh Athar, Sagnik Das, Dimitris Samaras |
WACV | 3 |
| 2024 | Controllable Dynamic Appearance for Neural 3D PortraitsabstractRecent advances in Neural Radiance Fields (NeRFs) have made it possible to reconstruct and reanimate dynamic portrait scenes with control over head-pose, facial expressions and viewing direction. However, training such models assumes photometric consistency over the deformed region e.g. the face must be evenly lit as it deforms with changing head-pose and facial expression. Such photometric consistency across frames of a video is hard to maintain, even in studio environments, thus making the created reanimatable neural portraits prone to artefacts during reanimation. In this work, we propose CoDyNeRF, a system that enables the creation of fully controllable 3D portraits in real-world capture conditions. CoDyNeRF learns to approximate illumination dependent effects via a dynamic appearance model in the canonical space that is conditioned on predicted surface normals and the facial expressions and head-pose deformations. The surface normals prediction is guided using 3DMM normals that act as a coarse prior for the normals of the human head, where direct prediction of normals is hard due to rigid and non-rigid deformations induced by head-pose and facial expression changes. Using only a smartphone-captured short video of a subject for training, we demonstrate the effectiveness of our method on free view synthesis of a portrait scene with explicit head pose and expression controls, and realistic lighting effects. Shahrukh Athar, Zhixin Shu, Zexiang Xu, Fujun Luan, Sai Bi, Kalyan Sunkavalli, Dimitris Samaras |
3DV | 1 |
| 2024 | Bridging the Gap: Studio-Like Avatar Creation from a Monocular Phone Capture
Shahrukh Athar, Shunsuke Saito, Zhengyu Yang 0003, Stanislav Pidhorskyi, Chen Cao 0001 |
ECCV (12) | 1 |
| 2024 | Boosting Image Quality Assessment Performance: Unsupervised Score Fusion by Deep Maximum a Posteriori EstimationabstractOver the past decades, numerous Image Quality Assessment (IQA) models have emerged, aiming to predict the perceptual quality of images. However, individual models are often biased toward certain types of image content or distortions, depending on the design principle and process. An intuitive idea is to harness the strengths and mitigate the weaknesses of each IQA model, by fusing the scores of multiple models into a stronger one. Here we make one of the first attempts to seek an optimal solution for the idea and propose a general framework for unsupervised IQA score fusion using deep Maximum a Posteriori (MAP) estimation. The proposed model conducts fine-grained uncertainty estimation at the score level to increase the accuracy and reduce the uncertainty in fused predictions. Comprehensive experiments demonstrate the superiority of the proposed model over individual IQA models and other fusion methods. It also exhibits an interesting capability of rejecting "bad" models in the fusion process. Zhongling Wang, Raymond Zhou, Shahrukh Athar, Wenbo Yang 0001, Zhou Wang 0001 |
ICASSP | 3 |
| 2023 | FLAME-in-NeRF: Neural control of Radiance Fields for Free View Face AnimationabstractThis paper presents a neural rendering method for controllable portrait video synthesis. Recent advances in volumetric neural rendering, such as neural radiance fields (NeRF), have enabled the photorealistic novel view synthesis of static scenes with impressive results. However, modeling dynamic and controllable objects as part of a scene with such scene representations is still challenging. In this work, we design a system that enables 1) novel view synthesis for portrait video, of both the human subject and the scene they are in and 2) explicit control of the facial expressions through a low-dimensional expression representation. We represent the distribution of human facial expressions using the expression parameters of a 3D Morphable Model (3DMM) and condition the NeRF volumetric function on them. In order to guide the network to learn disentangled control for static scene appearance and dynamic facial actions, we impose a spatial prior via 3DMM fitting. We show the effectiveness of our method on free view synthesis of portrait videos with expression controls. To train a scene, our method only requires a short video of a subject captured by a mobile device. Shahrukh Athar, Zhixin Shu, Dimitris Samaras |
FG | 1 |
| 2023 | LipNeRF: What is the right feature space to lip-sync a NeRF?abstractSynthesizing high-fidelity talking head videos of an arbitrary identity, lip-synced to a target speech segment, is a challenging problem. Recent GAN-based methods succeed by training a model on a large amount of videos, allowing the generator to learn a variety of audio-lip representations. However, they are unable to handle head pose changes. On the other hand, Neural Radiance Fields (NeRFs) model the 3D face geometry more accurately. Current audio-conditioned NeRFs are not as good in lip synchronization as GANs, since they are trained on limited video data of a single identity. In this work, we propose LipNeRF, a lip-syncing NeRF that bridges the gap between the accurate lip synchronization of GAN-based methods and the accurate 3D face modeling of NeRFs. LipNeRF is conditioned on the expression space of a 3DMM, instead of the audio feature space. We experimentally demonstrate that the expression space gives a better representation for the lip shape than the audio feature space. LipNeRF shows a significant improvement in lip-sync quality over the current state-of-the-art, especially in high-definition videos of cinematic content, with challenging pose, illumination and expression variations. Aggelina Chatziagapi, Shahrukh Athar, Rohith MV, Vimal Bhat, Dimitris Samaras |
FG | 2 |
| 2023 | Degraded Reference Image Quality AssessmentabstractIn practical media distribution systems, visual content usually undergoes multiple stages of quality degradation along the delivery chain, but the pristine source content is rarely available at most quality monitoring points along the chain to serve as a reference for quality assessment. As a result, full-reference (FR) and reduced-reference (RR) image quality assessment (IQA) methods are generally infeasible. Although no-reference (NR) methods are readily applicable, their performance is often not reliable. On the other hand, intermediate references of degraded quality are often available, e.g., at the input of video transcoders, but how to make the best use of them in proper ways has not been deeply investigated. Here we make one of the first attempts to establish a new paradigm named degraded-reference IQA (DR IQA). Specifically, by using a two-stage distortion pipeline we lay out the architectures of DR IQA and introduce a 6-bit code to denote the choices of configurations. We construct the first large-scale databases dedicated to DR IQA and have made them publicly available. We make novel observations on distortion behavior in multi-stage distortion pipelines by comprehensively analyzing five multiple distortion combinations. Based on these observations, we develop novel DR IQA models and make extensive comparisons with a series of baseline models derived from top-performing FR and NR models. The results suggest that DR IQA may offer significant performance improvement in multiple distortion environments, thereby establishing DR IQA as a valid IQA paradigm that is worth further exploration. Shahrukh Athar, Zhou Wang 0001 |
IEEE Trans. Image Process. | 1 |
| 2022 | RigNeRF: Fully Controllable Neural 3D PortraitsabstractVolumetric neural rendering methods, such as neural radiance fields (NeRFs), have enabled photo-realistic novel view synthesis. However, in their standard form, NeRFs do not support the editing of objects, such as a human head, within a scene. In this work, we propose RigNeRF, a system that goes beyond just novel view synthesis and enables full control of head pose and facial expressions learned from a single portrait video. We model changes in head pose and facial expressions using a deformation field that is guided by a 3D morphable face model (3DMM). The 3DMM effectively acts as a prior for RigNeRF that learns to predict only residuals to the 3DMM deformations and allows us to render novel (rigid) poses and (non-rigid) expressions that were not present in the input sequence. Using only a smartphone-captured short video of a subject for training, we demonstrate the effectiveness of our method on free view synthesis of a portrait scene with explicit head pose and expression controls. Shahrukh Athar, Zexiang Xu, Kalyan Sunkavalli, Eli Shechtman, Zhixin Shu |
CVPR | 1 |
| 2022 | Deep Image DebandingabstractBanding or false contour is an annoying visual artifact whose impact negatively degrades the perceptual quality of visual content. Since users are increasingly expecting better visual quality from such content and banding leads to deteriorated quality-of-experience, the area of banding removal or debanding has taken paramount importance. Existing debanding approaches are mostly knowledge-driven, while data-driven debanding approaches remain surprisingly missing. In this work, we construct a large-scale dataset of 51,490 pairs of corresponding pristine and banded image patches, which enables us to make one of the first attempts at developing a deep learning based banding artifact removal method for images that we name deep debanding network (deepDeband). We also develop a bilateral weighting scheme that fuses patch-level debanding results to full-size images. Extensive performance evaluation shows that deepDeband is successful at greatly reducing banding artifacts in images, outperforming existing methods both quantitatively and visually. The proposed algorithm and dataset are made publicly available.1 Raymond Zhou, Shahrukh Athar, Zhongling Wang, Zhou Wang 0001 |
ICIP | 2 |
| 2021 | SIDER: Single-Image Neural Optimization for Facial Geometric Detail RecoveryabstractWe present SIDER (Single-Image neural optimization for facial geometric DEtail Recovery), a novel photometric optimization method that recovers detailed facial geometry from a single image in an unsupervised manner. Inspired by classical techniques of coarse-to-fine optimization and recent advances in implicit neural representations of 3D shape, SIDER combines a geometry prior based on statistical models and Signed Distance Functions (SDFs) to recover facial details from single images. First, it estimates a coarse geometry using a morphable model represented as an SDF. Next, it reconstructs facial geometry details by optimizing a photometric loss with respect to the ground truth image. In contrast to prior work, SIDER does not rely on any dataset priors and does not require additional supervision from multiple views, lighting changes or ground truth 3D shape. Extensive qualitative and quantitative evaluation demonstrates that our method achieves state-of-the-art on facial geometric detail recovery, using only a single in the-wild image. Aggelina Chatziagapi, Shahrukh Athar, Francesc Moreno-Noguer, Dimitris Samaras |
3DV | 2 |
| 2021 | Variational Feature Disentangling for Fine-Grained Few-Shot ClassificationabstractData augmentation is an intuitive step towards solving the problem of few-shot classification. However, ensuring both discriminability and diversity in the augmented samples is challenging. To address this, we propose a feature disentanglement framework that allows us to augment features with randomly sampled intra-class variations while preserving their class-discriminative features. Specifically, we disentangle a feature representation into two components: one represents the intra-class variance and the other encodes the class-discriminative information. We assume that the intra-class variance induced by variations in poses, backgrounds, or illumination conditions is shared across all classes and can be modelled via a common distribution. Then we sample features repeatedly from the learned intra-class variability distribution and add them to the class-discriminative features to get the augmented features. Such a data augmentation scheme ensures that the augmented features inherit crucial class-discriminative features while exhibiting large intra-class variance. Our method significantly outperforms the state-of-the-art methods on multiple challenging fine-grained few-shot image classification benchmarks. Code is available at: https://github.com/cvlab-stonybrook/vfd-iccv21 Hieu Le 0001, Mingzhen Huang, Shahrukh Athar, Dimitris Samaras |
ICCV | 4 |
| 2020 | Self-supervised Deformation Modeling for Facial Expression EditingabstractDeep generative models have recently demonstrated impressive results in photo-realistic facial image synthesis and editing. Existing neural network-based approaches usually only rely on texture generation to edit expressions and largely neglect the motion information. However, facial expressions are inherently the result of muscle movement. In this work, we propose a novel end-to-end network that disentangles the task of facial editing into two steps: a “motionediting” step and a “texture-editing” step. In the “motion-editing” step, we explicitly model facial movement through an image deformation, warping the image into the desired expression. In the “texture-editing” step, we generate the necessary textures, such as teeth and shading effects, for a photorealistic result. Our physically-based task-disentanglement system design allows each step to learn a focused task, and thus need not generate texture to hallucinate motion. Our system is trained in a self-supervised manner, requiring no ground truth deformation annotation. Using Action Units [8] as the representation for facial expression, our method improves the state-of-the-art facial expression editing performance in both qualitative and quantitative evaluations. Shahrukh Athar, Zhixin Shu, Dimitris Samaras |
FG | 1 |
| 2020 | Nonverbal Behavioral Patterns Predict Social Rejection Elicited Aggression
Megan Quarmley, Zhibo Yang 0002, Shahrukh Athar, Gregory J. Zelinsky, Dimitris Samaras, Johanna M. Jarcho |
FG | 3 |
| 2019 | Perceptual Quality Assessment of UHD-HDR-WCG VideosabstractHigh Dynamic Range (HDR) Wide Color Gamut (WCG) Ultra High Definition (4K/UHD) content has become increasingly popular recently. Due to the increased data rate, novel video compression methods have been developed to maintain the quality of the videos being delivered to consumers under bandwidth constraints. This has led to new challenges for the development of objective Video Quality Assessment (VQA) models, which are traditionally designed without sufficient calibration and validation based on subjective quality assessment of UHD-HDR-WCG videos. The large performance variations between different consumer HDR TVs, and between consumer HDR TVs and professional HDR reference displays used for content production, further complicates the task of acquiring reliable subjective data that faithfully reflects the impact of compression on UHD-HDR-WCG videos. In this work, we construct a first-of-its-kind video database composed of PQ-encoded UHD-HDR-WCG content, which is subsequently compressed by H.264 and HEVC encoders. We carry out a subjective study on a professional 4K-HDR reference display in a controlled lab environment. We also benchmark representative Full Reference (FR) and No-Reference (NR) objective VQA models against the subjective data to evaluate their performance on compressed UHD-HDR-WCG video content. The database will be made available to the public, subject to content copyright constraints. Shahrukh Athar, Thilan Costa, Kai Zeng 0003, Zhou Wang 0001 |
ICIP | 1 |
| 2019 | Latent Convolutional Models
Shahrukh Athar, Evgeny Burnaev, Victor S. Lempitsky |
ICLR (Poster) | 1 |
| 2017 | Quality assessment of images undergoing multiple distortion stagesabstractIn practical media distribution systems, visual content often undergoes multiple stages of quality degradations along the delivery chain between the source and destination. By contrast, current image quality assessment (IQA) models are typically validated on image databases with a single distortion stage. In this work, we construct two large-scale image databases that are composed of more than 2 million images undergoing multiple stages of distortions and examine how state-of-the-art IQA algorithms behave over distortion stages. Our results suggest that the performance of existing IQA models degrades rapidly with distortion stages, especially when the distortion types of different stages vary. We also find that full-reference and no-reference frameworks, though both readily applicable, have major drawbacks at predicting the quality of images at middle distortion stages. However, when the quality level of the previous stage is accessible, significantly improved quality prediction performance may be achieved. This study points out a new avenue of degraded-reference IQA research that is both practically desirable and technically challenging. Shahrukh Athar, Abdul Rehman 0001, Zhou Wang 0001 |
ICIP | 1 |
| 2015 | Data rate and dynamic range compression of medical images: Which one goes first?abstractAdvances in the field of medical imaging have led to an immense increase in the volume of images being acquired. A fast growing application is to enable physicians to access image data remotely from any viewing device. This casts new challenges for data rate compression. Meanwhile medical images typically have High Dynamic Range (HDR), which needs to be transformed to Low Dynamic Range (LDR) through a so-called “windowing” operation in order for them to be viewed on standard displays to best visualize specific types of content such as tissues or bone structures. This leads to a basic question: Should data compression be performed before windowing or vice versa? Answering this question needs domain knowledge and also requires comparing HDR and LDR images in terms of objective measures, which has only recently become possible. In this paper, we compare the two alternative schemes by using a recently proposed structural fidelity measure. Our study suggests that data compression followed by windowing delivers better performance than the other alternative. Shahrukh Athar, Hojatollah Yeganeh, Zhou Wang 0001 |
ICIP | 1 |
| 2012 | Teaching and research in FPGA based Digital Signal Processing using Xilinx System GeneratorabstractThis paper presents an efficient approach for the implementation of typical DSP structures studied in class or conceived during research. This scheme is beneficial where the objective is to implement the physical working of complex DSP structures or algorithms without requiring detailed knowledge of hardware design and hardware description languages. The approach is based on the Xilinx System Generator for DSP tool, which integrates itself with the MATLAB based Simulink Graphics environment and relieves the user of the textual HDL programming. In addition to introducing this scheme for teaching of DSP, some useful examples based on Delta Sigma Modulators are also presented. These modulators are selected because they are an integral part of modern Analog to Digital Converters and encompass many important Signal Processing concepts. The advantage of this approach in DSP research in terms of reducing concept-to-Silicon design time and effort is also highlighted. Shahrukh Athar, Muhammad Ali Siddiqi, Shahid Masud |
ICASSP | 1 |
| 2010 | Design and FPGA Implementation of a 2nd Order Adaptive Delta Sigma Modulator with One Bit QuantizationabstractThis paper presents the design and FPGA implementation of a 2ndorder all-digital Adaptive Delta Sigma (ΔΣ) modulator with one bit quantization. It has a modulator stage and an adaptation stage. The adaptation stage produces a feedback signal that tracks the input signal and is subtracted from it. This difference signal is in a controlled and reduced range. It is given to the input of the modulator stage which has a 2ndorder ΔΣ modulator. This results in a reduction of quantization noise and an increase in the overall Signal to Quantization Noise Ratio (SQNR) of the modulator. The design was implemented on a Xilinx Spartan family FPGA using the Xilinx System Generator for DSP tool. The Hardware Co-Simulation mode of the System Generator was used which enables Simulink to run the FPGA directly, thus facilitating extensive testing. The spectral and SQNR analysis of the FPGA output was performed in MATLAB. The 2ndorder adaptive ΔΣ modulator presented here, exhibits an average SQNR improvement of 24.66 dB, 22.11 dB, 16.59 dB and 8.24 dB over the 2ndorder non-adaptive ΔΣ modulator at Over Sampling Ratios (OSRs) of 512, 256, 128 and 64 respectively in an input power range of -80 to 20 dB. It also exhibits an increased dynamic range of approximately 24 dB over the 2ndorder non-adaptive ΔΣ modulator. Shahrukh Athar, Muhammad Ali Siddiqi, Shahid Masud |
FPL | 1 |