VLDB 2026 Research / reviewers in the wild / expert
Naina Dhingra
dblp:241/4883
· DBLP profile ↗
8ranked-venue papers
5as first author
5since 2021 · last 2025
0000-0001-7546-1213ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Discontinuity-Aware Normal Integration for Generic Central Camera ModelsabstractRecovering a 3D surface from its surface normal map, a problem known as normal integration, is a key component for photometric shape reconstruction techniques such as shape-from-shading and photometric stereo. The vast majority of existing approaches for normal integration handle only implicitly the presence of depth discontinuities and are limited to orthographic or ideal pinhole cameras. In this paper, we propose a novel formulation that allows modeling discontinuities explicitly and handling generic central cameras. Our key idea is based on a local planarity assumption, that we model through constraints between surface normals and ray directions. Compared to existing methods, our approach more accurately approximates the relation between depth and surface normals, achieves state-of-the-art results on the standard normal integration benchmark, and is the first to directly handle generic central camera models. Francesco Milano 0001, Manuel Lopez-Antequera, Naina Dhingra, Roland Siegwart, Robert Thiel |
ICCV | 3 |
| 2023 | MMG-Ego4D: Multi-Modal Generalization in Egocentric Action RecognitionabstractIn this paper, we study a novel problem in egocentric action recognition, which we term as “Multimodal Generalization“ (MMG). MMG aims to study how systems can generalize when data from certain modalities is limited or even completely missing. We thoroughly investigate MMG in the context of standard supervised action recognition and the more challenging few-shot setting for learning new action categories. MMG consists of two novel scenarios, designed to support security, and efficiency considerations in real-world applications: (1) missing modality generalization where some modalities that were present during the train time are missing during the inference time, and (2) cross-modal zero-shot generalization, where the modalities present during the inference time and the training time are disjoint. To enable this investigation, we construct a new dataset MMG-Ego4D containing data points with video, audio, and inertial motion sensor (IMU) modalities. Our dataset is derived from Ego4D [27] dataset, but processed and thoroughly re-annotated by human experts to facilitate research in the MMG problem. We evaluate a diverse array of models on MMG-Ego4D and propose new methods with improved generalization ability. In particular, we introduce a new fusion module with modality dropout training, contrastive-based alignment training, and a novel cross-modal prototypical loss for better few-shot performance. We hope this study will serve as a benchmark and guide future research in multimodal generalization problems. The benchmark and code are available at https://github.com/facebookresearch/MMG_Ego4D Xinyu Gong, Sreyas Mohan, Naina Dhingra, Jean-Charles Bazin, Yilei Li, Zhangyang Wang |
CVPR | 3 |
| 2022 | Language-Attention Modular-Network for Relational Referring Expression Comprehension in VideosabstractReferring expression (RE) for video domain describes the video using a natural language expression. Relational RE comprehension in a video domain localizes an object in relation to a distinguishing context object. Unlike object grounding in videos using REs, not much work has been done in videos using relational REs. In this paper, we focus on (1) relational RE comprehension for videos, and (2) demonstrating the significance of attention for the task. We propose a novel modular network based approach for relational RE comprehension in highly ambiguous settings for videos. We show the significance of the language attention in modular approach by: (1) Using two different networks, i.e., modATN consisting of attention mechanism, visual modules, and a natural language expression input, and modSTR consisting of visual modules, and structured input (subject, subject adjective, object, object adjective, action, relation); (2) Introducing a new dataset having structured RE for relation RE comprehension task in modSTR. Finally, we propose an optimised modular network that outperforms and shows significant improvements over the baseline networks. Naina Dhingra, Shipra Jain |
ICPR | 1 |
| 2022 | LwPosr: Lightweight Efficient Fine Grained Head Pose EstimationabstractThis paper presents a lightweight network for head pose estimation (HPE) task. While previous approaches rely on convolutional neural networks, the proposed network LwPosr uses mixture of depthwise separable convolutional (DSC) and transformer encoder layers which are structured in two streams and three stages to provide fine-grained regression for predicting head poses. The quantitative and qualitative demonstration is provided to show that the proposed network is able to learn head poses efficiently while using less parameter space. Extensive ablations are conducted using three open-source datasets namely 300W-LP, AFLW2000, and BIWI datasets. To our knowledge, (1) LwPosr is the lightest network proposed for estimating head poses compared to both keypoints-based and keypoints-free approaches; (2) it sets a benchmark for both overperforming the previous lightweight network on mean absolute error and on reducing number of parameters; (3) it is first of its kind to use mixture of DSCs and transformer encoders for HPE. This approach is suitable for mobile devices which require lightweight networks. Naina Dhingra |
WACV | 1 |
| 2021 | HeadPosr: End-to-end Trainable Head Pose Estimation using Transformer EncodersabstractHead pose estimation (HPE) is the task of estimating head pose given an RGB image, video, or RGB-D data. In this paper, a network, HeadPosr is proposed to predict the head poses using a single RGB image. HeadPosr uses a novel architecture which includes a transformer encoder. In concrete, it consists of: (1) backbone; (2) connector; (3) transformer encoder; (4) prediction head. The significance of using a transformer encoder for HPE is studied. An extensive ablation study is performed on varying the (1) number of encoders; (2) number of heads; (3) different position embeddings; (4) different activations; (5) input channel size, in a transformer used in HeadPosr. Further studies on using: (1) different backbones, (2) using different learning rates are also shown. The elaborated experiments and ablations studies are conducted using three different open-source widely used datasets for HPE, i.e., 300W-LP, AFLW2000, and BIWI datasets. Experiments illustrate that HeadPosr outperforms all the state-of-art methods including both the landmark-free and the others based on using landmark or depth estimation on the AFLW2000 dataset and BIWI datasets when trained with 300W-LP. It also outperforms when averaging the results from the compared datasets, hence setting a benchmark for the problem of HPE, also demonstrating the effectiveness of using transformers over the state-of-the-art. Naina Dhingra |
FG | 1 |
| 2020 | Pointing Gesture Based User Interaction of Tool Supported Brainstorming MeetingsabstractAbstract This paper presents a brainstorming tool combined with pointing gestures to improve the brainstorming meeting experience for blind and visually impaired people (BVIP). In brainstorming meetings, BVIPs are not able to participate in the conversation as well as sighted users because of the unavailability of supporting tools for understanding the explicit and implicit meaning of the non-verbal communication (NVC). Therefore, the proposed system assists BVIP in interpreting pointing gestures which play an important role in non-verbal communication. Our system will help BVIP to access the contents of a Metaplan card, a team member in the brainstorming meeting is referring to by pointing. The prototype of our system shows that targets on the screen a user is pointing at can be detected with 80% accuracy. Naina Dhingra, Reinhard Koutny, Sebastian Günther 0001, Klaus Miesenberger, Max Mühlhäuser, Andreas M. Kunz |
ICCHP (2) | 1 |
| 2020 | Accessible Multimodal Tool Support for Brainstorming MeetingsabstractAbstract In recent years, assistive technology and digital accessibility for blind and visually impaired people (BVIP) has been significantly improved. Yet, group discussions, especially in a business context, are still challenging as non-verbal communication (NVC) is often depicted on digital whiteboards, including deictic gestures paired with visual artifacts. However, as NVC heavily relies on the visual perception, whichrepresents a large amount of detail, an adaptive approach is required that identifies the most relevant information for BVIP. Additionally, visual artifacts usually rely on spatial properties such as position, orientation, and dimensions to convey essential information such as hierarchy, cohesion, and importance that is often not accessible to the BVIP. In this paper, we investigate the requirements of BVIP during brainstorming sessions and, based on our findings, provide an accessible multimodal tool that uses non-verbal and spatial cues as an additional layer of information. Further, we contribute by presenting a set of input and output modalities that encode and decode information with respect to the individual demands of BVIP and the requirements of different use cases. Reinhard Koutny, Sebastian Günther 0001, Naina Dhingra, Andreas M. Kunz, Klaus Miesenberger, Max Mühlhäuser |
ICCHP (2) | 3 |
| 2019 | Res3ATN - Deep 3D Residual Attention Network for Hand Gesture Recognition in VideosabstractHand gesture recognition is a strenuous task to solve in videos. In this paper, we use a 3D residual attention network which is trained end to end for hand gesture recognition. Based on the stacked multiple attention blocks, we build a 3D network which generates different features at each attention block. Our 3D attention based residual network (Res3ATN) can be build and extended to very deep layers. Using this network, an extensive analysis is performed on other 3D networks based on three publicly available datasets. The Res3ATN network performance is compared to C3D, ResNet-10, and ResNext-101 networks. We also study and evaluate our baseline network with different number of attention blocks. The comparison shows that the 3D residual attention network with 3 attention blocks is robust in attention learning and is able to classify the gestures with a better accuracy, thus outperforming existing networks. Naina Dhingra, Andreas M. Kunz |
3DV | 1 |