Muskan Dosi

dblp:311/2256 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0001-7451-3317ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Security and privacy · 2 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2025 Harmonizing Geometry and Uncertainty: Diffusion with Hyperspheres
abstract
Do contemporary diffusion models preserve the class geometry of hyperspherical data? Standard diffusion models rely on isotropic Gaussian noise in the forward process, inherently favoring Euclidean spaces. However, many real-world problems involve non-Euclidean distributions, such as hyperspherical manifolds, where class-specific patterns are governed by angular geometry within hypercones. When modeled in Euclidean space, these angular subtleties are lost, leading to suboptimal generative performance. To address this limitation, we introduce \textbf{HyperSphereDiff} to align hyperspherical structures with directional noise, preserving class geometry and effectively capturing angular uncertainty. We demonstrate both theoretically and empirically that this approach aligns the generative process with the intrinsic geometry of hyperspherical data, resulting in more accurate and geometry-aware generative models. We evaluate our framework on four object datasets and two face datasets, showing that incorporating angular uncertainty better preserves the underlying hyperspherical manifold.
Muskan Dosi, Chiranjeev Chiranjeev, Kartik Thakral, Mayank Vatsa, Richa Singh 0001
ICML1
2025 Beyond shadows and light: Odyssey of face recognition for social good
Chiranjeev Chiranjeev, Muskan Dosi, Shivang Agarwal, Jyoti Chaudhary, Pranav Pant, Mayank Vatsa, Richa Singh 0001
Comput. Vis. Image Underst.2
2024 BirdCollect: A Comprehensive Benchmark for Analyzing Dense Bird Flock Attributes
abstract
Automatic recognition of bird behavior from long-term, un controlled outdoor imagery can contribute to conservation efforts by enabling large-scale monitoring of bird populations. Current techniques in AI-based wildlife monitoring have focused on short-term tracking and monitoring birds individually rather than in species-rich flocks. We present Bird-Collect, a comprehensive benchmark dataset for monitoring dense bird flock attributes. It includes a unique collection of more than 6,000 high-resolution images of Demoiselle Cranes (Anthropoides virgo) feeding and nesting in the vicinity of Khichan region of Rajasthan. Particularly, each image contains an average of 190 individual birds, illustrating the complex dynamics of densely populated bird flocks on a scale that has not previously been studied. In addition, a total of 433 distinct pictures captured at Keoladeo National Park, Bharatpur provide a comprehensive representation of 34 distinct bird species belonging to various taxonomic groups. These images offer details into the diversity and the behaviour of birds in vital natural ecosystem along the migratory flyways. Additionally, we provide a set of 2,500 point-annotated samples which serve as ground truth for benchmarking various computer vision tasks like crowd counting, density estimation, segmentation, and species classification. The benchmark performance for these tasks highlight the need for tailored approaches for specific wildlife applications, which include varied conditions including views, illumination, and resolutions. With around 46.2 GBs in size encompassing data collected from two distinct nesting ground sets, it is the largest birds dataset containing detailed annotations, showcasing a substantial leap in bird research possibilities. We intend to publicly release the dataset to the research community. The database is available at: https://iab-rubric.org/resources/wildlife-dataset/birdcollect
Kshitiz, Sonu Sreshtha, Bikash Dutta, Muskan Dosi, Mayank Vatsa, Richa Singh 0001, Saket Anand, Sudeep Sarkar, Sevaram Mali Parihar
AAAI4
2024 HyperSpaceX: Radial and Angular Exploration of HyperSpherical Dimensions
Chiranjeev Chiranjeev, Muskan Dosi, Kartik Thakral, Mayank Vatsa, Richa Singh 0001
ECCV (88)2
2024 Is Face Super Resolution Truly Pushing the Boundaries of Face Recognition?
abstract
With the improving efficacy of generative algorithms, the performance of face super-resolution algorithms is also increasing towards generating high-quality facial data. However, are these images useful for face recognition? This paper investigates whether these enhanced super-resolved facial images only improve the visual quality or they also aid in improving the recognizability of these images, thus contributing towards addressing the challenge of low-resolution face recognition. We conduct a comprehensive empirical and statistical analysis of human perception and face recognition tasks. Extensive experiments are performed using multiple state-of-the-art generative and face recognition models across six publicly available face datasets to assess whether face super-resolution algorithms are effective in recognizing individuals in low-resolution conditions. The results and supporting analysis indicate that the ability of super-resolution images to improve recognizability is limited, and further research is required to design generative AI algorithms that improve both visual appearance and recognizability of low-resolution images.
Muskan Dosi, Udaybhan Rathore, Chiranjeev Chiranjeev, Akshay Agarwal 0001, Richa Singh 0001, Mayank Vatsa
IJCB1
2023 UG-LDFace: Unified and Generalized Framework for Long-Range Disguised Face Recognition
abstract
Long-range, low-resolution videos have widespread applications in active monitoring, crowd counting, traffic analysis, and person verification/identification. The problem of analyzing faces in such an environment is exacerbated by the presence of disguise and occlusion. This research presents UG-LDFace, a novel face recognition model to address this arduous challenge. A single-stage unified framework is proposed that integrates two feature refinement techniques: feature enhancement for low-resolution data and feature selection for disguised faces. The proposed model also comprises a revised distribution technique to generalize UG-LDFace on unseen data. The proposed approach shows its efficacy on five different datasets, DroneSURF, SCface, D-LORD, DSIMF, and LFW, containing various levels of occlusion and low-resolution data.
Muskan Dosi, Chiranjeev Chiranjeev, Richa Singh 0001, Mayank Vatsa
IJCB1
2021 AECNet: Attentive EfficientNet For Crowd Counting
abstract
In the COVID pandemic situation, crowd counting became one of the tools to monitor if the social-distancing norms are being followed or not. However, in designing crowd counting algorithm, there are several challenges such as background noise, camera-to-objects distance, occlusion, and variations due to illumination, scale, and viewpoint. In this research, we propose a novel pipeline for density estimation in crowd counting. The proposed pipeline makes use of an encoder-decoder-based architecture in which we explore the family of EfficientN ets for the encoder architecture. For the decoder, we propose a deeper attention network to assist the model in a better distinction between foreground and background pixels. We empirically show that for a crowd counting dataset, the use of average pooling operation for any backbone architecture of encoder gives a significant improvement in performance. In terms of Mean Absolute Error, the proposed pipeline outperforms existing state-of-the-art techniques by a large margin on large-scale and small-scale counting datasets, UCF-QNRF and UCF _CC_50 dataset. We also achieve state-of-the-art results on the ShanghaiTech and Mall datasets. We additionally propose a crowd counting dataset captured using drones. We perform benchmark experiments on this dataset with existing and the proposed methods. The proposed dataset can be found at http://www.iab-rubric.org/resources/CrowdUAV.html.
Muskan Dosi, Kartik Thakral, Surbhi Mittal, Mayank Vatsa, Richa Singh 0001
FG1
2021 Dual Sensor Indian Masked Face Dataset
abstract
With the advancements in deep learning technologies, real-world applications like face detection, gender prediction, and face recognition have achieved human-level performance. However, the emergence of the COVID-19 pandemic brought new challenges to existing deep learning algorithms. People are forced to wear a mask to limit the spread of COVID-19. These face masks occlude a significant portion of the face, thereby posing multiple challenges to existing algorithms. Images captured using surveillance cameras have a low resolution which hinders the model performance. Along with this, skin tone, ethnicity and attire also play a significant role in detection and recognition performance. India is a large country with huge diversity in skin tone and attire of the people. To address the challenges due to masks in the Indian context, we propose a novel Dual Sensor Indian Masked Face (DS- IMF) dataset, which contains images captured in constrained environmental settings with a variety of masks and degrees of occlusion. Multiple experiments are performed on the DS- IMF dataset at different resolutions. Experimental results demonstrate the limitations of existing algorithms on low-resolution masked face images. The proposed dataset can be found at http://www.iab-rubric.org/resources/dsimf.html.
Shiksha Mishra, Puspita Majumdar, Muskan Dosi, Mayank Vatsa, Richa Singh 0001
FG3