Zahid Hasan 0001

dblp:21/8367-1 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0002-8495-0948ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Lorentz Entailment Cone for Semantic Segmentation
abstract
Semantic segmentation in hyperbolic space can capture hierarchical structure in low dimensions with uncertainty quantification. Existing approaches choose the Poincaré ball model for hyperbolic geometry, which suffers from numerical instabilities, optimization, and computational challenges. We propose a novel, tractable, architecture-agnostic semantic segmentation framework in the hyperbolic Lorentz model. We employ text embeddings with semantic and visual cues to guide hierarchical pixel-level representations in Lorentz space. This enables stable and efficient optimization without requiring a Riemannian optimizer, and easily integrates with existing Euclidean architectures. Beyond segmentation, our approach yields free uncertainty estimation, confidence map, boundary delineation, hierarchical and text-based retrieval, and zero-shot performance, reaching generalized flatter minima. We further introduce a novel uncertainty and confidence indicator in Lorentz cone embeddings. Extensive experiments on ADE20K, COCO-Stuff-164k, Pascal-VOC, and Cityscapes with state-of-the-art models (DeepLabV3 and SegFormer) validate the effectiveness and generality of our approach. Our results demonstrate the potential of hyperbolic Lorentz embeddings for robust and uncertainty-aware semantic segmentation. Code is available at https://github.com/mxahan/Lorentz_semantic_segmentation.
Zahid Hasan 0001, Masud Ahmed, Nirmalya Roy
WACV1
2024 Contactless Physiological Health Sensing: Challenges, Solutions & Opportunities
abstract
Contactless health monitoring techniques, such as Remote Photoplethysmography (rPPG) [1] and video-based Respiratory Rate (RR) [2] estimation, have emerged as promising methods utilizing regular camera sensors for capturing vital signs like Heart Rate (HR) and Respiratory Rate (RR). These approaches offer a cost-effective, widely applicable, and safe solution for regular and long-term health monitoring. However, the primary challenge lies in extracting accurate vital signals from the captured videos due to their low signal-to-noise ratios. Traditional signal processing analyzes temporal properties of spatial regions of interest across video frames to extract vital signs, while computer vision-based Deep Learning (DL) approaches leverage large-scale datasets and computational power for data-driven decision-making without heuristic reliance [3]. In this tutorial, we aim to provide a comprehensive background on rPPG and video-based heart rate estimation, covering signal acquisition and physics-inspired estimation principles. We will discuss signal-processing approaches and their limitations. We will delve into DL-based estimation systems, addressing the challenge of inherent aleatoric uncertainty (irreducible uncertainty from various stochastic factors like sensor variations and inter-subject differences) in the ground truth data annotation that hinders the development of generalized deep learning-based rPPG estimation system. To reinforce a robust DL model addressing these inherent uncertainties in rPPG data streams, we will discuss three novel deep-learning approaches. First, a multi-task learning method for rPPG estimation that learns the core rPPG features using shared embedding across noisy ground truth by separating individual targets. Second, a self-supervised learning method that leverages unlabeled rPPG data to learn the intrinsic rPPG properties and improve rPPG feature representation and signal reconstruction by incorporating the prior domain knowledge (HR frequency, phase, and temporal coherence). Third, a generative adversarial method that enables fine-grained rPPG learning without direct supervision using large-scale rPPG data. Finally, we will present the validation of the proposed methods across in-house MPSC-rPPG datasets, and multiple public datasets with open-source code for further exploration. Furthermore, we will demonstrate a real-time heart rate estimation system, RhythmEdge, using a low-cost camera sensor and describe the rPPG-specific pruning techniques to reduce rPPG model size for efficient edge implementation. We will conclude the tutorial with existing research gaps and potential directions in contactless rPPG and respiratory rate estimation research.
Nirmalya Roy, Zahid Hasan 0001
SMARTCOMP2
2023 NEV-NCD: Negative Learning, Entropy, and Variance Regularization Based Novel Action Categories Discovery
abstract
Novel Categories Discovery (NCD) facilitates learning from a partially annotated label space and enables deep learning (DL) models to operate in an open-world setting by identifying and differentiating instances of novel classes based on the labeled data notions. One of the primary assumptions of NCD is that the novel label space is perfectly disjoint and can be equipartitioned, but it is rarely realized by most NCD approaches in practice. To better align with this assumption, we propose a novel single-stage joint optimization-based NCD method, Negative learning, Entropy, and Variance regularization NCD (NEV-NCD). We demonstrate the efficacy of NEV-NCD in previously unexplored NCD applications of video action recognition (VAR) with the public UCF101 dataset and a curated in-house partial action-space annotated multi-view video dataset. Further, we perform a thorough ablation study by varying the composition of final joint loss and associated hyper-parameters. During our experiments with UCF101 and multi-view action dataset, NEV-NCD achieves ≈ 83% classification accuracy in test instances of labeled data. NEV-NCD achieves ≈ 70% clustering accuracy over unlabeled data outperforming both naive baselines and state-of-the-art pseudo-labeling-based approaches by ≈ 40% and ≈ 3.5% over both datasets.
Zahid Hasan 0001, Masud Ahmed, Abu Zaher Md Faridee, Sanjay Purushotham, Heesung Kwon, Hyungtae Lee, Nirmalya Roy
ICIP1
2023 An Online Continuous Semantic Segmentation Framework With Minimal Labeling Efforts
abstract
The annotation load for a new dataset has been greatly decreased using domain adaptation based semantic segmentation, which iteratively constructs pseudo labels on unlabeled target data and retrains the network. However, realistic segmentation datasets are often imbalanced, with pseudo-labels tending to favor certain "head" classes while neglecting other "tail" classes. This can lead to an inaccurate and noisy mask. To address this issue, we propose a novel hard sample mining strategy for an active domain adaptation based semantic segmentation network, with the aim of automatically selecting a small subset of labeled target data to fine-tune the network. By calculating class-wise entropy, we are able to rank the difficulty level of different samples. We use a fusion of focal loss and regional mutual information loss instead of cross-entropy loss for the domain adaptation based semantic segmentation network. Our entire framework has been implemented in real-time using the Robotics Operating System (ROS) with a server PC and a small Unmanned Ground Vehicle (UGV) known as the ROSbot2.0 Pro. This implementation allows ROSbot2.0 Pro to access any type of data at any time, enabling it to perform a variety of tasks with ease. Our approach has been thoroughly evaluated through a series of extensive experiments, which demonstrate its superior performance compared to existing state-of-the-art methods. Remarkably, by using just 20% of hard samples for fine-tuning, our network has achieved a level of performance that is comparable (≈88%) to that of a fully supervised approach, with mIOU scores of 60.51% in the In-house dataset.
Masud Ahmed, Zahid Hasan 0001, Tim Yingling, Eric O'Leary, Sanjay Purushotham, Suya You, Nirmalya Roy
SMARTCOMP2
2023 A Systematic Study on Object Recognition Using Millimeter-wave Radar
abstract
Millimeter-wave (MMW) radar is becoming an essential sensing technology in smart environments due to its light and weather-independent sensing capability. Such capabilities have been widely explored and integrated with intelligent vehicle systems, often deployed in industry-grade MMW radars. However, industry-grade MMW radars are often expensive and difficult to attain for deployable community-purpose smart environment applications. On the other hand, commercially available MMW radars pose hidden underpinning challenges that are yet to be well investigated for tasks such as recognizing objects, and activities, real-time person tracking, object localization, etc. Such tasks are frequently accompanied by image and video data, which are relatively easy for an individual to obtain, interpret, and annotate. However, image and video data are light and weather-dependent, vulnerable to the occlusion effect, and inherently raise privacy concerns for individuals. It is crucial to investigate the performance of an alternative sensing mechanism where commercially available MMW radars can be a viable alternative to eradicate the dependencies and preserve privacy issues. Before championing MMW radar, several questions need to be answered regarding MMW radar’s practical feasibility and performance under different operating environments. To answer the concerns, we have collected a dataset using commercially available MMW radar, Automotive mmWave Radar (AWR2944) from Texas Instruments, and reported the optimum experimental settings for object recognition performance using several deep learning algorithms in this study. Moreover, our robust data collection procedure allows us to systematically study and identify potential challenges in the object recognition task under a cross-ambience scenario. We have explored the potential approaches to overcome the underlying challenges and reported extensive experimental results.
Maloy Kumar Devnath, Avijoy Chakma, Mohammad Saeid Anwar, Emon Dey, Zahid Hasan 0001, Marc Conn, Biplab Pal, Nirmalya Roy
SMARTCOMP5
2023 SrPPG: Semi-Supervised Adversarial Learning for Remote Photoplethysmography with Noisy Data
abstract
Remote Photoplethysmography (rPPG) systems offer contactless, low-cost, and ubiquitous heart rate (HR) monitoring by leveraging the skin-tissue blood volumetric variation-induced reflection. However, collecting large-scale time-synchronized rPPG data is costly and impedes the development of generalized end-to-end deep learning (DL) rPPG models to perform under diverse scenarios. We formulate the rPPG estimation as a generative task of recovering time-series PPG from facial videos and propose SrPPG, a novel semi-supervised adversarial learning framework using heterogeneous, asynchronous, and noisy rPPG data. More specifically, we develop a novel encoder-decoder architecture, where rPPG features are learned from video in a self-supervised manner (encoder) to reconstruct the time-series PPG (decoder/generator) with physics-inspired novel temporal consistency regularization. The generated PPG is scrutinized against the real rPPG signals by a frequency-class conditioned discriminator, forming a generative adversarial network. Thus, SrPPG generates samples without point-wise supervision, alleviating the need for time-synchronized data collection. We experiment and validate SrPPG by amassing three public datasets in heterogeneous settings. SrPPG outperforms both supervised and self-supervised state-of-the-art methods in HR estimation across all datasets without any time-synchronous rPPG data. We also perform extensive experiments to study the optimal generative setting (architecture, joint optimization) and provide insight into the SrPPG behavior.
Zahid Hasan 0001, Abu Zaher Md Faridee, Masud Ahmed, Shibi Ayyanar, Nirmalya Roy
SMARTCOMP1
2022 CoDEm: Conditional Domain Embeddings for Scalable Human Activity Recognition
abstract
We explore the effect of auxiliary labels in improving the classification accuracy of wearable sensor-based human activity recognition (HAR) systems, which are primarily trained with the supervision of the activity labels (e.g. running, walking, jumping). Supplemental meta-data are often available during the data collection process such as body positions of the wearable sensors, subjects' demographic information (e.g. gender, age), and the type of wearable used (e.g. smartphone, smart-watch). This information, while not directly related to the activity classification task, can nonetheless provide auxiliary supervision and has the potential to significantly improve the HAR accuracy by providing extra guidance on how to handle the introduced sample heterogeneity from the change in domains (i.e positions, persons, or sensors), especially in the presence of limited activity labels. However, integrating such meta-data information in the classification pipeline is non-trivial - (i) the complex interaction between the activity and domain label space is hard to capture with a simple multi-task and/or adversarial learning setup, (ii) meta-data and activity labels might not be simultaneously available for all collected samples. To address these issues, we propose a novel framework Conditional Domain Embeddings (CoDEm). From the available unlabeled raw samples and their domain meta-data, we first learn a set of domain embeddings using a contrastive learning methodology to handle inter-domain variability and inter-domain similarity. To classify the activities, CoDEm then learns the label embeddings in a contrastive fashion, conditioned on domain embeddings with a novel attention mechanism, enforcing the model to learn the complex domain-activity relationships. We extensively evaluate CoDEm in three benchmark datasets against a number of multi-task and adversarial learning baselines and achieve state-of-the-art nerformance in each avenue.
Abu Zaher Md Faridee, Avijoy Chakma, Zahid Hasan 0001, Nirmalya Roy, Archan Misra
SMARTCOMP3
2022 RhythmEdge: Enabling Contactless Heart Rate Estimation on the Edge
abstract
The primary contribution of this paper is designing and prototyping a real-time edge computing system, RhythmEdge, that is capable of detecting changes in blood volume from facial videos (Remote Photoplethysmography; rPPG), enabling cardio-vascular health assessment instantly. The benefits of RhythmEdge include non-invasive measurement of cardiovascular activity, real-time system operation, inexpensive sensing components, and computing. RhythmEdge captures a short video of the skin using a camera and extracts rPPG features to estimate the Photoplethysmography (PPG) signal using a multi-task learning framework while offloading the edge computation. In addition, we intelligently apply a transfer learning approach to the multi-task learning framework to mitigate sensor heterogeneities to scale the RhythmEdge prototype to work with a range of commercially available sensing and computing devices. Besides, to further adapt the software stack for resource-constrained devices, we postulate novel pruning and quantization techniques (Quantization: FP32, FP16; Pruned-Quantized: FP32, FP16) that efficiently optimize the deep feature learning while minimizing the runtime, latency, memory, and power usage. We benchmark RhythmEdge prototype for three different cameras and edge computing platforms while evaluating it on three publicly available datasets and an in-house dataset collected under challenging environmental circumstances. Our analysis indicates that RhythmEdge performs on par with the existing contactless heart rate monitoring systems while utilizing only half of its available resources. Furthermore, we perform an ablation study with and without pruning and quantization to report the model size (87%) vs. inference time (70%) reduction. We attested the efficacy of RhythmEdge prototype with a maximum power of 8W and a memory usage of 290MB, with a minimal latency of 0.0625 seconds and a runtime of 0.64 seconds per 30 frames.
Zahid Hasan 0001, Emon Dey, Sreenivasan Ramasamy Ramamurthy, Nirmalya Roy, Archan Misra
SMARTCOMP1
2022 Demo: RhythmEdge: Enabling Contactless Heart Rate Estimation on the Edge
abstract
In this demo paper, we design and prototype RhythmEdge [1], a low-cost, deep-learning-based contact-less system for regular HR monitoring applications. RhythmEdge benefits over existing approaches by facilitating contact-less nature, real-time/offline operation, inexpensive and available sensing components, and computing devices. Our RhythmEdge system is portable and easily deployable for reliable HR estimation in moderately controlled indoor or outdoor environments. RhythmEdge measures HR via detecting changes in blood volume from facial videos (Remote Photoplethysmography; rPPG) and provides instant assessment using off-the-shelf commercially available resource-constrained edge platforms and video cameras. We demonstrate the scalability, flexibility, and compatibility of the RhythmEdge by deploying it on three resource-constrained platforms of differing architectures (NVIDIA Jetson Nano, Google Coral Development Board, Raspberry Pi) and three heterogeneous cameras of differing sensitivity, resolution, properties (web camera, action camera, and DSLR). RhythmEdge further stores longitudinal cardiovascular information and provides instant notification to the users. We thoroughly test the prototype stability, latency, and feasibility for three edge computing platforms by profiling their runtime, memory, and power usage.
Zahid Hasan 0001, Emon Dey, Sreenivasan Ramasamy Ramamurthy, Nirmalya Roy, Archan Misra
SMARTCOMP1