VLDB 2026 Research / reviewers in the wild / expert
Brejesh Lall
dblp:47/4820
· DBLP profile ↗
90ranked-venue papers
0as first author
46since 2021 · last 2026
0000-0003-2677-3071ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 39 · 21 since 2021Artificial intelligence and machine learning · 31 · 14 since 2021Computer networks · 14 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 7 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Performance Optimization Using Delay in Multirate Reconstruction Systems
Sandeep Patel, Ravindra Dhuli, Brejesh Lall |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2026 | DeepQ-Charge: A Decentralized Model-Free Framework for Real-Time Mobile EV Charger Dispatch Using Deep Reinforcement Learning
Taniya Manzoor, Ubaid Qureshi, Bijaya K. Panigrahi, Brejesh Lall |
IEEE Trans. Ind. Informatics | 4 |
| 2025 | DSFace : Conditional Diffusion Inpainting for Sketch-to-Face SynthesisabstractGenerating realistic human faces from monochromatic sketches is a challenging task with applications in forensic reconstruction, character design, and digital art. The limited semantic information in single-channel sketches makes this problem difficult, as they often lack fine details like expressions, skin tone and accessories. While GANs have shown promise, they suffer from unstable training and poor structural guidance. Diffusion models offer improved image generation, but struggle with monochrome inputs and high computational costs. To address these challenges, we propose DSFace, a latent diffusion-based framework that treats sketch-to-face generation as a conditional inpainting problem. DSFace utilises a frozen Paint-by-Example (PBE) inpainting diffusion model, conditioned with a ControlNet encoder, ensuring precise control over face synthesis. Our novel approach utilises a GAN-generated coarse image to compute DINO-V2 embeddings, which provide fine-grained details for improved facial and garment features. Trained on the CUFS dataset, DSFace achieves state-of-the-art performance, surpassing existing methods in visual realism, perceptual quality and structural alignment with the input sketches. Sanhita Pathak, Vinay Kaushik, Brejesh Lall |
ICIP | 3 |
| 2025 | CLEARSTR: Contextual Learning with Edge-guided and Adaptive-texture Reconstruction for Scene Text RemovalabstractScene text removal is a challenging task in computer vision, requiring the seamless restoration of text-masked regions while preserving the structural and aesthetic coherence of the background. Current methods often fail to achieve natural integration of the restored regions into the surrounding context, especially in complex scenes. We introduce structure guidance to the task of scene text removal utilising a novel framework that extends denoising diffusion probabilistic models (DDPMs) for text removal tasks. Our approach integrates depth-aware neighborhood estimation to identify regions with similar depth profiles near the text-masked area, providing spatial cues to guide the inpainting process. Additionally, the model leverages localized texture reconstruction, ensuring that the synthesized textures align with the intricate details of the surrounding image. We propose a unified approach for context-aware guidance that dynamically integrates both depth and spatial proximity constraints into a single, coherent neighborhood definition. To ensure semantic consistency in generated scene image we also propose context loss. We evaluate our approach on the SCUT-EnsText and SCUT-Syn datasets, demonstrating its ability to achieve superior text removal quality, combining high perceptual fidelity with robust quantitative performance. By incorporating structural depth information and context-aware texture generation, this work sets a new benchmark in scene text removal research. Sanhita Pathak, Vinay Kaushik, Brejesh Lall |
ICME | 3 |
| 2025 | Cricket Shot Analysis using Conditional Directed Spatio-Temporal Graph networksabstractRecent Use of Conditional Spatio-temporal Directed Graph Convolutional Networks(Cond ST-DGCN) [8] to represent human pose estimation has significantly helped in capturing varying non-local dependencies between limbs for different actions. This can be immensely helpful in Sports analytics where player pose plays key role in shot evaluation and can help in corrective action. In this article, we propose CondDGCN [8] based framework to explore use of spatial-temporal relation of batsman shot sequences (labelled and annotated 2D cricket dataset [1]) for Cricket shot action recognition by conditioning the graph network on batsman 2D poses. We achieve 97% accuracy for shot recognition and further explore visualization of conditional graph connections to establish importance of particular limbs for shots. The proposed framework uses fine-tuned 2D Pose estimator OpenPose [11] (fine-tuned for cricket dataset [1]) which in turn helps in easy adaptation of our solution to internet cricket videos for shot analytics. Tanu Aneja, Brejesh Lall, Harsha Malaviya |
IPAS | 2 |
| 2025 | Strategic Offloading of Grant-Based Traffic to Grant-Free Resources: A DRL Based Approach for Optimizing Bandwidth in 5G-Advanced NetworksabstractEfficient bandwidth utilization is essential for nextgeneration cellular networks such as 5G -Advanced and 6G, to support diverse services with varying quality-of-service (QoS) requirements. Grant-free (GF) communication, introduced in 5G NR, enables unscheduled data transmission, significantly reducing latency. However, it often leads to inefficiencies, such as underutilized GF bandwidth in scenarios involving sporadic status messages. In contrast, grant-based (GB) communication, which relies on scheduled transmissions, offers greater reliability. This paper proposes a novel deep reinforcement learning (DRL)-based mechanism to strategically offload GB traffic to underutilized GF resources. Termed “soft access class barring” (SACB), the approach dynamically offloads traffic from GB to GF bandwidth to optimize spectrum utilization while maintaining QoS, without requiring complex bandwidth reconfigurations. Simulation results demonstrate the potential of SACB to address challenges in hybrid traffic management, paving the way for its adoption in future network standards. Yukti Kaura, Brejesh Lall, Ranjan K. Mallik, Amit Singhal 0002 |
VTC2025-Spring | 2 |
| 2025 | COVITON : Consistency driven integration of TPS and flow for virtual tryon
Sanhita Pathak, Vinay Kaushik, Brejesh Lall |
Comput. Graph. | 3 |
| 2025 | Multi spectral visible-thermal IR image translation using improved u-net & conditional diffusion
Mahroosh Banday, Brejesh Lall |
Neurocomputing | 2 |
| 2025 | ThermalDiff: A diffusion architecture for thermal image synthesis
Tayeba Qazi, Brejesh Lall, Prerana Mukherjee |
J. Vis. Commun. Image Represent. | 2 |
| 2025 | Adaptive Scheduling of Shared Grant-Free Resources for Heterogeneous Massive Machine Type Communication in 5G and Beyond NetworksabstractMassive machine-type communication (mMTC) has been identified as a key service type in fifth-generation new radio (5G NR) communication systems. The third-generation partnership (3GPP) project, starting with 5G, has introduced grant-free (GF) or configured grant (CG) scheduling for uplink traffic with small data packets to reduce signaling and latency overheads as compared to prevalent grant-based (GB) schemes. However, when heterogeneous MTC devices compete for pre-configured, shared GF resources, the access results in collisions. No standardized methods exist for ensuring priority-based access in the shared GF scheduling scheme. In this work, we introduce novel methods which utilize both heuristic and multi-objective deep reinforcement learning (DRL) techniques for priority-enabled GF access. The proposed methods adaptively partition GF bandwidth resources per allocation interval for scheduling configured grants to heterogeneous MTC device groups in a way which improves their probability of successful transmission, thereby resulting in a lower average age of information and packet drop rate and simultaneously ensuring fairness. Through extensive simulations set in the context of cyber-physical systems (CPS) with diverse quality-of-service (QoS) requirements across various 5G NR numerology schemes, we exemplify that our proposed approach provides significantly better performance and resource utilization than conventional schemes. Yukti Kaura, Brejesh Lall, Ranjan K. Mallik, Amit Singhal 0002 |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2025 | Garment Recycle Training and Conditional Garment-Person Outline Attention-Guided Virtual TryonabstractVirtual try-on, a significant application in computer vision, aims to seamlessly simulate the appearance of clothing on a person from a single image. We propose a diffusion-based tryon approach, solving virtual tryon as a problem of conditional image inpainting. Our method introduces GarNet and OutlineNet as two learnable Stable Diffusion ControlNet encoders conditioned on the garment and person outline images, enhancing the controllability and realism of the generated try-on. We propose a two-stage garment diffusion recycling training strategy, utilizing \(x_{0}\) -parameterization. We estimate the initial clean image that is conditioned on the maximum noised input and feed the same to the same diffusion model again to estimate total noise. This reduces over-fitting and makes our model more generalized. We also introduce a zero garment-outline conditioning (ZGOC) block along with a Garment-Outline Cross Attention layer to optimize garment draping and ensure global consistency in the try-on results. The ZGOC block provides control and adaptability by prioritizing garment details that are most affected by body shape, ensuring precise garment alignment with the person’s outline. Our comprehensive experiments on the VITON-HD and Dresscode dataset demonstrate that our proposed approach achieves state-of-the-art realism and controllability in VITON, marking a significant advancement in virtual fashion experiences and online shopping applications. Sanhita Pathak, Vinay Kaushik, Brejesh Lall |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Thermal Videodiff (TVD): A Diffusion Architecture For Thermal Video SynthesisabstractMachine perception gathers information about its surroundings through sophisticated sensors. Thermal sensors offers scalable perception in conditions of low visibility such as night time, smoke or fog. However, integrating thermal data with deep learning algorithms is a challenge because of scarcity of thermal data due to high cost of thermal sensors. In this paper, we propose for the first time, a deep learning based framework for thermal video synthesis as an affordable alternative to purchasing costly thermal imaging devices. Here, we have introduced diffusion model for estimating thermal videos from videos in the visible spectrum. Results show that our Thermal VideoDiff (TVD) is capable of synthesizing high fidelity video samples and captures temperature variations from thermal data effectively. Our work addresses the challenge posed by the scarcity of thermal data, as well as brings deep learning to the domain of infrared video generation, enabling research and development in infrared domain. The implementation details of the model and fundamentals of thermal imaging are available at https://github.com/Tayeba/TVD. Tayeba Qazi, Brejesh Lall |
ICIP | 2 |
| 2024 | Single Stage Warped Cloth Learning and Semantic-Contextual Attention Feature Fusion for Virtual TryonabstractImage-based virtual try-on aims to fit an in-shop garment onto a clothed person image. Garment warping, which aligns the target garment with the corresponding body parts in the person image, is a crucial step in achieving this goal. Existing methods often use multi-stage frameworks to handle clothes warping, person body synthesis and tryon generation separately or rely on noisy intermediate parser-based labels. We propose a novel single-stage framework that implicitly learns the same without explicit multi-stage learning. Our approach utilizes a novel semantic-contextual fusion attention module for garment-person feature fusion, enabling efficient and realistic cloth warping and body synthesis from target pose keypoints. By introducing a lightweight linear attention framework that attends to garment regions and fuses multiple sampled flow fields, we also address misalignment and artifacts present in previous methods. To achieve simultaneous learning of warped garment and try-on results, we introduce a Warped Cloth Learning Module. Our proposed approach significantly improves the quality and efficiency of virtual try-on methods, providing users with a more reliable and realistic virtual try-on experience. Sanhita Pathak, Vinay Kaushik, Brejesh Lall |
ICME | 3 |
| 2024 | ICPR 2024 Leaf Inspect Competition: Leaf Instance Segmentation and Counting
Swati Bhugra, Prerana Mukherjee, Vinay Kaushik, Siddharth Srivastava 0004, Viswanathan Chinnusamy, Brejesh Lall, Santanu Chaudhary |
ICPR (34) | 7 |
| 2024 | SUPI-Rear: Privacy-Preserving Subscription Permanent Identification Strategy in 5G-AKA
K. Sowjanya, Pabitra Pal, Aman Verma, Bijoy Das, Dhiman Saha, Anand M. Baswade, Brejesh Lall |
SSS | 7 |
| 2024 | Semantic-Aided Image Transmission System with Unequal Error Protection for Next-Generation Communication NetworksabstractSemantic communication (SC) aims to convey the meaning of data instead of focusing on its bit-by-bit reconstruction. SC finds applications in beyond 5G and 6G networks for artificial intelligence-empowered multimedia content delivery. In this paper, we propose a novel semantic-aided autoencoder-based image transmission system that leverages semantic information in the form of the segmentation map of an image. We demonstrate up to 23% and 18% improvement (in terms of mean square error and peak signal-to-noise ratio, respectively) in the quality of the received image with only 2% extra bandwidth over a traditional autoencoder-based image transmission system. The study also explores channel coding strategies for our proposed system. We focus on the intrinsically robust nature of semantic data, as compared to traditional data, to design low-density parity check code, Hamming code, and polar code-based unequal error protection (UEP) schemes. Comparative evaluations between UEP and equal error protection schemes show that while both approaches yield similar performance, UEP schemes are more efficient. Nargis Fayaz, Aman Shreshtha, Smruti R. Sarangi, Ranjan K. Mallik, Brejesh Lall |
WCNC | 5 |
| 2024 | Multi-armed bandit based online model selection for concept-drift adaptationabstractAbstract Ensemble methods are among the most effective concept‐drift adaptation techniques due to their high learning performance and flexibility. However, they are computationally expensive and pose a challenge in applications involving high‐speed data streams. In this paper, we present a computationally efficient heterogeneous classifier ensemble entitled OMS‐MAB which uses online model selection for concept‐drift adaptation by posing it as a non‐stationary multi‐armed bandit (MAB) problem. We use a MAB to select a single adaptive learner within the ensemble for learning and prediction while systematically exploring promising alternatives. Each ensemble member is made drift resistant using explicit drift detection and is represented as an arm of the MAB. An exploration factor controls the trade‐off between predictive performance and computational resource requirements, eliminating the need to continuously train and evaluate all the ensemble members. A rigorous evaluation on 20 benchmark datasets and 9 algorithms indicates that the accuracy of OMS‐MAB is statistically at par with state‐of‐the‐art (SOTA) ensembles. Moreover, it offers a significant reduction in execution time and model size in comparison to several SOTA ensemble methods, making it a promising ensemble for resource constrained stream‐mining problems. Jobin Wilson, Santanu Chaudhury, Brejesh Lall |
Expert Syst. J. Knowl. Eng. | 3 |
| 2023 | Deep Learning based Diagnostic and Severity Assessment Framework for Lung Diseases using Chest RadiographsabstractComputer-aided diagnosis and prediction of the severity of lung diseases is a promising way to help overburdened medical experts in accelerating and improving their diagnosis. The objective of this work is to investigate the use of deep learning techniques to design a framework for the automatic diagnosis of lung diseases along with the prediction of severity using chest radiographs. We identified input chest radiographs as healthy or belonging to patients with lung disease along with the confidence score of prediction. The unhealthy chest radiograph is further examined to calculate clinical parameters considered in the severity prediction of lung diseases. We calculate clinical parameters such as the extent of lung involvement in disease manifestation, the type of abnormalities present in chest radiographs, and their location in terms of lung zones. We conduct experiments with our in-house Indian database and achieved an accuracy of 95.65% in the classification between healthy and unhealthy chest radiographs. We obtained average precision scores of 0.8128, 1.00, 0.8214, and 0.9650 for the detection of effusion, cavity, lymphadenopathy, and opacity respectively. Experimental results indicated that the proposed framework can be used to provide rapid and cost-effective screening in places where massive traditional testing is not feasible. Anushikha Singh, Brejesh Lall, Bijaya K. Panigrahi, Anjali Agrawal, Anurag Agrawal, Balamugesh Thangakunam, D. J. Christopher |
CBMS | 2 |
| 2023 | Spectral Grouping Driven Hyperspectral Super-ResolutionabstractConvolutional neural networks have proven to be proficient when extracting low-level concepts in an image. With the wonderful performance of transformers in exploiting the long-range correlations in an image, many methods have been explored where one exploit benefits of both the architectures. Therefore, in order to strengthen our network we add an important feature to transformers wherein single image super-resolution (SISR) is exploited using band grouping leveraging a simple CNN architecture. This paper aims to train a set of simple residual modelling architectures and then integrate them into a transformer architecture to solve super-resolution problem in HSI. We take a step forward to analyse how to adapt swinIR to fully exploit the information derived from band grouping for efficient SISR. Sadia Hussain, Brejesh Lall |
ICIP | 2 |
| 2023 | Hierarchical Multi-task Learning via Task Affinity GroupingsabstractMulti-task learning (MTL) permits joint task learning based on a shared deep learning architecture and multiple loss functions. Despite the recent advances in MTL, one loss often dominates the learning optimization in multiple unrelated tasks. This often results in poor performance compared to the corresponding single task learning. To overcome the aforementioned "negative transfer", we propose a novel hierarchical framework that leverages task relations via inter-task affinity to supervise multi-task learning. Specifically, the inter-task affinity generated task sets, with low-level task set and complex task set at the bottom and top layers respectively, enables iterative multi-task information sharing. In addition, it also alleviates simultaneous image annotations for multiple tasks. The proposed framework achieves state-of-the-art results on classification, detection, semantic segmentation and depth estimation across three standard benchmarks. Furthermore, with state of the results on two benchmarks for image retrieval task, we also demonstrate that the embeddings learned using such a framework provide good generalization and robust representation learning. Siddharth Srivastava 0004, Swati Bhugra, Vinay Kaushik, Brejesh Lall |
ICIP | 4 |
| 2023 | Weakly Supervised Few-Shot and Zero-Shot Semantic Segmentation with Mean Instance Aware Prompt LearningabstractThe existing Zero-Shot Segmentation (ZSS) and Few-Shot Segmentation (FSS) methods utilize fully supervised pixel-labeled seen classes to segment unseen classes. Pixel-level labels are hard to obtain, and using weak supervision in the form of inexpensive image labels is often more practical. To this end, we propose a novel unified weakly supervised Zero-Shot and Few-Shot semantic segmentation pipeline that can perform ZSS and FSS on novel classes without using pixel-level labels for either the base (seen) or the novel (unseen) classes. We propose Mean Instance Aware Prompt based Network (MIAPNet), a novel language-guided segmentation pipeline that i) learns context vectors with batch aggregates (mean) to map class prompts to image features and ii) decouples weak ZSS/FSS into weak semantic segmentation and Zero-Shot segmentation. MIAPNet beats existing methods for weak generalized ZSS and weak FSS by 39 and 3 mIOU points respectively on PASCAL VOC and weak FSS by 5 mIOU points on MS COCO. Prashant Pandey 0002, Mustafa Chasmai, Monish Natarajan, Brejesh Lall |
ICME | 4 |
| 2023 | KL Regularized Normalization Framework for Low Resource Tasks
Neeraj Kumar 0010, Ankur Narang, Brejesh Lall |
ECML/PKDD (4) | 3 |
| 2023 | Altering Backward Pass Gradients to Improve Convergence (S)abstractIn standard neural network training, the gradients in the backward pass are determined by the forward pass.As a result, the two stages are coupled.This is how most neural networks are trained hitherto.Gradient modification in the backward pass has seldom been studied in the literature.In this paper we explore decoupled training, where we alter the gradients in the backward pass.We propose a simple yet powerful method called PowerGrad Transform (PGT), that alters the gradients before the weight update in the backward pass and significantly enhances the predictive performance of a convolutional neural network.PGT trains networks to arrive at a better optima at convergence.It is computationally efficient, and adds no additional cost to either memory or compute, but results in improved final accuracies on both the training and test datasets.Power-Grad Transform is easy to integrate into existing training routines, requiring just a few lines of code.With decoupled training, our method improves baseline accuracies for ResNet-50 by 0.73%, for SE-ResNet-50 by 0.66% and by more than 1.0% for the non-normalized ResNet-18 network on the ImageNet classification task. Bishshoy Das, Milton Mondal, Brejesh Lall, Shiv Dutt Joshi, Sumantra Dutta Roy |
SEKE | 3 |
| 2023 | AnoLeaf: Unsupervised Leaf Disease Segmentation via Structurally Robust Generative InpaintingabstractPlant diseases severely limits agriculture production, necessitating the high-throughput monitoring of plant leaves. Currently, this is formulated as an automatic disease segmentation task addressed via deep learning frameworks. These deep leaning frameworks trained with leaf image data in a supervised paradigm have few limitations, mainly: (1) training datasets are heavily imbalanced towards healthy leaf images, (2) disease region annotation is labour-intensive and (3) due to the heterogeneity of disease symptoms, these frameworks lacks generalisability. In this paper, we reformulate disease segmentation as an anomaly localisation task. Specifically, we introduce a novel unsupervised framework (AnoLeaf) based on an edge-guided in-painting that optimises the learning of contextual attention on only healthy leaf images. The network utilisation on diseased leaf images results in reconstruction of its healthy counterparts, generating an inpainting error. The contextual attention maps reinforce the inpainting error to effectively localise the disease. Thus, AnoLeaf alleviates the acquisition and annotation of rare disease images. Additional experiments on MVTec anomaly detection dataset further demonstrate its generalisability. Swati Bhugra, Vinay Kaushik, Brejesh Lall, Santanu Chaudhury |
WACV | 4 |
| 2023 | On the Optimal Beamwidth of UAV-Assisted Networks Operating at Millimeter WavesabstractThe millimeter-wave (mm-wave) bands enable very large antenna arrays that can generate narrow beams for beamforming and spatial multiplexing. However, directionality introduces beam misalignment and leads to reduced energy efficiency. Thus, employing the narrowest possible beam in a cell may not necessarily imply maximum coverage. The objective of this work is to determine the optimal sector beamwidth for a cellular architecture served by an unmanned aerial vehicle (UAV) acting as a base station (BS). The users in a cell are assumed to be distributed according to a Poisson Point Process (PPP) with a given user density. We consider hybrid beamforming at the UAV, such that multiple concurrent beams serve all the sectors simultaneously. An optimization problem is formulated to maximize the sum rate over a given area while limiting the total power available to each sector. We observe that, for a given transmit power, the optimal sector beamwidth increases as the user density in a cell decreases, and varies based on the height of the UAV. Thus, we provide guidelines towards the optimal beamforming configurations for users in rural areas. Manishika Rawat, Marco Giordani, Brejesh Lall, Abdelaali Chaoub, Michele Zorzi |
WCNC | 3 |
| 2023 | Extending function mixture network for improved spectral super-resolution
Sadia Hussain, Brejesh Lall |
Comput. Vis. Image Underst. | 2 |
| 2023 | Feature independent Filter Pruning by Successive Layers analysis
Milton Mondal, Bishshoy Das, Brejesh Lall, Pushpendra Singh 0002, Sumantra Dutta Roy, Shiv Dutt Joshi |
Comput. Vis. Image Underst. | 3 |
| 2023 | Homogeneous-Heterogeneous Hybrid Ensemble for concept-drift adaptation
Jobin Wilson, Santanu Chaudhury, Brejesh Lall |
Neurocomputing | 3 |
| 2023 | A novel energy-efficient adaptive superframe structure for OWC-based real-time bio-sensor networks
Monica Bhutani, Brejesh Lall, Monika Agrawal 0002 |
Peer Peer Netw. Appl. | 2 |
| 2023 | Reconstruction of a signal from multirate observations: A recursive approach
Sandeep Patel, Ravindra Dhuli, Brejesh Lall |
Signal Process. | 3 |
| 2023 | Robust Prototypical Few-Shot Organ Segmentation With Regularized Neural-ODEsabstractDespite the tremendous progress made by deep learning models in image semantic segmentation, they typically require large annotated examples, and increasing attention is being diverted to problem settings like Few-Shot Learning (FSL) where only a small amount of annotation is needed for generalisation to novel classes. This is especially seen in medical domains where dense pixel-level annotations are expensive to obtain. In this paper, we propose Regularized Prototypical Neural Ordinary Differential Equation (R-PNODE), a method that leverages intrinsic properties of Neural-ODEs, assisted and enhanced by additional cluster and consistency losses to perform Few-Shot Segmentation (FSS) of organs. R-PNODE constrains support and query features from the same classes to lie closer in the representation space thereby improving the performance over the existing Convolutional Neural Network (CNN) based FSS methods. We further demonstrate that while many existing Deep CNN-based methods tend to be extremely vulnerable to adversarial attacks, R-PNODE exhibits increased adversarial robustness for a wide array of these attacks. We experiment with three publicly available multi-organ segmentation datasets in both in-domain and cross-domain FSS settings to demonstrate the efficacy of our method. In addition, we perform experiments with seven commonly used adversarial attacks in various settings to demonstrate R-PNODE's robustness. R-PNODE outperforms the baselines for FSS by significant margins and also shows superior performance for a wide array of attacks varying in intensity and design. Prashant Pandey 0002, Mustafa Chasmai, Tanuj Sur, Brejesh Lall |
IEEE Trans. Medical Imaging | 4 |
| 2022 | Adversarially Robust Prototypical Few-Shot Segmentation with Neural-ODEs
Prashant Pandey 0002, Aleti Vardhan, Mustafa Chasmai, Tanuj Sur, Brejesh Lall |
MICCAI (8) | 5 |
| 2022 | Adaptive CNN filter pruning using global importance metric
Milton Mondal, Bishshoy Das, Sumantra Dutta Roy, Pushpendra Singh 0002, Brejesh Lall, Shiv Dutt Joshi |
Comput. Vis. Image Underst. | 5 |
| 2022 | Depth analysis of kinect v2 sensor in different mediums
Aditi Bhateja, Adarsh Shrivastav, Himanshu Chaudhary, Brejesh Lall, Prem Kumar Kalra |
Multim. Tools Appl. | 4 |
| 2022 | Correction to: depth analysis of kinect v2 sensor in different mediums
Aditi Bhateja, Adarsh Shrivastav, Himanshu Chaudhary, Brejesh Lall, Prem Kumar Kalra |
Multim. Tools Appl. | 4 |
| 2022 | On impropriety for a large-sized discrete fourier transform of a real-valued stationary process
Ramanjit Ahuja, Brejesh Lall, Surendra Prasad |
Signal Process. | 2 |
| 2022 | Zero-Shot Normalization Driven Multi-Speaker Text to Speech SynthesisabstractText-to-speech (TTS) systems are designed to synthesize natural and expressive speech, adapt to an unseen voice, and capture the speaking style of an unseen speaker by converting text into speech. The introduction of an unseen speaker’s speaking style into a TTS system offers a wide range of application scenarios, including personal assistant, news broadcast, and audio navigation, among others. The style of the speech varies from person to person and every person exhibits his or her style of speaking that is determined by the language, demography, culture and other factors. Style is best captured by the prosody of a signal. It is an ongoing research area with numerous real-world applications that produces high-quality multi-speaker voice synthesis while taking into account prosody and in a zero-shot manner. Despite the fact that several efforts have been made in this area, it continues to be an interesting and difficult topic to solve. In this paper, we present a novel zero-shot multi-speaker speech synthesis approach (ZSM-SS) that leverages the normalization architecture and speaker encoder with non-autoregressive multi-head attention driven encoder-decoder architecture. Given an input text and a reference speech sample of an unseen person, ZSM-SS can generate speech in that person’s style in a zero-shot manner. Additionally, we demonstrate how the affine parameters of normalization help in capturing the prosodic features such as energy and fundamental frequency in a disentangled fashion and can be used to generate morphed speech output. We generate the 256 dimensional speaker embedding using a speaker encoder based on wav2vec2.0 based architecture. We demonstrate the efficacy of our proposed architecture on multi-speaker VCTK [1] and LibriTTS [2] datasets, using visualization of hessian of proposed model, multiple quantitative metrics that measure generated speech distortion and MOS, along with speaker embedding analysis of the proposed speaker encoder model. Neeraj Kumar 0010, Ankur Narang, Brejesh Lall |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Temporal Causal Modelling on Large Volume Enterprise DataabstractStructural Causal Modelling (SCM) with its intervention analysis is one of the promising modelling approach that assists in data driven decision making. SCM not only overcomes the black box modelling associated with most of the classification algorithms but also gives enterprises an opportunity to perform intervention analysis without having to perform randomized controlled experiments. But the large volume of enterprises' data pose challenges in learning the causal structure as existing algorithms are not suitable to learn from data present in Distributed File System (DFS). Hence algorithm presented in this paper, proposes a novel variation to PC-Stable algorithm to efficiently learn the causal structure from data present in DFS - thus enabling temporal causal modelling on large volume time-series data. The proposed learning algorithm is used to determine the causal story associated with churn in telecommunication industry and flight delay in airline industry. Our model identifies and quantifies the respective causal factors for unfavourable events churn and flight delay. Ram Mohan, Santanu Chaudhury, Brejesh Lall |
IEEE Trans. Big Data | 3 |
| 2022 | A Recursive Method for Estimating Missing Data in Spatio-Temporal ApplicationsabstractMissing data is a major data reliability problem in spatio-temporal (ST) applications. This article proposes an online method for estimating missing data in case of a network of$n$sensors. The true sensor value at a specific location is expressed using an integro-difference equation. The Karhunen–Loéve Expansion of the spatial process allows one to represent the ST field values at$n$locations in the form of a linear state-space model. The parameters of the model are identified using the maximum likelihood method. The parameters are updated in a rolling window approach. Whenever missing data are encountered, the algorithm predicts the missing observations based on the constrained solution of state evolution equation. The constrained solution is obtained by representing the optimal state as the orthogonal sum decomposition of a deterministic and a stochastic component. The utility of the algorithm is presented on two sensor network datasets. Abhishek Grover, Brejesh Lall |
IEEE Trans. Ind. Informatics | 2 |
| 2021 | Multi Modal Normalization
Neeraj Kumar 0010, Ankur Narang, Brejesh Lall, Srishti Goel |
ICONIP (1) | 3 |
| 2021 | Normalization Driven Zero-Shot Multi-Speaker Speech Synthesis
Neeraj Kumar 0010, Srishti Goel, Ankur Narang, Brejesh Lall |
Interspeech | 4 |
| 2021 | Automatic Quantification of Plant Disease from Field Image Data Using Deep LearningabstractPlant disease is a major factor in yield reduction. Thus, plant breeders currently rely on selecting disease-resistant plant cultivars, which involves disease severity rating of a large variety of cultivars. Traditional visual screening of these cultivars is an error-prone process, which necessitates the development of an automatic framework for disease quantification based on field-acquired images using unmanned aerial vehicles (UAVs) to augment the throughput. Since these images are impaired by complex backgrounds, uneven lighting, and densely overlapping leaves, state-of-the-art frameworks formulate the processing pipeline as a dichotomy problem (i.e. presence/absence of disease). However, additional information regarding accurate disease localization and quantification is crucial for breeders. This paper proposes a deep framework for simultaneous segmentation of individual leaf instances and corresponding diseased region using a unified feature map with a multi-task loss function for an end-to-end training. We test the framework on field maize dataset with Northern Leaf Blight (NLB) disease and the experimental results show a disease severity correlation of 73% with the manual ground truth data and run-time efficiency of 5fps. Kanish Garg, Swati Bhugra, Brejesh Lall |
WACV | 3 |
| 2021 | MAC layer performance modelling for IEEE 802.15.7 based on discrete-time Markov chainabstractAbstract Potential lack of enough wireless frequency spectrum has guided researchers towards visible light communication (VLC), emerging as a sturdy support to Wi‐Fi. Though the IEEE 802.15.7 standard has comprehensive medium access control (MAC) and physical (PHY) layer specifications, the emerging VLC technology still faces MAC challenges. Further, throughput is one of the major concerns for the VLC personal area network (VPAN) as it directly impacts the network speed. We propose a novel Markov chain model with two clear channel assessments, which proves to be a significant milestone in improving the network throughput for a star topology VPAN. This paper extends the already available analytical models for the MAC layer to efficiently plan and predict the VPAN performance. We also extensively evaluate other network metrics, like network collision probability and end‐to‐end epoch, to demonstrate the proposed model's applicability. We verify the model's analytical results with elaborate MATLAB simulations and obtain good estimation, especially for large network sizes. Monica Bhutani, Brejesh Lall, Abhishek Dixit 0002 |
IET Commun. | 2 |
| 2021 | An Ontology Representation Language for Multimedia Event ApplicationsabstractThis paper presents formalization of a new Multimedia Web Ontology Language (E-MOWL) to handle events with media depictions. The temporal, spatial and entity aspects that are implicitly linked to an event are represented through this language to model the context of events. The already existing Multimedia Web Ontology Language (MOWL) can be leveraged for perceptual modelling of a domain, where the concepts manifest into media patterns in the multimedia document and helps in semantic processing of the contents. The language E-MOWL provides a rich method for representing knowledge corresponding to a specific domain wherein the context specifies the intended meaning of each element of the domain of discourse; an element in different context may correspond to different functional role. The context information associated with an event ties the audiovisual data with event related aspects. All these aspects when considered altogether provide the evidence and contribute towards recognizing an event from multimedia documents. The language also enables reasoning with the uncertainty associated with the events and is organized in the form of Bayesian Network (BN). The media items that are semantically relevant can be assimilated together on the basis of their association with events. We have demonstrated the efficacy of our approach by utilizing an ontology for the entertainment category in news domain to offer an application \textit{news aggregation} and event-based book recommendations. Nisha Pahal, Brejesh Lall, Santanu Chaudhury |
J. Web Eng. | 2 |
| 2021 | Measure for degree of time variance and measure for degree of non-stationarity: application to discrete LPTV systems
Ayan Kumar Dutta, Shiv Dutt Joshi, Brejesh Lall |
Signal Process. | 3 |
| 2021 | A Data-Driven Framework for Deploying Sensors in Environment Sensing ApplicationabstractSensor deployment routines for environment sensing applications make various assumptions about the underlying spatiotemporal field. These assumptions render the deployment ineffective in a practical scenario. This article proposes a two-step process: initially, the sensors are deployed based on geographical covariates. Then, after a fixed period, the data collected from sensors are used to find optimal locations for sensors. The spatiotemporal representation of sensor values has been modeled as the sum of a systematic trend component and a residual process. The trend component is modeled as the sum of deterministic functions, and the residual component is modeled using support vector regression. The locations with maximum support vector count in the residual model are identified as optimal for the deployment of sensors. The method can be used for both static and dynamic deployments. The proposed strategy has been applied to a specific case study of air pollution dataset. Abhishek Grover, Brejesh Lall |
IEEE Trans. Ind. Informatics | 2 |
| 2020 | A Hierarchical Framework for Leaf Instance Segmentation: Application to Plant PhenotypingabstractImage based analysis of plants is a high-throughput and non-invasive approach to study plant traits. The quantitative estimation of many plant traits (leaf area index, biomass etc.) from plant images is primarily based on accurate segmentation of individual leaves. This is a challenging task due to the presence of overlapped leaves and lack of discernible boundaries between them. To overcome these limitations, state-of-the-art supervised deep learning algorithms have been recently employed. However, the annotations of individual leaf instances is time consuming, in addition the variability in leaf shapes and its arrangement among different plant species limits the broad utilisation of these algorithms. To relieve this bottleneck, we propose a novel framework that relies on a graph based formulation to extract leaf shape knowledge for the task of leaf instance segmentation. These shape priors are generated based on leaf shape characteristics independent of plant species. Evaluation of the proposed framework on multiple plant datasets i.e. Arabidopsis, Komatsuna and salad demonstrates its broad utility. Swati Bhugra, Kanish Garg, Santanu Chaudhury, Brejesh Lall |
ICPR | 4 |
| 2020 | Using Scene Graphs for Detecting Visual RelationshipsabstractIn this paper we solve the problem of detecting relationships between pairs of objects in an image. We develop spatially aware word embeddings using scene graphs and use joint feature representations containing visual, spatial and semantic embeddings from the input images to train a deep network on the task of relationship detection. Further, we propose to utilize context aligned scene graph embeddings from the train set, without requiring explicit availability of scene graphs at test time. We show that the proposed method outperforms the state-of-the-art methods for predicate detection and provides competing results on relationship detection. We also show the generalization ability of the proposed method by performing predictions under zero shot settings. Further, we also provide an exhaustive empirical evaluation on each component of the proposed network. Anurag Tripathi, Siddharth Srivastava 0004, Brejesh Lall, Santanu Chaudhury |
ICPR | 3 |
| 2020 | Improved Polynomial Neural Networks with Normalised ActivationsabstractPolynomials, which are widely used to study non-linear systems, have been shown to be extremely useful in analyzing neural networks (NNs). However, the existing methods for training neural networks with polynomial activation functions (PAFs), called as PNNs, are applicable for shallow networks and give a stable performance with quadratic PAFs only. This is due to the optimization issues encountered during training PNNs. We propose a working model for PAFs using a novel normalizing transformation which alleviates the problem of training PNNs with arbitrary degree. Our PAF can be directly used to train shallow PNNs in practice for degrees as high as ten. It can also be utilized to learn multivariate sparse polynomials of small degrees. We also propose a way to train deep CNNs with PAFs which achieve performance similar to deep CNNs with standard activations. Through rigorous experimentation on multiple data sets, we show that PNNs can be effectively trained in practice. This also highlights the potential of the proposed method to support the research on using polynomials to study deep learning. Mohit Goyal, Rajan Goyal, Brejesh Lall |
IJCNN | 3 |
| 2020 | Paradigm Shift in Public Warning Systems: A Two-tier Approach towards BroadcastingabstractWith the recent outbreak of COVID-19 and other pandemics, improving the public safety communication is essential for efficient communication in the 5G and beyond wireless communication networks. The key requirements shall be lower delays, improved coordination and efficient resource utilization, to achieve higher efficiency in the network performance in an emergency/pandemic situation. Since deployment costs and scarce resource availability are major constraints in the network functioning, looking forward to a new network solution, a heterogenous network (HetNet) architecture has been proposed in this paper, for an efficient broadcast network set up during emergency situations. This paper proposes a two-tier heterogenous network (HetNet) architecture, with the macro base station (MBS) tier being Tier 1 and the small cell tier (SCT) being Tier 2. Here the SCT is mostly involved in setting up of a public warning communication system. The HetNets also intend to promote device-to-device (D2D) communication links, in case of absence of connectivity to the user via the MBT or the SCT. Use of small cells and D2D links shall improve the overall system performance. Certain research challenges however persist, and are stated in the paper. Pimmy Gandotra, Vimal Bhatia, Brejesh Lall |
TENCON | 3 |
| 2020 | Conditional Random Field based salient proposal set generation and its application in content aware seam carving
Prerana Mukherjee, Brejesh Lall |
Signal Process. Image Commun. | 2 |
| 2020 | Riemannian Curvature of Deep Neural NetworksabstractWe analyze deep neural networks using the theory of Riemannian geometry and curvature. The objective is to gain insight into how Riemannian geometry can characterize and predict the trained behavior of neural networks. We define a method for calculating Riemann and Ricci curvature tensors, and Ricci scalar curvature values for a trained neural net, in such a way that the output classifier softmax values are related to the input transformations, through the curvature equations. We also measure these curvature tensors experimentally for different networks which are pretrained with stochastic gradient descent and offer a way of visualizing and understanding the measurements to gain insight into the effect curvature has on behavior the neural networks locally, and possibly predict their behavior for different transformations of the test data. We also analyze the effect of variation in depth of the neural networks as well as how it behaves for different choices of data set. Piyush Kaul, Brejesh Lall |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2019 | UnDispNet: Unsupervised Learning for Multi-Stage Monocular Depth PredictionabstractDespite advances in single view depth estimation, most existing techniques treat the task in a supervised manner. Recent approaches utilize the possibility of learning without ground truth depth, by minimizing the photometric error. In this paper, we propose a deep framework that refines predicted depth from a single image using a two-stage process, exploiting sub-pixel convolutions for depth super resolution. The first stage uses a pyramidal input and learns depth at 4 scales, utilizing depth super resolution. The second stage uses warp errors, reconstructed images, predicted depth along with original left input to refine the depth predicted by the first stage. We use data augmentation by varying color, scale and incorporating left-right flipped images in our data. We train our network in a completely unsupervised way on photoconsistency imposing occlusion, left right consistency and disparity smoothness constraints. We transform the learning process into optimally distributed steps, varying the combination of scales and losses to minimize over-regularization of depth maps. We evaluate our model for monocular inputs on KITTI driving benchmark. Our depth predictions surpass state-of-the art self-supervised approaches for monocular depth prediction. Vinay Kaushik, Brejesh Lall |
3DV | 2 |
| 2019 | Time-varying Spectral Entropy Based Analysis of Impulse NoisesabstractImpulse noises encountered in indoor wireline communications systems consist of a series of wideband sub-impulses embedded in strongly colored background noise. The temporal attributes and spectral contents of these sub-impulses vary with time exhibiting interesting patterns both in time and frequency domain. Conventional methods of spectral analysis such as those based on power spectral density estimation give information only about the average spectral contents of observed impulse noise and do not convey any information about their nature of distribution. Time-frequency analysis methods are only able to quantify the magnitude of changing spectral content w.r.t time. Thus, in this paper we propose time-varying spectral entropy based analysis to capture the changing distribution of spectral contents of impulse noise recordings. This approach complements the existing time-frequency domain based characterization techniques and thus provides deeper insight into the structure of the impulse noises by tracking the changing distribution of time-varying spectrum. Neelima Singh, Brejesh Lall |
PIMRC | 2 |
| 2019 | Guided Compositional Generative Adversarial NetworksabstractIn this paper, we propose to synthesize natural images from a set of input objects. The proposed technique generates a scene which has high correlation with the provided set of input objects while also maintaining the natural placement of objects within the scene. The technique constitutes of a generative adversarial network trained on a large corpus of objects and natural scenes. This is in contrast with earlier works where the objective was to generate a natural scene from a noise vector or conditioning the network over a variable. However, such methods have limitations in their ability to control the objects within the generated images. On the contrary, we show that by training a Generative Adversarial Network with raw image pixels as input, we can generate scenes which constitute the objects as well as generate the surrounding environment suitable for the combination of the input objects. We provide qualitative and quantitative results on challenging MS-COCO dataset to show the effectiveness of the proposed technique. Anurag Tripathi, Siddharth Srivastava 0004, Brejesh Lall, Santanu Chaudhury |
SMC | 3 |
| 2019 | Performance analysis of diffusive molecular timing channelsabstractIn this study, the authors consider a diffusive molecular communication channel where the information to be transmitted is the time of release of the information molecules. Using the truncated Lévy distribution to model the first passage time of a molecule, the authors develop some general statistics for a random variable described by a truncated Lévy distribution. The authors then consider a multi‐particle molecular communication system, where the average arrival time of the information particles can be represented by a truncated Lévy flight, and the number of molecules required for the cross‐over from the Lévy regime to the Gaussian regime to occur is calculated. Furthermore, they use these results to analyse single as well as multi‐particle molecular communication channels in terms of capacity bounds and symbol error probability. Nilay Pandey, Ranjan K. Mallik, Brejesh Lall |
IET Commun. | 3 |
| 2019 | Performance analysis of multi-hop multi-branch frequency-selective channel for DF relayingabstractIn this work, the authors have explored the average bit error rate (BER) and ergodic capacity (EC) of a frequency‐selective channel consisting of multi‐hop and multi‐branch (MHMB). Each hop consists of multiple paths and signal travels in successive hops using decode and forward (DF) relaying. The channel is assumed to be log‐normal distributed. Numerical integration (NI) formulations based on the Gauss‐quadrature rule representation of the moment generating function of log‐normal distribution have been derived for both average BER and EC for multi‐hop and MHMB cooperative communication framework using maximal ratio combining and selection combining (SC) techniques at receiver. Results are obtained for a variable number of hops and branches. Excellent matching between Monte Carlo simulation and NI plots validate the accuracy of the proposed formulations. The contribution of this work to the existing research scenario is two‐fold: (i) to the best of the authors' knowledge, this is a first attempt to study a generic MHMB framework of a log‐normal distributed frequency‐selective channel. (ii) DF which is a more practical relaying protocol has not been explored till date for analysis of a log‐normal distributed MHMB network. Both of these aspects make this work pertinent. Manishika Rawat, Brejesh Lall |
IET Commun. | 2 |
| 2019 | Memorability-based image compressionabstractThis study is concerned with achieving the image compression using the concept of memorability. The authors have used memorability of an image, as a perceptual measure while image coding. In the proposed approach, a region‐of‐interest‐based memorability preserving image compression algorithm which is accomplished via two sub‐processes namely, memorability prediction and image compression is introduced. The memorability of images is predicted using convolutional neural network and restricted Boltzmann machine features. Based on these features, the memorability score of individual patches in an image is calculated and these scores are used to generate the memorability map. These memorability map values are used for optimised image compression. In order to validate the results, an eye tracking experiment with human participants is performed. The comparative analysis shows that the memorability‐based compression outperforms the state‐of‐the‐art compression techniques. Meera Thapar Khanna, Chetan Ralekar, Anurika Goel, Santanu Chaudhury, Brejesh Lall |
IET Image Process. | 5 |
| 2019 | Multi-channel local ternary pattern for content-based image retrieval
Megha Agarwal, Amit Singhal 0002, Brejesh Lall |
Pattern Anal. Appl. | 3 |
| 2019 | DeepPoint3D: Learning discriminative local descriptors using deep metric learning on 3D point clouds
Siddharth Srivastava 0004, Brejesh Lall |
Pattern Recognit. Lett. | 2 |
| 2019 | Estimation of Source Wavelet From Seismic Traces Using Groebner BasesabstractAn accurate and effective seismic wavelet estimation technique has extreme significance in the seismic data processing for analyzing the earth's subsurface layer information. The seismic wavelet to be determined is modeled as a moving average (MA) process and assumed to be driven by a zero mean, non-Gaussian, statistically independent, and identically distributed (IID) process. In order to estimate the MA model parameter from the observed noisy seismic signal, we pose this as a blind system identification (BSI) problem. In the BSI, a set of multivariate polynomial equations is obtained by matching higher order cumulant of observed noisy data with a higher order moment of blind system's impulse response. The Groebner bases that form the solution to this set of equations are obtained using the proposed algorithm. Numerical results demonstrate that the proposed method has a lower estimation error as compared to the previously reported methods. Karthikeyan Elumalai, Brejesh Lall, Rakesh Kumar Patney |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2018 | Spectral Entropy Features Based Analysis of Impulse Noise Sources for PLC SystemsabstractImpulse noises generated by domestic electrical and electronic appliances impact the performance of indoor Power Line Communication (PLC) systems severely. To understand the nature of complexity of these noises, especially in frequency domain, we need to look beyond average Power Spectral Density (PSD) based attributes. Hence, in this paper we have proposed spectral entropy based analysis for quantifying complexity of impulse noises in frequency domain. It has been demonstrated in this paper that Shannon Spectral Entropy is capable of quantifying the degree of regularity present in the distribution of spectral components of the impulse noises under observation. Thus, spectral entropy based features may contribute significantly in reliable diagnostics or identification of impulse noise sources which are primary disturbers in PLC systems. Neelima Singh, Brejesh Lall |
GLOBECOM | 2 |
| 2018 | Automatic Quantification of Stomata for High-Throughput Plant PhenotypingabstractStomatal morphology is a key phenotypic trait for plants' response analysis under various environmental stresses (e.g. drought, salinity etc.). Stomata exhibit diverse characteristics with respect to orientation, size, shape and varying degree of papillae occlusion. Thus, the biologists currently rely on manual or semi-automatic approaches to accurately compute its morphological traits based on scanning electron microscopic (SEM) images of leaf surface. In contrast to these subjective and low-throughput methods, we propose a novel automated framework for stomata quantification. It is realized based on a hybrid approach where the candidate stomata region is first detected by a convolutional neural network (CNN) and the occlusion is dealt with an inpainting algorithm. In addition, we propose stomata segmentation based quantification framework to solve the problem of shape, scale and occlusion in an end-to-end manner. The performance of the proposed automated frameworks is evaluated by comparing the derived traits with manually computed morphological traits of stomata. With no prior information about its size and location, the hybrid and end-to-end machine learning frameworks shows a correlation of 0.94 and 0.93, respectively on rice stomata images. Furthermore, they successfully enable wheat stomata quantification showing generalizability in terms of cultivars. Swati Bhugra, Deepak Mishra 0003, Anupama Anupama, Santanu Chaudhury, Brejesh Lall, Archana Chugh |
ICPR | 5 |
| 2018 | Need of Time-domain Pre-processing of Impulse Noises for PLC/xDSL SystemsabstractAppropriate characterization of impulse noise sources is necessary for reliable diagnostics of disturbers to handle performance issues in broadband communication systems. The impulse noises originating from primary disturbers, i.e., electrical and electronic appliances, reach the modem through different paths (channels) and hence pluralities of such noises increase. Diverse topologies of household electrical wiring introduce different types of time and frequency domain ambiguities in impulse noise characteristics before the noises reach the receiver. Hence, we propose that, if impulse noises are being observed at receiver and not directly at source, then time domain characterization should be done only after minimizing the effects of the channels from the observed noise recordings. In this paper we have demonstrated that time-domain pre-processing such as deconvolution is necessary to achieve reliable time-domain characterization if impulse noises are being observed at receiver. Neelima Singh, Brejesh Lall |
ISNCC | 2 |
| 2018 | Joint Image Classification and Annotation Prediction Using Iterative Learning on Local NeighbourhoodabstractImage annotation (tag) and classification play a critical role in many computer vision applications, such as image retrieval, scene understanding, scene description etc. While, databases such as ImageNet have high quality labels for images, in real world, a large number of images have missing labels or tags that completely describe the contents of an image. To solve this problem, in this paper, we work on the hypothesis that class and tag information are correlated and propose a joint optimization for image classification and annotation. We construct a unified cost function to learn the class scoring vectors as well as tag scoring vectors. The proposed approach achieves state-of-the-art results on benchmark datasets for joint tag prediction and classification. Anurag Tripathi, Siddharth Srivastava 0004, Santanu Chaudhury, Brejesh Lall |
SMC | 4 |
| 2018 | Large Scale Novel Object Discovery in 3DabstractWe present a method for discovering never-seen-before objects in 3D point clouds obtained from sensors like Microsoft Kinect. We generate supervoxels directly from the point cloud data and use them with a Siamese network, built on a recently proposed 3D convolutional neural network architecture. We use known objects to train a non-linear embedding of supervoxels, by optimizing the criteria that supervoxels which fall on the same object should be closer than those which fall on different objects, in the embedding space. We test on unknown objects, which were not seen during training, and perform clustering in the learned embedding space of supervoxels to effectively perform novel object discovery. We validate the method with extensive experiments, quantitatively showing that it can discover numerous unseen objects while being trained on only a few dense 3D models. We also show very good qualitative results of object discovery in point cloud data when the test objects, either specific instances or even categories, were never seen during training. Siddharth Srivastava 0004, Gaurav Sharma 0004, Brejesh Lall |
WACV | 3 |
| 2018 | Clustering short temporal behaviour sequences for customer segmentation using LDAabstractAbstract Customer segmentation based on temporal variation of subscriber preferences is useful for communication service providers (CSPs) in applications such as targeted campaign design, churn prediction, and fraud detection. Traditional clustering algorithms are inadequate in this context, as a multidimensional feature vector represents a subscriber profile at an instant of time, and grouping of subscribers needs to consider variation of subscriber preferences across time. Clustering in this case usually requires complex multivariate time series analysis‐based models. Because conventional time series clustering models have limitations around scalability and ability to accurately represent temporal behaviour sequences (TBS) of users, that may be short, noisy, and non‐stationary, we propose a latent Dirichlet allocation (LDA) based model to represent temporal behaviour of mobile subscribers as compact and interpretable profiles. Our model makes use of the structural regularity within the observable data corresponding to a large number of user profiles and relaxes the strict temporal ordering of user preferences in TBS clustering. We use mean‐shift clustering to segment subscribers based on their discovered profiles. Further, we mine segment‐specific association rules from the discovered TBS clusters, to aid marketers in designing intelligent campaigns that match segment preferences. Our experiments on real world data collected from a popular Asian communication service provider gave encouraging results. Jobin Wilson, Santanu Chaudhury, Brejesh Lall |
Expert Syst. J. Knowl. Eng. | 3 |
| 2018 | Denoising of pre-stack seismic data using subspace estimation methodsabstractDenoising is one of the core steps in seismic data processing flow. The seismic gather consists of multiple traces captured at different receivers. A set of receivers observe waves which are reflected from the same reflection point. Those traces need to be grouped together as they contain the same information about the earth subsurface layers. This is done by finding a common mid‐point (CMP) between the source and geophones. The time delay between CMP gathered traces are corrected by the normal move out correction method but the individual traces are corrupted by noise. In this paper we, propose a method for denoising individual traces. The set of traces can be modelled as belonging to a low‐dimensional subspace of an ambient signal space. This allows for construction of sparse representations of each trace in terms of other traces in the CMP gather. The resulting sparse representations are subsequently utilised to construct approximations of individual traces and thus, noise is suppressed. We constructed, the approximations using orthogonal matching pursuit. We applied proposed method to synthetic and field seismic data, the proposed technique performs better on established benchmarks while capturing the true locations of weak reflections and effectively attenuating the random noise. Karthikeyan Elumalai, Brejesh Lall, Rakesh Kumar Patney |
IET Signal Process. | 3 |
| 2018 | 3D local ternary co-occurrence patterns for natural, texture, face and bio medical image retrieval
Megha Agarwal, Amit Singhal 0002, Brejesh Lall |
Neurocomputing | 3 |
| 2018 | Restricted affine motion compensation and estimation in video coding with particle filtering and importance sampling: a multi-resolution approach
Mithilesh Kumar Jha, Ravi Chaudhary, Sumantra Dutta Roy, Mona Mathur, Brejesh Lall |
Multim. Syst. | 5 |
| 2017 | Truncated Lévy Statistics for Diffusion Based Molecular CommunicationabstractIn this paper, we use the truncated Lévy distribution to model the first arrival time in a molecular communication channel where information on the release time of the molecules is modulated and the molecules have an exponentially distributed lifetime. The general statistics for a random variable described by a truncated Lévy distribution are developed. Considering systems where a large number of molecules is used so that the average arrival time can be represented by a truncated Lévy flight, we derive an expression for the number of molecules required for the cross-over from the Lévy regime to the Gaussian regime to occur. We also use these statistics to obtain bounds on the capacity of these channels in an information theoretic sense. Nilay Pandey, Ranjan K. Mallik, Brejesh Lall |
GLOBECOM | 3 |
| 2017 | A Noise-Resilient Super-Resolution Framework to Boost OCR PerformanceabstractRecognizing text from noisy low-resolution (LR) images is extremely challenging and is an open problem for the computer vision community. Super-resolving a noisy LR text image results in noisy High Resolution (HR) text image, as super-resolution (SR) leads to spatial correlation in the noise, and further cannot be de-noised successfully. Traditional noise-resilient text image super-resolution methods utilize a denoising algorithm prior to text SR but denoising process leads to loss of some high frequency details, and the output HR image has missing information (texture details and edges). This paper proposes a noise-resilient SR framework for text images and recognizes the text using a deep BLSTM network trained on high resolution images. The proposed end-to-end deep learning based framework for noise-resilient text image SR simultaneously perform image denoising and super-resolution as well as preserves missing details. Stacked sparse denoising auto-encoder (SSDA) is learned for LR text image denoising, and our proposed coupled deep convolutional auto-encoder (CDCA) is learned for text image super-resolution. The pretrained weights for both these networks serve as initial weights to the end-to-end framework during finetuning, and the network is jointly optimized for both the tasks. We tested on several Indian Language datasets and the OCR performance of the noise-resilient super-resolved images is at par with the original HR images. Anupama Ray, Santanu Chaudhury, Brejesh Lall |
ICDAR | 4 |
| 2017 | Salprop: Salient object proposals via aggregated edge cuesabstractIn this paper, we propose a novel object proposal generation scheme by formulating a graph-based salient edge classification framework that utilizes the edge context. In the proposed method, we construct a Bayesian probabilistic edge map to assign a saliency value to the edgelets by exploiting low level edge features. A Conditional Random Field is then learned to effectively combine these features for edge classification with object/non-object label. We propose an objectness score for the generated windows by analyzing the salient edge density inside the bounding box. Extensive experiments on PASCAL VOC 2007 dataset demonstrate that the proposed method gives competitive performance against 10 popular generic object detection techniques while using fewer number of proposals. Prerana Mukherjee, Brejesh Lall, Sarvaswa Tandon |
ICIP | 2 |
| 2017 | Deep learning based frameworks for image super-resolution and noise-resilient super-resolutionabstractOur paper is motivated from the advancement in deep learning algorithms for various computer vision problems. We are proposing a novel end-to-end deep learning based framework for image super-resolution. This framework simultaneously calculates the convolutional features of low-resolution (LR) and high-resolution (HR) image patches and learns the non-linear function that maps these convolutional features of LR image patches to their corresponding HR image patches convolutional features. Here, proposed deep learning based image super-resolution architecture is termed as coupled deep convolutional auto-encoder (CDCA) which provides state-of-the-art results. Super-resolution of a noisy/distorted LR images results in noisy/distorted HR images, as super-resolution process gives rise to spatial correlation in the noise, and further, it cannot be de-noised successfully. Traditional noise resilient image super-resolution methods utilize a de-noising algorithm prior to super-resolution but de-noising process gives rise to loss of some high-frequency information (edges and texture details) and super-resolution of the resultant image provides HR image with missing edges and texture information. We are also proposing a novel end-to-end deep learning based framework to obtain noise resilient image super-resolution. Proposed end-to-end deep learning based framework for noise resilient super-resolution simultaneously perform image de-noising and super-resolution as well as preserves textural details. First, stacked sparse de-noising auto-encoder (SSDA) was learned for LR image de-noising and proposed CDCA was learned for image superresolution. Then, both image de-noising and super-resolution networks were cascaded. This cascaded deep learning network was employed as one integral network where pre-trained weights were serving as initial weights. The integral network was end-to-end trained or fine-tuned on a database having noisy, LR image as an input and target as an HR image. In fine-tuning, all layers of the combined end-to-end network was jointly optimized to perform image de-noising and super-resolution simultaneously. Experimental results show that proposed noise resilient super-resolution framework outperforms the conventional and state-of-the-art approaches in terms of PSNR and SSIM metrics. Santanu Chaudhury, Brejesh Lall |
IJCNN | 3 |
| 2017 | Poster: DRIZY: Collaborative Driver Assistance Over Wireless NetworksabstractDriver assistance systems, that rely on vehicular sensors such as cameras, LIDAR and other on-board diagnostic sensors, have progressed rapidly in recent years to increase road safety. Road conditions in developing countries like India are chaotic where roads are not well maintained and thus vehicular sensors alone do not suffice in detecting impending collisions. In this paper, we investigate a collaborative driver assistance system "DRIZY: DRIve eaSY" for such scenarios where inference is drawn from on-board camera feed to alert drivers of obstacles ahead and the cloud uses GPS sensor data uploaded by all vehicles to alert drivers of vehicles in potential collision trajectory. Thus, we combine computer vision and vehicle-to-cloud communication to create comprehensive situational awareness. We prototype our system to consider two types of collisions: vehicle-to-vehicle collisions based on uploading GPS sensor data of vehicles to cloud and vehicle-to-pedestrian collisions based on detecting pedestrians from vehicle's dashboard camera feed. Sensor data processing in each vehicle occurs on smartphone for GPS values which are then uploaded to cloud and on raspberry pi3 for video feeds to make a cost-effective solution. Experiments over both 4G and wireless networks in India show that collaborative driver assistance is feasible in low traffic density within acceptable driver reaction time of <5 sec, but can be limited by the time to process compute-intensive video feeds in real-time. We investigate novel ways to optimize the processing to find an acceptable trade-off. Nakul Garg, Ishani Janveja, Divyansh Malhotra, Chetan Chawla, Pulkit Gupta, Harshil Bansal, Aakanksha Chowdhery, Prerana Mukherjee, Brejesh Lall |
MobiCom | 9 |
| 2017 | Saliency and KAZE features assisted object segmentation
Prerana Mukherjee, Brejesh Lall |
Image Vis. Comput. | 2 |
| 2017 | A Novel Hybrid Kinect-Variety-Based High-Quality Multiview Rendering Scheme for Glass-Free 3D DisplaysabstractThis paper presents a new hybrid Kinect-variety-based synthesis scheme that renders artifact-free multiple views for autostereoscopic/automultiscopic displays. The proposed approach does not explicitly require dense scene depth information for synthesizing novel views from arbitrary viewpoints. Instead, the integrated framework first constructs a consistent minimal image–space parameterization of the underlying 3D scene. The compact representation of scene structure is formed using only implicit sparse depth information of a few reference scene points extracted from raw RGB depth data. The views from arbitrary positions can be inferred by moving the novel camera in parameterized space by enforcing Euclidean constraints on reference scene images under a full-perspective projection model. Unlike the state-of-the-art depth image-based rendering (DIBR) methods, in which input depth map accuracy is crucial for high-quality output, our proposed algorithm does not depend on precise per-pixel geometry information. Therefore, it simply sidesteps to recover and refine the incomplete or noisy depth estimates with advanced filling or upscaling techniques. Our approach performs fairly well in unconstrained indoor/outdoor environments, where the performance of range sensors or dense depth-based algorithms could be seriously affected due to scene complex geometric conditions. We demonstrate that the proposed hybrid scheme provides guarantees on the completeness, optimality with respect to the inter-view consistency of the algorithm. In the experimental validation, we performed a quantitative evaluation as well as subjective assessment of the scene with complex geometric or surface properties. A comparison with the latest representative DIBR methods is additionally performed to demonstrate the superior performance of the proposed scheme. Santanu Chaudhury, Brejesh Lall |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2015 | Saliency map based improved segmentationabstractIn this paper we present a novel approach for refining segmentation using saliency map. To achieve this, we first develop a new saliency detection method based on cues at various levels. Initially preprocessing step is done using non-linear anisotropic diffusion filtering in order to preserve the edge information in the foreground salient objects and smoothen the background. Then we apply grab cut segmentation using saliency map as the input to get improved segmentation. Repeated application of the scheme is used for multi-object segmentation. The experimental results for the saliency technique show high precision and recall rates against the state-of-the-art methods. Prerana Mukherjee, Brejesh Lall, Archit Shah |
ICIP | 2 |
| 2015 | DEMD-based video coding for textured videos in an H.264/MPEG framework
Mithilesh Kumar Jha, Sumantra Dutta Roy, Brejesh Lall |
Pattern Recognit. Lett. | 3 |
| 2015 | Performance Analysis of Amplitude Modulation Schemes for Diffusion-Based Molecular CommunicationabstractIn this paper, we investigate modulation techniques for end-to-end communication between two nanomachines placed in a fluid medium. The information is encoded as the number of molecules transmitted leading to such schemes being aptly named as amplitude modulation schemes. The propagation of molecules obeys the laws of Brownian motion with a positive drift from the transmitter to the receiver nanomachine. The channel is characterized by two parameters of the fluid medium: the drift velocity and the diffusion coefficient. Assuming the molecules degrade over time, the life expectancy of the molecules also plays a significant role in such communication scenarios. We consider anM-ary modulation scheme and also propose an extended scheme, which is a slight variation of a binary modulation scheme. The received symbol is corrupted by interference from the previous symbols as well as other noise sources present in the medium. Considering maximum likelihood detection at the receiver, we derive analytical expressions for the end-to-end symbol error probability and the capacity for these modulation schemes. Numerical results bring out the impact of various parameters on the performance of the system. Our results show that these schemes offer a promising approach to set up molecular communication over diffusion-based channels. Amit Singhal 0002, Ranjan K. Mallik, Brejesh Lall |
IEEE Trans. Wirel. Commun. | 3 |
| 2014 | Kinect-Variety Fusion: A Novel Hybrid Approach for Artifacts-Free 3DTV Content GenerationabstractThis paper presents a novel low-cost hybrid Kinect-variety based content generation scheme for 3DTV displays. The integrated framework constructs an efficient consistent image-space parameterization of 3D scene structure using only sparse depth information of few reference scene points. Under full-perspective camera model, the enforced Euclidean constraints simplify the synthesis of high quality novel multiview content for distinct camera motions. The algorithm does not rely on complete precise scene geometry information, and are unaffected by scene complex geometric properties, unconstrained environmental variations and illumination conditions. It, therefore, performs fairly well under a wider set of operation condition where the 3D range sensors fail or reliability of depth-based algorithms are suspect. The robust integration of vision algorithm and visual sensing scheme complement each other's shortcomings. It opens new opportunities for envisioning vision-sensing applications in uncontrolled environments. We demonstrate that proposed robust integration provides guarantees on the completeness and consistency of the algorithm. This leads to improved reliability on an extensive set of experimental results. Santanu Chaudhury, Brejesh Lall |
ICPR | 3 |
| 2014 | Performance of amplitude modulation schemes for molecular communication over a fluid mediumabstractIn this paper, we consider three different amplitude modulation schemes for molecular communication through a fluid medium with a positive drift velocity. The channel is divided into time slots of equal duration and transmission is done at the beginning of every slot. The molecules, transmitted by a nanomachine, propagate via Brownian motion and reach the receiver nanomachine. The amplitude levels are represented by the number of molecules transmitted. The communication is corrupted by randomness due to diffusion as well as interference from the previous slots. Considering maximum likelihood detection at the receiver, the probability of error is analyzed for all the three schemes, and its variation with the parameters involved is studied. The results are also compared with previously proposed time modulation schemes, where the information is encoded as the transmit time of molecules. Amit Singhal 0002, Ranjan K. Mallik, Brejesh Lall |
PIMRC | 3 |
| 2014 | Molecular communication with Brownian motion and a positive drift: performance analysis of amplitude modulation schemesabstractIn this study, the authors consider molecular communication between two nanomachines placed in a fluid medium for three different amplitude modulation schemes. The number of molecules transmitted represents the amplitude levels for these schemes. Each molecule released by the transmitter travels with Brownian motion and a positive drift to reach the receiver nanomachine. They consider a time slotted channel, where the information in every slot is corrupted by stray molecules from the previous slots. The capacity of such a molecular communication channel is investigated for all the three modulation schemes. Analytical expressions for the end‐to‐end symbol error probability are derived, considering maximum likelihood detection at the receiver. Numerical results indicate that arbitrarily low probabilities of error can be achieved for high drift velocities. An increase in the slot length further improves the performance, albeit at the cost of data rate. The results also demonstrate the improvements offered by the amplitude modulation schemes over the previously proposed time modulation schemes. Amit Singhal 0002, Ranjan K. Mallik, Brejesh Lall |
IET Commun. | 3 |
| 2014 | A flexible architecture for multi-view 3DTV based on uncalibrated cameras
Santanu Chaudhury, Brejesh Lall, M. S. Venkatesh |
J. Vis. Commun. Image Represent. | 3 |
| 2014 | Design and analysis of matrix Wiener synthesis filter for multirate filter bank
Sandeep Patel, Ravindra Dhuli, Brejesh Lall |
Signal Process. | 3 |
| 2013 | Space-Time Parameterized Variety Manifolds: A Novel Approach for Arbitrary Multi-perspective 3D View GenerationabstractThis paper presents a novel image variety-Based approach that elegantly models the space of a broad class of perspective and non-perspective stereo varieties within a single, unified framework. The basic concept of parameterized variety presented earlier by Genc and Ponce [1] is extended to represent the nonlinear space of images. An efficient algebraic framework is constructed to parameterize the variety associated with full perspective cameras. The algorithm seeks the manifolds that constrain this space of six-dimensional variety to generate compelling multi-perspective 3D effects from arbitrary virtual viewpoints. Combining geometric space of multiple uncalibrated perspective views with appearance space in a globally optimized way leads to numerous potential applications, especially in content creation for multi-perspective 3DTV. The proposed approach works for uncalibrated static/dynamic scenes, containing parallax and unstructured object motion. It even seamlessly deals with images or video sequences that do not share a common origin, thus provides an effective tool for montaging, indexing and virtual navigation. Santanu Chaudhury, Brejesh Lall |
3DV | 3 |
| 2012 | Parameterized Variety Based View Synthesis Scheme for Multi-view 3DTV
Santanu Chaudhury, Brejesh Lall |
ACCV (4) | 3 |
| 2012 | A free viewpoint 3DTV system based on parameterized variety model
Santanu Chaudhury, Brejesh Lall |
ICPR | 3 |
| 2010 | A Liberal Carrier Sensing for Increased Spatial Reuse in Multi-Hop Wireless Ad Hoc NetworksabstractRecent experimental results have shown that the minimum signal-to-interference ratio required at a receiver (CPth) depends on the order of arrival of the overlapping frames. For a given sender-receiver distance, this differential capture capability of a receiver leads to two distinct interference ranges (ri) around the receiver, and its value is much smaller when the sender's frame arrives earlier. This feature also suggests a possibility of increased spatial reuse by allowing the (secondary) nodes outside the primary receiver's rito communicate concurrently once the primary receiver starts its DATA reception. In this paper, we propose a liberal carrier sensing (LCS) scheme wherein some already available information at an otherwise 'exposed' receiver are exploited to help decide when it is safe to respond to a secondary transmission request. The proposed modification in the carrier sensing approach results in a significantly improved spatial reuse, thereby increasing overall system throughput. Our simulation studies show that, compared to the conventional carrier sensing scheme with differential capture capable receivers, the end-to-end TCP throughput with LCS can be improved by more than 20% in regular topologies and up to about 9% in random topologies. Mayur M. Vegad, Swades De, Brejesh Lall |
ICC | 3 |
| 2010 | Alias reduction: Generalised pseudocirculant conditions
Ravindra Dhuli, Brejesh Lall |
Signal Process. | 2 |