EDBT 2026 Demo / reviewers in the wild / expert
Subhankar Ghosh
dblp:53/10604
· DBLP profile ↗
34ranked-venue papers
7as first author
30since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 5 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 3 first-author · 19 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Open Full-duplex Voice Agent with Speech-to-Speech Language ModelabstractWe present the system demonstration and opensource code release of a novel, data-efficient framework that converts any standard text Large Language Model (LLM) into a full-duplex end-to-end (E2E) speech-to-speech (S2S) model, for building conversational voice agents. Our new modeling method enables any LLMs to simultaneously listen and speak without requiring extensive speech-text pretraining. Moreover, we demonstrate how to put together a low-latency and full-duplex voice agent with open-source modeling, inference optimization, and serving solutions. This work significantly lowers the barrier to entry for developing low-latency, human-like voice agents by providing a generalizable, end-to-end solution built on open-source technologies. Edresson Casanova, Chen Chen 0075, Kevin Hu, Ankita Pasad, Elena Rastorgueva, Seelan Lakshmi Narasimhan, Slyne Deng, Ehsan Hosseini-Asl, Piotr Zelasko, Valentin Mendelev, Subhankar Ghosh, Yifan Peng 0003, Zhehuai Chen, Jason Li 0007, Jagadeesh Balam, Vitaly Lavrukhin, Boris Ginsburg |
ASRU | 11 |
| 2025 | Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free GuidanceabstractShehzeen Samarah Hussain, Paarth Neekhara, Xuesong Yang, Edresson Casanova, Subhankar Ghosh, Roy Fejgin, Mikyas T. Desta, Rafael Valle, Jason Li. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Shehzeen Hussain, Paarth Neekhara, Xuesong Yang, Edresson Casanova, Subhankar Ghosh, Roy Fejgin, Mikyas T. Desta, Rafael Valle, Jason Li 0007 |
EMNLP | 5 |
| 2025 | TTS-Transducer: End-to-End Speech Synthesis with Neural TransducerabstractThis work introduces TTS-Transducer – a novel architecture for text-to-speech, leveraging the strengths of audio codec models and neural transducers. Transducers, renowned for their superior quality and robustness in speech recognition, are employed to learn monotonic alignments and allow for avoiding using explicit duration predictors. Neural audio codecs efficiently compress audio into discrete codes, revealing the possibility of applying text modeling approaches to speech generation. However, the complexity of predicting multiple tokens per frame from several codebooks, as necessitated by audio codec models with residual quantizers, poses a significant challenge. The proposed system first uses a transducer architecture to learn monotonic alignments between tokenized text and speech codec tokens for the first codebook. Next, a non-autoregressive Transformer predicts the remaining codes using the alignment extracted from transducer loss. The proposed system is trained end-to-end. We show that TTS-Transducer is a competitive and robust alternative to contemporary TTS systems1. Vladimir Bataev, Subhankar Ghosh, Vitaly Lavrukhin, Jason Li 0007 |
ICASSP | 2 |
| 2025 | Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and InferenceabstractLarge language models (LLMs) have significantly advanced audio processing through audio codecs that convert audio into discrete tokens, enabling the application of language modeling techniques to audio data. However, audio codecs often operate at high frame rates, resulting in slow training and inference, especially for autoregressive models. To address this challenge, we present the Low Frame-rate Speech Codec (LFSC): a neural audio codec that leverages finite scalar quantization and adversarial training with large speech language models to achieve high-quality audio compression with a 1.89 kbps bitrate and 21.5 frames per second. We demonstrate that our novel codec can make the inference of LLM-based text-to-speech models around three times faster while improving intelligibility and producing quality comparable to previous models. Edresson Casanova, Ryan Langman, Paarth Neekhara, Shehzeen Hussain, Jason Li 0007, Subhankar Ghosh, Ante Jukic, Sang-gil Lee |
ICASSP | 6 |
| 2025 | NanoCodec: Towards High-Quality Ultra Fast Speech LLM Inference
Edresson Casanova, Paarth Neekhara, Ryan Langman, Shehzeen Hussain, Subhankar Ghosh, Xuesong Yang, Ante Jukic, Jason Li 0007, Boris Ginsburg |
INTERSPEECH | 5 |
| 2025 | Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model
Ehsan Hosseini-Asl, Chen Chen 0075, Edresson Casanova, Subhankar Ghosh, Piotr Zelasko, Zhehuai Chen, Jason Li 0007, Jagadeesh Balam, Boris Ginsburg |
INTERSPEECH | 5 |
| 2025 | FASTER: A Font-Agnostic Scene Text Editing and Rendering FrameworkabstractScene Text Editing (STE) is a challenging research prob-lem, that primarily aims towards modifying existing texts in an image while preserving the background and the font style of the original text. Despite its utility in numerous real-world applications, existing style-transfer-based approaches have shown sub-par editing performance due to (1) complex image backgrounds, (2) diverse font attributes, and (3) varying word lengths within the text. To address such limitations, in this paper, we propose a novel font-agnostic scene text editing and rendering framework, named FASTER, for simultaneously generating text in arbitrary styles and locations while preserving a natural and realistic appearance and structure. A combined fusion of target mask generation and style transfer units, with a cascaded self-attention mech-anism has been proposed to focus on multi-level text region edits to handle varying word lengths. Extensive evaluation on a real-world database withfurther subjective human eval-uation study indicates the superiority of FASTER in both scene text editing and rendering tasks, in terms of model per-formance and efficiency. The code and pre-trained models have been released in our Gi thub repo. Alloy Das, Sanket Biswas, Prasun Roy, Subhankar Ghosh, Umapada Pal 0001, Michael Blumenstein, Josep Lladós 0001, Saumik Bhattacharya |
WACV | 4 |
| 2025 | Climate smart computing: A perspective
Mingzhou Yang 0001, Bharat Jayaprakash, Subhankar Ghosh, Hyeonjung Tari Jung, Matthew Eagon, William F. Northrop, Shashi Shekhar 0001 |
Pervasive Mob. Comput. | 3 |
| 2024 | Towards Statistically Significant Taxonomy Aware Co-Location Pattern Detection (Short Paper)
Subhankar Ghosh, Arun Sharma 0006, Jayant Gupta, Shashi Shekhar 0001 |
COSIT | 1 |
| 2024 | Towards Kriging-informed Conditional Diffusion for Regional Sea-Level Data Downscaling: A Summary of ResultsabstractGiven coarser-resolution projections from global climate models or satellite data, the downscaling problem aims to estimate finer-resolution regional climate data, capturing fine-scale spatial patterns and variability. Downscaling is any method to derive high-resolution data from low-resolution variables, often to provide more detailed and local predictions and analyses. This problem is societally crucial for effective adaptation, mitigation, and resilience against significant risks from climate change. The challenge arises from spatial heterogeneity and the need to recover finer-scale features while ensuring model generalization. Most downscaling methods [21] fail to capture the spatial dependencies at finer scales and underperform on real-world climate datasets, such as sea-level rise. We propose a novel Kriging-informed Conditional Diffusion Probabilistic Model (Ki-CDPM) to capture spatial variability while preserving fine-scale features. Experimental results on climate data show that our proposed method is more accurate than state-of-the-art downscaling techniques. Subhankar Ghosh, Arun Sharma 0006, Jayant Gupta, Aneesh Subramanian, Shashi Shekhar 0001 |
SIGSPATIAL/GIS | 1 |
| 2024 | SALM: Speech-Augmented Language Model with in-Context Learning for Speech Recognition and TranslationabstractWe present a novel Speech Augmented Language Model (SALM) with multitask and in-context learning capabilities. SALM comprises a frozen text LLM, a audio encoder, a modality adapter module, and LoRA layers to accommodate speech input and associated task instructions. The unified SALM not only achieves performance on par with task-specific Conformer baselines for Automatic Speech Recognition (ASR) and Speech Translation (AST), but also exhibits zero-shot in-context learning capabilities, demonstrated through keyword-boosting task for ASR and AST. Moreover, speech supervised in-context training is proposed to bridge the gap between LLM training and downstream speech tasks, which further boosts the in-context learning ability of speech-to-text models. Proposed model is open-sourced via NeMo toolkit1. Zhehuai Chen, He Huang 0012, Andrei Andrusenko, Oleksii Hrinchuk, Krishna C. Puvvada, Jason Li 0007, Subhankar Ghosh, Jagadeesh Balam, Boris Ginsburg |
ICASSP | 7 |
| 2024 | λ-Color: Amplifying Long-Range Dependencies for Image Colorization
Subhankar Ghosh, Saumik Bhattacharya, Prasun Roy, Umapada Pal 0001, Michael Blumenstein |
ICPR (22) | 1 |
| 2024 | d-Sketch: Improving Visual Fidelity of Sketch-to-Image Translation with Pretrained Latent Diffusion Models without Retraining
Prasun Roy, Saumik Bhattacharya, Subhankar Ghosh, Umapada Pal 0001, Michael Blumenstein |
ICPR (25) | 3 |
| 2024 | Semantically Consistent Person Image Generation
Prasun Roy, Saumik Bhattacharya, Subhankar Ghosh, Umapada Pal 0001, Michael Blumenstein |
ICPR (25) | 3 |
| 2024 | Improving Robustness of LLM-based Speech Synthesis by Learning Monotonic AlignmentabstractLarge Language Model (LLM) based text-to-speech (TTS) systems have demonstrated remarkable capabilities in handling large speech datasets and generating natural speech for new speakers.However, LLM-based TTS models are not robust as the generated output can contain repeating words, missing words and mis-aligned speech (referred to as hallucinations or attention errors), especially when the text contains multiple occurrences of the same token.We examine these challenges in an encoder-decoder transformer model and find that certain cross-attention heads in such models implicitly learn the text and speech alignment when trained for predicting speech tokens for a given text.To make the alignment more robust, we propose techniques utilizing CTC loss and attention priors that encourage monotonic cross-attention over the text tokens.Our guided attention training technique does not introduce any new learnable parameters and significantly improves robustness of LLM-based TTS models. Paarth Neekhara, Shehzeen Hussain, Subhankar Ghosh, Jason Li 0007, Boris Ginsburg |
INTERSPEECH | 3 |
| 2024 | Conformal Prediction for Class-wise Coverage via Augmented Label Rank CalibrationabstractConformal prediction (CP) is an emerging uncertainty quantification framework that allows us to construct a prediction set to cover the true label with a pre-specified marginal or conditional probability.
Although the valid coverage guarantee has been extensively studied for classification problems, CP often produces large prediction sets which may not be practically useful.
This issue is exacerbated for the setting of class-conditional coverage on imbalanced classification tasks with many and/or imbalanced classes.
This paper proposes the Rank Calibrated Class-conditional CP (RC3P) algorithm to reduce the prediction set sizes to achieve class-conditional coverage, where the valid coverage holds for each class.
In contrast to the standard class-conditional CP (CCP) method that uniformly thresholds the class-wise conformity score for each class, the augmented label rank calibration step allows RC3P to selectively iterate this class-wise thresholding subroutine only for a subset of classes whose class-wise top-$k$ error is small.
We prove that agnostic to the classifier and data distribution, RC3P achieves class-wise coverage. We also show that RC3P reduces the size of prediction sets compared to the CCP method.
Comprehensive experiments on multiple real-world datasets demonstrate that RC3P achieves class-wise coverage and $26.25\\%$ $\downarrow$ reduction in prediction set sizes on average. Yuanjie Shi, Subhankar Ghosh, Taha Belkhouja, Janardhan Rao Doppa, Yan Yan 0006 |
NeurIPS | 2 |
| 2024 | TIC: text-guided image colorization using conditional generative modelabstractAbstract Image colorization is a well-known problem in computer vision. However, due to the ill-posed nature of the task, image colorization is inherently challenging. Though several attempts have been made by researchers to make the colorization pipeline automatic, these processes often produce unrealistic results due to a lack of conditioning. In this work, we attempt to integrate textual descriptions as an auxiliary condition, along with the grayscale image that is to be colorized, to improve the fidelity of the colorization process. To the best of our knowledge, this is one of the first attempts to incorporate textual conditioning in the colorization pipeline. To do so, a novel deep network has been proposed that takes two inputs (the grayscale image and the respective encoded text description) and tries to predict the relevant color gamut. As the respective textual descriptions contain color information of the objects present in the scene, the text encoding helps to improve the overall quality of the predicted colors. The proposed model has been evaluated using different metrics like SSIM, PSNR, LPISPS and achieved scores of 0.917, 23.27,0.223, respectively. These quantitative metrics have shown that the proposed method outperforms the SOTA techniques in most of the cases. Subhankar Ghosh, Prasun Roy, Saumik Bhattacharya, Umapada Pal 0001, Michael Blumenstein |
Multim. Tools Appl. | 1 |
| 2024 | Physics-Based Abnormal Trajectory Gap DetectionabstractGiven trajectories with gaps (i.e., missing data), we investigate algorithms to identify abnormal gaps in trajectories which occur when a given moving object did not report its location, but other moving objects in the same geographic region periodically did. The problem is important due to its societal applications, such as improving maritime safety and regulatory enforcement for global security concerns, such as illegal fishing, illegal oil transfers, and trans-shipments. The problem is challenging due to the difficulty of bounding the possible locations of the moving object during a trajectory gap, and the very high computational cost of detecting gaps in such a large volume of location data. The current literature on anomalous trajectory detection assumes linear interpolation within gaps, which may not be able to detect abnormal gaps since objects within a given region may have traveled away from their shortest path. In preliminary work, we introduced an abnormal gap measure that uses a classical space-time prism model to bound an object's possible movement during the trajectory gap and provided a scalable memoized gap detection algorithm (Memo-AGD). In this article, we propose a space time-aware gap detection (STAGD) approach to leverage space-time indexing and merging of trajectory gaps. We also incorporate a dynamic region merge-based (DRM) approach to efficiently compute gap abnormality scores. We provide theoretical proofs that both algorithms are correct and complete and also provide analysis of asymptotic time complexity. Experimental results on synthetic and real-world maritime trajectory data show that the proposed approach substantially improves computation time over the baseline technique. Arun Sharma 0006, Subhankar Ghosh, Shashi Shekhar 0001 |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2023 | Improving Uncertainty Quantification of Deep Classifiers via Neighborhood Conformal Prediction: Novel Algorithm and Theoretical AnalysisabstractSafe deployment of deep neural networks in high-stake real-world applications require theoretically sound uncertainty quantification. Conformal prediction (CP) is a principled framework for uncertainty quantification of deep models in the form of prediction set for classification tasks with a user-specified coverage (i.e., true class label is contained with high probability). This paper proposes a novel algorithm referred to as Neighborhood Conformal Prediction (NCP) to improve the efficiency of uncertainty quantification from CP for deep classifiers (i.e., reduce prediction set size). The key idea behind NCP is to use the learned representation of the neural network to identify k nearest-neighbor calibration examples for a given testing input and assign them importance weights proportional to their distance to create adaptive prediction sets. We theoretically show that if the learned data representation of the neural network satisfies some mild conditions, NCP will produce smaller prediction sets than traditional CP algorithms. Our comprehensive experiments on CIFAR-10, CIFAR-100, and ImageNet datasets using diverse deep neural networks strongly demonstrate that NCP leads to significant reduction in prediction set size over prior CP methods. Subhankar Ghosh, Taha Belkhouja, Yan Yan 0006, Janardhan Rao Doppa |
AAAI | 1 |
| 2023 | Vani: Very-Lightweight Accent-Controllable TTS for Native And Non-Native Speakers With Identity PreservationabstractWe introduce VANI, a very lightweight multi-lingual accent controllable speech synthesis system. Our model builds upon disentanglement strategies proposed in RADMMM[1] and supports explicit control of accent, language, speaker and fine-grained F0and energy features for speech synthesis. We utilize the Indic languages dataset, released for LIMMITS 2023 as part of ICASSP Signal Processing Grand Challenge, to synthesize speech in 3 different languages. Our model supports transferring the language of a speaker while retaining their voice and the native accent of the target language. We utilize the large-parameter RADMMM model for Track 1 and lightweight VANI model for Track 2 and 3 of the competition. Rohan Badlani, Akshit Arora, Subhankar Ghosh, Rafael Valle, Kevin J. Shih, João Felipe Santos, Boris Ginsburg, Bryan Catanzaro |
ICASSP | 3 |
| 2023 | Adapter-Based Extension of Multi-Speaker Text-To-Speech Model for New Speakers
Cheng-Ping Hsieh, Subhankar Ghosh, Boris Ginsburg |
INTERSPEECH | 2 |
| 2023 | Probabilistically robust conformal predictionabstractConformal prediction (CP) is a framework to quantify uncertainty of machine learning classifiers including deep neural networks. Given a testing example and a trained classifier, CP produces a prediction set of candidate labels with a user-specified coverage (i.e., true class label is contained with high probability). Almost all the existing work on CP assumes clean testing data and there is not much known about the robustness of CP algorithms w.r.t natural/adversarial perturbations to testing examples. This paper studies the problem of probabilistically robust conformal prediction (PRCP) which ensures robustness to most perturbations around clean input examples. PRCP generalizes the standard CP (cannot handle perturbations) and adversarially robust CP (ensures robustness w.r.t worst-case perturbations) to achieve better trade-offs between nominal performance and robustness. We propose a novel adaptive PRCP (aPRCP) algorithm to achieve probabilistically robust coverage. The key idea behind aPRCP is to determine two parallel thresholds, one for data samples and another one for the perturbations on data (aka "quantile-of-quantile” design). We provide theoretical analysis to show that aPRCP algorithm achieves robust coverage. Our experiments on CIFAR-10, CIFAR-100, and ImageNet datasets using deep neural networks demonstrate that aPRCP achieves better trade-offs than state-of-the-art CP and adversarially robust CP algorithms. Subhankar Ghosh, Yuanjie Shi, Taha Belkhouja, Yan Yan 0006, Janardhan Rao Doppa |
UAI | 1 |
| 2023 | Multi-scale attention guided pose transfer
Prasun Roy, Saumik Bhattacharya, Subhankar Ghosh, Umapada Pal 0001 |
Pattern Recognit. | 3 |
| 2022 | TIPS: Text-Induced Pose Synthesis
Prasun Roy, Subhankar Ghosh, Saumik Bhattacharya, Umapada Pal 0001, Michael Blumenstein |
ECCV (38) | 2 |
| 2022 | Towards a tighter bound on possible-rendezvous areas: preliminary resultsabstractGiven trajectories with gaps, we investigate methods to tighten spatial bounds on areas (e.g., nodes in a spatial network) where possible rendezvous activity could have occurred. The problem is important for reducing manual effort to post-process possible rendezvous areas using satellite imagery and has many societal applications to improve public safety, security, and health. The problem of rendezvous detection is challenging due to the difficulty of interpreting missing data within a trajectory gap and the very high cost of detecting gaps in such a large volume of location data. Most recent literature presents formal models, namely space-time prism, to track an object's rendezvous patterns within trajectory gaps on a spatial network. However, the bounds derived from the space-time prism are rather loose, resulting in unnecessarily extensive postprocessing manual effort. To address these limitations, we propose a Time Slicing-based Gap-Aware Rendezvous Detection (TGARD) algorithm to tighten the spatial bounds in spatial networks. We propose a Dual Convergence TGARD (DC-TGARD) algorithm to improve computational efficiency using a bi-directional pruning approach. Theoretical results show the proposed spatial bounds on the area of possible rendezvous are tighter than that from related work (space-time prism). Experimental results on synthetic and real-world spatial networks (e.g., road networks) show that the proposed DC-TGARD is more scalable than the TGARD algorithm. Arun Sharma 0006, Jayant Gupta, Subhankar Ghosh |
SIGSPATIAL/GIS | 3 |
| 2022 | Scene Aware Person Image Generation through Global Contextual ConditioningabstractPerson image generation is an intriguing yet challenging problem. However, this task becomes even more difficult under constrained situations. In this work, we propose a novel pipeline to generate and insert contextually relevant person images into an existing scene while preserving the global semantics. More specifically, we aim to insert a person such that the location, pose, and scale of the person being inserted blends in with the existing persons in the scene. Our method uses three individual networks in a sequential pipeline. At first, we predict the potential location and the skeletal structure of the new person by conditioning a Wasserstein Generative Adversarial Network (WGAN) on the existing human skeletons present in the scene. Next, the predicted skeleton is refined through a shallow linear network to achieve higher structural accuracy in the generated image. Finally, the target image is generated from the refined skeleton using another generative network conditioned on a given image of the target person. In our experiments, we achieve high-resolution photo-realistic generation results while preserving the general context of the scene. We conclude our paper with multiple qualitative and quantitative benchmarks on the results. Prasun Roy, Subhankar Ghosh, Saumik Bhattacharya, Umapada Pal 0001, Michael Blumenstein |
ICPR | 2 |
| 2022 | Virtual-reality-based digital twin of office spaces with social distance measurement featureabstractSocial distancing is an effective way to reduce the spread of the SARS-CoV-2 virus. Many students and researchers have already attempted to use computer vision technology to automatically detect human beings in the field of view of a camera and help enforce social distancing. However, because of the present lockdown measures in several countries, the validation of computer vision systems using large-scale datasets is a challenge. In this paper, a new method is proposed for generating customized datasets and validating deep-learning-based computer vision models using virtual reality (VR) technology. Using VR, we modeled a digital twin (DT) of an existing office space and used it to create a dataset of individuals in different postures, dresses, and locations. To test the proposed solution, we implemented a convolutional neural network (CNN) model for detecting people in a limited-sized dataset of real humans and a simulated dataset of humanoid figures. We detected the number of persons in both the real and synthetic datasets with more than 90% accuracy, and the actual and measured distances were significantly correlated (r=0.99). Finally, we used intermittent-layer- and heatmap-based data visualization techniques to explain the failure modes of a CNN. A new application of DTs is proposed to enhance workplace safety by measuring the social distance between individuals. The use of our proposed pipeline along with a DT of the shared space for visualizing both environmental and human behavior aspects preserves the privacy of individuals and improves the latency of such monitoring systems because only the extracted information is streamed. Abhishek Mukhopadhyay, G. S. Rajshekar Reddy, Kamalpreet Singh Saluja, Subhankar Ghosh, Anasol Peña-Ríos, Gokul Kumar Gopal, Pradipta Biswas |
Virtual Real. Intell. Hardw. | 4 |
| 2021 | Adversarial Training of Variational Auto-encoders for Continual Zero-shot Learning(A-CZSL)abstractMost existing artificial neural networks(ANNs) fail to learn continually due to catastrophic forgetting, while humans can do the same by maintaining previous tasks' performances. Although storing all the previous data can alleviate the problem, it takes a large memory, infeasible in real-world utilization. We propose a continual zero-shot learning model(A-CZSL) that is more suitable in real-case scenarios to address the issue that can learn sequentially and distinguish classes the model has not seen during training. Further, to enhance the reliability, we develop A -CZSL for a single head continual learning setting where task identity is revealed during the training process but not during the testing. We present a hybrid network that consists of a shared VAE module to hold information of all tasks and task-specific private VAE modules for each task. The model's size grows with each task to prevent catastrophic forgetting of task-specific skills, and it includes a replay approach to preserve shared skills. We demonstrate our hybrid model outperforms the baselines and is effective on several datasets, i.e., CUB, AWA1, AWA2, and aPY. We show our method is superior in class sequentially learning with ZSL(Zero-Shot Learning) and GZSL(Generalized Zero-Shot Learning). The code url is available at the arxiv paper. Subhankar Ghosh |
IJCNN | 1 |
| 2021 | Validating Social Distancing through Deep Learning and VR-Based Digital TwinsabstractThe Covid-19 pandemic resulted in a catastrophic loss to global economies, and social distancing was consistently found to be an effective means to curb the virus's spread. However, it is only as effective when every individual partakes in it with equal alacrity. Past literature outlined scenarios where computer vision was used to detect people and to enforce social distancing automatically. We have created a Digital Twin (DT) of an existing laboratory space for remote monitoring of room occupancy and automatically detecting violation of social distancing. To evaluate the proposed solution, we have implemented a Convolutional Neural Network (CNN) model for detecting people, both in a limited-sized dataset of real humans, and a synthetic dataset of humanoid figures. Our proposed computer vision models are validated for both real and synthetic data in terms of accurately detecting persons, posture, and intermediate distances among people. Abhishek Mukhopadhyay, G. S. Rajshekar Reddy, Subhankar Ghosh, L. R. D. Murthy, Pradipta Biswas |
VRST | 3 |
| 2021 | Deep neural network to detect COVID-19: one architecture for both CT Scans and Chest X-rays
Himadri Mukherjee, Subhankar Ghosh, Ankita Dhar, Sk Md Obaidullah, KC Santosh, Kaushik Roy 0004 |
Appl. Intell. | 2 |
| 2020 | STEFANN: Scene Text Editor Using Font Adaptive Neural NetworkabstractTextual information in a captured scene plays an important role in scene interpretation and decision making. Though there exist methods that can successfully detect and interpret complex text regions present in a scene, to the best of our knowledge, there is no significant prior work that aims to modify the textual information in an image. The ability to edit text directly on images has several advantages including error correction, text restoration and image reusability. In this paper, we propose a method to modify text in an image at character-level. We approach the problem in two stages. At first, the unobserved character (target) is generated from an observed character (source) being modified. We propose two different neural network architectures - (a) FANnet to achieve structural consistency with source font and (b) Colornet to preserve source color. Next, we replace the source character with the generated character maintaining both geometric and visual consistency with neighboring characters. Our method works as a unified platform for modifying text in images. We present the effectiveness of our method on COCO-Text and ICDAR datasets both qualitatively and quantitatively. Prasun Roy, Saumik Bhattacharya, Subhankar Ghosh, Umapada Pal 0001 |
CVPR | 3 |
| 2019 | Deep learning for spoken language identification: Can we visualize speech signal patterns?
Himadri Mukherjee, Subhankar Ghosh, Shibaprasad Sen, Sk Md Obaidullah, KC Santosh, Santanu Phadikar, Kaushik Roy 0004 |
Neural Comput. Appl. | 2 |
| 2018 | A CNN Based Framework for Unistroke Numeral Recognition in Air-WritingabstractAir-writing refers to virtually writing linguistic characters through hand gestures in three dimensional space with six degrees of freedom. In this paper a generic video camera dependent convolutional neural network (CNN) based air-writing framework has been proposed. Gestures are performed using a marker of fixed color in front of a generic video camera followed by color based segmentation to identify the marker and track the trajectory of marker tip. A pre-trained CNN is then used to classify the gesture. The recognition accuracy is further improved using transfer learning with the newly acquired data. The performance of the system varies greatly on the illumination condition due to color based segmentation. In a less fluctuating illumination condition the system is able to recognize isolated unistroke numerals of multiple languages. The proposed framework achieved 97.7%, 95.4% and 93.7% recognition rate in person independent evaluation over English, Bengali and Devanagari numerals, respectively. Prasun Roy, Subhankar Ghosh, Umapada Pal 0001 |
ICFHR | 2 |
| 2014 | Text summarization using Wikipedia
Yogesh Sankarasubramaniam, Krishnan Ramanathan, Subhankar Ghosh |
Inf. Process. Manag. | 3 |