Sohini Roychowdhury

dblp:145/3415 · DBLP profile ↗
← Back
18ranked-venue papers
12as first author
6since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 9 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 6 first-author · 2 since 2021Databases, data management, data science and information retrieval · 7 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2024 ERATTA: Extreme RAG for enterprise-Table To Answers with Large Language Models
abstract
Large language models (LLMs) with retrieval augmented-generation (RAG) have been the optimal choice for scalable generative AI solutions in the recent past. Although RAG implemented with AI agents (agentic-RAG) has been recently popularized, its suffers from unstable cost and unreliable performances for Enterprise-level data-practices. Most existing use-cases that incorporate RAG with LLMs have been either generic or extremely domain specific, thereby questioning the scalability and generalizability of RAG-LLM approaches. In this work, we propose a unique LLM-based system where multiple LLMs can be invoked to enable data authentication, user-query routing, data-retrieval and custom prompting for question-answering capabilities from Enterprise-level data tables on sustainability. The source tables here are highly fluctuating and large in size storing carbon footprint, energy and water usage at buildings in regional levels globally and the proposed framework enables structured responses in under 10 seconds per query. Additionally, we propose a five metric scoring module that detects and reports hallucinations in the LLM responses. Our proposed system and scoring metrics achieve >90% confidence scores across hundreds of user queries in the sustainability, financial health and social media domains. Extensions to the proposed extreme RAG architectures can enable heterogeneous source querying using LLMs.
Sohini Roychowdhury, Marko Krema, Anvar Mahammad, Arijit Mukherjee, Punit Prakashchandra
IEEE Big Data1
2024 Journey of Hallucination-minimized Generative AI Solutions for Financial Decision Makers
abstract
Generative AI has significantly reduced the entry barrier to the domain of AI owing to the ease of use and core capabilities of automation, translation, and intelligent actions in our day to day lives. Currently, Large language models (LLMs) that power such chatbots are being utilized primarily for their automation capabilities on a limited scope. One major limitation of the currently evolving family of LLMs is hallucinations, wherein inaccurate responses are reported as factual. Hallucinations are primarily caused by biased training data, ambiguous prompts and inaccurate LLM parameters, and they majorly occur while combining mathematical facts with language-based context. In this work we present the three major stages in the journey of designing hallucination-minimized LLM-based solutions that are specialized for the decision makers of the financial domain, namely: prototyping, scaling and LLM evolution using human feedback. These three stages and the novel data to answer generation modules presented in this work are necessary to ensure that the Generative AI products are reliable and high-quality to aid key decision-making processes.
Sohini Roychowdhury
WSDM1
2023 Hallucination-minimized Data-to-answer Framework for Financial Decision-makers
abstract
Large Language Models (LLMs) have been applied to build several automation and personalized question-answering prototypes so far. However, scaling such prototypes to robust products with minimized hallucinations or fake responses still remains an open challenge, especially in niche data-table heavy domains such as financial decision making. In this work, we present a novel Langchain-based framework that transforms data tables into hierarchical textual "data chunks" to enable a wide variety of actionable question answering. First, the user-queries are classified by intention followed by automated retrieval of the most relevant data chunks to generate customized LLM prompts per query. Next, the custom prompts and their responses undergo multi-metric scoring to assess for hallucinations and response confidence. The proposed system is optimized with user-query intention classification, advanced prompting, data scaling capabilities and it achieves over $ 90\%$ confidence scores for a variety of user-queries responses ranging from {What, Where, Why, How, predict, trend, anomalies, exceptions} that are crucial for financial decision making applications. The proposed data to answers framework can be extended to other analytical domains such as sales and payroll to ensure optimal hallucination control guardrails.
Sohini Roychowdhury, Marko Krema, Maria Paz Gelpi, Punit Agrawal, Federico Martin Rodriguez, Angel Rodriguez, Jose Ramon Cabrejas, Pablo Martinez Serrano, Arijit Mukherjee
IEEE Big Data1
2022 Semi-supervised and Deep learning Frameworks for Video Classification and Key-frame Identification
abstract
Automating video-based data and machine learning pipelines poses several challenges including metadata generation for efficient storage and retrieval and isolation of key-frames for scene understanding tasks. In this work, we present two semi-supervised approaches that automate this process of manual frame sifting in video streams by automatically classifying scenes for content and filtering frames for fine-tuning scene understanding tasks. The first rule-based method starts from a pre-trained object detector and it assigns scene type, uncertainty and lighting categories to each frame based on probability distributions of foreground objects. Next, frames with the highest uncertainty and structural dissimilarity are isolated as key-frames. The second method relies on the simCLR model for frame encoding followed by label-spreading from 20% of frame samples to label the remaining frames for scene and lighting categories. Also, clustering the video frames in the encoded feature space further isolates key-frames at cluster boundaries. The proposed methods achieve 64–93% accuracy for automated scene categorization for outdoor image videos from public domain datasets of JAAD and KITTI. Also, less than 10% of all input frames can be filtered as key-frames that can then be sent for annotation and fine tuning of machine vision algorithms. Thus, the proposed framework can be scaled to additional video data streams for automated training of perception-driven systems with minimal training images.
Sohini Roychowdhury
IJCNN1
2021 Cirrus: A Long-range Bi-pattern LiDAR Dataset
abstract
In this paper, we introduce Cirrus, a new long-range bi-pattern LiDAR public dataset for autonomous driving tasks such as 3D object detection, critical to highway driving and timely decision making. Our platform is equipped with a high-resolution video camera and a pair of LiDAR sensors with a 250-meter effective range, which is significantly longer than existing public datasets. We record paired point clouds simultaneously using both Gaussian and uniform scanning patterns. Point density varies significantly across such a long range, and different scanning patterns further diversify object representation in LiDAR. In Cirrus, eight categories of objects are exhaustively annotated in the LiDAR point clouds for the entire effective range. To illustrate the kind of studies supported by this new dataset, we introduce LiDAR model adaptation across different ranges, scanning patterns, and sensor devices. Promising results show the great potential of this new dataset to the robotics and computer vision communities.
Ze Wang 0008, Sihao Ding 0002, Ying Li 0139, Jonas Fenn, Sohini Roychowdhury, Andreas Wallin, Lane Martin, Scott Ryvola, Guillermo Sapiro, Qiang Qiu 0001
ICRA5
2021 OPAM: Online Purchasing-behavior Analysis using Machine learning
abstract
Customer purchasing behavior analysis plays a key role in developing insightful communication strategies between online vendors and their customers. To support the recent increase in online shopping trends, in this work, we present a customer purchasing behavior analysis system using supervised, unsupervised and semi-supervised learning methods. The proposed system analyzes session and user-journey level purchasing behaviors to identify customer categories/clusters that can be useful for targeted consumer insights at scale. We observe higher sensitivity to the design of online shopping portals for session-level purchasing prediction with accuracy/recall in range 91-98%/73-99%, respectively. The user-journey level analysis demonstrates five unique user clusters, wherein New Shoppers are most predictable and Impulsive Shoppers are most unique with low viewing and high carting behaviors for purchases. Further, cluster transformation metrics and partial label learning demonstrates the robustness of each user cluster to new/unlabelled events. Thus, customer clusters can aid strategic targeted nudge models.
Sohini Roychowdhury, Ebrahim Alareqi, Wenxi Li
IJCNN1
2020 Few Shot Learning Framework to Reduce Inter-observer Variability in Medical Images
abstract
Most computer aided pathology detection systems rely on large volumes of quality annotated data to aid diagnostics and follow up procedures. However, quality assuring large volumes of annotated medical image data can be subjective and expensive. In this work we present a novel standardization framework that implements three few-shot learning (FSL) models that can be iteratively trained by atmost 5 images per 3D stack to generate multiple regional proposals (RPs) per test image. These FSL models include a novel parallel echo state network (ParESN) framework and an augmented U-net model. Additionally, we propose a novel target label selection algorithm (TLSA) that measures relative agreeability between RPs and the manually annotated target labels to detect the “best” quality annotation per image. Using the FSL models, our system achieves 0.28-0.64 Dice coefficient across vendor image stacks for intra-retinal cyst segmentation. Additionally, the TLSA is capable of automatically classifying high quality target labels from their noisy counterparts for 60-97% of the images while ensuring manual supervision on remaining images. Also, the proposed framework with ParESN model minimizes manual annotation checking to 12-28% of the total number of images. The TLSA metrics further provide confidence scores for the automated annotation quality assurance. Thus, the proposed framework is flexible to extensions for quality image annotation curation of other image stacks as well.
Sohini Roychowdhury
ICPR1
2020 AV-SLAM: Autonomous Vehicle SLAM with Gravity Direction Initialization
abstract
Simultaneous localization and mapping (SLAM) algorithms aimed for autonomous vehicles (AVs) are required to utilize sensor redundancies specific to AVs and enable accurate, fast and repeatable estimations of pose and path trajectories. In this work, we present a combination of three SLAM algorithms that utilize a different subset of available sensors such as inertial measurement unit (IMU), a gray-scale mono-camera, and a Lidar. Also, we propose a novel acceleration-based gravity direction initialization (AGI) method for the visual-inertial SLAM algorithm. We analyze the SLAM algorithms and initialization methods for pose estimation accuracy, speed of convergence and repeatability on the KITTI odometry sequences. The proposed VI-SLAM with AGI method achieves relative pose errors less than 2%, convergence in half a minute or less and convergence time variability less than 3s, which makes it preferable for AVs.
Kaan Yilmaz, Baris Suslu, Sohini Roychowdhury, L. Srikar Muppirisetty
ICPR3
2019 Segmentation of Intra-Retinal Cysts From Optical Coherence Tomography Images Using a Fully Convolutional Neural Network Model
abstract
Optical coherence tomography (OCT) is an imaging modality that is used extensively for ophthalmic diagnosis, near-histological visualization, and quantification of retinal abnormalities such as cysts, exudates, retinal layer disorganization, etc. Intra-retinal cysts (IRCs) occur in several macular disorders such as, diabetic macular edema, retinal vascular disorders, age-related macular degeneration, and inflammatory disorders. Automated segmentation of IRCs poses challenges owing to variations in the acquisition system scan intensities, speckle noise, and imaging artifacts. Several segmentation methods have been proposed in the literature for IRC segmentation on vendor-specific OCT images that lack generalizability across imaging systems. In this paper, we propose a fully convolutional network (FCN) model for vendor-independent IRC segmentation. The proposed method counteracts image noise variabilities and trains FCN models on OCT sub-images from the OPTIMA cyst segmentation challenge dataset (with four different vendor-specific images, namely, Cirrus, Nidek, Spectralis, and Topcon). Further, optimal data augmentation and model hyperparametrization are shown to prevent over-fitting for IRC area segmentation. The proposed method is evaluated on the test dataset with a recall/precision rate of 0.66/0.79 across imaging vendors. The Dice correlation coefficient of the proposed method outperforms that of the published algorithms in the OPTIMA cyst segmentation challenge with a Dice rate of 0.71 across the vendors.
G. N. Girish, Bibhash Thakur, Sohini Roychowdhury, Abhishek R. Kothari, Jeny Rajan
IEEE J. Biomed. Health Informatics3
2018 Fast Proposals for Image and Video Annotation Using Modified Echo State Networks
abstract
Deep learning frameworks for computer-vision applications require fast and scalable annotation systems. Since manually annotated data for semantic segmentation tasks is time-consuming and tough to quality assure, accurate and automated region-based proposals can significantly aid high quality data annotation. In this work, we propose modified Echo State Network (ESN) models that iteratively learn from a small subset of data (20-30% images) and adapt to a variety of semantic segmentation goals without manual supervision on test images. We observe that the modified ESN model that relies on 3 x 3 pixel neighborhood features scales across segmentation tasks with mean segmentation F_scores in the range of 0.58-0.87 for complete foreground and specific foreground segmentation tasks, respectively. Thus, the proposed methods can be specifically useful for fast semantic proposal estimation to enhance the annotation resourcefulness for time sensitive applications in the automotive field.
Sohini Roychowdhury, L. Srikar Muppirisetty
ICMLA1
2018 Machine Learning Models for Road Surface and Friction Estimation using Front-Camera Images
abstract
Automotive active safety systems can significantly benefit from real-time road friction estimates (RFE) by adapting driving styles, specific to the road conditions. This work presents a 2-stage approach for indirect RFE estimation using front-view camera images captured from vehicles. In stage-1, convolutional neural network model architectures are implemented to learn region-specific features for road surface condition (RSC) classification. Texture-based features from the drivable surface, sky and surroundings are found to be separate regions of interest for dry, wet/water, slush and snow\ice RSC classification. In stage-2, a rule-based model that relies on domain-specific guidelines is implemented to segment the ego-lane drivable surface into [5×3] patches, followed by patch classification and quantization to separate images with high, medium and low RFE. The proposed method achieves average accuracy of 97% for RSC classification in stage-1 and 89% for RFE classification in stage-2, respectively. The 2-stage models are trained using publicly available data sets to enable benchmarking for future methodologies in the autonomous driving domain.
Sohini Roychowdhury, Minming Zhao, Andreas Wallin, Niklas Ohlsson, Mats Jonasson
IJCNN1
2016 Topic modeling for management sciences: A network-based approach
abstract
Big data mining and unsupervised pattern recognition from large corpus of text-based documents has been an active research topic over the past decade. This paper presents a novel sequence of network-based models for identifying high-dimensional clustering patterns between topics for quantitative and predictive modeling of trends in Management Science (an INFORMS Journal) papers over the past 54 years. The proposed methods extrapolate a new spatial dimension from publication records to identify and assess topic inter-dependence and clustering trends over time. First, the optimal number of topics for trend analysis is identified based on spatial clustering patterns using Self-Organizing-Maps (SOM). Next, topic models are used to construct weighted and unweighted complex networks. Based on spatio-temporal clustering trends in the complex networks, the influence, importance and uniqueness of topics are quantified. Finally, the dynamic trends in topic influence are modeled for predictive purposes using Hidden Markov Models (HMM). The proposed methods provide insights into topic type co-existence patterns, topic type rankings, identify ~40% topics as unique and predict topic importance with average accuracy per topic in the range of 79-84%. Thus, the proposed methods provide the apparatus to translate time-series text-intensive data sets to spatio-temporal models that can provide additional insights on data interdependencies and inter-data influences.
Max Menenberg, Surya Pathak, Hari P. Udyapuram, Srinagesh Gavirneni 0001, Sohini Roychowdhury
IEEE BigData5
2016 Application of big data analytics for automated estimation of CT image quality
abstract
With the increasing applications of Big Data analytics in medical image processing systems, there has been a growing need for quantitative medical image quality assessment techniques. Specifically for computed tomography (CT) images, quantitative image assessment can allow for benchmarking image processing methods and optimization of image acquisition parameters. In this work, large volumes of CT images from phantoms and patients are analyzed using 3 data models that vary in their implementation time complexities. The goal here is to identify the optimal method that scales across data set variabilities for predictive modeling of CT image quality (CTIQ). The first two models rely on spatial segmentation of regions-of-interest (ROIs) and estimate CTIQs in terms of segmented pixel variabilities. The third, convolutional neural network (CNN) model relies on error back-propagation from the training set of images to learn the regions indicative of CTIQ. We observe that for 70/30 data split, the average multi-class classification accuracies for CTIQ prediction using the 3 data models range from 73.6-100% and 50-100% for the phantom and patient CT images, respectively. Using variance of pixels within the segmented ROIs as a CTIQ classification parameter, the spatial segmentation data models are found to be more generalizable that the CNN model. However, the CNN model is found to be more suitable for CT image texture classification in the absence of structural variabilities. Our analysis demonstrates that spatial ROI segmentation data models are consistent CTIQ estimators while the CNN models are consistent identifiers of structural similarities for CT image data sets.
Maitham D. Naeemi, Johnny Ren, Nathan Hollcroft, Adam M. Alessio, Sohini Roychowdhury
IEEE BigData5
2016 Non-deep CNN for multi-modal image classification and feature learning: An Azure-based model
abstract
Convolutional Neural Networks (CNN) are useful methods for identification of previously unknown embedded patterns in images. Several object and facial recognition along with image segmentation tasks have benefited from the non-linear abstraction of hybrid features using CNN. This work presents a novel CNN model parametrization work-flow developed on the cloud-computing platform of Microsoft Azure Machine Learning Studio (MAMLS) that is capable of learning from the feature maps and classifying multi-modal images with different variabilities using one common flow. This two-step work-flow trains CNN models using 70/30 data split. First, the CNN layers are fixed and the optimal kernel and normalization parameters are identified that maximize classification accuracy on the test data. Next, using the optimal kernel and normalization parameters, the best CNN architecture that maximizes classification accuracy is detected. Finally, the activated feature maps (AFMs) from the optimally parameterized CNN model are analyzed to learn new features that can enhance image-based classification accuracies. The proposed flow achieves classification accuracies in the range of 92.5-99.2% that can be further enhanced by doubling the samples based on the features learned from the AFMs. The proposed non-deep CNN models in the MAMLS platform are capable of processing image data sets with 400-4 million samples using a common flow without exponential increase in the computation time. Thus, optimally parametrized non-deep CNN models are capable of identifying novel features that may enhance image-based classification accuracies.
Sohini Roychowdhury, Johnny Ren
IEEE BigData1
2016 Optic Disc Boundary and Vessel Origin Segmentation of Fundus Images
abstract
This paper presents a novel classification-based optic disc (OD) segmentation algorithm that detects the OD boundary and the location of vessel origin (VO) pixel. First, the green plane of each fundus image is resized and morphologically reconstructed using a circular structuring element. Bright regions are then extracted from the morphologically reconstructed image that lie in close vicinity of the major blood vessels. Next, the bright regions are classified as bright probable OD regions and non-OD regions using six region-based features and a Gaussian mixture model classifier. The classified bright probable OD region with maximum Vessel-Sum and Solidity is detected as the best candidate region for the OD. Other bright probable OD regions within 1-disc diameter from the centroid of the best candidate OD region are then detected as remaining candidate regions for the OD. A convex hull containing all the candidate OD regions is then estimated, and a best-fit ellipse across the convex hull becomes the segmented OD boundary. Finally, the centroid of major blood vessels within the segmented OD boundary is detected as the VO pixel location. The proposed algorithm has low computation time complexity and it is robust to variations in image illumination, imaging angles, and retinal abnormalities. This algorithm achieves 98.8%-100% OD segmentation success and OD segmentation overlap score in the range of 72%-84% on images from the six public datasets of DRIVE, DIARETDB1, DIARETDB0, CHASE_DB1, MESSIDOR, and STARE in less than 2.14 s per image. Thus, the proposed algorithm can be used for automated detection of retinal pathologies, such as glaucoma, diabetic retinopathy, and maculopathy.
Sohini Roychowdhury, Dara Koozekanani, Sam N. Kuchinka, Keshab K. Parhi
IEEE J. Biomed. Health Informatics1
2015 A generalized flow for multi-class and binary classification tasks: An Azure ML approach
abstract
The constant growth in the present day real-world databases pose computational challenges for a single computer. Cloud-based platforms, on the other hand, are capable of handling large volumes of information manipulation tasks, thereby necessitating their use for large real-world data set computations. This work focuses on creating a novel Generalized Flow within the cloud-based computing platform: Microsoft Azure Machine Learning Studio (MAMLS) that accepts multi-class and binary classification data sets alike and processes them to maximize the overall classification accuracy. First, each data set is split into training and testing data sets, respectively. Then, linear and nonlinear classification model parameters are estimated using the training data set. Data dimensionality reduction is then performed to maximize classification accuracy. For multi-class data sets, data-centric information is used to further improve overall classification accuracy by reducing the multi-class classification to a series of hierarchical binary classification tasks. Finally, the performance of optimized classification model thus achieved is evaluated and scored on the testing data set. The classification characteristics of the proposed flow are comparatively evaluated on 3 public data sets and a local data set with respect to existing state-of-the-art methods. On the 3 public data sets, the proposed flow achieves 78-97.5% classification accuracy. Also, the local data set, created using the information regarding presence of Diabetic Retinopathy lesions in fundus images, results in 85.3-95.7% average classification accuracy, which is higher than the existing methods. Thus, the proposed generalized flow can be useful for a wide range of application-oriented "big data sets".
Matthew Bihis, Sohini Roychowdhury
IEEE BigData2
2015 Blood Vessel Segmentation of Fundus Images by Major Vessel Extraction and Subimage Classification
abstract
This paper presents a novel three-stage blood vessel segmentation algorithm using fundus photographs. In the first stage, the green plane of a fundus image is preprocessed to extract a binary image after high-pass filtering, and another binary image from the morphologically reconstructed enhanced image for the vessel regions. Next, the regions common to both the binary images are extracted as the major vessels. In the second stage, all remaining pixels in the two binary images are classified using a Gaussian mixture model (GMM) classifier using a set of eight features that are extracted based on pixel neighborhood and first and second-order gradient images. In the third postprocessing stage, the major portions of the blood vessels are combined with the classified vessel pixels. The proposed algorithm is less dependent on training data, requires less segmentation time and achieves consistent vessel segmentation accuracy on normal images as well as images with pathology when compared to existing supervised segmentation methods. The proposed algorithm achieves a vessel segmentation accuracy of 95.2%, 95.15%, and 95.3% in an average of 3.1, 6.7, and 11.7 s on three public datasets DRIVE, STARE, and CHASE_DB1, respectively.
Sohini Roychowdhury, Dara Koozekanani, Keshab K. Parhi
IEEE J. Biomed. Health Informatics1
2014 DREAM: Diabetic Retinopathy Analysis Using Machine Learning
abstract
This paper presents a computer-aided screening system (DREAM) that analyzes fundus images with varying illumination and fields of view, and generates a severity grade for diabetic retinopathy (DR) using machine learning. Classifiers such as the Gaussian Mixture model (GMM), k-nearest neighbor (kNN), support vector machine (SVM), and AdaBoost are analyzed for classifying retinopathy lesions from nonlesions. GMM and kNN classifiers are found to be the best classifiers for bright and red lesion classification, respectively. A main contribution of this paper is the reduction in the number of features used for lesion classification by feature ranking using Adaboost where 30 top features are selected out of 78. A novel two-step hierarchical classification approach is proposed where the nonlesions or false positives are rejected in the first step. In the second step, the bright lesions are classified as hard exudates and cotton wool spots, and the red lesions are classified as hemorrhages and micro-aneurysms. This lesion classification problem deals with unbalanced datasets and SVM or combination classifiers derived from SVM using the Dempster-Shafer theory are found to incur more classification error than the GMM and kNN classifiers due to the data imbalance. The DR severity grading system is tested on 1200 images from the publicly available MESSIDOR dataset. The DREAM system achieves 100% sensitivity, 53.16% specificity, and 0.904 AUC, compared to the best reported 96% sensitivity, 51% specificity, and 0.875 AUC, for classifying images as with or without DR. The feature reduction further reduces the average computation time for DR severity per image from 59.54 to 3.46 s.
Sohini Roychowdhury, Dara Koozekanani, Keshab K. Parhi
IEEE J. Biomed. Health Informatics1