Aryya Gangopadhyay

dblp:g/AryyaGangopadhyay · DBLP profile ↗
← Back
69ranked-venue papers
4as first author
18since 2021 · last 2025
0000-0002-7553-7932ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 33 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 25 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 21 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 8 · 2 since 2021Security and privacy · 6Systems, architecture and hardware · 3 · 3 since 2021Software engineering, systems software and programming languages · 2Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Real-Time Object Tracking on the Edge
abstract
This paper presents a real-time detection, tracking, counting, and distance estimation framework deployed on the Boston Dynamics Spot robot, equipped with RGB and thermal cameras. Leveraging edge computing devices such as the NVIDIA Jetson Nano, the system autonomously processes data in dynamic terrains with minimal human intervention. A custom-trained YOLOv8 model, fine-tuned on a unique dataset tailored for military applications, is integrated with the StrongSORT algorithm for object tracking. Additionally, a novel geometric calculation methodology enables precise angle estimation and spatial mapping, enhancing situational awareness. The framework's capabilities include live visualization of detected objects, area-based counting, and mapping within an operational environment using ROS-based tools. Field demonstrations conducted at the Army Research Lab validate the system's effectiveness in processing complex thermal and RGB data in real-time. The proposed solution offers significant potential for enhancing autonomous robotic deployments in mission-critical applications.
Mohana Kundurthi, Tamilselvan Gurunathan, Md. Azim Khan, Muhammad Shehrose Raza, Aryya Gangopadhyay
ICFEC5
2024 A Progressive Meta-Algorithm for Large and Seamless Super Resolution Images
abstract
We present a progressive meta-algorithm for single image super resolution of very large and seamless images. Deep single image super resolution backbone networks are capable of upsampling local image patches but are susceptible to severe tiling artifacts when attempting to integrate the patches together into a single large image. For large images such as 4K images, the individual low-resolution patches also comprise a narrow receptive field that lacks context regarding the pixel information from the surrounding patches. Our Progressive meta-algorithm is inspired by prior works in progressive GANs and is designed to resolve inter-patch tiling artifact across varying scales while further incorporating context by expanding the receptive field to include the surrounding patches through multiple rounds of progressive upsampling. Our meta-algorithm is compatible with standard super-resolving backbones and enables super-resolution of very large images such as 4K images while overcoming practical GPU memory limitations with commodity graphics cards. We evaluate our Meta-Algorithm with the challenging task of 16x super resolution for 4096×4096 images using the ESRGAN super-resolving backbone and demonstrate significantly improved image quality versus a baseline patch-based approach as evaluated using the Learned Perceptual Image Patch Similarity (LPIPS) metric.
Jayalakshmi Mangalagiri, Aryya Gangopadhyay, David Chapman 0001
IEEE Big Data2
2024 Gradient Inversion Attacks on Acoustic Signals: Revealing Security Risks in Audio Recognition Systems
abstract
With a greater emphasis on data confidentiality and legislation, distributed training and collaborative machine learning algorithms are being developed to protect sensitive private data. Gradient exchange has become a widely used practice in those multi-node machine learning systems. But with the advent of gradient inversion attacks, it is already established that private training data can be revealed from the gradients. Gradient inversion attacks covertly spy on gradient updates and backtrack from the gradients to obtain information about the raw data. Although this attack has been widely studied in computer vision and natural language processing tasks, understanding the impact of this attack on acoustic signals still requires a comprehensive investigation. To the best of our knowledge, we are the first to explore gradient inversion attacks on acoustic signals by extracting the speakers’ voices from an audio recognition system. Here, we design a new application of gradient inversion attack to retrieve the audio signal used for training the deep learning model, irrespective of whether the audio has undergone conversion into mel-spectrogram or MFCC representations prior to feed to neural network. Experimental results demonstrate the capability of our attack method to extract the input vectors of the audio data from the gradients, which highlight the security risks in revealing the sensitive audio data from highly secured systems. We also discuss several possible strategies as countermeasures and their effectiveness to prevent the attack.
Pretom Roy Ovi, Aryya Gangopadhyay
ICASSP2
2024 GestRight: Understanding the Feasibility of Gesture-driven Tele-Operation in Human-Robot Teams
abstract
In this paper, we propose GestRight, a real-time system for gesture-based tele-operation of a mobile robot. For field use (e.g., smart factory settings, search and rescue missions, etc.), relying on tablet-based controls or joysticks are limiting which has led to the recent interest in hands-free operation of these assistive robots. In this work, we design three gesture-based schemes, namely, fist, touch, and wheel, represent three levels of precision–intuitiveness tradeoffs for low-level navigational control of mobile robots. GestRight includes a head-mounted device that captures hand joint data for accurate gesture recognition which is then translated to motion commands at an edge server. Through a user study involving seventeen participants, we present quantitative insights in comparison to traditional modes of control. Specifically, we evaluate GestRight in terms of the ease of navigational control, task time, and amount of errors/corrective actions required, run extensive statistical analyses, and provide a series of design recommendations for gesture-driven teleoperation systems. Our results show that gesture based schemes perform as well as traditional modes of control in contrast to participants’ self-reports on how successful they felt in controlling the robots.
Kevin Rippy, Aryya Gangopadhyay, Kasthuri Jayarajah
IROS2
2023 Flood-ResNet50: Optimized Deep Learning Model for Efficient Flood Detection on Edge Device
abstract
Floods are highly destructive natural disasters that result in significant economic losses and endanger human and wildlife lives. Efficiently monitoring Flooded areas through the utilization of deep learning models can contribute to mitigating these risks. This study focuses on the deployment of deep learning models specifically designed for classifying flooded and non-flooded in UAV images. In consideration of computational costs, we propose modified version of ResNet50 called Flood-ResNet50. By incorporating additional layers and leveraging transfer learning techniques, Flood-ResNet50 achieves comparable performance to larger models like VGG16/19, AlexNet, DenseNet161, EfficientNetB7, Swin(small), and vision transformer. Experimental results demonstrate that the proposed modification of ResNet50, incorporating additional layers, achieves a classification accuracy of 96.43%, F1 score of 86.36%, Recall of 81.11%, Precision of 92.41 %, model size 98MB and FLOPs 4.3 billions for the FloodNet dataset. When deployed on edge devices such as the Jetson Nano, our model demonstrates faster inference speed (820 ms), higher throughput (39.02 fps), and lower average power consumption (6.9 W) compared to larger ResNet101 and ResNet152 models.
Md. Azim Khan, Joyce Padela, Muhammad Shehrose Raza, Aryya Gangopadhyay, Jianwu Wang 0001, James R. Foulds, Carl E. Busart, Robert F. Erbacher
ICMLA5
2023 Vision Transformer-based Real-Time Camouflaged Object Detection System at Edge
abstract
Camouflaged object detection is a challenging task in computer vision that involves identifying objects that are intentionally or unintentionally hidden in their surrounding environment. Vision Transformer mechanisms play a critical role in improving the performance of deep learning models by focusing on the most relevant features that help object detection under camouflaged conditions. In this paper, we utilized a vision transformer (VT) in two phases, a) By integrating VT with a deep learning architecture for efficient monocular depth map generation for camouflaged objects and b) By embedding VT multiclass object detection model with multimodal feature input (RGB with RGB-D) that increases the visual cues and provides more representational information to the model for performance enhancement. Additionally, we performed an ablation study to understand the role of the vision transformer in camouflaged object detection and incorporated GRAD-CAM on top of the model to visualize the performance improvement achieved by embedding the VT in the model architecture. We deployed the model on resource-constrained edge devices for real-time object detection to realistically test the performance of the trained model.
Rohan Putatunda, Md. Azim Khan, Aryya Gangopadhyay, Jianwu Wang 0001, Carl E. Busart, Robert F. Erbacher
SMARTCOMP3
2023 Reproducible and Portable Big Data Analytics in the Cloud
abstract
Cloud computing has become a major approach to help reproduce computational experiments. Yet there are still two main difficulties in reproducing batch based Big Data analytics (including descriptive and predictive analytics) in the cloud. The first is how to automate end-to-end scalable execution of analytics including distributed environment provisioning, analytics pipeline description, parallel execution, and resource termination. The second is that an application developed for one cloud is difficult to be reproduced in another cloud, a.k.a. vendor lock-in problem. To tackle these problems, we leverage serverless computing and containerization techniques for automated scalable execution and reproducibility, and utilize the adapter design pattern to enable application portability and reproducibility across different clouds. We propose and develop an open-source toolkit that supports 1) fully automated end-to-end execution and reproduction via a single command, 2) automated data and configuration storage for each execution, 3) flexible client modes based on user preferences, 4) execution history query, and 5) simple reproduction of existing executions in the same environment or a different environment. We did extensive experiments on both AWS and Azure using four Big Data analytics applications that run on virtual CPU/GPU clusters. The experiments show our toolkit can achieve good execution performance, scalability, and efficient reproducibility for cloud-based Big Data analytics.
Xin Wang 0122, Pei Guo, Xingyan Li, Aryya Gangopadhyay, Carl E. Busart, Jade Freeman, Jianwu Wang 0001
IEEE Trans. Cloud Comput.4
2023 SAM-VQA: Supervised Attention-Based Visual Question Answering Model for Post-Disaster Damage Assessment on Remote Sensing Imagery
abstract
Each natural disaster leaves a trail of destruction and damage that must be effectively managed to reduce its negative impact on human life. Any delay in making proper decisions at the post-disaster managerial level can increase human suffering and waste resources. Proper managerial decisions after any natural disaster rely on an appropriate assessment of damages using data-driven approaches, which are needed to be efficient, fast, and interactive. The goal of this study is to incorporate a deep interactive data-driven framework for proper damage assessment to speed up the response and recovery phases after a natural disaster. Hence, this paper focuses on introducing and implementing the Visual Question Answering (VQA) framework for post-disaster damage assessment based on drone imagery, namely Supervised Attention-Based VQA (SAM-VQA). In visual question answering, query-based answers from images regarding the situation in disaster-affected areas can provide valuable information for decision-making. Unlike other computer vision tasks, visual question answering is more interactive and allows one to get instant and effective scene information by asking questions in natural language from images. In this work, we present a VQA dataset and propose a novel supervised attention-based VQA framework (SAM-VQA) for post-disaster damage assessment on remote sensing images. Our model outperforms state-of-the-art attention-based VQA techniques, including Stacked Attention Networks (SAN) [1] and Multi-modal Factorized Bilinear (MFB) with Co-Attention [2]. Furthermore, our proposed model can derive appropriate visual attention based on questions to predict answers, making our approach trustworthy.
Argho Sarkar, Tashnim Chowdhury, Robin R. Murphy, Aryya Gangopadhyay, Maryam Rahnemoonfar
IEEE Trans. Geosci. Remote. Sens.4
2022 GADAN: Generative Adversarial Domain Adaptation Network For Debris Detection Using Drone
abstract
In maritime, coastal, and riverine environments debris has become an abundant pollutant, posing a significant threat to aquatic life. Existing debris detection systems are mostly designed to detect debris for specific type of environment. Therefore, we propose GADAN, an autonomous debris detection model fusing RGB and thermal images, and experiment on both over land and aquatic environment. We collected data using the Parrot Anafi Thermal drone from different water streams, and construction sites near the aquatic environment. We employ a generative domain adaptation based architecture to generate the thermal image compatible with the RGB image. We postulate a two-stream network on these thermal and RGB pairs of images separately, and pass the concatenated features to object detecting YOLO model. We also demonstrated the performance of debris detection using only RGB and only thermal images. Our proposed GADAN network incorporating RGB and thermal images outperforms (with a mean average precision of 0.97) RCNN and single modality based models.
Masud Ahmed, Naima Khan, Pretom Roy Ovi, Nirmalya Roy, Sanjay Purushotham, Aryya Gangopadhyay, Suya You
DCOSS6
2022 Secure Federated Training: Detecting Compromised Nodes and Identifying the Type of Attacks
abstract
Federated learning (FL) allows a set of clients to collaboratively train a model without sharing private data. As a result, FL has limited control over the local data and corresponding training process. Therefore, it is susceptible to poisoning attacks in which malicious clients use malicious training data or local updates to poison the global model. In this work, we first studied the data level and model level poisoning attacks. We simulated model poisoning attacks by tampering the local model updates during each round of communication and data poisoning attacks by training a few clients on malicious data. And clients under such attacks carry faulty information to the server, poison the global model, and restrict it from convergence. Therefore, detecting clients under attacks as well as identifying the type of attacks are required to recover the clients from their malicious status. To address these issues, we proposed a way under federated framework that enables the detection of malicious clients and attack types while ensuring data privacy. Our clustering-based approach utilizes the neuron’s activations from the local models to identify the type of poisoning attacks. We also proposed to check the weight distribution of local model updates among the participating clients to detect malicious clients. Our experimental results validated the robustness of the proposed framework against the attacks mentioned above by successfully detecting the compromised clients and the attack types. Moreover, the global model trained on MNIST data couldn’t reach the optimal point even after 75 rounds because of malicious clients, whereas the proposed approach by detecting the malicious clients ensured convergence within only 30 rounds and 40 rounds in independent and identically distributed (IID) and non- independent and identically distributed (non-IID) setup respectively.
Pretom Roy Ovi, Aryya Gangopadhyay, Robert F. Erbacher, Carl E. Busart
ICMLA2
2022 Environmental Sound Classification for Flood Event Detection
abstract
Flood is one of the common natural disasters that can severely affect human life and properties. Early detection, therefore, is of paramount importance to provide help through an emergency response team. Robust flood detection techniques so far have been based on computer vision using images either from cameras, satellite imagery, remote sensing, or radar-based images. However, sound signal-based flood event detection has not been widely explored. In this work, we design an end-to-end architecture for a deep learning-based flood-related sound event detection model. We employ Mel-Spectrogram-based auditory signal analysis and deep learning models for sound event detection (SED). We evaluated four deep learning models under the following two categories: (i) Binary classification Flood/No Flood, vs. Windy vs. Non-Windy, and (ii) Multi-classification for more granular flood and wind events. The experimental results performed in these settings on the datasets collected from real deployment showed an accuracy of around 78%.
Bipendra Basnyat, Nirmalya Roy, Aryya Gangopadhyay, Adrienne Raglin
Intelligent Environments3
2022 CogAx: Early Assessment of Cognitive and Functional Impairment from Accelerometry
abstract
An individual’s cognitive and functional abilities are commonly assessed through physical and mental status examination, observational performance measures, surveys and proxy reports of symptoms. These strategies are not ideal for early impairment detection as the individual needs to be present physically at the clinic to avail the assessments, especially for older adults who require assistance from a caregiver, and experience mobility, cognitive and functional disabilities from neurodegenerative disorders. Moreover, these strategies rely on self-reporting and proxy reports for evaluation which often leads to under-reporting of symptoms and decrease the validity of these measures. We argue that an early assessment of functional, and cognitive health impairment can be obtained from the individual’s daily activities captured through accelerometry. In this work, we postulate to learn high-level motion related representations from accelerometer data to better correlate with underlying functional and cognitive health parameters of older adults using a contrastive and multi-task learning framework. In particular, we posit a novel indicator, Impairment Indicator using the proposed multi-task learning framework that can indicate functional or cognitive decline as neurodegenerative disease progresses. An extensive 24-hour data collection from 25 older adults with the clinician in-the-loop was carried out in a retirement community center with IRB approval. We collected the activity patterns using wearables in their homes in addition to survey-based assessments and observational performance measures recorded by a clinical evaluator to infer their current cognitive and functional impairment status. Our evaluation on the acquired dataset reveals that the representations learned using contrastive learning aids in improving the detection of activities, activity performance score, and stage of dementia to 92%, 97%, and 98%, respectively.
Sreenivasan Ramasamy Ramamurthy, Soumyajit Chatterjee, Elizabeth Galik, Aryya Gangopadhyay, Nirmalya Roy, Bivas Mitra, Sandip Chakraborty 0001
PerCom4
2022 An edge-cloud integrated framework for flexible and dynamic stream analytics
abstract
With the popularity of Internet of Things (IoT), edge computing and cloud computing, more and more stream analytics applications are being developed including real-time trend prediction and object detection on top of IoT sensing data. One popular type of stream analytics is the recurrent neural network (RNN) deep learning model based time series or sequence data prediction and forecasting. Different from traditional analytics that assumes data are available ahead of time and will not change, stream analytics deals with data that are being generated continuously and data trend/distribution could change (a.k.a. concept drift), which will cause prediction/forecasting accuracy to drop over time. One other challenge is to find the best resource provisioning for stream analytics to achieve good overall latency. In this paper, we study how to best leverage edge and cloud resources to achieve better accuracy and latency for stream analytics using a type of RNN model called long short-term memory (LSTM). We propose a novel edge–cloud integrated framework for hybrid stream analytics that supports low latency inference on the edge and high capacity training on the cloud. To achieve flexible deployment, we study different approaches of deploying our hybrid learning framework including edge-centric, cloud-centric and edge–cloud integrated. Further, our hybrid learning framework can dynamically combine inference results from an LSTM model pre-trained based on historical data and another LSTM model re-trained periodically based on the most recent data. Using real-world and simulated stream datasets, our experiments show the proposed edge–cloud deployment is the best among all three deployment types in terms of latency. For accuracy, the experiments show our dynamic learning approach performs the best among all learning approaches for all three concept drift scenarios.
Xin Wang 0122, Md. Azim Khan, Jianwu Wang 0001, Aryya Gangopadhyay, Carl E. Busart, Jade Freeman
Future Gener. Comput. Syst.4
2022 STAR-Lite: A light-weight scalable self-taught learning framework for older adults' activity recognition
Sreenivasan Ramasamy Ramamurthy, Indrajeet Ghosh, Aryya Gangopadhyay, Elizabeth Galik, Nirmalya Roy
Pervasive Mob. Comput.3
2022 Using Randomness to Improve Robustness of Tree-Based Models Against Evasion Attacks
abstract
Machine learning models have been widely used in security applications. However, it is well-known that adversaries can adapt their attacks to evade detection. There has been some work on making machine learning models more robust to such attacks. However, one simple but promising approach calledrandomizationis under-explored. In addition, most existing works focus on models with differentiable error functions while tree-based models do not have such error functions but are quite popular because they are easy to interpret. This paper proposes a novel randomization-based approach to improve robustness of tree-based models against evasion attacks. The proposed approach incorporates randomization into both model training time and model application time (meaning when the model is used to detect attacks). We also apply this approach to random forest, an existing ML method which already has incorporated randomness at training time but still often fails to generate robust models. We proposed a novel weighted-random-forest method to generate more robust models and a clustering method to add randomness at model application time. We also proposed a theoretical framework to provide a lower bound for adversaries’ effort. Experiments on intrusion detection and spam filtering data show that our approach further improves robustness of random-forest method.
Fan Yang 0124, Zhiyuan Chen 0003, Aryya Gangopadhyay
IEEE Trans. Knowl. Data Eng.3
2021 Classification of COVID-19 using Deep Learning and Radiomic Texture Features extracted from CT scans of Patients Lungs
abstract
COVID-19 is an air-borne viral infection, which infects the respiratory system in the human body, and it became a global pandemic in early March 2020. The damage caused by the COVID-19 disease in a human lung region can be identified using Computed Tomography (CT) scans. We present a novel approach in classifying COVID-19 infection and normal patients using a Random Forest (RF) model to train on a combination of Deep Learning (DL) features and Radiomic texture features extracted from CT scans of patient’s lungs. We developed and trained DL models using CNN architectures for extracting DL features. The Radiomic texture features are calculated using CT scans and its associated infection masks. In this work, we claim that the RFs classification using the DL features in conjunction with Radiomic texture features enhances prediction performance. The experiment results show that our proposed models achieve a higher True Positive rate with the average Area Under the Receiver Curve (AUC) of 0.9768, 95% Confidence Interval (CI) [0.9757, 0.9780].
Jayalakshmi Mangalagiri, Jones Sam Sugumar, Sumeet Menon, David Chapman 0001, Yaacov Yesha, Aryya Gangopadhyay, Yelena Yesha
IEEE BigData6
2021 ARIS: A Real Time Edge Computed Accident Risk Inference System
abstract
To deploy an intelligent transport system in urban environment, an effective and real-time accident risk prediction method is required that can help maintain road safety, provide adequate level of medical assistance and transport in case of an emergency. Reducing traffic accidents is an important problem for increasing public safety, so accident analysis and prediction have been a subject of extensive research in recent time. Even if a traffic hazard occurs, a readily deployable structure with an accurate prediction of accident can contribute to better management of rescue resources. But the significant shortcomings of current studies are the use of small-scale datasets with minimal scope, being based on extensive data sets, and not being applicable for real-time purposes. To overcome these challenges, we propose ARIS: a system for real-time traffic accident prediction built on a traffic accident dataset named ‘US-Accidents’ which covers 49 states of United States, collected from February 2016 to June 2020. Our approach is based on a deep neural network model that utilizes a variety of data characteristics, such as time-sensitive weather data, textual information, and discerning factors. We have tested ARIS against multiple baselines through a comprehensive series of experiments across several major cities of USA, and we have noticed significant improvement during inference especially in detecting accident classes. Additionally, to make our model edge-implementable we have compressed our model using a joint technique of magnitude-based weight pruning and model quantization. We have also demonstrated the inference results along with power consumption profiling after deploying the model on a resource constrained environment that consists of Intel Neural Compute Stick 2 (NCS2) with Raspberry Pi 4B (RPi4). Our investigation and observations indicate major improvements to predict unusual traffic accident event even after model compression and deployment. We have managed to reduce the model size and inference time by ≈ 6x, and ≈ 70 % respectively with insignificant drop in performance. Furthermore, to better understand the importance of each individual type of variables used in our analysis, we have showcased a comprehensive ablation study.
Pretom Roy Ovi, Emon Dey, Nirmalya Roy, Aryya Gangopadhyay
SMARTCOMP4
2021 STAR: A Scalable Self-taught Learning Framework for Older Adults' Activity Recognition
abstract
Activity Recognition (AR) in older adults living with Neurocognitive disorders caused by diseases such as Alzheimer’s is still a challenging research problem. The inherent natural variation in performing an activity increases while repeating the same activity for an older adult, let alone the variation introduced when another older adult performs the same activity. Moreover, the challenges in acquiring the labeled data while preserving the privacy, availability of annotators with domain knowledge, aversion towards cameras even for a minimal amount of time for ground truth data collection, and psychological and mental health status make AR for older adults challenging. In this paper, we postulate a self-taught learning-based approach that helps recognize activities with variations that are not being directly seen during the training phase. We hypothesize that the features extracted using deep architectures from unlabeled data instances can learn general underlying representations of activities efficiently and help improve activity classification in a supervised setting, although the data instances in labeled data do not follow the generative distribution of that of unlabeled data. We posit real data from a retirement community center using our in-house SenseBox infrastructure and survey-based assessments concurrently done by a clinical evaluator to study the relationship between activities and functional/behavioral health of older adults. We evaluate our proposed self-taught learning-based approach, STAR, using the presented in-house Alzheimer’s Activity Recognition (AAR) dataset acquired in a real-world deployment in 25 homes which outperforms the state-of-the-art algorithm by about 20%.
Sreenivasan Ramasamy Ramamurthy, Indrajeet Ghosh, Aryya Gangopadhyay, Elizabeth Galik, Nirmalya Roy
SMARTCOMP3
2020 Generating Realistic COVID-19 x-rays with a Mean Teacher + Transfer Learning GAN
abstract
COVID-19 is a novel infectious disease responsible for over 1.2 million deaths worldwide as of November 2020. The need for rapid testing is a high priority and alternative testing strategies including x-ray image classification are a promising area of research. However, at present, public datasets for COVID-19 x-ray images have low data volumes, making it challenging to develop accurate image classifiers. Several recent papers have made use of Generative Adversarial Networks (GANs) in order to increase the training data volumes. But realistic synthetic COVID-19 x-rays remain challenging to generate. We present a novel Mean Teacher + Transfer GAN (MTT-GAN) that generates COVID-19 chest x-ray images of high quality. In order to create a more accurate GAN, we employ transfer learning from the Kaggle pneumonia x-ray dataset, a highly relevant data source orders of magnitude larger than public COVID-19 datasets. Furthermore, we employ the Mean Teacher algorithm as a constraint to improve stability of training. Our qualitative analysis shows that the MTT-GAN generates x-ray images that are greatly superior to a baseline GAN and visually comparable to real x-rays. Although board-certified radiologists can distinguish MTT-GAN fakes from real COVID-19 x-rays, quantitative analysis shows that MTT-GAN greatly improves the accuracy of both a binary COVID-19 classifier as well as a multi-class pneumonia classifier as compared to a baseline GAN. Our classification accuracy is favorable as compared to recently reported results in the literature for similar binary and multi-class COVID-19 screening tasks.
Sumeet Menon, Joshua Galita, David Chapman 0001, Aryya Gangopadhyay, Jayalakshmi Mangalagiri, Yaacov Yesha, Yelena Yesha, Babak Saboury, Michael Morris
IEEE BigData4
2020 Image Segmentation for Dust Detection Using Semi-supervised Machine Learning
abstract
Dust plumes originating from the Earth's major arid and semi-arid areas can significantly affect the climate system and human health. Many existing methods have been developed to identify dust from non-dust pixels from a remote sensing point of view. However, these methods use empirical rules and therefore have difficulty detecting dust above or below the detectable thresholds. Supervised machine learning methods have also been applied to detect dust from satellite imagery, but these methods are limited especially when applying to areas outside the training data due to the inadequate amount of ground truth data. In this work, we proposed an automatic dust segmentation framework using semi-supervised machine learning, based on a collocated dataset using Visible Infrared Imaging Radiometer Suite (VIIRS) and Cloud-Aerosol Lidar and Infrared Pathfinder Satellite Observations (CALIPSO). The proposed method utilizes unsupervised machine learning for segmentation of VIIRS imagery, and leverages the guidance from the dust labels using the dust profile product of CALIPSO to determine the dust clusters as the final product. The dust clusters are determined based on the similarity of spectral signature from dust pixels along the CALIPSO tracks. Experiment results show that the accuracy of the proposed framework outperforms the traditional physical infrared method along CALIPSO tracks. In addition, the proposed method performs consistently over three different study areas, the North Atlantic Ocean, East Asia, and Northern Africa.
Manzhu Yu, Julie Bessac, Aryya Gangopadhyay, Yingxi Rona Shi, Jianwu Wang 0001
IEEE BigData4
2020 LASO: Exploiting Locomotive and Acoustic Signatures over the Edge to Annotate IMU Data for Human Activity Recognition
abstract
Annotated IMU sensor data from smart devices and wearables are essential for developing supervised models for fine-grained human activity recognition, albeit generating sufficient annotated data for diverse human activities under different environments is challenging. Existing approaches primarily use human-in-the-loop based techniques, including active learning; however, they are tedious, costly, and time-consuming. Leveraging the availability of acoustic data from embedded microphones over the data collection devices, in this paper, we propose LASO, a multimodal approach for automated data annotation from acoustic and locomotive information. LASO works over the edge device itself, ensuring that only the annotated IMU data is collected, discarding the acoustic data from the device itself, hence preserving the audio-privacy of the user. In the absence of any pre-existing labeling information, such an auto-annotation is challenging as the IMU data needs to be sessionized for different time-scaled activities in a completely unsupervised manner. We use a change-point detection technique while synchronizing the locomotive information from the IMU data with the acoustic data, and then use pre-trained audio-based activity recognition models for labeling the IMU data while handling the acoustic noises. LASO efficiently annotates IMU data, without any explicit human intervention, with a mean accuracy of $0.93$ ($\pm 0.04$) and $0.78$ ($\pm 0.05$) for two different real-life datasets from workshop and kitchen environments, respectively.
Soumyajit Chatterjee, Avijoy Chakma, Aryya Gangopadhyay, Nirmalya Roy, Bivas Mitra, Sandip Chakraborty 0001
ICMI3
2020 Vision Powered Conversational AI for Easy Human Dialogue Systems
abstract
In this paper, we propose an end to end goal-oriented conversational AI agent that can provide contextual information from a potential hazard site. We posit the conversational agent as a FloodBot capable of seeing, sensing, assessing hazard condition, and ultimately conversing about them. We present our domain-specific FloodBot design-solution and learning-experience from the real-time deployment in a flash flood devastated city that uses state-of-the-art deep learning models. We specifically used computer vision and pertinent natural language processing technologies to empower the conversation power of the FloodBot. To deliver such practical and usable AI, we chain multiple deep learning frameworks and create a human-friendly question-answer based dialogue system. We present our deployment details from the last five months and validate the results using ongoing COVID19's impact on the area as well.
Bipendra Basnyat, Neha Singh 0004, Nirmalya Roy, Aryya Gangopadhyay
MASS4
2020 AutoCogniSys: IoT Assisted Context-Aware Automatic Cognitive Health Assessment
abstract
Cognitive impairment has become epidemic in older adult population. The recent advent of tiny wearable and ambient devices, a.k.a Internet of Things (IoT) provides ample platforms for continuous functional and cognitive health assessment of older adults. In this paper, we design, implement and evaluate AutoCogniSys, a context-aware automated cognitive health assessment system, combining the sensing powers of wearable physiological (Electrodermal Activity, Photoplethysmography) and physical (Accelerometer, Object) sensors in conjunction with ambient sensors. We design appropriate signal processing and machine learning techniques, and develop an automatic cognitive health assessment system in a natural older adults living environment. We validate our approaches using two datasets: (i) a naturalistic sensor data streams related to Activities of Daily Living and mental arousal of 22 older adults recruited in a retirement community center, individually living in their own apartments using a customized inexpensive IoT system (IRB #HP-00064387) and (ii) a publicly available dataset for emotion detection. The performance of AutoCogniSys attests max. 93% of accuracy in assessing cognitive health of older adults.
Mohammad Arif Ul Alam, Nirmalya Roy, Sarah Holmes, Aryya Gangopadhyay, Elizabeth Galik
MobiQuitous4
2020 Flood Detection Framework Fusing The Physical Sensing & Social Sensing
abstract
We investigate the practical challenge of localized flood detection in real smart city environment using the fusion of physical sensor and social sensing models to depict a reliable and accurate flood monitoring and detection framework. Our proposed framework efficiently utilize the physical and social sensing models to provide the flood-related updates to the city officials. We deployed our flood monitoring system in Ellicott City, Maryland, USA and connect it to the social sensing module to perform the flood-related sensor and social data integration and analysis. Our ground-based sensor network model record and performs the predictive data analytic by forecasting the rise in water level (RMSE=0.2) that demonstrates the severity of upcoming flash floods whereas, our social sensing model helps collect and track the flood-related feeds from Twitter. We employ a pre-trained model and inductive transfer learning based approach to classify the flood-related tweets with 90% accuracy in the use of unseen target flood events. Finally our flood detection framework categorizes the flood relevant localized contextual details into more meaningful classes in order to help the emergency services and local authorities for effective decision making.
Neha Singh 0004, Bipendra Basnyat, Nirmalya Roy, Aryya Gangopadhyay
SMARTCOMP4
2020 Design and Deployment of a Flash Flood Monitoring IoT: Challenges and Opportunities
abstract
Successful implementation of the Internet of Thing (IoT) is precursory to a thriving smart city. However, the technical, physical, and environmental conditions can often pose challenges in their successful deployments. The deployment is further complicated if the time and location of implementation are amidst a natural disaster. In this work, we use flash flood detection as a natural hazard testbed and describe various IoT deployment, our progression, and first-hand experience from those implementations. We compare and contrast three IoTs and their performance in real-time execution. Next, we discuss systems architecture and their end-to-end design and present lessons learned from these heterogeneous deployments. Additionally, we evaluate and outline our observations, challenges, and opportunities for further improvement. We also formulate standard evaluation metrics for their scoring and document our deployment journey.
Bipendra Basnyat, Neha Singh 0004, Nirmalya Roy, Aryya Gangopadhyay
SMARTCOMP4
2020 Towards AI Conversing: FloodBot using Deep Learning Model Stacks
abstract
Talking to the electronic device and getting the required information at a minimal time has become today's norm. Although AI-powered conversational agents have percolated the commercial market, their use in a communal setting is still evolving. We postulate that the deployments of chatbots in disaster-prone areas can be beneficial to watch, monitor, and warn people during the crisis. Furthermore, the successful implementation of such technology can be life-saving. In this work, we discuss our deployment of a real-time flood monitoring chatbot called FloodBot. We collect, annotate and visually parse images from potentially hazardous areas. We detect the flood conditions and identify objects in harm's way by stacking deep learning models such as a convolutional neural network (CNN), single-shot multi-box object detection (SSD). We then feed the image contents to a knowledge base of our artificially intelligent FloodBot and explore its AI-Conversing power using end to end memory network. We also showcase the power of cross-domain transfer learning and model fusion techniques. In this work, we discuss our deployment of a real-time flood monitoring chatbot called FloodBot. We collect, annotate and visually parse images from potentially hazardous areas. We detect the flood conditions and identify objects in harm's way by stacking deep learning models such as a convolutional neural network (CNN), single-shot multi-box object detection (SSD). We then feed the image contents to a knowledge base of our artificially intelligent FloodBot and explore its AI-Conversing power using end to end memory network. We also showcase the power of cross-domain transfer learning and model fusion techniques.
Bipendra Basnyat, Nirmalya Roy, Aryya Gangopadhyay
SMARTCOMP3
2020 A Deep Learning Model for Detecting Dust in Earth's Atmosphere from Satellite Remote Sensing Data
abstract
In this paper we develop a deep learning model to distinguish dust from cloud and surface using satellite remote sensing image data. The occurrence of dust storms is increasing along with global climate change, especially in the arid and semi-arid regions. Originated from the soil, dust acts as a type of aerosol that causes significant impacts on the environment and human health. The dust and cloud data labels used in this paper are from CALIPSO (Cloud-Aerosol Lidar and Infrared Pathfinder Satellite Observation) satellite. The radiometric channels and geometric parameters from VIIRS (Visible Infrared Imaging Radiometer Suite) satellite sensor serve as features for our model. We trained and tested our deep learning model using 10,000 samples in March 2012. The developed model has five hidden layers and 512 neurons in each layer. The classification accuracy on the test set is 71.1%. In addition, we performed a shuffling procedure to identify the importance of features, which is calculated as the increase in the prediction error after we permute the feature's values. We also developed a method based on genetic algorithm to find the best subset of features for dust detection. The results show that the genetic algorithm can select a subset of features that have comparable performance as that of a model with all features. The shuffling procedure and the genetic algorithm both identify geometric information as important features for detecting mineral dust. The chosen subset will improve computational efficiency for dust detection and improve physical based methods.
Ping Hou, Pei Guo, Jianwu Wang 0001, Aryya Gangopadhyay
SMARTCOMP5
2019 A hybrid model using LSTM and decision tree for mortality prediction and its application in provider performance evaluation
abstract
The risk adjusted mortality rate, which is also called standardized mortality ratio (SMR), is one widely used quality measure to evaluate healthcare provider performance. Logistic regression and decision tree are two traditional risk models for mortality rate calculation. Though some machine learning based approaches could achieve higher accuracy, they are hard to interpret and may have poor calibration scores. In this paper, we evaluated multiple machine learning approaches with different formats of longitudinal data, and proposed a hybrid approach based on long short-term memory (LSTM) model and decision tree. The new hybrid method provides a comparable area under the receiver operating characteristic curve (AUC) performance as LSTM with a better calibration score. Using a set of 3,473 patients with 10 months of data from historical, large scale ESRD patient data, the LSTM with long format data approach achieved AUC for prediction of mortality of 0.772 compared to 0.758 for logistic regression and 0.726 for a decision tree model. The hybrid approach could reach 0.783, a little higher than both LSTM and decision tree model. The hybrid approach has the best calibration performance based on the Hosmer Lemeshow test.
Peichang Shi, Aryya Gangopadhyay, Carolyn Owens, Brenda Blunt, Christine Grogan
IEEE BigData2
2019 A Hybrid Algorithm for Mineral Dust Detection Using Satellite Data
abstract
Mineral dust, defined as aerosol originating from the soil, can have various harmful effects to the environment and human health. The detection of dust, and particularly incoming dust storms, may help prevent some of these negative impacts. In this paper, using satellite observations from Moderate Resolution Imaging Spectroradiometer (MODIS) and the Cloud-Aerosol Lidar and Infrared Pathfinder Satellite Observation Observation (CALIPSO), we compared several machine learning algorithms to traditional physical models and evaluated their performance regarding mineral dust detection. Based on the comparison results, we proposed a hybrid algorithm to integrate physical model with the data mining model, which achieved the best accuracy result among all the methods. Further, we identified the ranking of different channels of MODIS data based on the importance of the band wavelengths in dust detection. Our model also showed the quantitative relationships between the dust and the different band wavelengths.
Peichang Shi, Janita Patwardhan, Jianwu Wang 0001, Aryya Gangopadhyay
eScience6
2019 Identifying the Context of Hurricane Posts on Twitter using Wavelet Features
abstract
With the increase of natural disasters all over the world, we are in crucial need of innovative solutions with inexpensive implementations to assist the emergency response systems. Information collected through conventional sources (e.g., incident reports, 911 calls, physical volunteers, etc.) are proving to be insufficient [1]. Responsible organizations are now leaning towards research grounds that explore digital human connectivity and freely available sources of information. U.S. Geological Survey and Federal Emergency Management Agency (FEMA) introduced Critical Lifeline (CLL) s which identifies the most significant areas that require immediate attention in case of natural disasters. These organizations applied crowdsourcing by connecting digital volunteer networks to collect data on the critical lifelines from data sources including social media [3], [4], [5]. In the past couple of years, during some of the deadly hurricanes (e.g., Harvey, IRMA, Maria, Michael, Florence, etc.), people took on different social media platforms like never seen before, in search of help for rescue, shelter, and relief. Their posts reflect crisis updates and their real-time observations on the devastation that they witness. In this paper, we propose a methodology to build and analyze time-frequency features of words on social media to assist the volunteer networks in identifying the context before, during and after a natural disaster and distinguishing contexts connected to the critical lifelines. We employ Continuous Wavelet Transform to help create word features and propose two ways to reduce the dimensions which we use to create word clusters to identify themes of conversations associated with stages of a disaster and these lifelines. We compare two different methodologies of wavelet features and word clusters both qualitatively and quantitatively, to show that wavelet features can identify and separate context without using semantic information as inputs.
Amrita Anam, Aryya Gangopadhyay, Nirmalya Roy
SMARTCOMP2
2019 Analyzing Rhabdomyosarcoma using the Multimodal Clustering Approach (DReiM)
abstract
Rhabdomyosarcoma is a type of cancer that is connected to soft tissue, connective tissue, or bone. Every year 350 children are diagnosed with Rhabdomyosarcoma. Majority of the kids diagnosed with this disease are under ten years of age. Though the intensive conventional approach to treatment exists, patients are still at high risks to this aggressive disease. The need to gain understanding and insight into this disease can help the design of therapeutic agents. We utilized a multimodal network approach to gain an understanding of this mechanism. Protein phosphorylation has been mostly studied as a post-translational modification in eukaryotes. They play a significant role in various cellular processes. The mechanism of its oncogenes is not well known with various levels of signaling protein dysregulation. In the Phosphosite database, some proteins phosphorylate to cause Rhabdomyosarcoma. Our method utilizes a co-clustering approach using multimodal networks to analyze the rhabdomyosarcoma network. Rhabdomyosarcoma network in general consists of several heterogeneous networks that include gene-pathway, pathway-drug, and gene-drug. We reconstruct the phosphorylation network for Rhabdomyosarcoma, by creating a network that consists of different types of nodes. The goal is to implement this clustering approach to identify a potential candidate for Rhabdomyosarcoma. We applied network centrality measures to find the most influential nodes first and foremost and then used the clustering approach stated above towards drug repositioning.
Iyanuoluwa Emmanuel Odebode, Aryya Gangopadhyay
SMARTCOMP2
2019 Analyzing The Emotions of Crowd For Improving The Emergency Response Services
abstract
Twitter is an extremely popular micro-blogging social platform with millions of users, generating thousands of tweets per second. The huge amount of Twitter data inspire the researchers to explore the trending topics, event detection and event tracking which help to postulate the fine-grained details and situation awareness. Obtaining situational awareness of any event is crucial in various application domains such as natural calamities, man-made disaster and emergency responses. In this paper, we advocate that data analytics on Twitter feeds can help improve the planning and rescue operations and services as provided by the emergency personnel in the event of unusual circumstances. We take an emotional change detection approach and focus on the users’ emotions, concerns and feelings expressed in tweets during the emergency situations, and analyze those feelings and perceptions in the community involved during the events to provide appropriate feedback to emergency responders and local authorities. We employ improved emotion analysis and change point detection techniques to process, discover and infer the spatiotemporal sentiments of the users. We analyze the tweets from recent Las Vegas shooting (Oct. 2017) and note that the changes in the polarity of the sentiments and articulation of the emotional expressions, if captured successfully can be employed as an informative tool for providing feedback to EMS.
Neha Singh 0004, Nirmalya Roy, Aryya Gangopadhyay
Pervasive Mob. Comput.3
2018 Enhancing Interest in Cybersecurity Careers: A Peer Mentoring Perspective
abstract
The focus of this paper is an evaluation of our peer mentoring framework designed to encourage more students to seek cybersecurity career pathways through providing peer interactions. We present and compare results from two years (Spring 2016 and 2017) of interaction between students in an introductory Information Systems class (IS 300: Management of Information Systems) and an upper-level elective Cybersecurity course (IS 471: Data Analytics for Cybersecurity). Our results show a continuation of the general trend observed in the 2016 study. The students who receive peer mentoring show more interest in cybersecurity issues and careers and gain more overall knowledge throughout the semester, than those who don't. This is reflected by the results of an anonymous survey and overall grade improvements. These students show more variations regarding their choice of cybersecurity as a career compared to students who did not receive any mentoring, demonstrating that they are able to make more informed decisions. Female students exhibit more pronounced responses to peer mentoring in contrast to their male counterparts.
Vandana Pursnani Janeja, Abu Zaher Md Faridee, Aryya Gangopadhyay, Carolyn B. Seaman, Amy Everhart
SIGCSE3
2018 Evaluating Disaster Time-Line from Social Media with Wavelet Analysis
abstract
For over a decade, social media has proved to be a functional and convenient data source in the Internet of things. Social platforms such as Facebook, Twitter, Instagram, and Reddit have their own styles and purposes. Twitter, among them, has become the most popular platform in the research community due to its nature of attracting people to write brief posts about current and unexpected events (e.g., natural disasters). The immense popularity of such sites has opened a new horizon in `social sensing' to manage disaster response. Sensing through social media platforms can be used to track and analyze natural disasters and evaluate the overall response (e.g., resource allocation, relief, cost and damage estimation). In this paper, we propose a two-step methodology: i) wavelet analysis and ii) predictive modeling to track the progression of a disaster aftermath and predict the time-line. We demonstrate that wavelet features can preserve text semantics and predict the total duration for localized small scale disasters. The experimental results and observations on two real data traces (flash flood in Cummins Falls state park and Arizona swimming hole) showcase that the wavelet features can predict disaster time-line with an error lower than 20% with less than 50% of the data when compared to ground truth.
Amrita Anam, Aryya Gangopadhyay, Nirmalya Roy
SMARTCOMP2
2018 A Flash Flood Categorization System Using Scene-Text Recognition
abstract
Detecting flash floods in real-time and taking rapid actions are of utmost importance to save human lives, loss of infrastructures, and personal properties in a smart city. In this paper, we develop a low-cost low-power cyber-physical System prototype using a Raspberry Pi camera to detect the rising water level. We deployed the system in the real word and collected data in different environmental conditions (early morning in the presence of fog, sunny afternoon, late afternoon with sunsetting). We employ image processing and text recognition techniques to detect the rising water level and articulate several challenges in deploying such a system in the real environment. We envision this prototype design will pave the way for mass deployment of the flash flood detection system with minimal human intervention.
Bipendra Basnyat, Nirmalya Roy, Aryya Gangopadhyay
SMARTCOMP3
2018 Analyzing the Sentiment of Crowd for Improving the Emergency Response Services
abstract
Twitter is an extremely popular micro-blogging social platform with millions of users, generating thousands of tweets per second. The huge amount of Twitter data inspire the researchers to explore the trending topics, event detection and event tracking which help to postulate the fine-grained details and situation awareness. Obtaining situational awareness of any event is crucial in various application domains such as natural calamities, man made disaster and emergency responses. In this paper, we advocate that data analytics on Twitter feeds can help improve the planning and rescue operations and services as provided by the emergency personnel in the event of unusual circumstances. We take a different approach and focus on the users' emotions, concerns and feelings expressed in tweets during the emergency situations, and analyze those feelings and perceptions in the community involved during the events to provide appropriate feedback to emergency responders and local authorities. We employ sentiment analysis and change point detection techniques to process, discover and infer the spatiotemporal sentiments of the users. We analyze the tweets from recent Las Vegas shooting (Oct. 2017) and note that the changes in the polarity of the sentiments and articulation of the emotional expressions, if captured successfully can be employed as an informative tool for providing feedback to EMS.
Neha Singh 0004, Nirmalya Roy, Aryya Gangopadhyay
SMARTCOMP3
2017 Analyzing Social Media Texts and Images to Assess the Impact of Flash Floods in Cities
abstract
Computer Vision and Image Processing are emerging research paradigms. The increasing popularity of social media, micro- blogging services and ubiquitous availability of high-resolution smartphone cameras with pervasive connectivity are propelling our digital footprints and cyber activities. Such online human footprints related with an event-of-interest, if mined appropriately, can provide meaningful information to analyze the current course and pre- and post- impact leading to the organizational planning of various real-time smart city applications. In this paper, we investigate the narrative (texts) and visual (images) components of Twitter feeds to improve the results of queries by exploiting the deep contexts of each data modality. We employ Latent Semantic Analysis (LSA)-based techniques to analyze the texts and Discrete Cosine Transformation (DCT) to analyze the images which help establish the cross-correlations between the textual and image dimensions of a query. While each of the data dimensions helps improve the results of a specific query on its own, the contributions from the dual modalities can potentially provide insights that are greater than what can be obtained from the individual modalities. We validate our proposed approach using real Twitter feeds from a recent devastating flash flood in Ellicott City near the University of Maryland campus. Our results show that the images and texts can be classified with 67\% and 94\% accuracies respectively
Bipendra Basnyat, Amrita Anam, Neha Singh 0004, Aryya Gangopadhyay, Nirmalya Roy
SMARTCOMP4
2017 A smart segmentation technique towards improved infrequent non-speech gestural activity recognition model
Mohammad Arif Ul Alam, Nirmalya Roy, Aryya Gangopadhyay, Elizabeth Galik
Pervasive Mob. Comput.3
2017 Target-Based, Privacy Preserving, and Incremental Association Rule Mining
abstract
We consider a special case in association rule mining where mining is conducted by a third party over data located at a central location that is updated from several source locations. The data at the central location is at rest while that flowing in through source locations is in motion. We impose some limitations on the source locations, so that the central target location tracks and privatizes changes and a third party mines the data incrementally. Our results show high efficiency, privacy and accuracy of rules for small to moderate updates in large volumes of data. We believe that the framework we develop is therefore applicable and valuable for securely mining big data.
Madhu Ahluwalia, Aryya Gangopadhyay, Zhiyuan Chen 0003, Yelena Yesha
IEEE Trans. Serv. Comput.2
2016 Cybersecurity workforce development: A peer mentoring approach
abstract
In this paper we present a peer mentoring framework for students in an introductory Information Systems class (IS 300: Management of Information Systems) to interact with peers in an upper level elective course in cybersecurity (IS 471: Data Analytics for Cybersecurity). The purpose of this framework is to encourage more students to explore cybersecurity careers through peer led cybersecurity discussions. We discuss preliminary results from a peer interaction for the spring 2016 semester. Our initial results show an increased interest in cybersecurity for the IS 300 students. Moreover, we see that post the interactions both groups have a clearer interest and understanding of cybersecurity.
Vandana Pursnani Janeja, Carolyn B. Seaman, Kerrie Kephart, Aryya Gangopadhyay, Amy Everhart
ISI4
2016 Health Care Fraud Detection with Community Detection Algorithms
abstract
Fraud detection is interesting research topic and it not only needs data mining techniques but also needs a lot of inputs from domain experts. In health care claims, relationships between physicians and patients form complex communities structures and these communities could lead to potential fraud discoveries. Traditionally, researchers have focused on clustering physicians and patients and tried to find the suspicious communities. In this paper, we studied and discussed different types of relationships and focus on small but exclusive relationships that are suspicious and may indicate potential health care frauds. We developed two algorithms to detect these small and exclusive communities. These algorithms can be applied to larger dataset and are highly scalable. We tested these algorithms with a set of synthesized datasets. These synthesized datasets were created to resemble the real health care claims datasets and used to test the fraud detection algorithms. The test results show the these algorithms are very efficient and can evaluate the communities structures of 50,000 providers in about 1 minute.
Aryya Gangopadhyay
SMARTCOMP1
2016 Using Crowd Sourcing to Analyze Consumers' Response to Privacy Policies of Online Social Network and Financial Institutions at Micro Level
abstract
As it becomes easy and inexpensive to store huge amount of data, concerns about privacy are increasing as well. Although service providers have privacy policies, research shows that users rarely read privacy policies. As a result, there has been little work done on how consumers respond to individual segments of privacy policies, which is important for organizations when designing privacy policies. In this study, the authors break down privacy policies of two well-known social network companies (Facebook, Twitter) and financial institution (Bank of America) into simple segments. They then use crowd sourcing to analyze consumers' response to these policy segments. The authors ask questions on users' awareness, expectations, familiarity, and privacy concerns of these policy segments. The relationships between various factors such as demographic factors, data type, data flow and consumers' privacy concerns were also investigated. The authors conclude with guidelines and suggestions for improvement and ways to increase users' awareness of privacy policies.
Shaikha Alduaij, Zhiyuan Chen 0003, Aryya Gangopadhyay
Int. J. Inf. Secur. Priv.3
2015 A Graph-Based Method for Analyzing Electronic Medical Records
abstract
The recent years have seen a surge in the implementation of electronic health care records. These patient records contain valuable medical information including patient information, diagnosis, treatment methods, and eventual patient outcomes. It is important to analyze patterns within these records in order to more efficiently treat individuals. In this paper, we present a method for automatically discovering underlying themes and patterns within patient data. Our methodology includes the creation of the main themes or patterns in the data and linking the themes back to the corpus from which they were generated. In our research, we partitioned graphs from terms gathered from electronic health records. Modularity was used as the quality function, which strives to measure how well a given partition of a network compartmentalizes its communities. We have compared our method with probabilistic topic modeling algorithms, specifically LDA (Latent Dirichlet Allocation). Finally, recall and precision measures were used to evaluate the validity of our final results.
Rose Yesha, Aryya Gangopadhyay, Eliot L. Siegel
ASONAM2
2015 Acquisition of diabetes-related biological associations using a motif based network: Preliminary results
abstract
Diabetes remains the 7th leading cause of death in the US in 2010. There is an urgent need to find effective solutions to prevent and treat diabetes. With the advance in computational technology, computer-aided approaches has attracted much interest given the promising findings been generated accordingly. In this study, we introduced an approach that applied network analysis to reveal novel biological associations for diabetes from a large biological interaction network. In particular, network motif discovery has been performed to identify diabetes relevant motif-based network, from where the diabetes relevant biological associations can be detected via perturbation. The associations identified from this study will not only illustrate possible biological mechanism for diabetes, but also support biomedical application development ultimately, such as drug repositioning.
Iyanuoluwa Emmanuel Odebode, Aryya Gangopadhyay, Qian Zhu 0003
BIBM2
2015 STenSr: Spatio-temporal tensor streams for anomaly detection and pattern discovery
Aryya Gangopadhyay, Vandana Pursnani Janeja
Knowl. Inf. Syst.2
2014 Mining trajectories of moving dynamic spatio-temporal regions in sensor datasets
Michael P. McGuire, Vandana Pursnani Janeja, Aryya Gangopadhyay
Data Min. Knowl. Discov.3
2014 A generic and distributed privacy preserving classification method with a worst-case privacy guarantee
Madhushri Banerjee, Zhiyuan Chen 0003, Aryya Gangopadhyay
Distributed Parallel Databases3
2014 Discovery of temporal neighborhoods through discretization methods
abstract
Neighborhood discovery is a precursor to knowledge discovery in complex and large datasets such as Temporal data, which is a sequence of data tuples measured at successive time instances. Hence instead of mining the entire data, we are interested in
Sandipan Dey, Vandana Pursnani Janeja, Aryya Gangopadhyay
Intell. Data Anal.3
2014 A data recipient centered de-identification method to retain statistical attributes
Tamas S. Gal, Thomas C. Tucker, Aryya Gangopadhyay, Zhiyuan Chen 0003
J. Biomed. Informatics3
2012 Optimizing Privacy-Accuracy Tradeoff for Privacy Preserving Distance-Based Classification
abstract
Privacy concerns often prevent organizations from sharing data for data mining purposes. There has been a rich literature on privacy preserving data mining techniques that can protect privacy and still allow accurate mining. Many such techniques have some parameters that need to be set correctly to achieve the desired balance between privacy protection and quality of mining results. However, there has been little research on how to tune these parameters effectively. This paper studies the problem of tuning the group size parameter for a popular privacy preserving distance-based mining technique: the condensation method. The contributions include: 1) a class-wise condensation method that selects an appropriate group size based on heuristics and avoids generating groups with mixed classes, 2) a rule-based approach that uses binary search and several rules to further optimize the setting for the group size parameter. The experimental results demonstrate the effectiveness of the authors’ approach.
Zhiyuan Chen 0003, Aryya Gangopadhyay
Int. J. Inf. Secur. Priv.3
2011 Characterizing sensor datasets with multi-granular spatio-temporal intervals
abstract
Data from sensors and sensor networks are being collected at astronomical rates. This results in a massive dataset that is increasingly difficult to navigate to find interesting time periods where the spatial pattern of a process changes. The ability to navigate to such areas can lead to new knowledge about the factors that contribute to a spatio-temporal process. This paper proposes a method to automatically characterize sensor datasets based on a measure of spatial change over time resulting in a set of multi-granular spatio-temporal intervals. The resulting intervals can be used to focus knowledge discovery tasks at multiple temporal granularities within the dataset. Furthermore, the intervals enable a drill-down-style analysis where events of varying magnitudes can be identified within each granularity. Experiments were performed on a real-world dataset measuring NEXRAD precipitation accumulation. The results show that the multi-granular spatio-temporal intervals identify interesting time periods in the dataset as evidenced by naturally occurring events.
Michael P. McGuire, Vandana Pursnani Janeja, Aryya Gangopadhyay
GIS3
2009 Temporal Neighborhood Discovery Using Markov Models
abstract
Temporal data, which is a sequence of data tuples measured at successive time instances, is typically very large. Hence instead of mining the entire data, we are interested in dividing the huge data into several smaller intervals of interest which we call temporal neighborhoods. In this paper we propose an approach to generate temporal neighborhoods through unequal depth discretization. We describe two novel algorithms (a) similarity based merging (SMerg) and, (b) stationary distribution based merging (StMerg). These algorithms are based on the robust framework of Markov models and the Markov stationary distribution respectively. We identify temporal neighborhoods with distinct demarcations based on unequal depth discretization of the data. We discuss detailed experimental results in both synthetic and real world data. Specifically we show (i) the efficacy of our approach through precision and recall of labeled bins, (ii) the ground truth validation in real world datasets and, (iii) knowledge discovery in the temporal neighborhoods such as global anomalies. Our results indicate that we are able to identify valuable knowledge based on our ground truth validation from real world traffic data.
Sandipan Dey, Vandana Pursnani Janeja, Aryya Gangopadhyay
ICDM3
2009 Discretized Spatio-Temporal Scan Window
abstract
The focus of this paper is the discovery of anomalous spatio-temporal windows. We propose a Discretized Spatio-Temporal Scan Window approach to address the question of how we can treat Space and Time together without compromising on the properties of each and their impact on each other. In doing so we discover anomalous Spatio-Temporal windows, identify at what point in time the window changes, identify the spatial patterns of change over time and identify a spatial extent in time which is completely deviant with respect to the rest of the anomalous spatio-temporal windows. None of the current approaches address all these issues in combination. Subsequently we perform experiments on several real world datasets to validate our approach while comparing with the established approach of discovering a cylindrical spatio-temporal Scan window.
Seyed H. Mohammadi, Vandana Pursnani Janeja, Aryya Gangopadhyay
SDM3
2009 Preserving Privacy in Mining Quantitative Associations Rules
abstract
Association rule mining is an important data mining method that has been studied extensively by the academic community and has been applied in practice. In the context of association rule mining, the state-of-the-art in privacy preserving data mining provides solutions for categorical and Boolean association rules but not for quantitative association rules. This article fills this gap by describing a method based on discrete wavelet transform (DWT) to protect input data privacy while preserving data mining patterns for association rules. A comparison with an existing kd-tree based transform shows that the DWT-based method fares better in terms of efficiency, preserving patterns, and privacy.
Madhu Ahluwalia, Aryya Gangopadhyay, Zhiyuan Chen 0003
Int. J. Inf. Secur. Priv.2
2008 A privacy preserving technique for distance-based classification with worst case privacy guarantees
Shibnath Mukherjee, Madhushri Banerjee, Zhiyuan Chen 0003, Aryya Gangopadhyay
Data Knowl. Eng.4
2008 Measuring interestingness of discovered skewed patterns in data cubes
Aryya Gangopadhyay, Sanjay Bapna, George Karabatis, Zhiyuan Chen 0003
Decis. Support Syst.2
2008 Assisting decision making in the event-driven enterprise using wavelets
Stephen Russell 0002, Aryya Gangopadhyay, Victoria Y. Yoon
Decis. Support Syst.2
2008 A fuzzy programming approach for data reduction and privacy in distance-based mining
abstract
With the explosive growth of data and its distributed sources, there are increasing needs for secure cooperative data analysis. The issue of data reduction to decrease communication overheads and the issue of preservation of privacy of the shared data are becoming important. However, existing privacy preserving techniques do not work well for distance-based mining because they do not preserve distances. Besides, most of them either do not reduce data or are tied to very specific mining algorithms. Using the unitarity and energy compaction property of Fourier transforms, this paper proposes a novel framework to preserve privacy and reduce data size, yet preserve Euclidian distances. A fuzzy programming approach for selection of Fourier coefficients is proposed to optimise the objective of preserving Euclidean distances and obtaining privacy and data reduction through coefficient suppression. Experiments demonstrate the superiority of the proposed approach over the existing ones.
Shibnath Mukherjee, Zhiyuan Chen 0003, Aryya Gangopadhyay
Int. J. Inf. Comput. Secur.3
2008 A Privacy Protection Model for Patient Data with Multiple Sensitive Attributes
abstract
The identity of patients must be protected when patient data are shared. The two most commonly used models to protect identity of patients are L-diversity and K-anonymity. However, existing work mainly considers data sets with a single sensitive attribute, while patient data often contain multiple sensitive attributes (e.g., diagnosis and treatment). This article shows that although the K-anonymity model can be trivially extended to multiple sensitive attributes, the L-diversity model cannot. The reason is that achieving L-diversity for each individual sensitive attribute does not guarantee L-diversity over all sensitive attributes. We propose a new model that extends L-diversity and K-anonymity to multiple sensitive attributes and propose a practical method to implement this model. Experimental results demonstrate the effectiveness of our approach.
Tamas S. Gal, Zhiyuan Chen 0003, Aryya Gangopadhyay
Int. J. Inf. Secur. Priv.3
2007 Supporting mobile decision making with association rules and multi-layered caching
Aryya Gangopadhyay, George Karabatis
Decis. Support Syst.2
2007 Managing uncertainty in location services using rough set and evidence theory
Iftikhar U. Sikder, Aryya Gangopadhyay
Expert Syst. Appl.2
2007 Semantic Integration and Knowledge Discovery for Environmental Research
abstract
Environmental research and knowledge discovery both require extensive use of data stored in various sources and created in different ways for diverse purposes. We describe a new metadata approach to elicit semantic information from environmental data and implement semantic-based techniques to assist users in integrating, navigating, and mining multiple environmental data sources. Our system contains specifications of various environmental data sources and the relationships that are formed among them. User requests are augmented with semantically related data sources and automatically presented as a visual semantic network. In addition, we present a methodology for data navigation and pattern discovery using multi-resolution browsing and data mining. The data semantics are captured and utilized in terms of their patterns and trends at multiple levels of resolution. We present the efficacy of our methodology through experimental results.
Zhiyuan Chen 0003, Aryya Gangopadhyay, George Karabatis, Michael P. McGuire, Claire Welty
J. Database Manag.2
2006 A privacy-preserving technique for Euclidean distance-based mining algorithms using Fourier-related transforms
Shibnath Mukherjee, Zhiyuan Chen 0003, Aryya Gangopadhyay
VLDB J.3
2001 Conceptual modeling from natural language functional specifications
Aryya Gangopadhyay
Artif. Intell. Eng.1
2001 An image-based system for electronic retailing
Aryya Gangopadhyay
Decis. Support Syst.1
2000 On-Line User Interaction with Electronic Catalogs: Language Preferences Among Global Users
abstract
In this paper we study the behavior and performance of bilingual users in using an electronic catalog. The purpose of this research is to further the knowledge required for building electronic commerce systems that operate in multiple languages in global settings. We describe a bilingual electronic catalog that can be used by online retailers for selling products and/or services to customers interacting in either English or Chinese. We investigate into the nature of user interactions in multilingual electronic catalogs. We have defined three different groups of users: only Chinese speaking, only English speaking, and bilingual. We are specifically interested in investigating into the language preferences of the third group of users. In order to test language preferences, we have selected two types of products: office supplies and ethnic food. We hypothesize that bilingual users will exhibit differential language preferences for the type of products and the tasks performed in using the electronic catalog. Furthermore, learning curves and interaction effects are also tested. Three different task categories have been designed: browsing, directed search, and exact matches. In the first case, the user is a general browser who is looking for what is available in the catalog. In the second case, the user is looking for a class of products but is unsure of the exact item. In the third case the user knows exactly what item he/she is looking for. We propose to test the efficiency of usage by measuring the time as well as studying the path followed by the user in retrieving product information. This research will shed light on the important issue of designing multilingual electronic catalogs for both local and global applications.
Aryya Gangopadhyay, Zhensen Huang
J. Glob. Inf. Manag.1
1997 A Form-Based Natural Language Front-End to a CIM Database
abstract
The paper presents a methodology for developing a user interface that combines fourth generation interface tools (SQL forms) with a natural language processor for a database management system. The natural language processor consists of an index, a lexicon and a parser. The index is used to uniquely identify each form in the system through a conceptual representation of its purpose. The form fields specify database or nondatabase fields whose values are either entered by the user (user-defined) or are derived by the form (system-defined) in response to user input. A set of grammar rules are associated with each form. The lexicon consists of all words recognized by the system, their grammatical categories, roots, their associations (if any) with database objects and forms. The parser scans, a natural language query to identify a form in a bottom-up fashion. The information requested in the user query is determined in a top-down manner by parsing, through the grammar rules associated with the identified form. Extragrammatical inputs with limited deviations from the grammar rules are supported. Combining a natural language processor with SQL forms allows processing data modification tasks without violating any database integrity constraint, having duplicate records, or entering invalid data. A prototype natural language interface is described as a front-end to an ORACLE database for a computer integrated manufacturing system.
Nabil R. Adam, Aryya Gangopadhyay
IEEE Trans. Knowl. Data Eng.2
1993 Integrating Functional and Data Modeling in a Computer Integrated Manufacturing System
abstract
A structured methodology for linking data modeling with functional modeling in a computer integrated manufacturing system is presented. The target application, the functional and data models, and a method for developing the data model starting from the functional model are described. This approach ensures that the data model is complete and non-redundant with respect to the functional model. A scheme that enables various functions in the functional model to be linked with the data elements of the data model is also presented. Such a linkage makes it possible to determine the impact of a change in the functional model on the data model and vice versa.>
Nabil R. Adam, Aryya Gangopadhyay
ICDE2
1993 Design and Implementation of a Knowledge-Based Query Processor
abstract
This paper deals with query processing using semantic knowledge in relational databases. The Select-Project-Join (SPJ) conjunctive class of queries are dealt with in this paper. We propose to optimize highly repetitive queries by using semantic transformations in addition to syntactic transformations. Thus, we generate a set of pre-optimized queries. This set contains queries that are semantically equivalent to, syntactically different from, and more efficient to process than the user queries that we started with. The issues we address in this paper are: how to map a user query to a query that is in the set of pre-optimized and already optimized queries, how to search efficiently through the set of pre-optimized queries and set of semantic rules, and how to incorporate new queries to the set of pre-optimized queries, so that the number of queries that can be optimized using this method increases with the passage of time. Furthermore, we suggest some ideas of handling queries that do not have any semantically equivalent counterpart in the set of pre-optimized queries. We have tested the performance of the proposed method. An algorithm for mapping is implemented in Prolog. A database schema is implemented in the INGRES database management system. We have adopted a database schema that is widely used for measuring performance in the semantic query optimization literature.
Nabil R. Adam, Aryya Gangopadhyay, James Geller
Int. J. Cooperative Inf. Syst.2