VLDB 2026 Research / reviewers in the wild / expert
Yelena Yesha
dblp:y/YelenaYesha
· DBLP profile ↗
83ranked-venue papers
0as first author
5since 2021 · last 2024
0000-0001-8746-1157ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 36 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 24 · 5 since 2021Artificial intelligence and machine learning · 19 · 2 since 2021Computer networks · 14Software engineering, systems software and programming languages · 7Human-computer interaction and ubiquitous computing · 6Systems, architecture and hardware · 5Theory of computation · 4Security and privacy · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Improving VTE Identification through Language Models from Radiology Reports: A Comparative Study of Mamba, Phi-3 Mini, and BERTabstractVenous thromboembolism (VTE) is a critical cardio-vascular condition, encompassing deep vein thrombosis (DVT) and pulmonary embolism (PE). Accurate and timely identification of VTE is essential for effective medical care. This study builds upon our previous work, which addressed VTE detection using deep learning methods for DVT and a hybrid approach combining deep learning and rule-based classification for PE. Our earlier approaches, while effective, had two major limitations: they were complex and required expert involvement for feature engineering of the rule set. To overcome these challenges, we utilize the Mamba architecture-based classifier. This model achieves remarkable results, with a 97% accuracy and F1 score on the DVT dataset and a 98% accuracy and F1 score on the PE dataset. In contrast to the previous hybrid method on PE identification, the Mamba classifier eliminates the need for hand-engineered rules, significantly reducing model complexity while maintaining comparable performance. Additionally, we evaluated a lightweight Large Language Model (LLM), Phi-3 Mini, in detecting VTE. While this model delivers competitive results, outperforming the baseline BERT models, it proves to be computationally intensive due to its larger parameter set. Our evaluation shows that the Mamba-based model demonstrates superior performance and efficiency in VTE identification, offering an effective solution to the limitations of previous approaches. Jamie Deng, Yusen Wu 0001, Yelena Yesha |
BIBM | 3 |
| 2024 | Blockchain-Based Fine-Grained Access Control for Space Resource Sharing and ManagementabstractWith the expansion of space-based observation and exploration initiatives, the necessity for secure and efficient data exchange among interconnected space resources has become more critical. This encompasses sharing information about material characteristics, geographic positioning, and the availability of resources, all of which are integral to cybersecurity measures in the space sector. This paper analyzes these challenges in-depth, focusing on developing a robust framework for secure information exchange. To address these challenges, we explore implementing advanced technologies, such as permissioned blockchain systems. Our approach aims to enhance data communication’s security, efficiency, and reliability in the space industry, making it a valuable contribution to space resource management and cybersecurity. This work is particularly relevant for organizations engaged in collaborative space missions and those handling sensitive data, offering a novel perspective on protecting critical information in the evolving landscape of space exploration and observation. Yusen Wu 0001, Alex Pissinou Makki, Kevin Padron, Stephen Dennis, Yelena Yesha |
IGARSS | 6 |
| 2023 | Improving VTE Identification through Adaptive NLP Model Selection and Clinical Expert Rule-based Classifier from Radiology ReportsabstractRapid and accurate identification of Venous thromboembolism (VTE), a severe cardiovascular condition including deep vein thrombosis (DVT) and pulmonary embolism (PE), is important for effective treatment. Leveraging Natural Language Processing (NLP) on radiology reports, automated methods have shown promising advancements in identifying VTE events from retrospective data cohorts or aiding clinical experts in identifying VTE events from radiology reports. However, effectively training Deep Learning (DL) and the NLP models is challenging due to limited labeled medical text data, the complexity and heterogeneity of radiology reports, and data imbalance. This study proposes novel method combinations of DL methods, along with data augmentation, adaptive pre-trained NLP model selection, and a clinical expert NLP rule-based classifier, to improve the accuracy of VTE identification in unstructured (free-text) radiology reports. Our experimental results demonstrate the model’s efficacy, achieving an impressive 97% accuracy and 97% F1 score in predicting DVT, and an outstanding 98.3% accuracy and 98.4% F1 score in predicting PE. These findings emphasize the model’s robustness and its potential to significantly contribute to VTE research. Jamie Deng, Yusen Wu 0001, Hilary Hayssen, Brian Englum, Aman Kankaria, Minerva Mayorga-Carlin, Shalini Sahoo, John Sorkin, Brajesh Lal, Yelena Yesha |
BIBM | 10 |
| 2021 | Classification of COVID-19 using Deep Learning and Radiomic Texture Features extracted from CT scans of Patients LungsabstractCOVID-19 is an air-borne viral infection, which infects the respiratory system in the human body, and it became a global pandemic in early March 2020. The damage caused by the COVID-19 disease in a human lung region can be identified using Computed Tomography (CT) scans. We present a novel approach in classifying COVID-19 infection and normal patients using a Random Forest (RF) model to train on a combination of Deep Learning (DL) features and Radiomic texture features extracted from CT scans of patient’s lungs. We developed and trained DL models using CNN architectures for extracting DL features. The Radiomic texture features are calculated using CT scans and its associated infection masks. In this work, we claim that the RFs classification using the DL features in conjunction with Radiomic texture features enhances prediction performance. The experiment results show that our proposed models achieve a higher True Positive rate with the average Area Under the Receiver Curve (AUC) of 0.9768, 95% Confidence Interval (CI) [0.9757, 0.9780]. Jayalakshmi Mangalagiri, Jones Sam Sugumar, Sumeet Menon, David Chapman 0001, Yaacov Yesha, Aryya Gangopadhyay, Yelena Yesha |
IEEE BigData | 7 |
| 2021 | Tolerating Adversarial Attacks and Byzantine Faults in Distributed Machine LearningabstractAdversarial attacks attempt to disrupt the training, retraining, and utilizing of artificial intelligence (AI) and machine learning models in large-scale distributed machine learning systems. This causes security risks on its prediction outcome. For example, attackers attempt to poison the model by either presenting inaccurate misrepresentative data or altering the models' parameters. In addition, Byzantine faults including software, hardware, network issues occur in distributed systems which also lead to a negative impact on the prediction outcome. In this paper, we propose a novel distributed training algorithm, partial synchronous stochastic gradient descent (ParSGD), which defends adversarial attacks and/or tolerates Byzantine faults. We demonstrate the effectiveness of our algorithm under three common adversarial attacks again the ML models and a Byzantine fault during the training phase. Our results show that using ParSGD, ML models can still produce accurate predictions as if it is not being attacked nor having failures at all when almost half of the nodes are being compromised or failed. We will report the experimental evaluations of ParSGD in comparison with other algorithms. Yusen Wu 0001, Hao Chen 0068, Xin Wang 0122, Chao Liu 0039, Yelena Yesha |
IEEE BigData | 6 |
| 2020 | Automatic Tuning of Hyperparameters for Neural Networks in Serverless CloudabstractDeep Neural Networks are used to solve the most challenging world problems. In spite of the numerous advancements in the field, most of the models are being tuned manually. Experienced Data Scientists have to manually optimize hyperparameters, such as dropout rate, learning rate or number of neurons for Big Data applications. We have implemented a flexible automatic real-time hyperparameter tuning methodology. It works for arbitrary models written in Python and Keras. We also utilized state of the art Cloud services such as trigger based serverless computing (Lambda), and advanced GPU instances to implement automation, reliability and scalability.The existing tuning libraries, such as hyperopt, Scikit-Optimize or SageMaker, require developers to provide a list of hyperparameters and the range of their values manually. Our novel approach detects potential hyperparameters automatically from the source code, updates the original model to tune the parameters, runs the evaluation in the Cloud on spot instances, finds the optimal hyperparameters, and saves the results in the No-SQL database. The methodology can be applied to numerous Big Data Machine Learning systems. Alex Kaplunovich, Yelena Yesha |
IEEE BigData | 2 |
| 2020 | Generating Realistic COVID-19 x-rays with a Mean Teacher + Transfer Learning GANabstractCOVID-19 is a novel infectious disease responsible for over 1.2 million deaths worldwide as of November 2020. The need for rapid testing is a high priority and alternative testing strategies including x-ray image classification are a promising area of research. However, at present, public datasets for COVID-19 x-ray images have low data volumes, making it challenging to develop accurate image classifiers. Several recent papers have made use of Generative Adversarial Networks (GANs) in order to increase the training data volumes. But realistic synthetic COVID-19 x-rays remain challenging to generate. We present a novel Mean Teacher + Transfer GAN (MTT-GAN) that generates COVID-19 chest x-ray images of high quality. In order to create a more accurate GAN, we employ transfer learning from the Kaggle pneumonia x-ray dataset, a highly relevant data source orders of magnitude larger than public COVID-19 datasets. Furthermore, we employ the Mean Teacher algorithm as a constraint to improve stability of training. Our qualitative analysis shows that the MTT-GAN generates x-ray images that are greatly superior to a baseline GAN and visually comparable to real x-rays. Although board-certified radiologists can distinguish MTT-GAN fakes from real COVID-19 x-rays, quantitative analysis shows that MTT-GAN greatly improves the accuracy of both a binary COVID-19 classifier as well as a multi-class pneumonia classifier as compared to a baseline GAN. Our classification accuracy is favorable as compared to recently reported results in the literature for similar binary and multi-class COVID-19 screening tasks. Sumeet Menon, Joshua Galita, David Chapman 0001, Aryya Gangopadhyay, Jayalakshmi Mangalagiri, Yaacov Yesha, Yelena Yesha, Babak Saboury, Michael Morris |
IEEE BigData | 8 |
| 2020 | Intrusion-Tolerant and Confidentiality-Preserving Publish/Subscribe MessagingabstractWe present Chios, an intrusion-tolerant publish/subscribe system which protects against Byzantine failures. Chios is the first publish/subscribe system achieving decentralized confidentiality with fine-grained access control and strong publication order guarantees. This is in contrast to existing publish/subscribe systems achieving much weaker security and reliability properties. Chios is flexible and modular, consisting of four fully-fledged publish/subscribe configurations (each designed to meet different goals). We have deployed and evaluated our system on Amazon EC2. We compare Chios with various publish/subscribe systems. Chios is as efficient as an unreplicated, single-broker publish/subscribe implementation, only marginally slower than Kafka and Kafka with passive replication, and at least an order of magnitude faster than all Hyperledger Fabric modules and publish/subscribe systems using Fabric. Sisi Duan, Chao Liu 0039, Xin Wang 0122, Yusen Wu 0001, Yelena Yesha |
SRDS | 6 |
| 2019 | Scalability Analysis of Blockchain on a Serverless CloudabstractWhile adopting Blockchain technologies to automate their enterprise functionality, organizations are recognizing the challenges of scalability and manual configuration that the state of art present. Scalability of Hyperledger Fabric is an open challenge recognized by the research community. We have automated many of the configuration steps of installing Hyperledger Fabric Blockchain on AWS infrastructure and have benchmarked the scalability of that system. We have used the UCR (University of California Riverside) Time Series Archive with 128 timeseries datasets containing over 191,177 rows of data totaling 76,453,742 numbers. Using an automated Serverless approach, we have loaded this dataset, by chunks, into different AWS instances, triggering the load by SQS messaging. In this paper, we present the results of this benchmarking study and describe the approach we took to automate the Hyperledger Fabric processes using serverless Lambda functions and SQS triggering. We will also discuss what is needed to make the Blockchain technology more robust and scalable. Alex Kaplunovich, Karuna P. Joshi, Yelena Yesha |
IEEE BigData | 3 |
| 2018 | Consolidating billions of Taxi rides with AWS EMR and Spark in the Cloud : Tuning, Analytics and Best PracticesabstractSaving nature using Big Data Analytics is a very noble goal. Using New York taxi rides data, we decided to learn how many rides could be consolidated. It was a journey we would like to share. First, we had to choose the platform for calculation between Amazon Athena, Serverless Microservices, SQL or NoSql databases, Hadoop and Spark. Then, we had to find an optimal solution for the platform using assorted tuning and optimization techniques. Although the problem seems to be straight forward, it turned out that the solution is quite challenging because of the input size, data quality, calculation complexities and numerous EMR/Spark tuning options. We have been using New York taxi data from 2009 to 2017 to quantify the rides that can be joined together. The taxi rides were consolidated based on pickup location, pickup time and drop-off location. We have been calculating the percentage of taxi rides that can be joined. The benchmark originally set was rides within five minutes with a pickup and drop-off locations within half a kilometer. Then we started experimenting with different times and locations. We have been using parquet format, parallel Scala collections, compression, filtering, new column introduction, tuning parameters, I/O overhead tuning, bucketing, timeouts and partitioning. Over 1.2 billion rides were processed using Amazon EMR with Spark. We have been optimizing calculation time and processing price. Spark has hundreds of parameters, EMR has over fifty instances to choose from. It was challenging to process our data within reasonable time. We were able to find the optimal Spark queries (plans), tested different types of joins and compared their performances. Also, we were able to compare I/O and in-memory operations during partitioning and large files manipulation (the input file sizes were hundreds of Gigabytes). The results were amazing - we could consolidate around thirty five percent of total rides, saving tons of gas and improving environment and traffic in New York City. Alex Kaplunovich, Yelena Yesha |
IEEE BigData | 2 |
| 2018 | Collaborative data mining for clinical trial analyticsabstractClinical research and drug development trials generate large amounts of data. Due to the dispersed nature of clinical trial data across multiple sites and heterogeneous databases, it remains a challenge to harness these trial data for analytics to gain more understanding about the implementation of studies as well as disease processes. Moreover, the veracity of the results from analytics is difficult to establish in such datasets. We make a two-fold contribution in this paper: First, we provide a mechanism to extract task-relevant data using Master Data Management (MDM) from a clinical trial database with data spread over several domain datasets. Second, we provide a method for validating findings by collaborative utilization of multiple data mining techniques, namely: classification, clustering, and association rule mining. Overall, our approach aims at extracting useful knowledge from data collected during clinical trials to enable the development of faster and cheaper clinical trials that more accurate and impactful. For a demonstration of the efficacy of our proposed methods, we utilized the following datasets: (1) the National Institute on Drug Abuse (NIDA) data share repository and (2) the data from the Osteoarthritis initiative (OAI), where we found real-world implications in validating the findings using multiple data mining methods in a collaborative manner. The comparative results with existing state of the art techniques show the usefulness and high accuracy of our methods. Vandana Pursnani Janeja, Jay Gholap, Prathamesh Walkikar, Yelena Yesha, Naphtali Rishe, Michael A. Grasso |
Intell. Data Anal. | 4 |
| 2017 | Cloud big data decision support system for machine learning on AWS: Analytics of analyticsabstractMachine Learning algorithms on large datasets can be executed in the Cloud. Amazon Web Services (AWS) provides over 60 different On-Demand EC2 instances [1]. The instance prices range from $0.0059 (t2.nano) to $14.4 (p2.16xlarge) per hour. We decided to build an automatic recommendation system to choose the best instance for a dataset and a machine learning algorithm to optimize time and money spent. After running multiple algorithms for different Big Data sets on assorted AWS instances and collecting the results in the NoSQL DynamoDB database, we have trained machine learning models to predict time and cost using assorted regression ML methods. Alex Kaplunovich, Yelena Yesha |
IEEE BigData | 2 |
| 2017 | Target-Based, Privacy Preserving, and Incremental Association Rule MiningabstractWe consider a special case in association rule mining where mining is conducted by a third party over data located at a central location that is updated from several source locations. The data at the central location is at rest while that flowing in through source locations is in motion. We impose some limitations on the source locations, so that the central target location tracks and privatizes changes and a third party mines the data incrementally. Our results show high efficiency, privacy and accuracy of rules for small to moderate updates in large volumes of data. We believe that the framework we develop is therefore applicable and valuable for securely mining big data. Madhu Ahluwalia, Aryya Gangopadhyay, Zhiyuan Chen 0003, Yelena Yesha |
IEEE Trans. Serv. Comput. | 4 |
| 2016 | YinMem: A distributed parallel indexed in-memory computation system for large scale data analyticsabstractMachine learning and graph analytics typically process data in an iterative way, reading the same data multiple times and sharing intermediate results across the worker nodes in cluster. Hadoop MapReduce and Spark are two popular open source cluster compute frameworks for large scale data analytics. Apache Spark is currently the state-of-the-art in-memory computation model extending MapReduce by transforming data into RDDs stored in memory. One limitation of Spark, however, lies in the fact that data transformation and distribution is implicitly managed by HDFS. Data locality is not guaranteed for iterative machine learning algorithms which read the same data multiple times. For example, data needed for operations to one worker node might reside in RDDs stored in other worker nodes. The resulting data shuffling becomes a bottleneck when iteratively reading such RDDs. We propose YinMem, a parallel distributed indexed in-memory computation system, bridging the gap between Hadoop ecosystem and HPC by replacing MapReduce with MPI while obtaining the advantage of the distributed data storage. YinMem achieves fair load balancing prior to computation for large sparse matrix by scheduling and distributing indexed data from NoSQL database to the RAM of working nodes. YinMem explores Alluxio as the in-memory storage system and enables efficient data sharing of intermediate results. Preliminary results show that YinMem has achieved 3× speedup to Spark, for computing eigenvalue and eigenvectors of a 16-million scale sparse matrix. Yelena Yesha, Milton Halem, Yaacov Yesha, Shujia Zhou |
IEEE BigData | 2 |
| 2016 | Iterative unified clustering in big dataabstractWe propose a novel iterative unified clustering algorithm for data with both continuous and categorical variables, in the big data environment. Clustering is a well-studied problem and finds several applications. However, none of the big data clustering works discuss the challenge of mixed attribute datasets, with both categorical and continuous attributes. We study an application in the health care domain namely Case Based Reasoning (CBR), which refers to solving new problems based on solutions to similar past problems. This is particularly useful when there is a large set of clinical records with several types of attributes, from which similar patients need to be identified. We go one step further and include the genomic components of patient records to enhance the CBR discovery. Thus, our contributions in this paper spans across the big data algorithmic research and a key contribution to the domain of heath care information technology research. First, our clustering algorithm deals with both continuous and categorical variables in the data; second, our clustering algorithm is iterative where it finds the clusters which are not well formed and iteratively drills down to form well defined clusters at the end of the process; third we provide a novel approach to CBR across clinical and genomic data. Our research has implications for clinical trials and facilitating precision diagnostics in large and heterogeneous patient records. We present extensive experimental results to show the efficacy or our approach. Vasundhara Misal, Vandana Pursnani Janeja, Sai C. Pallaprolu, Yelena Yesha, Raghu Chintalapati |
IEEE BigData | 4 |
| 2015 | Using Big Data to Evaluate the Association between Periodontal Disease and Rheumatoid Arthritis
Michael A. Grasso, Angela C. Comer, Dana DiRenzo, Yelena Yesha, Naphtali Rishe |
AMIA | 4 |
| 2015 | Clinico-genomic Decision Support System for Precision Diagnostics and Management
Anuja Kench, Vandana Pursnani Janeja, Amanda Niskar, Yelena Yesha, Naphtali Rishe |
AMIA | 4 |
| 2015 | Collaborative data mining for clinical trial analyticsabstractThis paper proposes a collaborative data mining technique to provide multi-level analysis from clinical trials data. Clinical trials for clinical research and drug development generate large amount of data. Due to dispersed nature of clinical trial data, it remains a challenge to harness this data for analytics. In this paper, we propose a novel method using master data management (MDM) for analyzing clinical trial data, scattered across multiple databases, through collaborative data mining. Our aim is to validate findings by collaboratively utilizing multiple data mining techniques such as classification, clustering, and association rule mining. We complement our results with the help of interactive visualizations. The paper also demonstrates use of data stratification for identifying disparities between various subgroups of clinical trial participants. Overall, our approach aims at extracting useful knowledge from clinical trial data in order to improve design of clinical trials by gaining confidence in the outcomes using multi-level analysis. We provide experimental results in drug abuse clinical trial data. Jay Gholap, Vandana Pursnani Janeja, Yelena Yesha, Raghu Chintalapati, Harsh Marwaha, Kunal Modi |
BIBM | 3 |
| 2015 | Unified framework for clinical data analytics (U-CDA)abstractIn spite of significant progress in the area of data management and integration, heterogeneous nature of clinical data makes it challenging to develop a unified view of clinical data. Therefore, a central question we are trying to address is how we can utilize data analytics to discover insightful knowledge from the scattered & large amount of clinical data to simplify clinical decision making. We propose a Unified Framework for Clinical Data Analytics (U-CDA) for mining large amounts of heterogeneous data to build enhanced clinical data analytics system. The proposed framework (U-CDA) in this paper integrates relevant clinical data from structured and unstructured data sources such as electronic health records, legacy health information system databases, clinical notes, public registries, and genomic datasets after applying necessary cleansing and transformations. It further uses intelligent, versatile data analytics engine to analyze clinical data. Jay Gholap, Vandana Pursnani Janeja, Yelena Yesha |
IEEE BigData | 3 |
| 2015 | SQL-like big data environments: Case study in clinical trial analyticsabstractBig Data deals with enormous volumes of complex and exponentially growing data sets from multiple sources. With rapid growth in technology, we are now able to generate immense amount of data in almost any field imaginable including physical, biological and biomedical sciences. With the diversity and amount of data in health care industry there is an increasing need to evaluate the components in big data frameworks and gauge their adaptability to analytics techniques. However, a key step in adapting big data tools is the portability of relational databases to big data environment. Since SQL is considered to be the de-facto language for interactive queries, in this paper, we evaluate the performance of SQL-like big data solutions for the portability of existing relational databases. Our work focuses on benchmarking multiple SQL-like big data technologies over Hadoop based distributed file system (HDFS) for Study Data Tabulation Model (SDTM) used in clinical trial databases for improving the efficiency of research in clinical trials. We use publically available clinical trial data (from National Institute on Drug Abuse (NIDA)), which follows SDTM, as a test bed to measure key parameters like usability, adaptability, modularity, robustness and efficiency of these solutions. With the intention to demonstrate how current clinical trial functionality can be replicated on a big data backend with high SQL-like functionality, we evaluate several types of ad-hoc SQL queries. Akshay Grover, Jay Gholap, Vandana Pursnani Janeja, Yelena Yesha, Raghu Chintalapati, Harsh Marwaha, Kunal Modi |
IEEE BigData | 4 |
| 2015 | A database-based distributed computation architecture with Accumulo and D4M: An application of eigensolver for large sparse matrixabstractNoSQL distributed databases have been devised to tackle the challenges resulting from volume, velocity and variety of big data. Graph representation of datasets requires efficient distributed linear algebra operations for large sparse matrix constructed from big data. Storing the transformed matrix into the database not only speeds up the big data analysis process but also facilitates the computation because of indexing. The Hadoop based approach does not natively support iterative algorithms due to data shuffling during each iteration. This paper presents a novel database-based distributed computation architecture bridging the gap between Hadoop and HPC. The novelty results from exploring the indexing capability of D4M (Dynamic Distributed Dimensional Data Model) to support linear algebra operations in a distributed computation environment. The idea is to store input data and intermediate results in associative array format inside Accumulo table to facilitate the data sharing among working nodes. pMatlab is deployed as the parallel computation engine. Our proposed architecture is proved to be lighter, easier and faster than MapReduce based approach. One example application is calculating top k eigenvalues and eigenvectors for large sparse matrix. Experiments on Graph500 benchmark datasets demonstrate 2X speedup of our architecture as compared to HEIGEN (An eigensolver for billion-scale matrices using MapReduce). Yelena Yesha, Shujia Zhou |
IEEE BigData | 2 |
| 2014 | A Scalable System for Community Discovery in Twitter During Hurricane SandyabstractThe wide use of micro bloggers such as Twitter offers a valuable and reliable source of information during natural disasters. The big volume of Twitter data calls for a scalable data management system whereas the semi-structured data analysis requires full-text searching function. As a result, it becomes challenging yet essential for disaster response agencies to take full advantage of social media data for decision making in a near-real-time fashion. In this work, we use Lucene to empower HBase with full-text searching ability to build a scalable social media data analytics system for observing and analyzing human behaviors during the Hurricane Sandy disaster. Experiments show the scalability and efficiency of the system. Furthermore, the discovery of communities has the benefit of identifying influential users and tracking the topical changes as the disaster unfolds. We develop a novel approach to discover communities in Twitter by applying spectral clustering algorithm to retweet graph. The topics and influential users of each community are also analyzed and demonstrated using Latent Semantic Indexing (LSI). Han Dong, Yelena Yesha, Shujia Zhou |
CCGRID | 3 |
| 2014 | Automating Cloud Services Life Cycle through Semantic TechnologiesabstractManaging virtualized services efficiently over the cloud is an open challenge. Traditional models of software development are not appropriate for the cloud computing domain, where software (and other) services are acquired on demand. In this paper, we describe a new integrated methodology for the life cycle of IT services delivered on the cloud and demonstrate how it can be used to represent and reason about services and service requirements and so automate service acquisition and consumption from the cloud. We have divided the IT service life cycle into five phases of requirements, discovery, negotiation, composition, and consumption. We detail each phase and describe the ontologies that we have developed to represent the concepts and relationships for each phase. To show how this life cycle can automate the usage of cloud services, we describe a cloud storage prototype that we have developed. This methodology complements previous work on ontologies for service descriptions in that it is focused on supporting negotiation for the particulars of a service and going beyond simple matchmaking. Karuna P. Joshi, Yelena Yesha, Tim Finin |
IEEE Trans. Serv. Comput. | 2 |
| 2013 | SksOpen: Efficient Indexing, Querying, and Visualization of Geo-spatial Big DataabstractWith the fast growing use of web-based map services, the performance of indexing and querying of location-based data is becoming a critical quality of service aspect. Spatial indexing is typically time-consuming and is not available to end-users. To address this challenge, we have developed and open-sourced an Online Indexing and Querying System for Big Geospatial Data, sksOpen. Integrated with the TerraFly Geospatial database [1], TerraFly sksOpen is an efficient indexing and query engine for processing Top-k Spatial Boolean Queries. Further, we provide ergonomic visualization of query results on interactive maps to facilitate the user's data analysis. Yun Lu 0001, Mingjin Zhang, Shonda Witherspoon, Yelena Yesha, Yaacov Yesha, Naphtali Rishe |
ICMLA (2) | 4 |
| 2013 | Epidemiological Data Analysis in TerraFly Geo-spatial CloudabstractGIS systems and online services are growing at a very fast pace, however, there are few online services for the analysis of geospatial epidemiology and their functionality is limited. We present a geospatial epidemiology analysis system on the TerraFly Geo-spatial Cloud platform. The system provides comprehensive spatial analysis methods and visualization. In this system, the user is not required to program in order to employ the functionality. All the datasets are stored in the Geo-spatial Cloud. This system is accessible at http://terrafly.fiu.edu/GeoCloud/. The system API algorithms adapted to geospatial epidemiology. The application utilizes the GeoCloud distributed storage system for the Big Data to be analyzed, it utilizes an interactive mapping API to display results. Huibo Wang, Yun Lu 0001, Yudong Guang, Erik Edrosa, Mingjin Zhang, Raul Camarca, Yelena Yesha, Tajana Lucic, Naphtali Rishe |
ICMLA (2) | 7 |
| 2013 | Improving Word Similarity by Augmenting PMI with Estimates of Word PolysemyabstractPointwise mutual information (PMI) is a widely used word similarity measure, but it lacks a clear explanation of how it works. We explore how PMI differs from distributional similarity, and we introduce a novel metric, PMImax, that augments PMI with information about a word's number of senses. The coefficients of PMImaxare determined empirically by maximizing a utility function based on the performance of automatic thesaurus generation. We show that it outperforms traditional PMI in the application of automatic thesaurus generation and in two word similarity benchmark tasks: human similarity ratings and TOEFL synonym questions. PMImaxachieves a correlation coefficient comparable to the best knowledge-based approaches on the Miller-Charles similarity rating data set. Lushan Han, Tim Finin, Paul McNamee, Anupam Joshi, Yelena Yesha |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2012 | Special Issue: Exploring the frontiers of computing science and technology: efficiently utilizing multicore and many-core processorsabstractThe goal of this special issue is to address such issues by assembling some of the latest researches on efficiently utilizing multicore and many-core processors in real-world applications, and their strategies for coping with those challenges. Multicore (e.g., Intel Westmere and IBM Power7) and many-core (e.g., NVIDIA Tesla and AMD FireStream graphics processing units (GPUs)) microprocessors are enabling more compute-intensive and data-intensive computation in desktop computers, clusters, and leadership supercomputers. However, efficient utilization of these microprocessors is still a very challenging issue. Their differing architectures require significantly different programming paradigms when adapting real-world applications. The actual porting costs are actively debated, as well as the relative performance between GPUs and CPUs. The invited papers in this special issue represent amplified works originally presented at the Frontiers of Multicore Computing Conference 2010, held at the University of Maryland, Baltimore County in August 2010. The selected papers cover representative research addressing the issues above. There are four papers addressing the issues related to CPUs and two on GPUs. Seelam et al. report their experiences in building and scaling enterprise business analytics benchmark, report generation, and rendering, on an IBM Power7 multicore system with eight Power7 cores and 32 hardware threads 1. The paper by Tracy and Brown presents a multithreaded physics software design to eliminate overhead associated with bodies at rest and consequently accelerate physics simulation in large, continuous virtual environments on Intel multicore processors 2. The paper by Hammond et al. presents multilevel performance analysis for the computational chemistry software, NW Chem, in Blue Gene/P and on two large-scale clusters 3. The paper by Simon et al. provides a performance evaluation and investigation of the astrophysics code, FLASH, for a variety of Intel multicore processors 4. The paper by Blattner and Yang presents the key steps in porting one data assimilation algorithm to GPU 5. The paper by Malik et al. examines the programming paradigms of Compute Unified Device Architecture (CUDA), OpenCL, The Portland Group Accelerator Compiler (PGI), and MATLAB through developing kernels from the Numerical Aerodynamics Simulation (NAS) parallel benchmarking suite 6. The guest editors of this special issue would like to express their deep gratitude to all authors, external reviewers, and Geoffrey Fox for their efforts in making this issue possible. Shujia Zhou, Yelena Yesha, Milton Halem |
Concurr. Comput. Pract. Exp. | 2 |
| 2009 | Assured Information Sharing Life CycleabstractThis paper describes our approach to assured information sharing. The research is being carried out under a MURI 9Multiuniversiyt Research Initiative) project funded by the Air Force Office of Scientific Research (AFOSR). The main objective of our project is: define, design and develop an Assured Information Sharing Lifecycle (AISL) that realizes the DoD's information sharing value chain. In this paper we describe the problem faced by the Department of Defense and our solution to developing an AISL System. Tim Finin, Anupam Joshi, Hillol Kargupta, Yelena Yesha, Joel Sachs, Elisa Bertino, Ninghui Li 0001, Chris Clifton, Eugene H. Spafford, Bhavani Thuraisingham, Murat Kantarcioglu, Alain Bensoussan 0001, Nathan Berg, Latifur Khan, Jiawei Han 0001, ChengXiang Zhai, Ravi S. Sandhu, Shouhuai Xu, Jim Massaro, Lada A. Adamic |
ISI | 4 |
| 2009 | Special Issue: Exploring the Frontiers of Computing Science and Technology: Adapting Emerging Multi- and Many-core ProcessorsabstractRecent trends in computer microprocessor development have shifted from a single powerful core to multi- and many-cores. As a result, the continuous improvement in computing power fueled by the exponentially increasing speed of a single processor could be over. With applications such as Earth and space sciences, higher resolutions and more sophisticated treatments of physical processes make models even more computationally intensive. Moreover, the drastic increase of data collected by various instruments requires a significant increase in computing power for data processing and analysis. Therefore, it is crucial for the computational science and technology community to evaluate the impacts of this shift on computationally intensive modeling and data processing applications and to develop appropriate solutions. It is known that the computing power of conventional processors is limited by memory bandwidth. Adding more cores to the processors worsens the problem. It is necessary to adapt computing algorithms to effectively utilize the computing power of those conventional multi- and many-core processors. In the last two years, there have emerged few unconventional processors: IBM's Cell Broadband Engine (hereafter referred to as Cell) and NVIDIA's Graphics Processing Unit (GPU). Intel and AMD are also developing competing Cell- or GPU-like processors, in addition to conventional multi- and many-core processors. It has been demonstrated that certain computationally intensive applications with moderate data communication can benefit from both Cell and GPU with a significant performance improvement. However, these emerging processors require new programming paradigms, which increase the porting costs and impede their effective utilization. To address such issues with the unconventional multi- and many-core processors, this special issue assembles some of the latest research on assessing the impacts of multi- and many-core processors, developing appropriate solutions, and taking advantage of the abundant computing power. The invited papers in this special issue represent augmented works originally presented at the Frontiers of Multicore Computing Conference 2008, held at the University of Maryland, Baltimore County in August 2008. Of course, it is not possible for a single issue to include all the topics addressed by the conference. However, the selected papers cover representative research addressing the issues above. There are four and two papers addressing the issues related to Cell and multi-core processors, respectively. The paper by Germann et al. presents their approaches and optimization methods in porting a short-range parallel molecular dynamics code to the first petaflops computer, Roadrunner, which uses Cells as accelerators 1. The paper by Woodward et al. reports their initial experience with porting and optimizing gas dynamics simulation on Roadrunner 2. The paper by Zhou et al. presents a study on the impact of Cell on the programming paradigm for climate and weather models 3. The paper by Brown et al. presents a contemporary artwork, the ‘Scalable City,’ accelerated with Cell 4. The paper by Wang and JaJa presents a study on interactive direct volume rendering with Intel Clovertown quad-core processors 5. The paper by Simmon and McGalliard provides benchmark results for a variety of high-performance computing applications on the Cray XT3 and XT4, which use dual- and quad-core AMD Opteron processors, and discusses and analyzes multi-core effects 6. The guest editors of this special issue would like to express their deep gratitude to all authors, external reviewers, and Geoffrey Fox for their efforts in making this issue possible. Shujia Zhou, Yelena Yesha, Milton Halem |
Concurr. Comput. Pract. Exp. | 2 |
| 2009 | Service-Oriented Atmospheric Radiances (SOAR): Gridding and Analysis Services for Multisensor Aqua IR Radiance Data for Climate StudiesabstractThe Aqua spacecraft, launched on May 4, 2002, carries two well-calibrated independent infrared (IR) grating spectrometers Atmospheric Infrared Sounder (AIRS) and Moderate Resolution Imaging Spectrometer (MODIS), which have been continuously returning upwelling IR spectral radiance measurements for over five years. Based on an Aqua Sr. Project Review, estimates of available flight fuel, power, and orbital projections assess the life span of the Aqua satellite, and these two instruments, to be reliable to 2013. Since launch, these instruments have generated petabytes of data, which are managed and made available by the Goddard Space Flight Center (GSFC) Earth Science Data and Information Services Center and GSFC MODAPS. Agencies such as NOAA, DOD, EPA, and USGS use the AIRS data mostly for weather-related applications, whereas MODIS data are used, in addition to some climate-related studies, for studies of weather, oceans, and land processes, aerosols, natural and man-made disasters, and earth ecology. The Science Investigator-led Processing Systems (SIPS) teams have made many of the desired products derived from these data sets available either as level 2 products and/or level 3 gridded product fields. However, no gridded level 3 data products of radiances, either averaged for a grid element, max, min, or as brightness temperatures (BTs), are provided directly by the SIPS. Thus, one impediment that the general community faces in accessing these MODIS produced petabytes of data is storing such large data sets, interpreting the multiformatted data, and transforming it into helpful data sets for climate-research needs. The Service-Oriented Atmospheric Radiance (SOAR) system has been designed to bridge these gaps and overcome the challenges of bringing this rich data source to the science community, by delivering applications that process these valuable radiance data into standard spatial–temporal grids as well as user-defined criteria on demand. SOAR can serve this community with aggregated, enriched, and thinned gridded data sets provided with access to the data on demand, with query and subsetting capabilities across many dimensions. In addition, SOAR provides online user-specified visualization and analysis requests, all accessible via a Web browser. The utility of SOAR is exposed via Web-service routines, using the Simple Object Access Protocol. The Web-service library and supporting technologies (Axis, PostgreSQL, and Tomcat) reside on a University of Maryland Baltimore Campus client server, which interfaces to and invokes algorithms on the process server, a high-performance computer cluster and storage system. These servers are connected to the sensor data stores at the GSFC via a high-speed fiber-optic network connection [10 Gb/s], providing reliable and fast on-demand access to a vast online library of AIRS and current monthly MODIS source data. Milton Halem, Neal Most, Curt Tilmes, Kevin Stewart, Yelena Yesha, David Chapman 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2008 | Second Space: A Generative Model for the Blogosphere
Amit Karandikar, Akshay Java, Anupam Joshi, Tim Finin, Yaacov Yesha, Yelena Yesha |
ICWSM | 6 |
| 2008 | Providing Gridded Atmospheric Radiance Products and Services from MODIS and AIRS Instruments on NASA's Aqua SatelliteabstractAtmospheric temperature and moisture profile sounding data spanning more than three decades comprise one of the longest continuous US polar orbiting satellite data records available today. NASA's Aqua satellite launched on May 4, 2002 carries two well calibrated independent IR grating spectrometers, AIRS and MODIS. Both AIRS and MODIS have been used for extensive research and are truly flagship instruments for the Earth Science community, but neither offers an operational product with gridded calibrated radiances. Radiance data are very useful and have been used for operational weather forecasting models, and to determine observed global trends in temperature, ozone, albedo, desertification, aerosols, etc. The radiances are only available in their native orbital swath form at full resolution that can be awkward and bulky to work with. Their co-location on Aqua provides a unique opportunity for common observation of the same air mass with two different instruments over a long time period. Milton Halem, Curt Tilmes, Yelena Yesha, David Chapman 0001 |
IGARSS (4) | 3 |
| 2008 | Threshold-based intrusion detection in ad hoc networks and secure AODV
Anand Patwardhan, Jim Parker 0002, Michaela Iorga, Anupam Joshi, Tom Karygiannis, Yelena Yesha |
Ad Hoc Networks | 6 |
| 2007 | On the Structure, Properties and Utility of Internal Corporate Blogs
Pranam Kolari, Tim Finin, Kelly A. Lyons, Yelena Yesha, Yaacov Yesha, Stephen G. Perelgut, Jen Hawkins |
ICWSM | 4 |
| 2007 | A Ubiquitous Context-Aware Environment for Surgical TrainingabstractThe age of technology has changed the way that surgeons are being trained. Traditional methodologies for training can include lecturing, shadowing, apprenticing, and developing skills within live clinical situations. Computerized tools which simulate surgical procedures and/or experiences can allow for "virtual" experiences to enhance the traditional training procedures that can dramatically improve upon the older methods. However, such systems do not to adapt to the training context. We describe a ubiquitous computing system that tracks low-level events in the surgical training room (e.g. student locations, lessons completed, learning tasks assigned, and performance metrics) and from these derive the training context. This can be used to create an adaptive training system. Patricia Ordóñez 0002, Palani Kodeswaran, Vlad Korolev, Wenjia Li, Onkar Walavalkar, Ben Elgamil, Anupam Joshi, Tim Finin, Yelena Yesha, I. George |
MobiQuitous | 9 |
| 2007 | Service Oriented Atmospheric Radiances (SOAR) - A Web Service Research Tool for the Gridding and Synthesis of Multi-Sensor Satellite Radiance Data for Weather and Climate Studies
Milton Halem, Curt Tilmes, Yelena Yesha, Sharon Shen, Mitchell D. Goldberg, L. H. Zhou |
WEBIST (1) | 3 |
| 2007 | A Pervasive Computing System for the Operating Room of the Future
Sheetal K. Agarwal, Anupam Joshi, Tim Finin, Yelena Yesha, Tim Ganous |
Mob. Networks Appl. | 4 |
| 2006 | A Data Intensive Reputation Management Scheme for Vehicular Ad Hoc NetworksabstractIn vehicular ad hoc networks individual vehicles can help each other locate resources and establish trustworthiness under highly dynamic conditions, lacking any centralized trust authority. To ascertain the accuracy and reliability of data aggregated in a distributed manner, we present a reputation management system for such networks that enables devices to quickly adapt to changing local conditions and provides a bootstrapping method for establishing trust relationships where only a few may exist a priori. Our scheme considers cooperativeness and accuracy of peer-provided data as two aspects of trust when evolving trust relationships and managing reputations. We use an epidemic data exchange protocol that incorporates reputation and agreement to ensure high reliability of data and stimulate proactive collaboration above and beyond stipulation, to enhance availability and reliability of data. We present preliminary simulation results which demonstrate the effectiveness of our data intensive reputation management scheme Anand Patwardhan, Anupam Joshi, Tim Finin, Yelena Yesha |
MobiQuitous | 4 |
| 2006 | Integrating service discovery with routing and session management for ad-hoc networks
Dipanjan Chakraborty 0001, Anupam Joshi, Yelena Yesha |
Ad Hoc Networks | 3 |
| 2006 | A framework for specification and performance evaluation of service discovery protocols in mobile ad-hoc networks
Avinash Shenoi, Yelena Yesha, Yaacov Yesha, Anupam Joshi |
Ad Hoc Networks | 2 |
| 2006 | Toward Distributed Service Discovery in Pervasive Computing EnvironmentsabstractThe paper proposes a novel distributed service discovery protocol for pervasive environments. The protocol is based on the concepts of peer-to-peer caching of service advertisements and group-based intelligent forwarding of service requests. It does not require a service to be registered with a registry or lookup server. Services are described using the Web Ontology Language (OWL). We exploit the semantic class/subClass hierarchy of OWL to describe service groups and use this semantic information to selectively forward service requests. OWL-based service description also enables increased flexibility in service matching. We present simulation results that show that our protocol achieves increased efficiency in discovering services (compared to traditional broadcast-based mechanisms) by efficiently utilizing bandwidth via controlled forwarding of service requests. Dipanjan Chakraborty 0001, Anupam Joshi, Yelena Yesha, Tim Finin |
IEEE Trans. Mob. Comput. | 3 |
| 2005 | Active collaborations for trustworthy data management in ad hoc networksabstractWe propose a trust-based data management framework for enabling individual devices to harness the potential power of distributed computation, storage, and sensory resources available in pervasive computing environments. Available resources include those currently present in the fixed surrounding infrastructure as well as those resources made available by other mobile devices in the vicinity. We take a holistic approach that considers trust, security, and privacy issues of data management in these environments. We focus on collaborative mechanisms to provide a platform for trustworthy data management for devices in ad hoc networks. A fundamental aspect of our framework is a pack formation mechanism for enabling collaborative peer interaction in the pervasive computing environments based on context information and landmarks. A pack provides a routing substrate for enabling devices to find reliable sources of information. A pack also provides a platform for coordinated pro-active and reactive mechanisms that can detect and respond to malicious activity. Consequently, a pack can be used for providing a foundation to distributed trust management and data intensive interactions. We describe our proposed data management framework with an emphasis on forming packs in mobile ad-hoc networks and present preliminary results from our simulation of collaborative data management using packs Anand Patwardhan, Filip Perich, Anupam Joshi, Tim Finin, Yelena Yesha |
MASS | 5 |
| 2005 | Service Composition for Mobile Environments
Dipanjan Chakraborty 0001, Anupam Joshi, Tim Finin, Yelena Yesha |
Mob. Networks Appl. | 4 |
| 2005 | Collaborative joins in a pervasive computing environment
Filip Perich, Anupam Joshi, Yelena Yesha, Tim Finin |
VLDB J. | 3 |
| 2004 | Automatic video summarization for wireless and mobile environmentsabstractIn this paper, we propose a novel video summarization technique using which we can automatically generate high quality video summaries suitable for wireless and mobile environments. The significant contribution of this paper lies in the proposed clustering scheme. We use Delaunay diagrams to cluster multidimensional point data corresponding to the frame contents of the video. In contrast to the existing clustering techniques used for summarization, our clustering algorithm is fully automatic and well suited for batch processing. We illustrate the quality of our clustering and summarization scheme in an experiment using several video clips. Yong Rao, Padmavathi Mundur, Yelena Yesha |
ICC | 3 |
| 2004 | In Reputation We Believe: Query Processing in Mobile Ad-Hoc NetworksabstractResearch on data management in mobile ad-hoc networks focuses on discovering sources and acquiring information. Mobile devices assume answers to be correct and do not verify the veracity of the information or the providers. This assumption is suitable for most client-server environments; however, peer-to-peer environments lack the intrinsic stability of "anchored" sources. In mobile ad-hoc networks, sources may provide faulty information, which can lead to incorrect conclusions. Consequently, devices need a mechanism to evaluate the integrity of their peers and the accuracy of peer provided information. To address this problem we propose a query processing model that relies on distributed trust and belief. Each device maintains and shares beliefs regarding the degree of trust it has for its peers - where trust is determined by experience and reputation. Additionally, each device associates a value indicating its belief in the accuracy of the information it holds. This knowledge is used by devices to determine the reliability of query responses. We implement our model in GloMoSim and provide experimental results for different combinations of trust and accuracy algorithms. Filip Perich, Jeffrey Undercoffer, Lalana Kagal, Anupam Joshi, Tim Finin, Yelena Yesha |
MobiQuitous | 6 |
| 2004 | A distributed service composition protocol for pervasive environmentsabstractService composition in pervasive environments enables users to utilize services in the environment to solve complex queries. Current work in development of service composition architectures focuses on wired-networked environments where solutions are centralized and tailored towards a reliable network and fixed service topology. In this paper, we present an alternate and novel design architecture of a broker-based distributed service composition protocol for pervasive environments. We present simulation results by comparing our protocol to a centralized architecture for composition. Results show that our distributed broker-based composition architecture perform better than the centralized solution in terms of composition efficiency, broker arbitration efficiency and composition radius. Dipanjan Chakraborty 0001, Yelena Yesha, Anupam Joshi |
WCNC | 2 |
| 2004 | Service Discovery in Agent-Based Pervasive Computing Environments
Olga Ratsimor, Dipanjan Chakraborty 0001, Anupam Joshi, Tim Finin, Yelena Yesha |
Mob. Networks Appl. | 5 |
| 2004 | On Data Management in Pervasive Computing EnvironmentsabstractThis paper presents a framework to address new data management challenges introduced by data-intensive, pervasive computing environments. These challenges include a spatio-temporal variation of data and data source availability, lack of a global catalog and schema, and no guarantee of reconnection among peers due to the serendipitous nature of the environment. An important aspect of our solution is to treat devices as semi-autonomous peers guided in their interactions by profiles and context. The profiles are grounded in a semantically rich language and represent information about users, devices and data described in terms of “beliefs”, “desires”, and “intentions”. We present a prototype implementation of this framework over combined Bluetooth and Ad-Hoc 802.11 networks, and present experimental and simulation results that validate our approach and measure system performance. Filip Perich, Anupam Joshi, Tim Finin, Yelena Yesha |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2003 | eNcentive: a framework for intelligent marketing in mobile peer-to-peer environmentsabstractIn recent years, the growth of Mobile Computing, Electronic Commerce and Mobile Electronic Commerce has created a new concept of Mobile Electronic Marketing. New marketing models are being developed and used to target mobile users. Mobile environments introduces new challenges that need to be overcome by these marketing models in order to be successful and effective. This paper proposes a framework, called eNcentive, which addresses many of the issues that are characteristic of mobile environments. eNcentive facilitates peer-to-peer electronic marketing in mobile ad hoc environments. Our framework employs a intelligent marketing scheme, by providing users the capability to collect information like sales promotions and discounts. Users can propagate this marketing information to other users in the network. Participating users benefit from such circulation since businesses that originally created the promotions reward the active distributors with additional promotions and other compensations. Olga Ratsimor, Tim Finin, Anupam Joshi, Yelena Yesha |
ICEC | 4 |
| 2003 | Neighborhood-Consistent Transaction Management for Pervasive Computing Environments
Filip Perich, Anupam Joshi, Yelena Yesha, Tim Finin |
DEXA | 3 |
| 2003 | Using Peer-to-Peer Data Routing for Infrastructure-Based Wireless NetworksabstractA mobile ad-hoc network is an autonomous system of mobile routers that are self-organizing and completely decentralized with no requirements for dedicated infrastructure support. Wireless infrastructure in terms of base stations is often available in many popular areas offering highspeed data connectivity to a wired network. In this paper we describe an approach where infrastructure components utilize passing by mobile nodes to route data to other devices that are out of range. In our scheme, base stations track user mobility and determine data usage patterns of users as they pass by. Based on this, base stations predict the future data needs for a passing mobile device. These base stations then collaborate (over the wired network) to identify other mobile devices with spare capacity whose routes intersect that of a needy device and use these carriers to transport the needed data. When such a carrier meets a needy device, they form ad hoc peer-to-peer communities to transfer this data. In this paper, we describe the motivation behind our approach and the different component interactions. We present the results of simulation work that we have done to validate the viability of our approach. We also describe, Numi, our framework for supporting collaborative infrastructure and ad hoc computing along with a sample application built on top of this highlighting the benefits of our proposed approach. Sethuram Balaji Kodeswaran, Olga Ratsimor, Anupam Joshi, Tim Finin, Yelena Yesha |
PerCom | 5 |
| 2003 | On Using a Warehouse to Analyze Web Logs
Karuna P. Joshi, Anupam Joshi, Yelena Yesha |
Distributed Parallel Databases | 3 |
| 2003 | Guest editorial
Vijayalakshmi Atluri, Anupam Joshi, Yelena Yesha |
VLDB J. | 3 |
| 2002 | Profile Driven Data Management for Pervasive Environments
Filip Perich, Sasikanth Avancha, Dipanjan Chakraborty 0001, Anupam Joshi, Yelena Yesha |
DEXA | 5 |
| 2002 | On experiments with a transport protocol for pervasive computing environments
Sasikanth Avancha, Vlad Korolev, Anupam Joshi, Tim Finin, Yelena Yesha |
Comput. Networks | 5 |
| 2002 | Centaurus: An Infrastructure for Service Management in Ubiquitous Computing Environments
Lalana Kagal, Vlad Korolev, Sasikanth Avancha, Anupam Joshi, Tim Finin, Yelena Yesha |
Wirel. Networks | 6 |
| 2000 | Strategies for maximizing seller's profit under unknown buyer's valuations
Bella Belegradek, Konstantinos Kalpakis, Yelena Yesha |
Inf. Sci. | 3 |
| 2000 | Relational Transducers for Electronic Commerce
Serge Abiteboul, Victor Vianu, Bradley S. Fordham, Yelena Yesha |
J. Comput. Syst. Sci. | 4 |
| 1999 | Updating and Querying Databases that Track Mobile Units
Ouri Wolfson, A. Prasad Sistla, Sam Chamberlain, Yelena Yesha |
Distributed Parallel Databases | 4 |
| 1999 | Footprint handover rerouting protocol for low Earth orbit satellite networks
Hüseyin Uzunalioglu, Ian F. Akyildiz, Yelena Yesha, Wei Yen |
Wirel. Networks | 3 |
| 1998 | Relational Transducers for Electronic CommerceabstractElectronic commerce is emerging as one of the major Websupported applications requiring database support. We introduce and study high-level declarative specifications of business models, using an approach in the spirit of active databases. More precisely, business models are specified as relational transducers that map sequences of input relations into sequences of output relations. The semantically meaningful trace of an input-output exchange is kept as a sequence of log relations. We consider problems motivated by electronic commerce applications, such as log validation, verifying temporal properties of transducers, and comparing two relational transducers. Positive results are obtained for a restricted class of relational transducers called Spocus transducers (for semi-positive outputs and cumulative state). We argue that despite the restrictions, these capture a wide range of practically significant business models. 1 Introduction Electronic commerce is emerging as a major Web-s... Serge Abiteboul, Victor Vianu, Bradley S. Fordham, Yelena Yesha |
PODS | 4 |
| 1998 | Electronic Commerce: TutorialabstractAs we embark on the information age the use of electronic information is spreading through all sectors of society, both nationally and internationally. As a result, commercial organizations, educational institutions and government agencies are finding it essential to be linked by world wide networks, and commercial Internet usage is growing at an accelerating pace. Nabil R. Adam, Yelena Yesha |
SIGMOD Conference | 2 |
| 1998 | Towards a Theory of Cost Management for Digital Libraries and Electronic CommerceabstractOne of the features that distinguishes digital libraries from traditional databases is new cost models for client access to intellectual property. Clients will pay for accessing data items in digital libraries, and we believe that optimizing these costs will be as important as optimizing performance in traditional databases. In this article we discuss cost models and protocols for accessing digital libraries, with the objective of determining the minimum cost protocol for each model. We expect that in the future information appliances will come equipped with a cost optimizer, in the same way that computers today come with a built-in operating system. This article makes the initial steps towards a thery and practice of intellectual property cost management. A. Prasad Sistla, Ouri Wolfson, Yelena Yesha, Robert H. Sloan |
ACM Trans. Database Syst. | 3 |
| 1997 | Evolving Databases: An Application to Electronic CommerceabstractMany complex and dynamic database applications such as product modeling and negotiation monitoring require a number of features that have been adopted in semantic models and databases such as active rules, constraints, inheritance, etc. Unfortunately, each feature has largely been considered in isolation. Furthermore, in a commercial negotiation, participants staking their financial well beings will never accept a system they cannot gain a precise behavioral understanding of. We attack these problems with a rich and extensible database model, evolving databases, with a clear and precise semantics based on evolving algebras (E. Borger, 1994). We also briefly describe a prototype implementation of the model (B. Fordham et al.). Bradley S. Fordham, Serge Abiteboul, Yelena Yesha |
IDEAS | 3 |
| 1997 | Pythia and Pythia/WK: Tools for the Performance Analysis of Mass Storage SystemsabstractThe constant growth on the demands imposed on hierarchical mass storage systems creates a need for frequent reconfiguration and upgrading to ensure that the response times and other performance metrics are within the desired service levels. This paper describes the design and operation of two tools, Pythia and Pythia/WK, that assist system managers and integrators in making cost-effective procurement decisions. Pythia automatically buids and solves an analytic model of a mass storage system based on a graphical description of the architecture of the system, and on a description of the workload imposed on the system. The use of a modeling wizard to perform this conversion from a graphical description of a mass storage system to an analytic model makes Pythia unique among analytic performance tools. Pythia/WK uses clustering algorithms to characterize the workload from the log files of the mass storage system. The resulting workload characterization is used as input to Pythia. © 1997 John Wiley & Sons, Ltd. Odysseas I. Pentakalos, Daniel A. Menascé, Yelena Yesha |
Softw. Pract. Exp. | 3 |
| 1997 | Analytical Performance Modeling of Hierarchical Mass Storage SystemsabstractMass storage systems are finding greater use in scientific computing research environments for retrieving and archiving the large volumes of data generated and manipulated by scientific computations. This paper presents a queuing network model that can be used to carry out capacity planning studies of hierarchical mass storage systems. Measurements taken on a Unitree mass storage system and a detailed workload characterization provided the workload intensity and resource demand parameters for the various types of read and write requests. The performance model developed here is based on approximations to multiclass Mean Value Analysis of queuing networks. The approximations were validated through the use of discrete event simulation and the complete model was validated and calibrated through measurements. The resulting model was used to analyze three different scenarios: effect of workload intensity increase, use of file compression at the server and client, and use of file abstractions. Odysseas I. Pentakalos, Daniel A. Menascé, Milton Halem, Yelena Yesha |
IEEE Trans. Computers | 4 |
| 1996 | An Analytic Model of Hierachical Mass Storage Systems with Network-Attached Storage DevicesabstractNetwork attached storage devices improve I/O performance by separating control and data paths and eliminating host intervention during data transfer. Devices are attached to a high speed network for data transfer and to a slower network for control messages. Hierarchical mass storage systems use disks to cache the most recently used files and tapes (robotic and manually mounted) to store the bulk of the files in the file system. This paper shows how queuing network models can be used to assess the performance of hierarchical mass storage systems that use network attached storage devices. The analytic model validated through simulation was used to analyze many different scenarios. Daniel A. Menascé, Odysseas I. Pentakalos, Yelena Yesha |
SIGMETRICS | 3 |
| 1996 | Scheduling Tree DAGs on Parallel Architectures
Konstantinos Kalpakis, Yelena Yesha |
Algorithmica | 2 |
| 1996 | Minimizing message complexity of partially replicated data on hypercubesabstractWithin the framework of distributed and parallel computing, we consider partially replicated data on a hypercube. We address the problem of placing copies on the hypercube in order to minimize message complexity. With realistic restrictions on the read/write ratio and the number of copies, we find the unique optimal configuration of copies. We compute the communication cost of this configuration. The optimal configuration is a linear array satisfying certain properties. © 1996 John Wiley & Sons, Inc. Keith E. Humenik, Peter Matthews, A. B. Stephens, Yelena Yesha |
Networks | 4 |
| 1996 | Guest Editors' Introduction: Special Section on Digital Libraries
Nabil R. Adam, Yelena Yesha |
IEEE Trans. Knowl. Data Eng. | 2 |
| 1995 | A Lower Bound on the Probability of Conflict Under Nonuniform Access in Database Systems
Keith E. Humenik, Peter Matthews, A. B. Stephens, Yelena Yesha |
Algorithmica | 4 |
| 1994 | The Role of the Database Community in the National Information InfrastructureabstractThe computer science community is increasingly focusing its efforts on tasks related to the National Information Infrastructure. Hardware and software advances are being sought to make wide-area networks and technological resources in general more useful to mainstream society. With this transition comes a set of unavoidable political and social issues that have never been satisfactorily dealt with in the past. Grand challenges face the computer science community on both the technical and the socio-political sides of this “major upgrade.” This paper discusses various aspects of the enormous coordination problem that faces all of us and considers what the database community can do to help. David W. Flater, Yelena Yesha |
CIKM | 2 |
| 1994 | Managing Read-Only Data on Arbitrary Networks with Fully Distributed CachingabstractIn a large information system, the amount of cache space available to store read-only replicas of data may be limited. Since acquiring these data from their sources may be an expensive and time-consuming operation, it is essential to make efficient use of the available cache space. This cache space may be unevenly distributed over a large number of loosely coupled sites. An intelligent caching strategy is needed to insure that replicas are created often enough that they can be inexpensively reached when necessary, but not so often that important data are forced out to make room. We present such a strategy, which we have developed for use in ALIBI, a networked resource discovery and information retrieval system. The TCF Strategy, as it is called, allows individual sites to adjust their level of cache turnover to provide better overall performance. This novel approach could no doubt be beneficially applied in other distributed systems which use caching. We include discussion and simulation results supporting the efficiency of the TCF Strategy. David W. Flater, Yelena Yesha |
Int. J. Cooperative Inf. Syst. | 2 |
| 1994 | Probalistic Analysis of Transaction Blocking under Arbitrary Data Access Distribution in Database System
Mukesh Singhal, Ming T. Liu, Yelena Yesha |
Inf. Sci. | 3 |
| 1994 | Guest editors' corner
Keith E. Humenik, Yelena Yesha |
J. Syst. Softw. | 2 |
| 1994 | Optimal Allocation for Partially Replicated Database Systems on Ring NetworksabstractConsiders a distributed database with partial replication of data objects located on a ring network. Certain placements of replicated objects are shown to optimize the probability of read-only success and the probability of write-only success. We also obtain optimal placements for k-terminal reliability and expected minimal path length for read-only and write-only operations.> A. B. Stephens, Yelena Yesha, Keith E. Humenik |
IEEE Trans. Knowl. Data Eng. | 2 |
| 1994 | On a Unified Framework for the Evaluation of Distributed Quorum Attainment ProtocolsabstractQuorum attainment protocols are an important part of many mutual exclusion algorithms. Assessing the performance of such protocols in terms of number of messages, as is usually done, may be less significant than being able to compute the delay in attaining the quorum. Some protocols achieve higher reliability at the expense of increased message cost or delay. A unified analytical model which takes into account the network delay and its effect on the time needed to obtain a quorum is presented. A combined performability metric, which takes into account both availability and delay, is defined, and expressions to calculate its value are derived for two different reliable quorum attainment protocols: D. Agrawal and A. El Abbadi's (1991) and Majority Consensus algorithms (R.H. Thomas, 1979). Expressions for the primary site approach are also given as upper bound on performability and lower bound on delay. A parallel version of the Agrawal and El Abbadi protocol is introduced and evaluated. This new algorithm is shown to exhibit lower delay at the expense of a negligible increase in the number of messages exchanged. Numerical results derived from the model are discussed.> Daniel A. Menascé, Yelena Yesha, Konstantinos Kalpakis |
IEEE Trans. Software Eng. | 2 |
| 1993 | Properties of Networked Information Retrieval with ALIBIabstractArticle Properties of networked information retrieval with ALIBI Share on Authors: David W. Flater University of Maryland, Baltimore County, Baltimore, MD University of Maryland, Baltimore County, Baltimore, MDView Profile , Yelena Yesha University of Maryland, Baltimore County, Baltimore, MD University of Maryland, Baltimore County, Baltimore, MDView Profile Authors Info & Claims CIKM '93: Proceedings of the second international conference on Information and knowledge managementDecember 1993 Pages 31–38https://doi.org/10.1145/170088.170098Online:01 December 1993Publication History 0citation199DownloadsMetricsTotal Citations0Total Downloads199Last 12 Months1Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access David W. Flater, Yelena Yesha |
CIKM | 2 |
| 1993 | An Efficient Management of Read-Only Data in a Distributed Information SystemabstractWhen dealing with massive amounts of primarily read-only data, significant improvements can be made over distributed DBMS for making these data available to a large network. This paper outlines some methods for heuristic query routing and cooperative caching which manage read-only replicas of data in a fully distributed manner on a network of arbitrary topology. These methods insure that query throughout increases steadily as nodes are added to the network while maintaining good response time. The resulting system is capable of providing automatic resource discovery and information retrieval over a wide area network without relying on resource directories. David W. Flater, Yelena Yesha |
Int. J. Cooperative Inf. Syst. | 2 |
| 1993 | Information and Knowledge Management: Guest Editors' Introduction
Charles K. Nicholas, Yelena Yesha |
Int. J. Cooperative Inf. Syst. | 2 |
| 1993 | Extensions to the C Programming Language for Enhanced Fault DetectionabstractAbstract The acceptance of the C programming language by academia and industry is partially responsible for the ‘software crisis’. The simple, trusting semantics of C mask many common faults, such as range violations, which would be detected and reported at run‐time by programs coded in a robust language such as Ada. Ada is a registered trademark of the U.S. Government (Ada Joint Program Office) This needlessly complicates the debugging of C programs. Although the assert macro lets programmers add run‐time consistency checks to their programs, the number of instantiations of this macro needed to make a C program robust makes it highly unlikely that any programmer could correctly perform the task. We make some unobtrusive extensions to the C language which support the efficient detection of faults at run‐time without reducing the readability of the source code. Examples of the extensions are automatic checking of error codes returned by library routines, constrained subtypes and detection of references to uninitialized and/or non‐existent array elements. David W. Flater, Yelena Yesha, E. K. Park |
Softw. Pract. Exp. | 2 |
| 1988 | A Polynomial Algorithm for Computation of the Probability of Conflicts in a Database Under Arbitrary Data Access Distribution
Mukesh Singhal, Yelena Yesha |
Inf. Process. Lett. | 2 |